upvote
Speed up loop was how we called this trick a long time ago. Guess this time we call it intelligence loop.

https://thedailywtf.com/articles/The-Speedup-Loop

reply
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.

reply
I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
reply
Really!? Glm5.3 is my daily driver and I feel im having the most productive experience with agentic collaborations so far, by a lot. Using pi with tons of custom extensions, that to be fair I developed since making the jump off of codex and claude about 12 weeks ago. I primarily do not write code for a living. I do a lot of modeling and commercial analysis and a lot of math (related to differentiable simulation)
reply
Qwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context.

Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?

I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.

I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.

So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?

reply
if the answer to 'the model is bad at X' is "you're over-reliant on it" - then yes, the model is bad at X in comparison to alternatives.
reply
Last time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.
reply
I have 2 x ChatGPT Pro 20x, Claude Max 20x, and Kimi Vivace. It's about ~12 months payback for two units and the cable.

The problem is they can't fit any frontier level open models.

reply
Is Kimi really competitive enough to have it in your mix?
reply
Last time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC.

Where'd you get 30 years from? Show your work.

reply
I would like to subscribe to your newsletter.
reply
flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans
reply
Serving compute is their main value prop

Yet… even Altman called out Anthropic for serving dumbed down models.

Shits weird man

reply
> Yet… even Altman called out Anthropic for serving dumbed down models.

Even Altman called out Anthropic? Isn't Anthropic the biggest competitor Sam Altman has?

reply
>Also I will likely save some money on subscriptions.

Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.

reply
I wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.
reply
I max out my Claude Max plan every week, and I can measure the output, and for me it's stayed fairly constant, subject to the various "bonuses" whenever Anthropic is feeling the competitive pressure.
reply
There could be gym logic at play. Hundreds signed up, 20 people actually exercising. Though it's probably more likely in the lower tiers.
reply
> to create a perceived improvement when in reality there isn’t really one?

This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.

reply
This is a good point I hadn’t considered, thank you.

Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?

Sorry if it’s a dumb question, I don’t really know much about the topic.

reply
Overfitting to benchmarks. And puff, you have the exact same effect.
reply
Far more likely it's about reducing costs.
reply
Also, in a world where there are several models competing with each other for public perception of which is best, that seems like an extremely bad move.
reply
but there is a gap between benchmarks and user feel.

Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.

reply
* Release new model that scores an arbitrary 100 on a benchmark

* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.

* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.

* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.

* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.

Bis repetitae.

reply
> * Tune it down over time so that you end up only scoring 75 on the benchmark

Where?

I see so many accusations of this happening and it's so easy to check, but nobody ever proves it.

reply
Why can't the models be benchmarked again after a few weeks/months to confirm this (likely true) theory?

I imagine some people have their own personal in depth benchmarks they could do this for.

reply
>gaslight them into thinking it never changed or that it's just a harness problem

Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).

This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.

reply
There must be some benefit if all the providers are doing it independently.

GPT5.6-Sol on Max thinking just became regarded as of a few days ago.

The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).

The cycle repeats.

reply
deleted
reply
Again, I’m out of my element here, but isn’t the entire industry dependent on “new better releases frequently”? If so, and if no one has made any meaningful breakthrough, might they all pursue this kind of deception just to stay afloat/“competitive”/relevant?

Thanks for your insight

reply
Kinda. Off the top of my head, DeepSeek and their thinking model was pretty new and interesting. Multi input models are also newish (combined input of text, image, video, audio, etc). Then there's Jev, a recently release that has a lot of people talking. It isn't really an LLM, but also is one.

Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models.

Overall models have become cheaper to run and smarter per token.

reply
deleted
reply
This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason
reply
That's not a rumour, there are countless recorded examples.
reply
You can serve Fable from a cloud vendor (like AWS, Azure). They have frozen versions of the models, so likely this should not be an issue?

I would do a test to verify my suspicions.

reply
In my experience, the API versions are as good as ever; it's the subscriptions that are severely degraded.
reply
Sounds like a good smoke test.

I’m actually so far removed from this tech that I couldn’t run such a test myself lol

reply
The Opus 4-6,4-8,5 arc is exactly this. As one person commented in here, opus 5 is a terrorist. This is undeniable. Opus 4-6 was awesome. 4-8 was worse behaviorally but produced better code.

Fable seems to be following the same enshittification arc of other Anthropic models.

Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.

reply
I’m so behind on this topic but I find it interesting how quickly things change. I feel like just yesterday I way hearing how anthropic is far and away better than OAI, and now this.

I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments

reply
> Making your product worse over time rarely plays well with customers.

On the other hand, New Coke was a resounding success. Well, it, itself wasn't, but in the aftermath, Coke outsold Pepsi 2:1.

reply
Yep, that's what they've been doing for a long while now. Also the amount of tokens you get per sub varies drastically from month to month. Needs to be regulated.
reply
The Shepard tone of "progress"
reply
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Not unless your competitors do the same, or else you will only be perceived as falling behind others.

reply
Yes that makes sense. In my hypothetical, the industry frontier is stagnating, meaning no one is making big breakthroughs, so they all resort to this.

If one lab makes a breakthrough, can the other labs just distill to bear parity anyways and then set a new baseline industry wide.

I’m quite ignorant on this topic, so if any of this sounds moronic, forgive me

reply
This would only provides a benefit if we're approaching some sort of theoretical limit of how good LLMs can be with the current approaches and data.

Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so on, all working together?

So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.

reply
planned obsolescence
reply
Artificial obsolescence!
reply
I have not been attributing it so much to malice, just that all the major cloud vendors seem to be running at full capacity, and can't build new datacenters fast enough. I just kind of assumed that as they got busy training newer models, that they allocated less resources to handle the existing systems, because they aren't able to get more capacity right now.
reply
I’m not sure why this point keeps coming up — if your service/product is so popular that it’s capacity-constrained, then the answer is to raise prices, not degrade service, because the demand should be inelastic.
reply
Raising prices also has second order effects, like consumer and business expectations around how widespread the tech can be. Valuations depend on it being reasonably affordable to roll out on a much more massive scale than today. If people get the impression that it seems too limited to very rich people (200 is affordable for a North American / Western European professional), the impression about the trajectory will change.
reply
This really depends where the load shedding point is.

A very small raise in prices may cause a very large loss in customers that you risk never getting back.

For example if customers figure out that the Chinese models are just as good, they are gone because they are so much cheaper.

reply
Exactly, it’s not a good business to be in if they’re capacity-constrained and can’t raise prices.
reply
Every SOTA model I've used at launch uses deeper, longer inference then gradually turns down over time, until the next model comes out which seems to be trained on some new data, but mostly performance due to deeper longer inference for another period.
reply
...releasing a new model that’s marginally if at all better than the original...

This isn't what we see in benchmarks.

reply
you mean like a Shepards Tone (https://en.wikipedia.org/wiki/Shepard_tone); i wouldn't doubt they slowly tweak quants to try to eke out.

there's also probably load balancers that downgrade models during high use.

reply
Nah its because they cache and preprocess requests by dumb models and send them too often to another dumb models instead of the top tier model.
reply
> For an industry that’s stagnant in progress

Yes, the AI technology is known primarily for how stagant it is.

reply
Yes, I freely admitted I was entertaining a pure hypothetical I pulled out of my butt.

I have no idea, just had a thought and put it out there

reply