Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?
I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.
I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.
So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?
The problem is they can't fit any frontier level open models.
Where'd you get 30 years from? Show your work.
Yet… even Altman called out Anthropic for serving dumbed down models.
Shits weird man
Even Altman called out Anthropic? Isn't Anthropic the biggest competitor Sam Altman has?
Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?
Sorry if it’s a dumb question, I don’t really know much about the topic.
Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
Bis repetitae.
Where?
I see so many accusations of this happening and it's so easy to check, but nobody ever proves it.
I imagine some people have their own personal in depth benchmarks they could do this for.
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.
GPT5.6-Sol on Max thinking just became regarded as of a few days ago.
The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).
The cycle repeats.
Thanks for your insight
Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models.
Overall models have become cheaper to run and smarter per token.
I would do a test to verify my suspicions.
I’m actually so far removed from this tech that I couldn’t run such a test myself lol
Fable seems to be following the same enshittification arc of other Anthropic models.
Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.
I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments
On the other hand, New Coke was a resounding success. Well, it, itself wasn't, but in the aftermath, Coke outsold Pepsi 2:1.
Not unless your competitors do the same, or else you will only be perceived as falling behind others.
If one lab makes a breakthrough, can the other labs just distill to bear parity anyways and then set a new baseline industry wide.
I’m quite ignorant on this topic, so if any of this sounds moronic, forgive me
Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so on, all working together?
So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.
A very small raise in prices may cause a very large loss in customers that you risk never getting back.
For example if customers figure out that the Chinese models are just as good, they are gone because they are so much cheaper.
This isn't what we see in benchmarks.
there's also probably load balancers that downgrade models during high use.
Yes, the AI technology is known primarily for how stagant it is.
I have no idea, just had a thought and put it out there