On what task? By who? On what benchmark? How do you measure in you own workflow the “betterness” or “more goodness” of these or any models? If you don’t say those things you’re just writing a bad ad copy.
> It's plausible that open models are 6 - 12 months behind, and there is no "good enough".
Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough. It depends on the task you use them for and how you measure the output. For most tasks you really do not need frontier capability. Also how do we know how much of these “big improvements” come from the harness and tooling rather than the raw capability of the model?
It's always some sort of "I don't notice the difference".
And honestly, if you don't see a difference between the SOTA from 6 months ago, which would be GPT 5.4, and today's Opus 5.5, you would have to be downright blind. Not sure what else to say - the results are obviously different for any kind of meaningful output.
> Also how do we know how much of these “big improvements” come from the harness and tooling rather than the raw capability of the model?
By simply running the old models in the latest harness. Which none of the people who argue "it's all the harness" ever do.
The difference between GPT 5.4 and Opus 5.5 is obvious.
What do you do, where apparently you cannot see a difference?
I honestly can't imagine, unless it's like sorting your emails.
Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough.
>It depends on the task you use them for and how you measure the output. For most tasks you really do not need frontier capability.
And on that specific kind of software, ultimately a big CRUD, there really isn't that much of a difference between GLM5.3 and Opus/OpenAI.
You see the differences when you get to different class of software.
I also have data entry applications that use LLM to actually parse documents, it's all Chinese models self hosted because the economic calculus beated a hosted API by about 5x
For mobile apps, I find that nowadays with Opus 5.5 the UI looks better, the UX is better, it can implement more tricky animations and gestures, and it can do all of that with far fewer iterations and feedback than eg. GPT 5.5 would have required.
Also vision capabilities were improved significantly with GPT 6 Astra or Opus 5.5, even compared to GPT 5.6 Sol.
There was no way the old models such as GPT 5.4 would have done a comparable job when asked to align an implementation to a visual reference.
Even for basic websites with no interactive functionality, this should make a significant difference.
I think use cases are the real reason why people have such different experiences, I too find that Opus5.5/Astra/6 are better for UI/UX now, it wasn't the case a year ago, at some point Gemini pro 3.5 was the best one at that.
That's also why I use all of them and try to not be locked to a single harness as well.
If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?
I think there would be a market and I think it would be large.
So, I agree.
Unless you're doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.
There's people at my job that get so much more done than everyone else using Fable/Opus/Astra. and all they use is the fastest cheapest models. I'd say the people who are using sota models for everything are doing it just because they prefer to be lazy.
You simply do not need these frontier models, they outgrew most people's needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.
However I do actually have a project where I need the frontier models--I'm working on a deep learning project of moderate complexity (something novel/state of the art within its domain, adapting a known approach from published research in a related domain). The difference from Opus 5 -> Opus 5.5 was huge for my project. Opus 5 was struggling, Opus 5.5 is doing really well.
I think the demand for frontier models will continue to be there, at least for a subset of tasks, although I agree that it is probably going to shrink as the non-frontier becomes more and more capable.
Certainly, there's a real market for it too, with people who would actually use its advanced capabilities, and see the 100x price as worth it.
But sure, even a mostly-FOMO market is still a market. If people are paying, people are paying.
99% of everything is CRUD LoB apps.
Not asking to be mean, I just genuinely dont know why you'd need the frontier for basic applications.
I cannot trust current models to find all the necessary context, or to make what I consider to be good trade offs. A much more capable model would be able to see my existing patterns (or at least not have context rot make them blind to my convention docs) and make trade offs I agree with much more consistently, and I'd be able to do more with my time.
I've actually found models to be pretty poor at driving things I don't know well, so I generally don't do that unless its general design/product exploration and the end product code is throw-away.
Except being priced out.
The big labs' financials are based on their products being used widely by a lot of the general public. If it turns out that they're actually selling a premium product to premium-product consumers at a premium price point (while everyone else buys DeepSeek-like cheaper/worse products), that's a big issue for them.
If a consumer computer hardware company launched by promising investors that it'd be the next Dell/HP and it turned out to be the next Apple (talking Macs here, not phones or apps/services), that'd be an issue for them too.
In actual day to day development the differences are a lot harder to spot. Maybe deepseek is worse, but I asked it to run until it was able to launch itself and verify it worked as expected, and it did. Maybe it wasted some turns, idk, but when it said it was done, it was done.
I have no doubt there's things it's worse at, but what percentage of development is truly novel?
Was a night and day difference going directly to deepseek api
even their harnesses are far surpassed by pi and opencode at this point
also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days
Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.
I've seen different benchmarks come to different conclusions
Benchmarking these models must be an incredibly complex and difficult problem
How can a lay person know which benchmarks actually have good signal?