It's always some sort of "I don't notice the difference".
And honestly, if you don't see a difference between the SOTA from 6 months ago, which would be GPT 5.4, and today's Opus 5.5, you would have to be downright blind. Not sure what else to say - the results are obviously different for any kind of meaningful output.
> Also how do we know how much of these “big improvements” come from the harness and tooling rather than the raw capability of the model?
By simply running the old models in the latest harness. Which none of the people who argue "it's all the harness" ever do.
The difference between GPT 5.4 and Opus 5.5 is obvious.
What do you do, where apparently you cannot see a difference?
I honestly can't imagine, unless it's like sorting your emails.
Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough.
>It depends on the task you use them for and how you measure the output. For most tasks you really do not need frontier capability.
And on that specific kind of software, ultimately a big CRUD, there really isn't that much of a difference between GLM5.3 and Opus/OpenAI.
You see the differences when you get to different class of software.
I also have data entry applications that use LLM to actually parse documents, it's all Chinese models self hosted because the economic calculus beated a hosted API by about 5x
For mobile apps, I find that nowadays with Opus 5.5 the UI looks better, the UX is better, it can implement more tricky animations and gestures, and it can do all of that with far fewer iterations and feedback than eg. GPT 5.5 would have required.
Also vision capabilities were improved significantly with GPT 6 Astra or Opus 5.5, even compared to GPT 5.6 Sol.
There was no way the old models such as GPT 5.4 would have done a comparable job when asked to align an implementation to a visual reference.
Even for basic websites with no interactive functionality, this should make a significant difference.
I think use cases are the real reason why people have such different experiences, I too find that Opus5.5/Astra/6 are better for UI/UX now, it wasn't the case a year ago, at some point Gemini pro 3.5 was the best one at that.
That's also why I use all of them and try to not be locked to a single harness as well.