upvote
What kind of evidence would satisfy you. Referring to benchmarks is apparently not sufficient because "you won't notice the difference in your everyday tasks", but saying "I do notice the difference in my everyday tasks" is just vibes.
reply
I have 500+ hours of experience building native mobile apps with AI between GPT 5.4 half a year ago and today, in addition to my regular software engineering job.

The difference between GPT 5.4 and Opus 5.5 is obvious.

What do you do, where apparently you cannot see a difference?

I honestly can't imagine, unless it's like sorting your emails.

reply
I struggle to see how a software engineer fails to see that anecdotal evidence does not matter here at all. I can say I built seven fully functioning operating systems last month, or that I am the fastest runner in the world or that my daddy is the strongest man in the world. None of it matters without data. I can say it is warmer in the living room and you can say that no, it is much warmer in the bedroon, without having something definite to measure and something accurate to measure it with, the whole discussion is pointless - just vibes. You haven't even said how you have anecdotally experienced the difference between the older or the open weight models. So excuse me, but then you will fail to convince people on your claims, especially when the whole discussion seems to be in the middle of some bloody information warfare at the moment.
reply
This you?

Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough.

reply
The next sentence:

>It depends on the task you use them for and how you measure the output. For most tasks you really do not need frontier capability.

reply
It sounds like you measure it with vibes, but demand others produce benchmarks (which are easily found if you actually care).
reply
I have similar experience than you, with a big caveat, opus 5.5 would have still broken badly a custom ERP.

And on that specific kind of software, ultimately a big CRUD, there really isn't that much of a difference between GLM5.3 and Opus/OpenAI.

You see the differences when you get to different class of software.

I also have data entry applications that use LLM to actually parse documents, it's all Chinese models self hosted because the economic calculus beated a hosted API by about 5x

reply
That's fair enough.

For mobile apps, I find that nowadays with Opus 5.5 the UI looks better, the UX is better, it can implement more tricky animations and gestures, and it can do all of that with far fewer iterations and feedback than eg. GPT 5.5 would have required.

Also vision capabilities were improved significantly with GPT 6 Astra or Opus 5.5, even compared to GPT 5.6 Sol.

There was no way the old models such as GPT 5.4 would have done a comparable job when asked to align an implementation to a visual reference.

Even for basic websites with no interactive functionality, this should make a significant difference.

reply
I even find that for what amounts to a web search Gemini 3.8 flash is better than the big models, stuff like the links to updated tax codes in different countries, or whatever MS is doing with Azure/o365 or AWS with their plethora of products.

I think use cases are the real reason why people have such different experiences, I too find that Opus5.5/Astra/6 are better for UI/UX now, it wasn't the case a year ago, at some point Gemini pro 3.5 was the best one at that.

That's also why I use all of them and try to not be locked to a single harness as well.

reply
[dead]
reply