upvote
> These open models still did not beat February's Mythos / Fable 5.

On what task? By who? On what benchmark? How do you measure in you own workflow the “betterness” or “more goodness” of these or any models? If you don’t say those things you’re just writing a bad ad copy.

> It's plausible that open models are 6 - 12 months behind, and there is no "good enough".

Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough. It depends on the task you use them for and how you measure the output. For most tasks you really do not need frontier capability. Also how do we know how much of these “big improvements” come from the harness and tooling rather than the raw capability of the model?

reply
Did you see any person claiming an open model was more intelligent than Fable 5?

It's always some sort of "I don't notice the difference".

And honestly, if you don't see a difference between the SOTA from 6 months ago, which would be GPT 5.4, and today's Opus 5.5, you would have to be downright blind. Not sure what else to say - the results are obviously different for any kind of meaningful output.

> Also how do we know how much of these “big improvements” come from the harness and tooling rather than the raw capability of the model?

By simply running the old models in the latest harness. Which none of the people who argue "it's all the harness" ever do.

reply
Ok so you have nothing, just vibes. Fair enough.
reply
What kind of evidence would satisfy you. Referring to benchmarks is apparently not sufficient because "you won't notice the difference in your everyday tasks", but saying "I do notice the difference in my everyday tasks" is just vibes.
reply
I have 500+ hours of experience building native mobile apps with AI between GPT 5.4 half a year ago and today, in addition to my regular software engineering job.

The difference between GPT 5.4 and Opus 5.5 is obvious.

What do you do, where apparently you cannot see a difference?

I honestly can't imagine, unless it's like sorting your emails.

reply
I struggle to see how a software engineer fails to see that anecdotal evidence does not matter here at all. I can say I built seven fully functioning operating systems last month, or that I am the fastest runner in the world or that my daddy is the strongest man in the world. None of it matters without data. I can say it is warmer in the living room and you can say that no, it is much warmer in the bedroon, without having something definite to measure and something accurate to measure it with, the whole discussion is pointless - just vibes. You haven't even said how you have anecdotally experienced the difference between the older or the open weight models. So excuse me, but then you will fail to convince people on your claims, especially when the whole discussion seems to be in the middle of some bloody information warfare at the moment.
reply
This you?

Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough.

reply
The next sentence:

>It depends on the task you use them for and how you measure the output. For most tasks you really do not need frontier capability.

reply
It sounds like you measure it with vibes, but demand others produce benchmarks (which are easily found if you actually care).
reply
I have similar experience than you, with a big caveat, opus 5.5 would have still broken badly a custom ERP.

And on that specific kind of software, ultimately a big CRUD, there really isn't that much of a difference between GLM5.3 and Opus/OpenAI.

You see the differences when you get to different class of software.

I also have data entry applications that use LLM to actually parse documents, it's all Chinese models self hosted because the economic calculus beated a hosted API by about 5x

reply
That's fair enough.

For mobile apps, I find that nowadays with Opus 5.5 the UI looks better, the UX is better, it can implement more tricky animations and gestures, and it can do all of that with far fewer iterations and feedback than eg. GPT 5.5 would have required.

Also vision capabilities were improved significantly with GPT 6 Astra or Opus 5.5, even compared to GPT 5.6 Sol.

There was no way the old models such as GPT 5.4 would have done a comparable job when asked to align an implementation to a visual reference.

Even for basic websites with no interactive functionality, this should make a significant difference.

reply
I even find that for what amounts to a web search Gemini 3.8 flash is better than the big models, stuff like the links to updated tax codes in different countries, or whatever MS is doing with Azure/o365 or AWS with their plethora of products.

I think use cases are the real reason why people have such different experiences, I too find that Opus5.5/Astra/6 are better for UI/UX now, it wasn't the case a year ago, at some point Gemini pro 3.5 was the best one at that.

That's also why I use all of them and try to not be locked to a single harness as well.

reply
[dead]
reply
I run my own benchmarks at the interface of math and coding (ML, math-opt, stochastic models), and keep my own set of benchmarks, i.e. somewhat niche linear algebra and numerical recipes. DS-4.1 is the first open model that passes most of these benchmarks, and GPT-5.6-Sol is definitely not ahead. Even if it's all distillation, then DS does a formidable job at it.
reply
I was thinking about this earlier today and I came to the following question:

If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?

I think there would be a market and I think it would be large.

So, I agree.

reply
Many people would, and you'll find that they're building crappy webapps where you dont need SoTA. Like seriously who needs these frontier models?

Unless you're doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.

There's people at my job that get so much more done than everyone else using Fable/Opus/Astra. and all they use is the fastest cheapest models. I'd say the people who are using sota models for everything are doing it just because they prefer to be lazy.

You simply do not need these frontier models, they outgrew most people's needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.

reply
I agree that for average web dev tasks the open models are already good enough. I've had good experiences with both DeepSeek and GLM. And these models are just better for anything security related since they don't throw massive hissy fits.

However I do actually have a project where I need the frontier models--I'm working on a deep learning project of moderate complexity (something novel/state of the art within its domain, adapting a known approach from published research in a related domain). The difference from Opus 5 -> Opus 5.5 was huge for my project. Opus 5 was struggling, Opus 5.5 is doing really well.

I think the demand for frontier models will continue to be there, at least for a subset of tasks, although I agree that it is probably going to shrink as the non-frontier becomes more and more capable.

reply
I think the market would be huge, especially if it's the "can complete a large task in 3 turns instead of 15" kind of smart. Lots of people and companies would pay for quality + speed.
reply
I think there would be a market, but it would mostly be a FOMO market. That is, people would be doing tasks on it that the "regular" model is more than capable of handling, because they're afraid they're leaving something on the table by not using the absolute best option.

Certainly, there's a real market for it too, with people who would actually use its advanced capabilities, and see the 100x price as worth it.

But sure, even a mostly-FOMO market is still a market. If people are paying, people are paying.

reply
What does 10x more capable look like now? Surely at some point we will reach an asymptote of what can be done purely digitally: all useful coding tasks can be automated, most math research, etc. At some point the physical world becomes the dyke holding back the singularity; until these genius models can scale their investigation into physical experiments and manufacturing, the future will have arrived only in the digital world.
reply
I think you're making a mistake in thinking the digital world is the only one reachable to AI. Robotics and sensing would be opened up by a sufficiently capable AI.
reply
I agree: large market. I think the future of these frontier labs is selling exceptionally powerful and exceptionally expensive models. They'll be used for precision, high value tasks. The rest of us will be happy with good enough and cheap models.
reply
Doing what? How many jobs involve solving Millennium Prize math challenges?

99% of everything is CRUD LoB apps.

reply
I am coding CRUD apps with a mix of astra, sol 6.1, fable and opus 5.5. A more capable model would still benefit me imo. Being able to follow high level guidance better, and being able to harness other models for each task would be a big improvement.
reply
Do you know how what you're doing, or do you find yourself working on things you dont understand and need the best model because it's the only way to push your own capabilities (because you're avoiding learning how to do the thing yourself)?

Not asking to be mean, I just genuinely dont know why you'd need the frontier for basic applications.

reply
It's a matter of bandwidth. The more I can offload onto the model, the more I can accomplish. For example, I had to do a lot of security work over the last 2 weeks to get ready for an event. This requires handholding current models on many fronts, like: 1) Do they actually implement the security fixes correctly. 2) Do their fixes create any new edge cases. 3) Do their fixes compromise existing interfaces or API surfaces.

I cannot trust current models to find all the necessary context, or to make what I consider to be good trade offs. A much more capable model would be able to see my existing patterns (or at least not have context rot make them blind to my convention docs) and make trade offs I agree with much more consistently, and I'd be able to do more with my time.

I've actually found models to be pretty poor at driving things I don't know well, so I generally don't do that unless its general design/product exploration and the end product code is throw-away.

reply
> leading labs have nothing to fear.

Except being priced out.

The big labs' financials are based on their products being used widely by a lot of the general public. If it turns out that they're actually selling a premium product to premium-product consumers at a premium price point (while everyone else buys DeepSeek-like cheaper/worse products), that's a big issue for them.

If a consumer computer hardware company launched by promising investors that it'd be the next Dell/HP and it turned out to be the next Apple (talking Macs here, not phones or apps/services), that'd be an issue for them too.

reply
Perhaps on certain benchmarks and for certain work, but anecdotally I've not been able to see a difference between it and Opus on a lot of dev work (web, Go, iOS/AppleTV native, scripting, general tasks)
reply
I think it's fair to say that an open model has not yet surpassed Fable 5. But I don't think it's fair to count the period prior to Fable's public release. In which case it's only been 4 months since the release of a model that's clearly ahead of today's top open models.
reply
Imo deepseek 4.1 (and a lot of the cheaper models, Luna is similar) show the issues in benchmarks. At this point.

In actual day to day development the differences are a lot harder to spot. Maybe deepseek is worse, but I asked it to run until it was able to launch itself and verify it worked as expected, and it did. Maybe it wasted some turns, idk, but when it said it was done, it was done.

I have no doubt there's things it's worse at, but what percentage of development is truly novel?

reply
Personally I find this shocking. I don't think our application is that complicated, (typescript full stack graphql reactnative etc) but deepseek 4.1 flash is a bumbling fool, junior-level at best, who takes a very long time to make a very big mess. Opus 5.5 one shots truly impressive code in 5 minutes, while deepseek 4.1 flash takes 20 minutes to do horribly. I simply don't understand how folks claim they get good engineering out of it. Maybe we still care enough about the fundamentals to notice the mess ...
reply
Are you using OpenRouter? I’m honestly surprised open model labs haven’t been calling them out, but heaps of providers either silently serve heavily quant versions, or don’t have inference set up correctly and don’t run the model properly.

Was a night and day difference going directly to deepseek api

reply
In this case, I was using Hugging Face inference set to one of a few American providers. For my professional work, I do not use Chinese inference.
reply
The benchmarks don’t mean shit. Opus 5 was a terrible model and yet it had very impressive benchmarks. All the labs are benchmaxxed to the tits, only open models’ benchmarks are even worth paying attention to because they literally cannot cheat.
reply
I have a different experience and find them just as good as the latest anthropic and openai models. But then, we probably can just have opinions on this as benchmarks are probably used for marketing to a large degree. I find it hard to find truly independent benchmarks that don't have any ties to the cash flow of openai and anthropic, can not be trained up on etc.
reply
Agree, will see after the dilution and CoT hack fixed, will they keep the pace now. MiMo had some good numbers recently because it's discovered that the post evaluation RL directly exposes answers to models, so RL and evaluation is runied.
reply
There absolutely is "good enough" and I agree with this author: DeepSeek 4.1 Flash is plenty good enough for all the things I would trust an AI to do at my job.
reply
Yes, trust is exactly the point. I trust the frontier models to tackle bigger chunks of work than the open weight models, and get it right.
reply
that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol

even their harnesses are far surpassed by pi and opencode at this point

also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days

reply
>> that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol

Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.

reply
I see people throw around these benchmark numbers

I've seen different benchmarks come to different conclusions

Benchmarking these models must be an incredibly complex and difficult problem

How can a lay person know which benchmarks actually have good signal?

reply
[flagged]
reply
Got it. Thanks for your input I guess.
reply
[flagged]
reply
[dead]
reply