I'm shocked anyone could conclude this. This year it became common for people to entirely delegate coding to AI (I know many competent programmers/researchers who do this now). Progress in math has just been insane. An internal model at OAI just resolved one of the most celebrated open problems in mathematics. If anything, progress has accelerated.
This has been the case for around 2 years now, more reliably - a year. We've mostly stayed there since then.
Saying that more people started doing it isn't indicative of significant improvement. Some people just started doing it later.
I can't speak about math because I haven't used AI for that application, but I know that there hasn't been any significant advancement in coding in this year on base models. There has been more RL work, more harness work, more tools, they all expanded some capabilities like cyber or orchestration or tool use, but raw intelligence of base models is no longer where the main focus is.
I have to disagree with this pretty strongly. Opus 4.5 needed a lot of handholding not to work itself into a corner pretty quickly. Fable I basically never need to correct, and I've most become a data source.
But models have been somewhat stagnant since Opus 4.6/7.
And in some regards there were even regressions like Claudeisms that are load bearing.
2 years ago a model could barely solve the AMC, 1 year ago it got IMO gold, and this year models have solved multiple millenium problems.
Even 1 year ago ai code was just unusable (claude code only became available 1.5 years ago!) and now basically everyone I know from independent shops all the way to faang and anthropic/openai themselves exclusively use some AI agent to code.
Why does HN continue to delude itself that "models are not improving?" Maybe for the simple things they care about its "roughly the same," but they are _clearly_ improving.
Have they? That seems like quite a claim given the last 6 months, particularly for cybersecurity.
That is a much easier catch up game. GLM 5.3 and DeepSeek flash 4.1 also demonstrate significant jump in cyber capabilities. So yeah, it is a slowdown in the place where it matters. RL has been around for ages, there's no moat there if you already have a good enough pretrain.
OpenAI just solved Navier-Stokes.
Seems like the US is on a takeoff ramp to me.
Astra can confidently one-shot 500k lines of slop, with 800k lines of tests covering it, without testing a single intended product requirement, and none of it actually working.
All models require hand holding. Fable and Astra are no exceptions. The difference is only in the amount of hand holding required, and there's essentially no gap here anymore between American and Chinese models.
I only use Chinese models sparingly because American models are so much cheaper with subscriptions, that it doesn't make economic sense to not use them. If/when that changes, I could simply route to cheapest model that's available at the moment and I wouldn't notice much difference in most applications.
1. Navier Stokes was plagiarism
2. All benchmarks were misleading wrong and incorrect
3. All other mathematical advances were again hype
4. HF incident was marketting ploy jointly coordinated by HF, METR and OpenAI (and also Anthropic)
5. Anthropic's HF like incident was again a marketing ploy [1]
Nothing ever happens. This whole thing is a scam. Everything is done to fool you and you have fallen for it. Congrats.
[1] https://www.anthropic.com/research/investigating-incidents-c...
This is obviously untrue… do you use any of them?
Yes, I do use them, quite heavily. The only difference at this point is in benchmarks that can't be trusted (see: artificial analysis on Astra), or in the way models communicate.
Most gains are now from RL, which for some reason is hyperfocused on improving cyber capabilities, and harnesses. Raw intelligence gains of base models is absolutely slowing down.
AI in general is just hype and unprofitable and all these companies are playing marketing tricks before the IPO after which they will cash out and let the economy crash.
This is legit what a lot of people think. To continue this narrative, they have to keep up the charade of "things are not improving".Downplaying the latest models capabilities is frankly insane considering what we’ve seen what OpenAI’s models have done without safeguards. That wasn’t possible before this latest generation.
And yet, they historically did agree on the existence of AI risk, since before OpenAI was even founded.
I mean it is literally economy 101: some capitalists getting on the top using free market, and then try to use government to remove free market so their top position were secured from any competitors.
This seems like a case of "save me from my own mistakes/ambition"
Duh! It’s called collusion. They want to try and hoard the technology for themselves if possible!