upvote
> This year it became common for people to entirely delegate coding to AI

This has been the case for around 2 years now, more reliably - a year. We've mostly stayed there since then.

Saying that more people started doing it isn't indicative of significant improvement. Some people just started doing it later.

I can't speak about math because I haven't used AI for that application, but I know that there hasn't been any significant advancement in coding in this year on base models. There has been more RL work, more harness work, more tools, they all expanded some capabilities like cyber or orchestration or tool use, but raw intelligence of base models is no longer where the main focus is.

reply
> This has been the case for around 2 years now, more reliably - a year.

I have to disagree with this pretty strongly. Opus 4.5 needed a lot of handholding not to work itself into a corner pretty quickly. Fable I basically never need to correct, and I've most become a data source.

reply
What are you talking about - I feel like we’re living in parallel realities. If I had to go back to opus 4.5 tomorrow I’d be hugely upset and significantly slowed down
reply
I'm not. Yes we normalized 1m context window and models tend to hallucinate less.

But models have been somewhat stagnant since Opus 4.6/7.

And in some regards there were even regressions like Claudeisms that are load bearing.

reply
Yes these guys are completely delusional.

2 years ago a model could barely solve the AMC, 1 year ago it got IMO gold, and this year models have solved multiple millenium problems.

Even 1 year ago ai code was just unusable (claude code only became available 1.5 years ago!) and now basically everyone I know from independent shops all the way to faang and anthropic/openai themselves exclusively use some AI agent to code.

Why does HN continue to delude itself that "models are not improving?" Maybe for the simple things they care about its "roughly the same," but they are _clearly_ improving.

reply