I summarized each into new fable 5.1 sessions, and both seem to have arrived at reasonable solutions that only need a few nits revised before they are commit worthy.
But yes, we might end up hitting the issue of "most jobs aren't solving hard problems" increasingly. The bigger potential benefit is higher trustworthiness, reliability/thoroughness, and squishy human things; people will likely continue to pay large premiums for those. "Solve it well and save time, long term". Those can be harder to see on a benchmark.
I get your point, but we can only have groundbreaking leaps once in a blue moon. That doesn’t mean incremental improvements aren’t useful.
The problem frontier LLMs face is that they are hundreds of billions to trillions of dollars short of finding that market that's big enough to sustain capex commitments and further product development. If they don't find something groundbreaking, they are going to have a very painful year next year, maybe even starting this year for some of them and their data center partners.
Anthropic and OpenAI can't afford to live in a world where LLMs are at or near the top of their S curve.