upvote
The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
reply
> The magnitude of improvement in unverifiable domains is small,

What makes you say that? What is an example of a domain where the improvement is small?

I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.

reply
My argument is that if you were to compare "analyze XYZ geopolitical situation" or "explain the ramifications of XYZ law" from Opus 3.5, 4.5 and 5.5, the difference would be marginal, at least for 4.5 - 5.5. Almost all the crazy capabilities newer models have is from RLVR variants, whereas capabilities driven by RLHF are inching along.
reply
How are the models making politics better? I don't count AI attack ads as an improvement.
reply
Is this a serious question?

Improvement in this context means "better quality results".

You can use better quality models to do worse things with.

I'm not making any claim about second order effects like that.

reply
Yeah it's a serious question. What are the better quality results in politics from AI?
reply