upvote
Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

It does work out to be a similar cost per task though

reply
You should probably look at the cost/score graph by effort level instead:

https://artificialanalysis.ai/models/claude-opus-5-5#intelli...

It is most of the pareto frontier.

reply
Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort
reply
so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)
reply
Is verboseness the only measure of token efficiency towards overall task completion?
reply
Disagree. Our internal company tests showed a cost per task drop from 0.35usd to 0.16usd . Opus 5low vs opus 5.5 low
reply
Very fast you were.
reply
Even created an account to tell us just that.
reply
I don't think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:

https://artificialanalysis.ai/models/claude-opus-5-5?models=...

reply
I tested it with Claude Code, and I can confirm it's way cheaper, better, faster and less verbose than Opus 5.
reply
parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
reply