upvote
Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.

But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.

reply
at that point, you can switch to dirt cheap open models
reply
At low and medium effort it is 1/3 cheaper, at high it’s a step above Opus/low. It only looks worse at xhigh.
reply
That makes sense. I'm interested in seeing where Haiku 5.5 comes in then when it gets released. It feels like the low intelligence / fast niche will be covered there.
reply
i'd love to see them re-enter that space but given haiku 5 never happened I wouldn't bet on it

i think they see what openai charges for luna and just don't want to try and compete

reply
They did mention in the Opus 5.5 announcement blogpost that Sonnet and Haiku 5.5 will follow soon.
reply
But they literally stated that they would release Sonnet 5.5 and Haiku 5.5 after Opus 5.5 was released
reply
Haiku 5.5 is DOA without a massive price cut. Luna is literally 10x cheaper at current pricing
reply
It depends entirely on its capabilities. If it is significantly smarter than Luna, which frankly is quite likely, then a lot of people won't mind paying more.
reply
they've already said in both the opus 5.5 and sonnet 5.5 blog posts that haiku 5.5 is coming
reply
There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.

Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.

reply
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?

This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.

reply
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.

With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.

reply
I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.

Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.

I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.

reply
just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.
reply
t/s maybe? IDK, because their token speed comparison was against Sonnet 5.
reply
It literally does not?
reply
[dead]
reply