upvote
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
reply
You're right about its real world performance, and I worded my original comment wrongly.

I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.

reply
Claude, is that you?
reply
you're absolutely right to push back
reply
Your clarification makes sense. The distinction between overall benchmark performance and why Terminal-Bench is an outlier is important
reply
> You're right about its real world performance, and I worded my original comment wrongly.

Damn, HN commenters starting to talk in claudisms now

reply
this is human writing...

this is claude writing...

corporate needs you to find the difference

reply
Well presumably now it’ll fall back to Sonnet 5.5 lol
reply
I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.

Anthropic made it that way, and I'd say the lower score is accurate.

reply
I believe you meant to cite the Opus 5.5 System Card which states:

> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).

> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...

I cannot find a Sonnet 5.5 system card.

reply
And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!
reply
it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

reply
So I suppose the easy fix for Anthropic would be to have Opus 5.5 now fall back to Sonnet 5.5, right?
reply
That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?

I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?

reply
Fallback was usually Opus 4.8
reply
Isn't that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?
reply
this "feature" is one of the primary that caused me to cancel and move to exclusively open weight based systems
reply