upvote
You're right about its real world performance, and I worded my original comment wrongly.

I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.

reply
Claude, is that you?
reply
you're absolutely right to push back
reply
Your clarification makes sense. The distinction between overall benchmark performance and why Terminal-Bench is an outlier is important
reply
> You're right about its real world performance, and I worded my original comment wrongly.

Damn, HN commenters starting to talk in claudisms now

reply
this is human writing...

this is claude writing...

corporate needs you to find the difference

reply
Well presumably now it’ll fall back to Sonnet 5.5 lol
reply