upvote
The implicit point being adding this type of safeguards to Fable dumbs down the model in measured performance even though it is not fundamentally different.

Note it may not even be actual performance, typically in most benchmarks the model would be scored zero for refusing a task just the same as not completing it, so it could just be the Fable's stronger safeguards is just making it refuse more or perhaps even drop down to Opus.

reply
The model cannot complete that task, for one reason or another, and therefore it scores lower.
reply
Artificial Analysis at least reports the results with fallback to an inferior model. So presumably Opus 5, and the score should be between Mythos 5.1 and that other model.
reply
Maybe they do that opaque degradation trick that whenever it's asked something questionable, it'll route to a worse model instead.
reply
Makes more sense if you recognize that Anthropic intentionally degrades outputs for most customers. Vetted customers get excluded from that practice.
reply