upvote
From my limited testing of just 2 hours, reasoning output of 5.1-max is at least 7x of 5-max, on the same project and comparable prompts.

It reasoned for ~2 minutes trying to figure out an appropriate directory name. I've never seen 5-max do that. Could be a misconfiguration though.

reply
I haven't dug deep but I burned through 30% of my weekly usage in a few hours which shocked me at first.
reply
Interesting, even if we were to ignore the cache-hits, reads and output, the reasoning cost (aka test time compute) per task should remain a fully comparable metric - it went from $1.25 (Fable5) to $1.48 (+18.4%) for an improvement significantly lower than 18%.
reply
I would expect the benchmark scores to be nonlinear near the top, as the easier tasks get solved and the harder ones are left over. So going from 10 to 15 would be easier than going from 60 to 65.

I only take the Intelligence Index value roughly though. Considering they put Opus 5 (High) at the same level as Fable 5 (Max), I don't trust it that much.

reply
> Considering they put Opus 5 (High) at the same level as Fable 5 (Max), I don't trust it that much.

Have you used both? I’ve never experienced any seemingly greater level of intelligence from Fable 5 over Opus.

reply