upvote
> Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.

I have seen people claim the exact opposite. If there was so much progress, then you wouldn't have endless disagreements with people championing their own favourite model as the strongest.

reply
Yes, and good senior software engineer is ahead of fable, but benchmarks can't capture that either.

We already know from testing humans that test scores don't correlate that well with how effective a person is at work. Same applies here, we just aren't that great at making good tests.

reply
If most models were getting 100% on the test it would be an inadequate benchmarks.

What were seeing is all models failing to ace these tests.

"Benchmark Saturation" is term that promotes lowering the bar.

reply
deleted
reply
The two times I tried to use fable I had it attempt something I had already had Opus 4.6 do with no issues. It blasted through 10% of my weekly allowance on a 100$ a month sub and produced something broken and nonsensical.
reply