upvote
Idk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.
reply
They're gaming benchmarks HTH
reply
This benchmark gives the same intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high)? That seems very wrong to me, unless I'm misinterpreting the visualizations.
reply
The most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1.

And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.

reply
I don't think this is quite true. We have other examples.

Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while Fable is more intelligent, Opus is more capable. I'd still opt for Fable in nearly every case if tokens were free.

So it can be true that the "smarter" model is perhaps not the smartest in every single niche dimension that its cousins have been fine-tuned for (yet!).

reply