upvote
If you look at the score distribution, Grok and Opus 5 especially stand out for doing consistently well and rarely ever scoring under 50%. Basically they always at give you something that at least works.

Most others, especially and famously Fable 5.1, seem to have a fair chance of completely failing, despite also sometimes excelling.

reply
We run each model multiple times against each challenge and take the average score. We include the variance below the score in the leaderboard.

GPT 5.5: 42.3±10.1 GPT 5.6 sol: 39.4±8.7

We were also surprised by the low sol score but it seems consistent with our experience in using it in the field in atopile as agent in our harness. In general OpenAI models didn't do too well on electronics, which seems to change now with GPT-6 Astra. Results are in soon!

reply
yeah noticed the same. I wonder if this will be a recurring theme for model releases: each release specializes on a set of headline benchmarks, along with regression in benchmarks that are less of a priority
reply