ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
(I coauthored the linked blog post)
It is described in their methodology: https://arcprize.org/policy
It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
Which LLMs participate on private set? Open weight LLMs only?
Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371
That's not clear. Need to see independent benchmarks first.
Still below Fable 5, let alone Fable 5.1.
EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
GPT 5.6 is also 61 like Astra.
With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.