upvote
Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

(I coauthored the linked blog post)

reply
Thanks for destroying all hope of a better life in my lifetime. I hope you enjoy your millions.
reply
Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
reply
Those people haven't verified their results against the private set: https://arcprize.org/leaderboard
reply
Astra also not verified using private set, but on "semi-private" set
reply
if that is true then why is astra on the official ARC leaderboard now ?
reply
ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.

It is described in their methodology: https://arcprize.org/policy

It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.

reply
Where are results for private data?

Which LLMs participate on private set? Open weight LLMs only?

reply
Yes, they run competitions once a year amongst open weight models
reply
yes it is.
reply
Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
reply
It is about memory retention. No heavy lifting done on the reasoning side so I hardly see anything misleading here.

Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371

reply