Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**].
[**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and it’s a big part of why the industry is so skeptical of the capabilities insiders keep claiming it has. I’d be skeptical too if I hadn’t tested it myself.
We lost access to Mythos Preview when Anthropic forced us onto Mythos 5 some weeks ago, which is garbage by comparison. I’ve already switched to GPT-5.5 and I’m working on adapting my harness(es) to less restrictive open-weight models. I don’t see any other way forward at this point.
That graph gives a good perspective of what models they've tested, and roughly what "subject" each step covers. It is on that task that they note this:
> Kimi K3 reaches step 17 on average, compared with step 11 for GLM-5.2.
[1] - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...