>The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”
>The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?
Obviously the result of OpenAI's investigation was that no usage data has interacted with the system after that date.
What else do you expect them to investigate?
If Buckmaster and co. provide their chats, OpenAI could potentially search for them in the anonymized opted-in usage data. Then they could say if any data has been used.
By all accounts individual usage data does not have the direct impact on the model most here fantasize about. To prove this, OpenAI would need to do new training runs to replicate the system used minus the particular usage data in question, if it exists, and then benchmark this on the problem again.
Potentially multiple times, in order to reach a conclusion.
The cost might be in the hundreds of millions.
It's comparable to Magnus Carlson saying that if he wanted to cheat, all he would need would be for someone to tell him to spend more time thinking about a specific move (just a wink would be enough) as an indication that a computer had found something interesting.
It's as-if after OpenAI first failing on Navier-Stokes (which OpenAI had just tweeted about 2 days earlier!), someone winked at them and said "you might want to try a little harder ...".
(TIL: paltering: exact and technically correct statement usage to create misleading impression)
1) OpenAI by their own admission, only re-tackled Navier-Stokes because they heard it had already been solved (but not yet published). This isn't advancing science or helping the mathematical community, this is just being a dick.
2) OpenAI, specifically Sebastien Brubeck, then threaten to "not be nice" and "ruin the career" of one of the mathematicians whose work they had succeeded in duplicating, unless he agreed (which he refused to do) that his collaborator, an Anthropic employee, was not named. This is not only against mathematical norms of credit assignment, it is also being a pathetic human being.
OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute, and the assistance of a whole team of people at OpenAI, to replicate (then exceed) the work that just took two people, with some academic grants as an AI spending budget to achieve (a few $100K - listed below).
https://cims.nyu.edu/~tristanb/
I'd say advantage humans this time. Better luck next time OpenAI - and if you don't want unfavorable comparisons then maybe choose to work on problems that have not been solved yet, and that humans are NOT making nice progress on.
I think you have to work pretty hard to minimize what OpenAI achieved here like this.
The Navier-Stokes equations have been around since 1850. The smoothness problem has been well known for over a hundred years and has only gained importance. It's been a Millennium Problem since 2000.
Levent Alpöge and Tristan Buckmaster did great work to solve the related Euler problem, but didn't solve the Navier-Stokes smoothness problem.
The Navier-Stokes smoothness problem has previously had significant resources working on it. Computational fluid dynamics is one of the most important tools in modern engineering and is closely related.
You speak of 10,000 agents as though it is somehow extreme, and yet within the past month I've had a single task that used over 100 agents on a mere Anthropic team plan. I think two orders of magnitude more compute to solve one of the greatest unsolved physics problems[1] is nothing.
I don't excuse Brubeck behavior because of this, but that doesn't minimize the achievement here.
[1] Wikipedia quote: In particular, solutions of the Navier–Stokes equations often include turbulence, which remains one of the greatest unsolved problems in physics, despite its immense importance in science and engineering. https://en.wikipedia.org/wiki/Navier%E2%80%93Stokes_existenc...
2. Yes Brubeck's comments were weird at face value. That said, Open AI's proof isn't a duplication of anything. Not only is Tristan's work a sub problem but the methods are different. And what OpenAI didn't want was Levant on the paper OpenAI authored not whatever they were working on (Euler). It's petty sure but it's fair enough. Tristan and Levant didn't have anything to do with the Navier Stokes solution, so it's really their call if they didn't want to collaborate on their own paper with the Anthropic employee.
>OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute,
$20M in approximated API prices doesn't mean they spent $20M worth of compute. The real number would obviously be substantially less.
>and the assistance of a whole team of people at OpenAI
You can't eat your cake and have it. What sort of guidance do you think is happening in a 10k agent, 320b token, 88 hour run ? AI did this one.
>I'd say advantage humans this time....to work on problems that have not been solved yet, and that humans are NOT making nice progress on.
Interesting way to frame progress that didn't move along till an LLM generated proof.
If you read the PDF release by Buckmaster, apparently the initial claim from Brubeck was that there as very little human input involved, then as the call progressed more and more people popped up that has been involved with it.
Does this aspect really matter? Not really, other than OpenAI wanting to present this as all the work of their model.
**
https://cims.nyu.edu/~tristanb/statement.pdf
I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.
I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
As it seems and as they tell it, they started the run modestly and diverted more resources towards it as it looked more and more promising. The run didn't start with 10k agents for instance. The point is there isn't anything humans are doing in this timeframe against all this text that would count more than "little human output". It's still a fair assessment I would say.
This part of your argument is totally wrong. The OpenAI approach begins with the B/L work. The belief / knowledge that their approach would pan out is worth a lot - it means essentially “depth-first” search in this direction will be more fruitful than a general search.
Unless you are counting the B/L work as LLM generated. Is that your argument? Even if you do consider it that way, to me racing in for a scoop isn’t a good look.
This is an odd way to gloss over threats.
Given Buckmaster's telling, this seems beyond "poor choice of words"... It was a veiled threat, that he then doubled down on with his "If you don’t want me to be nice, then I don’t have to be nice." follow-up.
**
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
**
FWIW there are also other people on Twitter, such as this DeepMind researcher, saying this is a pattern for Brubeck.
https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=2...
I see none, and consequently Brubeck's explanation makes more sense to me. I understand he meant these words, which he supposedly retracted on the spot, in a "why would you ruin your career with this behavior / turning down the opportunity I am offering" way.
You don't think that is the likely explanation?
Well it looks like they will announce at least one other millenium solution soon. In the same link they say they have "made substantial progress" on another millenium problem. The rumor mill before that statement was Hodge is done and Birch and Swinnerton-Dyer is on its way out.
They've pretty much said their own work was heavily agent driven. Levent is in a particularly bad place here because while he probably had a lot of background in the Jacobian Conjecture problem, he made the solution to that one sound like someone asked the question and he just fed it to Fable during the world cup. Whether that nonchalantness was to just seem hip or was to promote Anthropic, which he has stock in, or was just the truth I don't know though. But it makes this one seem similar, when they might have had really had nearly a year of very valuable feedback to the models.
Terrance Tao has lamented this practice as being unhelpful for mathematics, and likely to lead to humans working in private to avoid this.
Tao has also noted that many of these AI math proofs don't really help mathematics (nor does it seem they are intended to), since for many of them the proof was never the point, it was the math expected to be needed to be developed along the way, which the AI solutions don't provide.
A related point is that the actual solution approach is never revealed. What was the role of humans guiding the agents ? was it fully autonomous ? etc. It is in the incentive of the AI labs to trump the powers of the LLM, but in practice it is humans guiding the agents on the overall approach, This is never admitted. For example, in the announcement on NS there was only an output artifact given but no indication of how it was arrived at, and not even a writeup. This is what disappointed many folks as it was done purely for one-upmanship. As other have noted, the benefit is in the journey or process and not in arriving magically at a destination.
You can train on a sequence of outputs. In the end, OpenAI outputs are OpenAI's property.
You can learn a lot from a single side of a conversation.
Why should we trust them?
The only way the chat could have been used would be for Open AI to baldly violate their policies.
That said, sometimes it take very little information to point someone in a given direction, "I'm working on Navier-Stokes" said by someone with a given specialization might itself be very useful information.
OpenAI's statement says that they began training their new model on August 28.
edit: ffsm8 makes a great point below, it doesn't matter. I'm not great with dates, sorry.