upvote
What else can they declare really? Yeah the model has training data from previous attempts. Alpöge and Buckmaster also similarly benefited from attempts before theirs.
reply
I don't think OAI should be given the benefit of doubt. They are doing the research equivalent of front-running. Knowing where to look is one of the main challenges in research. Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.

"deidentified data" isn't much to go by. Say I prompted the internal model this way - "Hey there's a solution to a unsolved problem X. The solution uses a less known Method Y so don't bother wasting time with the usual methods. Take papers A, B and C as references. Oh btw, here's the last year's worth of data of all prompt sessions that mention this problem. Pay special attention to the ones that mention Method Y and sub-keywords Z,W".

This is obviously all speculation but the timing is very suspect. If OAI actually did this (and I suspect whatever they did is pretty much close to this), I think it is highly unethical.

reply
> Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.

This is a bad argument. This is clearly not how it works. And unless Tristan is some truly alien-like savant (and maybe he is), what's necessary to initiate the AI's work already exists in countless published research papers and not exclusively in his head or notes. AIs can survey the sum total of all prior work on a problem and discern reasonable paths for inquiry.

Tristan is acting as if he's working off of an outdated model of AI, similar to primitive chess-playing models that winnowed the search space much more deterministically. If someone this intelligent truly doesn't get that this is not at all what AI is anymore, then maybe there's no hope that we ever understand it.

But I think he does realize this and he's flailing about for counterarguments from a place of bitterness and dejection, accepting even those that are too weak to be defensible. And that is very human and even forgivable.

reply
Are you saying there is no search space intractable to LLMs? That wouldn't be possible. AIs are statistical pattern-matchers on steroids. The prompt is key to getting anything useful out of them. They are incredibly useful and major game changers but ultimately that does not alter this fact. People (including OAI) have already tried to solve Millenium Problems with it. That OAI woke up last week and suddenly decided that throwing their researchers armed with millions of compute on one particular idea to a problem is highly suspicious in itself.

Even if OAI had zero data from Buckmaster's sessions, this is in very poor taste and highly unethical. You are front running a researcher just to be able to say you did it first? Tao is right - OAI is treating math results like oil. This is the like Exxon getting a whiff of a massive oil field and racing to the punch by deploying their full crew.

reply
> Even if OAI had zero data from Buckmaster's sessions, this is in very poor taste and highly unethical. You are front running a researcher just to be able to say you did it first? Tao is right - OAI is treating math results like oil. This is the like Exxon getting a whiff of a massive oil field and racing to the punch by deploying their full crew.

I want to be clear that I agree with this view and with Tao more generally. But we're all just yelling at the wind now.

reply
> What else can they declare really?

Oh I don't know, maybe something like this?

"Given how seriously this would violate the most fundamental of academic standards, as well as taint the claimed capability behind this result, we take this issue very seriously, and we're launching a probe into identifying whether any of their research artifacts have entered our training set. We have further begun making changes to our UI/UX on all our surfaces, so that it is always clear whether any particular chat, or other user artifact, is eligible for being trained on."

reply
In OpenAI's case, if they were genuinely unsure, they wouldn't have said anything. "We cannot rule out" means they absolutely 100% for-sure did look at the existing prompts and bootstrapped from that, and they are trying to get ahead of the disclosure with this weasel-wording.
reply
Also possible: we're 99.999% sure, but a lawyer said to be safe and strictly accurate, we should stick in a sentence in saying we can't be perfectly sure, since it's infeasible for us to prove it.

I promise you that if we took their work from ChatGPT and stuck in a bunch of weasel words to give the opposite impression while remaining technically true, I would quit on the spot.

(I work at OpenAI.)

reply
While you can't necessarily prove it, you can say whether the data was in the training set at all.

You can also do something like a release of a GPT-OSS v2, where you actually release training data and checkpoints, and do an experiment where you have some held out math problem dataset, then demonstrate how much training it takes on solutions (or partial solutions) to that dataset before the model saturates that test. While of course that would be a test on a much smaller model, it would cost a tiny fraction of the training on your big model, and it could be used to demonstrate just how much effect data contaminaiton like this could have, especially if you did the same experiment on a few different sized of model to show the scaling laws involved.

reply
Nice damage control bud, too bad the veil's lifting and everyone's seeing what you sociopaths at OpenAI are really like
reply
Where did the veil lift? This feels like a witch-hunt to me.
reply
They could have thought about the problem for like 2 minutes and not done this! I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.
reply
> I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.

Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.

reply
Being scooped is not a new phenomenon, but the scooper's story is almost always that they were working on the problem independently or had some independent insight into it. By OpenAI's own account, they were inspired to start working on this by rumors that there might be Millennium Prize solutions to scoop.
reply
Research projects don't start in a vacuum.

It never happens that you wake up one morning and start working on a new problem that came to you in a dream (*unless you are Ramanujan).

This is business as usual for academia, it's amusing to the discussion over it.

reply
Apologies if this is against the rules, but could I ask if you have some background in scientific research (maybe you could elaborate lightly on topics you've worked on)?

From my perspective, this practice is quite bad mannered, unusual, and heavily frowned upon, but I recognize it's possible that these stories might be more common in other fields. Still, I'd appreciate a strong sign that you aren't making these statements up based on secondhand accounts of what 'academia is usually like'.

reply
CS PhD, used to be a professor, have worked at several industry research labs since.

> this practice is quite bad mannered, unusual, and heavily frowned upon,

Yes, it is.

Maybe you missed my point?

The fact that is "quite bad mannered, unusual, and heavily frowned upon" does not stop it from happening, and it is common for all high profile inventions and discoveries.

This does not disqualify the person doing the scooping, history remembers them as having the credit, and quietly forgets the person who was scooped.

reply
My initial read of this

> Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.

was that you were trying to suggest that it wasn't. It seems we actually share similar views here then? I cannot say anything about how common it is, since I've been pretty lucky when collaborating I guess.

reply
I meant that it's very common when it comes to high profile discoveries and inventions.

If you have made one, and there was never any dishonest competition, you have indeed been very lucky.

I haven't, but the history of science is rather nasty.

reply
If it's all business as usual and being scooped is no big deal, why was OpenAI in such a rush? They didn't have to launch this effort on the very day they heard the rumor, run "on the order of 10,000 concurrent agents", or try to coordinate announcement scheduling with Buckmaster in the middle of a long weekend. It seems to me that they understood very well this was not a "business as usual" announcement, and devoted huge amounts of money and focus to maximize the chance that they were first.
reply
I think you're making the opposite conclusion than what I intended?

Being scooped is a big deal for the one getting scooped.

It has never been a big deal for the one doing the scooping. History is full of math and science results being scooped. For example, we keep calling it Pythagoras' theorem a few thousand years later.

reply
right, surely they could've waited or even reached out? It reads as desperation to get there for marketing purposes
reply
They did reach out.

> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.

reply
We've heard from Buckmaster, who says that they demanded a condition of cutting Alpöge of all credit. If true, it doesn't make them look too good.
reply
It seems like this is going to be a PR nightmare, because they are now competing with their own customers. If you're using an LLM to help with your bright idea to cure cancer, you're going to have second thoughts about relying on OpenAI.
reply
This is desperate. They were expressly operating within a program. OpenAI isn't going to recover from this
reply
Would that be more or less unlikely than accidentally hacking another company? More or less unlikely than colonizing an obscure wiki?

Highly persistent agents + vibe-coded security seems like a problem.

reply
"Unlikely" lmao if it's in the corpus, it's gonna be brought up immediately.

This is no different than scooping them.

reply
It's not massively different from a certain President's teleprompter operator making bets on speech content. A moral hazard a mile wide which I don't think OpenAI can so easily wave away as they are apparently trying here, especially since they've spent something like $15e6 to keep $1e6 out of academic researchers' hands, right?
reply
Research equivalent of front-running.
reply
They'll probably claim a rogue AI agent accessed it accidentally!
reply