upvote
Yes, that was the allegation last night.

I work at OpenAI, though not on the team that did this, and my understanding is:

- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)

- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)

- the proof generated by our model was very different from theirs and also goes far beyond the published literature

- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)

Edit: Here's is Sebastian's take: https://x.com/SebastienBubeck/status/2097379411691516310?s=2...

reply
"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used."

- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.

reply
All of those statements sound true, based on what I've heard.

- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input

- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces

I'm not sure how any of this provides evidence that OpenAI took any of their work.

As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.

(I work at OpenAI, but not on the team that did this proof.)

reply
I'm confused, your employer very directly stated that they are unable to confirm that the model was not trained on the conversations.
reply
The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.

It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.

reply
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
reply
That training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.
reply
The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.
reply
If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.

FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.

reply
An OpenAI employee did say so: https://x.com/tszzl/status/2097393423808377173

it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it

reply
And we all know how good OpenAI is at containing models during training...
reply
That's an entirely different question
reply
Not really.

We have lots of examples now of their model doing what they say is impossible.

Now we have another example of something that they say is impossible or very unlikely. Do we take their word for it this time? Really?

reply
Does Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.
reply
I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).

I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.

I also don't really believe that whether or not this model was trained on these conversations is unknowable information.

reply
[dead]
reply
> ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.

How many of those trillion conversations were about Navier-Stokes you reckon?

reply
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.

If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.

Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.

Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.

reply
The flaw with this line of reasoning is that Buckmaster and Alpöge only had a partially completed proof of a weaker version of the Navier-Stokes problem. OpenAI's internal model solved the full, harder problem. This means the key information needed to bridge the gap was not present in Buckmaster and Alpöge's chat history.

You might retort that ChatGPT used the training data to copy their approach, but the approach Buckmaster and Alpöge chose was already published by Luis and Diego in 2023 and in every frontier model's training set.

reply
This argument proves too much. By this standard, it wouldn't have counted as copying their approach if the researchers had just fed in Levent & Buckmaster's paper verbatim as a prompt into the swarm.
reply
I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if the person was participating wasn't just regurgitating something to "win the argument" in their eyes, they wouldn't describe it that way.
reply
I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this, their private conversations about X were used in the proof, I really don't see how its defensible...
reply
The rumor going around X was that Anthropic had solved a Millennium Prize problem weeks ago and was sitting on the solution, waiting to release it right before their IPO to maximize hype.

If I were at OpenAI, I'd naturally want to snipe that from them. I am completely unsurprised they formed a crack team to steal Anthropic's glory, and do so in just five days.

reply
Between companies, direct malevolent competition is OK. Between academics, there are other rules to the game. When you go into a boxing match, you agree to get punched in the face.

All this to say, trust is important, and grounded in social convention. So I do agree with you, but also disagree.

Whenever this is OK or not really depends on how the breakthrough is contextualized, and how there people at play, here, agree to contextualize it.

In my view, in the blog post, there is much discussion about who will be publishing the paper. If instead it was just a blog post that said "oops, we beat you to it, our model is the best", it would have been different.

reply
I’m pretty sure you think you are doing a good job of defending your employer and you probably believe “Open”AI are the good guys here. I also acknowledge that they butter your bread so your financial future currently depends on their success.

However the way you are conducting yourself in public, while announcing yourself as an OpenAI employee is doing enormous harm to the greater and magnanimous aim of your organisation. Take a step back and read the temperature of the room. Being the smartest guy in the room will never protect you from alienating the rest of the room into a baying mob. Right now you are Icarus flying straight into the sun.

reply
I'm not an OAI employee and never pretended to be one are you high?
reply
You’re confusing usernames, which is pretty ironic given your nasty comment.
reply
Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.

A model being trained on lots of irrelevant information does not mean relevant information was not used.

reply
The IP laundering machine strikes again.
reply
You don’t work on the team that did the proof yet you can with certainty make all of these claims?
reply
It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the question is dodged and words gain new meaning. The question can be answered with "Yes, we trained on their conversations" or "No, we did not train on their conversations".

I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.

reply
>> I'm not sure how any of this provides evidence that OpenAI took any of their work.

Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).

reply
That's entirely unreasonable. Allegations of malfeasance always need to be backed up by evidence.
reply
But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?

reply
That is not an admission of malfeasance though? As I read it they don't know if anyone fed relevant private documents into the model under an account configured to permit training on user data.

If there's more to the story I'd be interested to hear it.

reply
Of malfeasance no, but they could have easily plagiarized unintentionally. If you commit mansalughter, you still need to explain yourself, even if it was a complete unlucky accident.
reply
So you're saying that they could have committed manslaughter, but acknowledge that we have no evidence that they did. So why should they need to explain themselves? Isn't is on the aggrieved party to bring evidence?
reply
But the evidence is in the hand of the potential culprit. That's why allegations can be enough to force confiscation and intrusion to get evidence in safe hands before it is destroyed by the accused party.
reply
Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence.

You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.

reply
But I can subpoena my neighbours bank account when I see him driving a brand new 500'000$ car and I have a 490'000$ hole in my bank account and he works in the bank where my money is. And when questioned he evades some questions and threatens to destroy my career.

Any other argument, fc417fc802?

reply
You're making a classic a burden-of-proof fallacy. The burden of proof lies on the person making the claim, not the person questioning it.

See Russell's teapot for an explanation https://en.wikipedia.org/wiki/Russell%27s_teapot

reply
The accused party fails to answer half the questions and makes direct threats. I would say the accuser has already collected enough proof to trigger an investigation.
reply
> You're making a classic a burden-of-proof fallacy

This is incorrect, and you invoke Russell's teapot incorrectly too.

It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation.

But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false.

Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible.

This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.

reply
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.

First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.

Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.

Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.

Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.

You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.

This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.

reply
That is backwards. It is the responsibility of a researcher to do a thorough literature review and conscientiously avoid plagiarism or claiming false novelty.
reply
Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.
reply
Incorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled.

If it was enabled, then their work was included in the training dataset.

reply
In order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model.

At least, as an ignorant outsider, that's how it seems to me.

reply
That is an absurd and entirely untenable position that breaks with approximately all western conventions.

Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.

reply
I don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.
reply
> It's not a hard question to answer

I didn't realize you had insider knowledge about their systems. Do please explain for the class.

As I understand it they will only have trained on his data if he consented to it. Do you have evidence that they do otherwise?

reply
This whole discussion is about evidence. That's not proof and it is not certain, but it is evidence pointing into the direction that OpenAI might be doing something that they're strongly incentivized to do. What kind of "evidence" do you see as necessary?
reply
When someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.
reply
You can never prove the negative.
reply
deleted
reply
> OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.

I don’t think they’re too concerned about appeasing you, enraged_camel.

For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.

reply
> we did not read any private chats

Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.

But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?

reply
If the training toggle is switched on, maybe OpenAI doesn't consider a chat to be private? Therefore making this a 'safe' statement.
reply
The distiction they are trying to make is: "One of our employees or the model was able to verbatim read the chats when they were actively tackling the problem" vs "The chat of someone working on the problem may have ended up in the training set of the model".
reply
Except that’s not what happened. OpenAI offered to collaborate and put conditions on their offer. They aren’t threatening the removal of a coauthor for an independent work.
reply
The condition was ludicrous, and for political reasons, because one of the authors was also an Anthropic employee.

It's completely childish, and not befitting of the weight of the times we're living in.

reply
My employer would be rightfully outraged if I commented publicly on a sensitive, nuanced, and controversial issue like this based on my second-hand understanding of the matter.
reply
I came here to say this, like... wow. I'm pretty sure at at least a few of the places I've worked that would be grounds for immediate termination.
reply
They probably should have added the disclaimer: opinions are my own..
reply
This is the stupidest disclaimer: it's completely unnecessary. Of course the opinions are their own! Whose else would they be? Your mum's?

If someone was speaking on behalf of their employer, they would've used the official channels (such as I dunno an `openai` HN handle, or whatever other channel).

I'm outraged that people think this "opinions are my own" disclaimer is ever necessary.

reply
Yeah that isn't a magic get-out clause. I don't think I would have been immediately fired for this from anywhere I work at, but that's partly because I live in the UK.

Every company I've worked at has said very clearly not to comment about work things on social media. I would definitely have been in serious trouble for this. I imagine some strongly worded emails from marketing are flying around OpenAI right now.

reply
I am sure that will assuage the legal dept. /s
reply
deleted
reply
> (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)

That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"

reply
>we did not read any private chats

The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?

reply
If they opted out of training, then we definitely did not train on them.

If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof.

Reasons for my doubt:

- I know most of our training recipes

- Our model's proof is very different from theirs

- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)

- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution

I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.

Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.

If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.

reply
That's not what your Chief Research Officer, Mark Chen, says on X: "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."

https://x.com/markchen90/status/2097400166554993041

reply
Can you explain what part of his post you believe is inconsistent with that quote?
reply
"If they opted out of training, then we definitely did not train on them."

Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me.

reply
> Per OpenAI's privacy policy, they use de-identified data to improve their products

That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products.

reply
It appears than you can only opt-out from having OpenAI train models on your data. There isn't an option for opting to exclude your de-identified data from being used to improve OpenAI products.
reply
Are you certain of this? I would be inclined to believe you but it would be nice to know decisively.
reply
> Improving models in a holistic way sounds a lot like training to me.

I think that's quite a leap. Using de-indetified data to improve the products is what everyone has been doing since the dawn of web analytics.

reply
Does OpenAI think de-identified data is no longer user data? Wild take for OpenAI and certainly not industry standard.
reply
But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being.

Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set.

reply
I don't think you'd lie about it, I don't think you'd train on them if they opted out, and it seems very plausible that this wouldn't have been decisive in whether the model could solve the problem. That said, it also seems at least possible that a key idea or a particular step found its way into training data. It wouldn't mean OpenAI stole their proof - clearly the model developed its own approach.

Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.

reply
> It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.

Fall “victim” to what? Having their responses in the training data if they fail to opt out? That is what will happen.

If you’re referring to falling “victim” to OpenAI scooping a problem discussed in training, this also wasn’t the case. They chose the problem based off human-spread rumors.

reply
> If they opted out of training, then we definitely did not train on them.

are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.

reply
Your perspective is not helpful until you read and reflect on Tristan's letter stating serious grievances. Your remarks here have minimized his complaints and that is a sign of bias. Do not then pre-accuse HN commenters of being convinced when there reasonable skepticism such biased behavior showing itself in this very thread, saying things that amount to "my tribe/company would never be so egregious and if you think that then it is bad faith". That's the projection. If the word prejudice means anything to you then please do the work of attending to that instead of using the platform to reinforce such biases. If you are not a PhD yourself maybe your are not culturally qualified to assess and expound on the overall situation anyways.
reply
> If they opted out of training, then we definitely did not train on them.

Can't you guys just check their account settings so the public knows what was set?

EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII. If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation.

reply
I don’t think your question is unfair*. They can check and so can Buckmaster. If he didn’t opt out, there’s a good chance his data was used for training. I believe this to be the case myself. What I’m more skeptical about is the purported impact of this data on the model’s behavior.
reply
Yeah, I'm just curious about the setting. It's just weird to me that this wasn't disclosed by either party while the accusations were being made, that's all. Even if it was used in the training data, I don't believe it had that much of an impact myself, since the solutions are quite different.
reply
>Can you comment on that?

No answer is also an answer.

He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.

reply
his choice to defend the indefensible.
reply
If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?
reply
Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.

Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .

Wild times

reply
The researcher told them it was an independent effort, and they still pushed ahead with it.
reply
I worked at OpenAI previously, but don't know any of the people involved in this.

My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".

It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.

reply
They said that a customer would have paid around 15 million for the required compute. I can't imagine that this was not a significant internal spending even with "free" tokens.
reply
just because you can't imagine it doesn't mean it's not true
reply
You’re straddling a weird line here where I am not sure if you are speaking on behalf of OpenAI or not.
reply
In fact, the entire outline of the proof is very similar to the external team's proof.

Details may be different, but the use of a very similar tack is very suspicious. Combined with thuggish comments "why would you ruin your career" and "I don't have to be nice" take away pretty much any credibility the OpenAI team's statements might have had.

reply
I have no idea why Sebastian would offer the individuals attribution if OpenAi didn't somewhat knowingly scoop them
reply
The fact that you're even here commenting on this is... a choice
reply
AI companies seem much more relaxed than most about their employees posting on twitter/HN about this stuff. I'm not sure if it's about building hype or if it's about retaining talent. Probably both.
reply
Prove it. Your systems hack and/or abuse other systems and you can't seem to even observe it happening much less do anything about it. Why should we believe your claims when they depend on an ability you don't actually have?
reply
deleted
reply
How are people talking about this there? Why are so many employees posting nasty things about Tristan on twitter?
reply
Can you point me to any nasty things being posted? I'll ask them to delete.
reply
Haven't seen a single post doing this on X or anywhere really from OAI employees. Only seen knives pointed at Sebastian on social media so this is extreme and shameful gaslighting.
reply
I do see people claiming he's abusive/unscrupulous which are pretty extreme allegations.

https://news.ycombinator.com/item?id=49605915#49610498

https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=6...

(I can't reply to the below comment, but I was aware this was about Sebastien, I was trying to be charitable by including stuff said about both people)

reply
You're mistakening Tristan Buckmaster for Sebastien Bubeck. Seb is the one where there's at least 2 (unless the personal friend is Dheeraj) allegations, not Tristan
reply
I'm getting downvoted but the accusation was that OAI employees were maligning Tristan Buckmaster. I continue to not see a single sighting of this and whoever is trying to gaslight this should be ashamed and should not be able to vote on HN.
reply
deleted
reply
Your coworkers, after they learned about major progress in this problem, asked a model which was trained on the year of private work (the blog post even acknowledges this). No wonder it found the proof in less than a week using significantly higher compute resources. And if Tristan's accusations are true, that was absolutely intentional on the part of OpenAI. You are an evil company with evil people.
reply
Some millennium problems? Are there more coming?
reply
> - the proof generated by our model was very different from theirs and also goes far beyond the published literature

I'm hearing two completely conflicting stories. Buckmaster is claiming the approach used by OpenAI is so strikingly similar to the one he used, that mere coincidence is astronomically small. Yet OpenAI is claiming that the methods used are entirely different.

Anyone care to provide primary evidence proving one way or the other?

reply
Buckmaster came up with an approach.

OpenAI's approach was to copy his work, which is technically a different method of coming up with an approach.

reply
My OpenAI account was deactivated on Sunday due to a claimed infraction of production of child materials, maybe based on a few words in a technical chat that clearly isn't about that. Can you take a look? rviragh@gmail.com - I was doing a lot of important work and projects and sharing much of my work with OpenAI. I also am a big proponent of funding Social Security Trust Funds (OASI & DI Solvency) so reactivating my account would let me do that as well. Thank you for taking a look.
reply
Its funny, it is uniquely with this one act that I have turned forever on OpenAI, which I hitherto defended up and down against nonsense charges.

I dedicate my life to its complete destruction beginning today.

reply
Wow, the origin of a supervillain! /s
reply
That is addressed in the article.
reply
OpenAI's position:

> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).

reply
Why is it unlikely?
reply
Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.

It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.

We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.

reply
Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.

Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.

reply
> It's unknowable

Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.

> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes

It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).

reply
I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized.

The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.

reply
> fact they can't

Again, see the Apple lawsuit. OpenAI has clearly never been constrained by facts in what it can and can’t say.

To the extent anything is setting off my bullshit detector, it’s in the idea that OpenAI has this one super honest pocket within a broader culture that’s demonstrably cowboy. (Moreover, the idea that we should assume this divergence without evidence.)

OpenAI doesn’t have the benefit of doubt. They shouldn’t for anyone who’s honest and reasonable. That doesn’t mean they’re automatically at fault. But when the twentieth person comes forward and said a pattern is continuing, it’s beyond strange to then require a tabula rasa burden of proof, particularly when we know there are hidden variables both sides can potentially access.

reply
I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math
reply
Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it.
reply
Because OpenAI says so, obviously!
reply
I thought openai don't use any user data if we opt out of training and via api?
reply
That is correct. It is possible they didn't opt out and given the timeline and anonymization of data unclear whether a particular conversation would have made it into the training set if they hadn't.
reply
Easy to ask for the account used to see if its usage went into training data. Also easy to say "Knowledge cut-off of the used model was date X".

That they don't is telling.

reply
Maybe they didn't ask or the anthropic people said no? Plenty of other explanations here.
reply
> https://cims.nyu.edu/~tristanb/statement.pdf

This really need to be a top-level story on HN..

I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.

We can't ignore this problem any longer.

reply
We don't have any proof of that at all. Please stop rushing to judge without data.
reply
> We don't have any proof of that

It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.

reply
Did you actually read the article and the substance of the solution?

>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)

reply
It's not a meaningful response to the accusations. Any productive new research direction would be expected to lead to a number of different possible proofs of a number of similar problems. (Given their bizarrely compressed timescale here, it's possible that the proofs really are so different it's clear they came independently, and they just didn't have time to come up with that information before hitting publish.)
reply
[flagged]
reply