upvote
If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.

If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.

reply
The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools

The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem

reply
In that it isn't able to genuinely solve problems

Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.

This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.

reply
I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach
reply
You are confusing ideas here. No one except OpenAI had a solution to Navier–Stokes. Buckmaster and Alpöge had a solution for the forced Euler problem, which they arrived at largely using LLMs (Claude and Codex). Buckmaster implies (but does not explicitly accuse, since he has no evidence) that training on his prompts had some influence on OpenAI's result. This seems unlikely to me but is not impossible. However, in either case, the solution was found due to an LLM. Of course the LLM built on past human work, but "plagiarism" is not sufficient to account for the distance between the papers of Martínez-Zoroa, or the prompts of Buckmaster, and the final resolution.
reply
I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants?
reply
If it could do this while properly crediting the researchers (the “giants”) it would be a different matter.
reply
Do OpenAI’s T&Cs that users accept not allow them to train on prompts people enter into it?
reply
OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct
reply
Why would you suspect that? Stealing children is illegal, and involves violating the rights of unwilling parties, whereas prompting openAI (or any LLM) is a business transaction, in which the transfer of money and data is legal.
reply
Terms and conditions are almost entirely about the company doing things that would otherwise be illegal.
reply
Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?

Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.

reply
I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data?
reply
All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.

That's it. The rest appears to be wild speculation.

reply
Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.

Never ever touch those requests. If you get a side by side comparison just resend the prompt.

reply
even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"?
reply
Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it.

Publication, though? Slimy is right.

But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.

reply
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.

reply
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,

I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.

reply