upvote
As far as I understand it, nobody is disputing the correctness of the Lean proof, or that it proves the conjecture it actually claims to prove. That's sufficient to consider the problem "solved". The natural language proof is a "nice to have".
reply
The claim in TFA is that the formalization(in Lean) of the problem does not correspond to the natural language statement of the problem, such that the statement proven is not the conjecture for which proof is required for the problem to be considered "solved".
reply
>the statement proven is not the conjecture for which proof is required for the problem to be considered "solved".

that's not the claim. the formal statement of the problem for the NS proof was written by humans not autoformalized.

reply
That's not the claim made in TFA. See the sibling comments, in particular about the DeepMind formalization.
reply
deleted
reply
[flagged]
reply
Please make your substantive points without swipes. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.
reply
Does that contradict what I said? In that quote, it says that the NL proof does not correspond to the Lean proof. However, the statement of the theorem in Lean is independent from the NL proof. It comes from a DeepMind repository, which as far as I'm aware has been accepted by the community as a valid formalization of the original Clay Institute statement.

https://github.com/google-deepmind/formal-conjectures/blob/8...

reply
[dead]
reply
Both proofs may be correct, and the problem may indeed be solved. My point is that it should not be assumed.
reply
> Maybe read the original article before replying, at a minimum.

Maybe read the comment before replying, at a minimum.

reply
>The idea that an AI company is beyond peer review is harmful.

i havent seen this sentiment expressed anywhere, have you?

isn't this comment chain on a submission about openai's claims being reviewed?

reply
OpenAI has expressed this sentiment by not submitting to or saying they will submit their results to peer reviewed journals.
reply
I would say it's released in the spirit of open source. "Peer review" in the narrow sense exists primarily to assign prestige in academia; but there's nothing stopping anyone from "peer reviewing" the GitHub repository.
reply
I would say it's released in the spirit of machine learning's competitive landscape (which is the culture this emerged from).
reply
deleted
reply
So they could also dump a 100 quadrillion line proof in Bourbaki notation and call it a day?

The proof was released in the spirit of being first at all costs without any attempt to clean it up. I doubt that OpenAI mathematicians could give a coherent talk about it, certainly not using a blackboard.

reply
Sure, why not? They can publish whatever they want, then the public can choose to ignore it, criticize it, or accept it.
reply
[dead]
reply
What kind of prestige? Peer review is anonymous unpaid work.

A good review does not merely check the correctness of logical arguments, it gives suggestions for the exposition, citing the correct references, putting everything in the right context, etc.

reply
> What kind of prestige? Peer review is anonymous unpaid work.

Prestige to the reviewed, not to the reviewer.

reply
All of the reasons you listed for peer review are valid. The broader point is that peer review can happen outside of academic journals, and nobody has an incentive to submit to them who isn't trying to play the academic prestige game. For a significant example, see the history of Perelman's proof of the Poincaré conjecture.
reply
> journals are not the arbiter of truth and getting published in them is something only academics have an incentive to do

Good, I just wanted to point out that peer review isn't primarily an arbitrage of truth, it is also to make sure the exposition is nice to read. When you get a reviewer who actually cares, you receive lots of feedback that isn't related to the correctness of Lemma 3.14.15 and stuff like that.

reply
This is an equivalent of a company producing security software, open sourcing their code, and then claiming that since no one has found any serious bugs, their software is secure.

No. The way to build confidence that your software is well made, you do a proper external security audit and obtain the requisite certificate from a proper auditing firm.

It's also incorrect to think peer review in mathematics is low quality (like it is in some other fields). Certainly, when major results are in place, editors ensure that high quality peer reviewers are recruited and do their job properly. Like all human processes this fails sometimes, but not enough to not do it.

reply
>then claiming that since no one has found any serious bugs, their software is secure.

which specific openai statements does this part of your analogy map to?

in the "sharing ai progress in mathematics" blog, openai simply says "results", and never once claims that all of them are unquestionably true. instead, they state they want to evaluate the results. their github states that the results are "different stages of verification" and also says "Some of the unformalized results could have issues"

that is the opposite of "claiming [...] their software is secure", to use your analogy.

reply
I didn't say peer review is low quality; just that it's not necessary or sufficient to determine the truth. Ultimately the OpenAI proof stands or falls on things that have been audited externally, namely the formalization of the problem in Lean and the correctness of the Lean software. There's no incentive for OpenAI to submit to a peer-reviewed journal when they don't need to play the academic prestige game. TFA is an example of peer review in action: they're analyzing the proof and finding points to criticize.
reply
not submitting to whatever journal is quite different than saying they are "beyond peer review"

are people not reviewing openai claims right now?

reply
People described the problems as solved the minute they were made public.
reply
this happens in approximately every scientific field. ive never heard it described as "idea that they are beyond peer review".

openai themselves specifically call out that there may be issues with their results. journalists and laypeople just happen to skip that part, like they do with ~all physics, health, astronomy, etc results.

reply
You are confusing two levels of indirection here.

Peer review is a proxy for correctness.

Peer review journal is a proxy for quality peer review, or at least it was, once upon a time.

reply
Because they want to release everything on github so everyone can peer review it themselves

This is far more efficient and they’re telling the academic industry to grow up

Sister comments are saying that academics dont like the Lean programming language and see a lack of human language described proof. Doesn’t sound like something I should care about but I’m watching for a better human language description of the problem as this discussion evolves

reply
"not interested in" != "beyond"
reply
I've seen a lot of breathless reporting about various mathematical things being "proven" on the basis of the LLM-generated Lean formulation compiling. We probably wouldn't declare that for a human-written proof until peers had checked the proof for errors
reply
This. The proof of Fermat’s Last Theorem took 15+ months to check. It’s absurd to see the media reporting that these big problems are solved based off of a news release and a hastily and mostly AI-written manuscript, and OpenAI et al. are all too happy to run with said breathless reporting.
reply
Wiles' proof was informal and couldn't be checked by a computer. In this case, the experts need to check 300 lines of Lean code (mostly comments) and confirm that it formalizes the problem statement correctly. There are papers building on the solution and analyzing it for more general versions of the problem, which suggests that the PDE community has already accepted it and moved on.
reply
deleted
reply
there's breathless reporting of just about everything scientific. physics, astronomy, archaeology, etc. have this sort of thing all the time.

yet i have never seen anyone say "the idea that physicists are beyond peer review is harmful" because some mainstream news articles published a piece about dark energy or whatever.

reply
Exactly. Coverage here is "OpenAI has solved problem X", not "OpenAI has claimed to solve problem X."
reply
Anyone who doesn't understand peer review (its intended workings, its negative effects by implementation flaws, etc.) automatically assumes expression is beyond academic peer review, so thats potentially a lot of people...
reply