upvote
That reminds me of this paper:

Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320

> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:

> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””

reply
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.

In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different?

(Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)

reply
There's also another (that I find more concerning) aspect to it.

As AIs become smarter and smarter, there will be no amount of clarity that will make more complex proofs understandable to humans - this is an inevitable effect of the cognitive capacity gap.

Complaining about bad style can make some sense now (I disagree anyway), but it's an argument that will be dead shortly.

reply
Arguably no human understands 100% of how a smartphone is produced, and, it doesn't matter?

Maybe no human will fully understand a future proof, but they could fully understand a little piece of it. And many humans in aggregate could understand it, each with their own little piece.

reply
> A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.

https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)

reply
Not a counterpoint, actually case in point, because:

>Mochizuki and a few other mathematicians claim that the theory indeed yields such a proof but this has so far not been accepted by the mathematical community.

Proof can't be understood, proof doesn't matter.

reply
I wait for AI to say something about this :)

Someone at OpenAI, please, work on this.

reply
Isn't this controversial, to say the least?
reply
If intelligence is compression, and these models are a different form of lesser intelligence than human, but being scaled up to brute force problems, then it makes sense the artifacts that produce (the proofs) would have worse compression than a human proof would.

In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.

I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.

In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.

It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.

reply
> I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.

Isn't it fairly established that (generally [0]) manually written / optimized skill files perform a lot better than generated ones? Meaning that yes, this likely does hold.

[0] or to be specific, that the pecking order is: ai generated < human co/written < hyperoptimized for the specific model via some convergence process

reply
This is just the first cut. I have no doubt that they will polish their proofs over time.
reply
idk if i'd even say they're "lesser", just very different. so they look like gods/babies depending on what they're doing because we anthropomorphise them.
reply
The post says there's a Lean certificate for this and other proofs ("some [...] not all of them").

> This looks like an AI IPO PR powerplay,

Interestingly, the post has actually also an argument for this:

> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.

> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”

reply
It's on OpenAI and Anthropic to prove that they obtained these results legitimately and credited all researchers who deserve credit. They do not get the benefit of the doubt.
reply
Have you actually read the article? It's been actually written, among the other things, because the author's wife has been trying to solve one of the problems for her whole life.
reply
But if you think they got them illegitimately, then how did they get them? And why are mathematicians reacting to this as a sudden explosion of new results that have resisted sustained effort? Where is the sudden productivity rise coming from?
reply
I'd say it comes from the same mathematicians that were strongly encouraged to use the machine to solve their problems for the last 2 years or so.

There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.

Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.

reply
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.

Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?

reply
From the entity that is producing these proofs, obviously.

As the old saying: great claims require great evidence.

reply
Ask away. They've dropped the mic, as far as they're concerned; they're not going to worry about what you do with it, or if you don't understand it.
reply
That's the best definition of slop I've ever read.
reply
Surely by the time of the IPO we will know whether the main results are correct, if only because a different AI will have produced a lean proof or found a logical flaw (the second case would be hard to verify but probably not impossible).

Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.

reply
Why should that make material difference to the IPO? What is the economic value of those results?

The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?

reply
The market works in mysterious ways. What companies do for marketing is often irrational, what companies do to attract investors is likewise often irrational, and why investors invest in companies is also often irrational.

Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.

reply
We should consider the possibility that at some abstraction levels, we can safely stop chasing "clarity" or "coherence" which is circularly defined in such a way that it's capped by human processing power.

Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.

reply
CS got that idea from mathematics. Theorems (with the definitions required to state them) are supposed to be self-contained units. Once the general consensus is that a theorem has been proven correct, people can use it without understanding the proof. Of course, people still want to understand how things work, and it often makes sense to understand them a couple of layers below the one you usually work at. But at some point, you should stop distracting yourself with irrelevant details and focus on your actual work.
reply
Thing is that many of the theorems here are not useful work in and of themselves, but were posed as research problems because it wasn't clear how they could be resolved with current techniques, implying that the process of trying to find a proof might result in new techniques. It's those new techniques that are the actual goal, but if they can't be easily extracted because the proof isn't structured to enable this, that's a bit of a headache.
reply