upvote
What do we do about the problems that don't require many millions of dollars in resources?

Last weekend I spun up a small agent swarm and pointed it at a field of math I have some affinity towards. Within four hours I had settled three conjectures, one of which is rather famous (for the field, not in general). It cost me about four hundred dollars.

I am at a loss about what to do with these results. On one hand I feel like the mathematicians working on these should know about them, but on the other I feel a bit like a barbarian who suddenly finds themselves sacking Rome.

reply
I'm not a mathematician but it seems to me that if all it took to solve the problem was an enthusiast level understanding of the domain and a few hundred dollars of tokens then the result probably isn't that valuable. Even assuming you are the first person in the world to solve it, these kinds of LLM-friendly problems that are now easy to solve and easy to verify will almost certainly be picked off by one person or another in the near future.

It's also possible that the result is already known and you just weren't aware of it. It's easy for someone outside of a field, or even one steeped in it, to not be aware of certain solutions.

reply
sounds a bit like cope. Crouzeix's conjecture was solved exactly under these circumstances and I wouldn't describe it as "not that valuable". In any case, AI capabilities will increase dramatically over the next few years while human math capability will not. That means that there's a fixed target regarding whatever is currently considered a "serious math problem" and its difficulty level. Soon the average problem solved with a few hundred dollars of compute will be at that bar.
reply
> while human math capability will not.

I actually partially disagree with this. What happened to all the excitement about Intelligence Augmentation (IA)? Now it's AI instead of IA. I think there's so much untapped potential for augmenting our intellect with the likes of https://dynamicland.org and https://folk.computer, as well as the work that's been going on in college math education, things like Lean, etc. I think the only reason human math capabilities haven't expanded that much is a failure of our imagination, not our potential.

reply
> I am at a loss about what to do with these results.

I would recommend publishing them to Palomar (https://palomar-registry.org/) - I have no affiliation, this is an online registry of Lean-verified proofs created by Terrence Tao.

I have submitted a proof there that's also minorly important in an extremely niche field.

Anyway, I feel like it's a good place to dump AI slop lean proofs because the main point of the registry is that it verifies that: 1) your Lean challenge statement is the same as what you informally state you're trying to prove; 2) your Lean proof actually compiles.

This could be useful to future AI slop researchers who want to know if a given result has already been formalized, and they may be able to mine some lemmas from your work. Also, it's good to know for the field in general what has been proven.

I'm fairly certain you can set your publishing name to be whatever you want, so you could set it to be just the word "Anonymous", or the name of the model you used.

reply
deleted
reply
It is interesting isn't it.

You asked for a painting. A robot made the painting. You looked at it and said, "well, I guess it's good. Should I put it online or something? Dunno. Hey Fred, what do you think of this?"

Meanwhile, your next door neighbor spends their entire life developing their understanding of life through art. They "understand" (maybe not in a way they can articulate) art. You go next door, you look at their painting and say, "well I guess it's good." But you also understand that your neighbor is just like you, and maybe you are a painter in another way.

I find it strange that, people can't see that, we don't need to solve hunger and poverty and work balance, and etc, by a round-about make-super-intelligent-AI. We could just solve it. It's pretty obvious how to, as well.

We can all be painters, if we put restrictions on the psychopaths.

reply
what about cancer?
reply
How do you know the proofs are correct?
reply
The numerical results are trivial to check. I wrote the analytical results in Lean by hand before asking a former professor to confirm after asking him to keep this private.

They're valid.

reply
It's really tough to come to terms with it but making academic contributions in general from the outside (with or without AI tbh) is not often welcome and the whole process feels very gate-kept.
reply
(academic.) Unfortunately the overwhelming majority of outsiders are missing core knowledge (or are cranks), so the optimal prior from a time management perspective is to ignore them. AI just makes engaging more costly because there’s more volume and it’s harder to get signal on whether they know what they are talking about.
reply
If that's the case, that makes me much less sympathetic, even though I can understand how it's very disturbing to see the field suddenly changing like this.
reply
I mean, let's say you spun up a swarm of agents to rewrite a large component of a well used open source library to be memory safe. You could dump it in a big PR and walk away (we all know how that would go), or you could try engaging, see if they're interested, write something up and see where it goes.

The biggest problem is, IMO, drivebys uninterested in actual results, just getting a check mark, and the equivalent of dropping a 200k line PR on people and expecting them to be interested and do the work for you. These are things many on HN are familiar with and know how to do better :)

reply
> and the equivalent of dropping a 200k line PR on people

I can understand why the community is pissed. So now, lean proofs can be churned out at scale, and the community is left to decipher all of that slop into human understanding. There are bad actors with misaligned incentives coming in with drive-by proofs upending what the community holds dear which is to practice and propagate the art. I applaud them for this declaration.

To re-align incentives the following could happen. AI slop lean proofs are dumped unceremoniously into a lean dumpster, and what gets rewarded are results that could digested into human understanding - via the already followed human review process. Prizes are not given to lean proofs since anyone with sufficient compute can churn them out.

reply
It depends.

If you are trying to understand better the field, then do a good write up of the proofs so that people can learn from it.

If you want to earn the respect of people because you found interesting proofs. Then do a good write up of the proofs so thst people can learn from it.

If you want to plant flags and pollute peoples minds. Then please publish it anonimously, no one wants to correct LLM slop for you.

Probably we should build a repository of AI slop proofs that are only allowed to be publish anonimously. That way people may be more inclined to work on it because they would feel like they are cleaning your house for free.

reply
Maybe I'll end up doing the write ups pseudonymously. I have taken care to make sure the results can meaningfully contribute to field but I don't want to plant flags or really receive credit of any kind. I just think they are interesting.

I like my current life and don't want to get dragged into the current fracas surrounding the use of AI in math.

reply
If you have taken care, then share it with the community. People will be thankful and happy.

The problem is with people that may do it without contributing to the community.

reply
If this was limited to just three results, I would agree with you. But those three are just the ones I've managed to verify myself. The list of unverified results is quite a bit larger.

The field I've been investigating is not large. Even if I were to take the time and care to beat the interesting results into something meaningful, I'm afraid the pace at which I'm able to produce these results would not be well received.

reply
That solution (stop pouring resources in to proofs, stay in your lane) works today. How does it work 5, 10, 20 years from now? The software and hardware advances will continue.
reply
> primarily as a marketing exercise

Like the mathematicians working on famous problems in private until they could claim full credit for something interesting wasn't also a marketing exercise for their own careers. The commercial value (or lack thereof) of a proof doesn't depend on whether it was done by a human or a machine.

reply
These mathematicians dedicated their life to math and were working for a long time to achieve the pinnacle of their careers.

OpenAI just burned millions of dollars over a weekend after hearing that someone else was close to solving the problems. Their interest was in their AI system more than the actual math problems.

Don’t you see how that’s different?

reply
I see how it can be devastating to their ego, but no, I don't see a particular difference in a company spending money for clout vs. a person spending time for clout. The underlying motivation is the same.
reply
Do you really think people dedicate their life to mathematics for clout?

If clout was the goal I don’t think becoming a lifelong mathematics academic would be the first step

reply
for many fairly strong mathematicians, the career calculus is fame and status (mild though it may be) through mathematics or anonymity but financial reward in tech or finance. Yes, clout by becoming a lifelong academic is in fact rational for some and part of their motivation. Of course they really like what they do as well, but earning the respect of the peers they know are also respected by a large swathe of society is very important.
reply
I think you’re right that some of the angst here is due to a fairly insular community having its culture disturbed by commercial forces.
reply
Yeah that’s totally fair!

I guess that type of “clout” feels different to me.

Wanting to be validated by peers for your talents in a niche field vs. using millions to try to solve a math problem that you don’t really care about with AI to market the gigantic company you work for.

reply
Yeah bro, same for probably the majority of this site that write software for a living and/or hobby.
reply
Sorry I’m not entirely sure what you’re trying to say
reply
How good were you at those “A is to B as C is to D” SAT questions?
reply
No need to be rude. I didn’t want to misinterpret when responding.

I can see how what is happening to mathematicians is similar to what is happening to coding.

I’m not sure what your broader point is? Mathematicians shouldn’t be upset? Coders should? Something else?

reply
The point is that nearly every white collar worker is in the same boat right now, and the majority of them probably have already had much more profound impacts to their fields. So yeah, we can certainly imagine what it’s like for mathematicians.
reply
Okay. I agree with all that you said in your latest comment.

From your first comment it seemed like you were disagreeing with me but I’m not sure how.

That’s why I asked for clarification.

Edit: BTW I’m a web developer and designer, Not a mathematician

reply
I suspect that beyond just marketing, these pursuits yield plenty of useful information about model design that will likely lead to model improvements and optimizations for both mathematics and general reasoning going forward.
reply
Are they claiming that the only value in solving these problems was for their field's personal development process? I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology. It would be insane to demand that people avoid making progress on technology that can save lives or improve general quality of life, just to protect the sanctity of your karate belt system. Perhaps in lieu of open problems left to solve, mathematicians should be welcome to take up chess or sudoku to keep their minds spry.
reply
No, there's most likely zero practical benefit of having found a pathological edge case in which the Navier-Stokes equations do not work. We're most assuredly not talking about "saving lives" here. Unless advanced aliens show up and tell us they'll destroy Earth unless a counterexample to the N-S equations is provided within 24 hours.
reply
> I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology.

This is the core misunderstanding that the open letter is attempting to correct.

Developing a better understanding of the Navier-Stokes equations could have a number of implications for useful technology. They're fundamental to fluid dynamics, and turbulence in particular is something that many people feel we could work with more effectively if we better understood how and why it's generated. The Navier-Stokes smoothness problem is an interesting and long-standing benchmark for this understanding; we don't know why it should be so hard to answer, so we hoped that the process of developing a proof to the problem would produce more understanding. (We may still be able to extract this understanding after the fact, if OpenAI's proof is fully human-comprehensible.)

Simply knowing that there exists a finite-time blowup is not practically useful. We know that fluids in the real world don't produce random singularities, so the result can't really have much physical meaning. What it illustrates is that the Navier-Stokes equations fail to model physical fluids in some yet to be characterized way.

reply
> There is no near term commercial value

How much do you think other AI companies would offer to get access to the transcripts of the generation that led to the proof? No doubt OpenAI will include it in their training data somehow and use it to build the next generation.

There is already economic value.

reply
Humanity is better off for knowing these proofs. This strikes me as academic NIMBYism.
reply
Do you "know", in any meaningful sense, any of OpenAI's recently publicized proofs? Do you suppose that there is any large community of non-academics that does?

One of the points the parent makes, along with the TFA, is that academia -- or more specifically, the "mathematical community"-- is a setting primarily for creating and ingesting mathematical knowledge, and disseminating it to the next generation and to other fields. Humans absorb this material slowly, through lots of discussion and collaboration -- it is necessarily a slow process. Facilitating this is one of the important functions of academia. Your usage of academic as a slur here is a bit silly for this exact reason.

I don't claim it is perfect, and we can argue about pedagogy in elementary courses till the cows come home. That's not really material. But this is one of the only settings in which such knowledge is broadly valued for its own sake, and in which there is a semblance of incentive to help others "know" this stuff as well, be they future generations of mathematicians, science and math educators and communicators, practitioners in other fields, or genuinely curious amateurs.

reply
Your point is largely addressed in the article, did you try reading it? "In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align."

The point is that these proofs are largely useless without the insights. The value of a proof is largely in the travel, not so much in the destination.

reply
Different person here, I read the article and they are all wrong. Hope that helps.

Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing. And instead of them - and nobody - spending millions of dollars to solve the problem, successfully, they want every problem of their academic industry to persist because even though they never solve the problem, they synthesize and solve lots of other problems nobody asked for. And get to boost their egos?

Yeah, stop that. Actual alignment is on the humans themselves, if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades and don't worry about the narcissistic elements that slow their industry down.

reply
Mathematical breakthroughs with commercial relevance are few and far between, and often depend on dusting off old results which were, at the time of discovery, "solutions nobody asked for."

The NS counterxample is actually, by any market measure, a "problem nobody asked for" in the sense that its existence doesn't have any commercial relevance (beyond juicing OpenAI's IPO). So the only long-term value solving it could have is by virtue of whatever reusable theory/insights were generated along the way to the counterexample itself. The letter is absolutely right on that point.

It's not actually clear that those insights will come faster from reverse engineering this LLM proof vs. humans building theory to solve the problem themselves. So what you're saying may or may not even be an efficient way of operating. Also, it implicitly depends on mathematicians to do the hard work of creating problems and then deciphering LLM hieroglyphics for essentially free while the only immediately profitable component gets outsourced to a frontier lab. In what world is that model going to work?

Reading between the lines, it seems like maybe you have a personal grudge for some reason and simply think the technology will advance enough to where we won't need academics at all. But you should say that in the first place.

reply
What irks me is the ego

My stance is that solving the problem is aligned with humankind

the rest is just hypothesizing a way that academics fit in this world at all

reply
Their claim is only indirectly related to the motivations of the people using their models. What they're saying is that doing math in this way does not produce the same value as traditional mathematical research, and the people using these AI models aren't concerned about that because their marketing objectives don't depend on whether their results produce mathematical value. If people doing valuable work are made irrelevant by people doing a larger volume of non-valuable work, that's not a positive outcome.
reply
But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.

I actually found this to be the case with some basic linear algebra notes I was recently doing in Lean (without using mathlib). The model could generate working proofs, but they obscure the basic ideas (actually I wonder somewhat if this is because the Lean code that's out there to train on doesn't make a huge effort to read like textbook proofs, which was my motivation in the first place). I give it a skeleton of a couple lines of `calc`, letting it fill in the reasoning for each line, and it does much better. Then ask it about making some macros to simplify "trivial" or "obvious" things, and it does even better. etc.

I suspect there's a good workflow where a big SOTA model makes an impenetrable proof (or code) and then a human works with a FIM model to simplify it (with the larger gnarly proof right there in context for FIM), but unfortunately everyone seems to only care about agents right now.

reply
>> Humans can then work on clarifying why it's true.

Presumably you're a human. Are you going to do that?

reply
There are vanishingly few research mathematician positions and it's one of the most competitive fields, so no. But I'm not sure how that's relevant. As the OP says, usually the value of a proof is not the knowledge that something is true per se, but the reasoning techniques to understand why. How can it be anything other than helpful then to have a truth oracle as you try to figure out why things are true?
reply
> But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.

To me this analogy points in the complete opposite direction. Imagine somebody takes a half-completed project design you're trying to figure out, vibecodes a rough prototype of it, emails your manager to announce that the project just launched in alpha, and then dumps it back on your lap for approvals and testing and productionization. Would you say that they've added value to this process? Or did they just strip away all the hard parts of the problem so they could claim credit for the easy part?

reply
One of the awesome things about LLMs is they make it quick and easy to make PoCs, so yes. Proving that an approach will work before spending a bunch of deep design effort is absolutely valuable. Your exact scenario is something I've literally done: give a half-completed design to a team member and asked them to vibecode a PoC to prove the approach will work and figure out some of the details, explore scaling and failure characteristics, etc. Or I do the PoC vibecoding myself too. LLMs have been a gamechanger here.
reply
It's valuable for you, the person who's going to spend a bunch of deep design effort, to make POCs. Is it valuable for someone else to drive by, dump some POCs on your lap, and then leave you to do the deep design effort while they run away to study AI?

If that person then runs around telling people that they're the real author of your project, because they generated the original POC, would you consider that an accurate assessment?

reply
Your analogy is far enough away from the way that the real world works that I'm not sure that I can really even strain my experiences to fit within it. Sure, I guess that would be annoying?

But mathematicians define their field. They're smart people. They're capable of recognizing when someone just did a vibecoded throwaway PoC and when someone has a well structured proof. Actually even before LLMs they'd publish new, clearer or more elegant proofs of old results. They can say that inscrutable proofs are exactly as valuable as they are, and that the first explanation people can actually understand carries its own prestige.

reply
Yes, mathematicians will clearly need to rewrite the qualifying criteria for prizes to better align with the actual goals and value they were hoping to get from a solved problem. The field as a whole assumed good faith actors and collaboration, not expecting a few trillion dollar companies to walk in and start turning in piles of Lean no human understands to be able to claim "first".

This letter includes someone like Terrance Tao who publicly expressed a lot of optimism about AI for solving novel math like with the Erdos problems. It's not sour grapes but the first steps to define those new expectations for the future to reduce the perverse incentives.

And yet, predictably, people are accusing him of "gatekeeping" and ignoring the arguments he has made here and elsewhere about the benefits vs. harm in different ways of using AI.

reply
It's a real example that's happened to me twice in the past year, so I'm not sure what to make of the idea that it's far away from how the real world works.

I'm also not sure I understand what you're objecting to if we agree that mathematicians define their field. The source link is a declaration from 25 Fields Medallists with precisely that goal. They believe/define/declare that the type of AI-generated proofs we've seen are vibecoded throwaway PoCs; they feel that a well-structured proof must include factors such as "a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and the success criterion is not a true/false conclusion but rather "development and integration into the mathematical canon".

reply
deleted
reply
> What the mathematicians are saying is stop pouring resources that most mathematicians can only dream of accessing into projects that are actively damaging to their field. They face a massive challenge of figuring out how maths can evolve in the face of this new technology, and this is not helping.

Is it reasonable for any field to make such demands? If this were doctors objecting to AI becoming good at medical practice would you have the same concerns?

While any idea of OpenAI spying on people to pursue their goals is disgusting, the rest of this is par for the course, as Kasparov experienced with IBM in the 90s. Humans still play chess after all.

reply
It's not a simple matter of "becoming good at", and yes, I could very well have similar concerns, depending on how it impacts the field. The statement itself mentions that such concerns exist in many other fields.

I doubt OpenAI will take such a combative stance and accuse these mathematicians of "demanding" things, as you do. As I said, the purpose of this is marketing and the statement simultaneously undermines the value of that marketing (showing these projects as irresponsible) and gives these companies an even better piece of marketing in its place: "our AI got so good at maths the mathematicians begged us to stop". It's entirely possible they will stop pouring millions into these projects.

reply
So let’s say OpenAI cure cancer and put every cancer researcher out of work depriving them of intellectual satisfaction, this would also be a problem? It would certainly impact the field.

The fact is these fields are supported by society because of the benefits to everyone else. Once the same results can be achieved in a cheaper and faster way that is what will be done. We should mourn this in the same way we do buggy whip manufacturers. Again people still ride horses.

reply
> So let’s say OpenAI cure cancer and put every cancer researcher out of work depriving them of intellectual satisfaction, this would also be a problem? It would certainly impact the field.

Would you expect Fields medalists to cure cancer if you moved them from the math department to a medical research lab? This is precisely the fallacy that the frontier labs are counting on to inflate their valuation as their IPO approaches. They want to use headline-grabbing problems in pure maths to make their models look "smart" in the public eye. But what does "smartness" in mathematics really mean in terms of economic value? It is not at all obvious whether success in abstract mathematics should translate to successes and, more importantly, profitability, in more grounded endeavors.

Look at OpenAI's job postings (https://openai.com/careers/search/). Those roles involve far more pedestrian yet profitable duties than research mathematics. So why isn't OpenAI automating them with their vaunted models? Success in one field, no matter how "difficult", does not predict results in another field.

reply
> The fact is these fields are supported by society because of the benefits to everyone else. Once the same results can be achieved in a cheaper and faster way that is what will be done.

Maybe it's worth double checking that you know how these fields benefit everyone else? Proving the blowup of the Navier Stokes equations in 3D isn't going to make your gas cheaper or make harvesting food easier or make drones easier to protect against. Maybe consider the deeper effects at work?

reply
Math people absolutely are important to things like the SpaceX landing control systems, hypersonic gliders, 5G networks and so on.

If you can make breakthroughs on such areas as fluid dynamics, control theory or information theory with AI then that absolutely is a big deal with real technological implications.

reply
A long complex proof to something we have been able to simulate in fidelity forever doesn't do any of that though.
reply
Why don't we first assume OpenAI can create a cold fusion reactor and then give them a trillion dollars to get it done?
reply
Chess is even more mainstream and accessible now!
reply
deleted
reply
yeah beacuse no one trust any of their benchmark results now they are scrambling to find a signal thats undeniable
reply
I expected this. They prove a millenium result, but it doesn't count because they are bad people.

It's this sort of thing that motivates people to burn down the institution you might be trying to defend.

reply
> It's this sort of thing that motivates people to burn down the institution you might be trying to defend.

lol, yes, this sort of thing is what many people who voted for Trump were saying, and things are going great for them.

> They prove a millenium result, but it doesn't count because they are bad people.

OpenAI has only themselves to blame for this, and they know it. They could have handled this so much better. I'd bet there's more meeting time right now going into how to unveil future math results than on meeting about the actual math research.

reply