I think your suspicions are warranted and your explanation seems plausible.
If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than humans can. This is great!
But there's so much vested interest in the AI companies to be opaque about all this, to hype up their models and avoid giving credit to people whose data made everything possible, that they would never tell us this fact if it were true.
I feel like so much of the AI hype cycle is like this. The models develop extremely useful capabilities, but it's hard to understand what they really are through the hype. The lies and obfuscation by their owners who have vested interests in capturing the value they provide makes it impossible to take anything they say at face value.
It's perhaps great in the short term although it's not very clear who it's great for. I'm not sure mathematicians find it all so great, I mean.
In the long term, if this contrives to destroy the tradition of human mathematics the whole endeavour is self-defeating. In time, there will be nobody left with the knowledge and skills to produce mathematics to train AI to do mathematics.
And then we'll be left with no mathematics at all: we'll have no human mathematicians and no AI that can do mathematics, either.
I don't know what you should do because I'm not a mathematician. But superintelligence schmuperintelligence. We didn't stop running because we have cars or playing chess or Go because there's chess and Go engines. Even more so than chess there's no point in maths unless it's people doing it, for other people. AI maths makes no sense, like AI art makes no sense, because those are things that people enjoy and can do pretty damn well ourselves so there's no point to automate them away. We gotta stop that bullshit, and we can stop it. And if we don't, if we just sit around and wait for OpenAI and Anthropic to destroy society then that's not their fault but ours.
Sorry, I'm not great at pep talks. Those are brave men. Let's go kill them!
If this is not possible it does make me question whether mathematics ever had any value except for economic or industrial reasons. I do believe it does however, so it must be possible.
And like - I think there’s a presumption you could make that AI models could overfit to asymptote towards just the capabilities and knowledge we currently have.
And that would be amazing! And crazy useful. And there are probably a whole world of complex problems that remain unsolved because they’re adjacent to knowledge we have but they haven’t been invested in.
But can a human reliably tell the difference between “can do 99.999% of the things we currently know how to do which includes a small subset of things we didn’t know we had the capacity to do” and “super intelligent math and science research pushing the frontier of what we know”
A physicist that knows all the things we currently know in excruciating detail feels like it should be able to make the leap beyond the frontier.
But since these are computer models it might just be that it can ride that line extraordinarily well while the line remains firm.
I've been thinking along exactly these lines... they very well could have a 21st century Mechanical Turk and its real superpower is getting people to "collaborate" asynchronously but it's just stealing their ideas and laundering them.
I don't think it's purely that, of course... but "consult other clients' transcripts" would be an easy tool to write.
Sam Altman knows what he’s doing. He will happily screw these folks to one-up his competition.
1. Systems that OpenAI is able to use (either public or private) are improving rapidly at open problems, even if they are still extraordinarily expensive
2. Researchers will inadvertently speed up the rate at which the AIs improve by feeding them valuable training data
This is pretty much the definition of a data flywheel.
Maybe I'm failing to read that graph properly but the y axis says "pass rate" and it only goes up to 0.5. That would mean every single problem is at most half-solved.
I don't know what that means though. What is "0.5 pass rate" in the context of "open math problems" (as in the graph title)?
This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.
Plus it is an unfair standard since so many scientists in the past have been caught unethically using the work of others without attribution (and so many more have been accused).
In history we also repeatedly see the phenomenon of multiple discovery or simultaneous invention. If that happens to AI because the topic is pregnant, would you call it "plagiarism" just to disparage AI? https://en.wikipedia.org/wiki/Multiple_discovery
How is it an unfair standard. OpenAI stole the work of others to build the AI. That's not different than scientists stealing from other works as their own, or artists copying others work as their own, etc. It's all plagarism. I'm applying the same standard for everybody.
As for multiple discovery, this is a thing, but I don't think the AI did a parallel discovery any more than Ray Kroc made the parallel discovery of the MacDonald brother's speedee service system.
I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability.
No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is.
I’m very pro AI long term btw but I’m not blinded.
Brute force would have been solving Navier-Stokes in 88 hours after plagiarizing all known 20th century math
When it needs to snoop live on what the actual mathematicians are working on that’s something else
Trust me I've seen it happen to myself. I no longer trust ChatGPT.
I can see right through his act. Altman is one devious f8k.
But when the building-block ideas are still being formed, I'm not sure that AI is good at forming them.
There are many actions being performed today that can be nicely packaged.
Im already working on such a project.
Just another rich man’s trick
Perhaps the last one before they destroy that world and try to hide away as people forget and history is rewritten again. I don’t think they’ll succeed this time.
It seems to me the academics are upset that AI scooped them. But scooping is a time-honored tradition between researchers. First to print and all that. In a nutshell, they are upset that they lost out on a publication.
I will also point out for those unaware that any mathematics that is produced is automatically part of the public domain and can be used freely in derivative works. It is not a protected intellectual class like other works of art.
Provided that it's properly accredited. And definitely not for others' unpublished work -- that's despised upon if not an academic integrity issue.
People even point out that you should add a reference to certain papers during the peer review process.