upvote
> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

See:

> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)

reply
Sure and if I make a half court shot after an hour of trying, the result only took 1 second.
reply
Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.

They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.

[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...

reply
That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.
reply
In a couple short years we've moved from "AIs can't do anything useful" to "they're lying about the actual cost of the innovative breakthroughs!".

I know that the former and the latter may be discrete subsets of the anti-AI crowd, but come on.

reply
It’s very typical in human math that explaining the final result after years of searching looks very simple too.
reply
Fourth, these hundreds of solved problems are the result of OpenAI attempting tens of thousands of problems and failing. When you hear claims that the average result took about 3 hours of model time, I simply do not believe it. If you account for all the time spend properly, it's probably orders of magnitude more.
reply
I think the announcement says they report the amount of problems attempted somewhere.

Edit: "Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above."

reply