See:
> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)
They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.
[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...
I know that the former and the latter may be discrete subsets of the anti-AI crowd, but come on.
Edit: "Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above."