And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.
The post that started this sub-thread asked:
> 1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
I think it's an extremely relevant question to ask, because it helps us better understand the current state of AI being able to handle math, for exactly the reasons I outlined. I was arguing against the idea this is just a reactionary anti-AI kind of question to ask. It's not! You can be very impressed by what AI is capable of in math (I am) and still think those are really interesting things for OpenAI to disclose (I do).
OpenAI specifically called out a $2000 per problem average, which implies something that's probably not true ("if you throw $2k at us we'll solve an open problem for you"). It would be cool to know what the actual number is.
If these 10 problems were solved by humans, it would be pretty impressive, even if it took a large number of researchers! Yet when AI does it, HN commenters suddenly feel the urge to play accountant.
But that's the start of math research, not the end.
The point is to get practice and experience doing research.
Did ChatGPT learn anything from these proofs, that it can build on?
Part of what's annoying people is that ChatGPT is churning though problems that are meant to be motivating. They are problems that aren't worth the effort of human professionals (usually because they are incredibly computation-hevy, so better suited for a computer than a human), so they are good for students to work on.
I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."
By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?
I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?
To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations
I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though
That's not normally how people act when they're confident in their product