upvote
It is comical at this point. Some people just can not stand the thought of AI actually delivering and are trying to find whatever ways to discredit it.
reply
This isn't really about delivering - it's more about helping to understand the shape of problems that AI can solve right now. If they took 1000 problems and threw the model at it and it solved these ten, is there something we learn about these ten problems and the kinds of things that current AI is good at? That's very different from picking ten problems _at random_ and solving all of them successfully, which would suggest a much less bumpy capability surface. It's interesting and it would be good science to release it.
reply
That's totally disjointed from anything in this thread. The main accusation is that openai is cherrypicking math problems and we should be against these results. As if a mathematical proof stops being provably correct because it was cherry picked

And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.

reply
It's disjointed?

The post that started this sub-thread asked:

> 1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?

I think it's an extremely relevant question to ask, because it helps us better understand the current state of AI being able to handle math, for exactly the reasons I outlined. I was arguing against the idea this is just a reactionary anti-AI kind of question to ask. It's not! You can be very impressed by what AI is capable of in math (I am) and still think those are really interesting things for OpenAI to disclose (I do).

OpenAI specifically called out a $2000 per problem average, which implies something that's probably not true ("if you throw $2k at us we'll solve an open problem for you"). It would be cool to know what the actual number is.

reply
It just feels silly to haggle about the price here. It doesn't even matter because it's going to drop by an OOM quickly.

If these 10 problems were solved by humans, it would be pretty impressive, even if it took a large number of researchers! Yet when AI does it, HN commenters suddenly feel the urge to play accountant.

reply
> In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it).

But that's the start of math research, not the end.

The point is to get practice and experience doing research.

Did ChatGPT learn anything from these proofs, that it can build on?

Part of what's annoying people is that ChatGPT is churning though problems that are meant to be motivating. They are problems that aren't worth the effort of human professionals (usually because they are incredibly computation-hevy, so better suited for a computer than a human), so they are good for students to work on.

reply
It's not about discrediting AI. We know LLM is a commodity technology like electricity at this point. If somebody in 1900 claimed they had a setup at home where they feed in electricity and cool air comes out the other end (meaning they invented AC), obviously people would want to know what the setup is, so everybody can have AC.
reply
Entirely comical too that some people can not stand the thought of people poking very big comulent holes in the claims of AI delivering what it claims to deliver. As if having skepticism is some how a way to discredit a person.
reply
Sure, but I don't really understand what the argument is to _not_ be transparent about methodology, since if the models are so powerful, then doing so would easily support the claims and put these concerns to rest. People are right to be skeptical given what is being implied and the orientation of the narrative

I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."

By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?

I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?

To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations

I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though

reply
deleted
reply
Yes, but if their machine god really is as good as they say it is, why are they constantly resorting to statistical sleight of hand at best and outright lies at worst with every public statement?

That's not normally how people act when they're confident in their product

reply
You need to bring both cost and benefit into the argument, and it's not necessarily an obvious win for either side. There are a few complicating factors here.

The cost of running a model is not only $/token, but the salaries of the people managing/orchestrating the models, deciding what theorems to try, etc. Once we factor that in, how much are we really paying per theorem?

The other factor is the subjective component of the value of a theorem. Not all theorems are created equal, and the only way to really measure the value is to ask professional mathematicians for their opinion, or publish the results and look at citations over months/years.

Once we have both of these nailed down, then we can start to do the cost/benefit analysis. To be fair, we should actually compare three groups: human experts, hybrid agent/human expert teams, and fully autonomous agents.

reply
it's still important. not everyone has access to 1 million USD. saying it "only" coat 2000 USD is highly misleading for the discussion and future. the concentration of power is a huge problem with AI.
reply
If you told them this was the problem and they would still have a job if they failed probably. The reasons people don't go head on these problems is career incentives and psychology.
reply
It would still provide better context to see the numbers that the parent proposes, though.
reply
Mentioning cost is fine, comparing may not be.
reply
Grad students on zero pay solve problems like this everyday. What exactly is your point here?
reply
Everyday? Which 10 problems were solved by mathematics grad students in the past 10 days?

OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?

reply
[dead]
reply
Even if they can solve problems like this every day, you still have a very limited number of grad students who can solve them. With model capabilities like this, you can have the equivalent of millions of grad students who can solve problems like this.
reply
Grad students do not solve problems such as the existence of non-sofic groups every day.
reply
Plenty of 'advances in mathematics' done pre-llm, no?
reply
Give the grad students these resources and they can do even more!!!
reply
Zero pay? These would be PhD candidates; surely they have a stipend?
reply
Nopes, often times especially in math they get paid due to teaching duties (at least in the US). So technically for the math research part they are not getting any stipend.
reply