upvote
I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.

Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

reply
Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.
reply
I'm not denying that, but I'd still like to know what that cost.
reply
They did say that. "3 hours of ChatGPT Pro thinking compute"
reply
Yes, what does that mean?
reply
It means the level of effort that a ChatGPT Pro plan summons when thinking. For 3 hours.

If folks are going to analyze this claim with a critical eye, I'd be zeroing in on "average" rather than acting like this measurement is somehow unclear.

reply
>OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]

That doesn't sound right

reply
Oops, edited.
reply
Oh cool, we will all now have a math genius on our computer.
reply
It was using their internal math model, so not yet for us
reply
I used future tense. It was implied this will be available.
reply
Well, on their computers. But you can rent them for a price.
reply
An open source model will reproduce it 6 months later
reply
which you an run IF you have the hardware. who knows how heavy these models are.
reply
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
reply
It doesn't imply that, it's just measuring the amount of compute.
reply
But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?
reply
Not necessarily, could be agent swarm with low N
reply
At some point that stops being a "swarm" and just a handful of subagents. E.g. 6 agents running for 30 minutes each isn't really a swarm in my eyes.
reply
That estimate is obviously going to conveniently ignore all the failed runs.
reply
deleted
reply