upvote
This seems to imply that training will ever be done? But yeah, I think the idea is that the appetite for thinking-on-tap will be enormous.

Even for smaller models, I think they’ve found that training an enormous, inefficient model and then distilling it internally to something much more efficient to serve is the way to go.

reply
> Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done?

I have no real data to back this up, but that has always been my assumption.

Claude says that K3 can be assumed to have required 10-100M GPU hours. If you have 100k GPUs that would mean like 6 weeks of training. 100k GPU's can serve 3-30 trillion tokens of K3 per day. Google apparently serves ≈100 trillion per day [0].

The big labs probably want to have capacity to fairly quickly train / post train different SOTA models continuously + being able to serve peak inference demand in valuable markets (US daytime?).

[0]: https://blog.google/innovation-and-ai/sundar-pichai-io-2026

reply
It's unclear how big a role distillation plays, but it may be a big one. There's also a law of diminishing returns. To get a meaningful increase in model quality you seemingly need an exponentially larger model. And most people don't think K3 is actually on par with top closed models.
reply
Google says 1 million blackwell gpus are being delivered monthly I'm really curious if it's companies just hoarding chips/memory/servers awaiting to be deployed in data centers not ready yet for months or years, or everything built is actually deployed upon delivery. Plus google and amazon have there own chips in the mix.
reply
Your just missing many things and so have an incorrect picture of the situation - k3 is not on par with fable or astra. Closed models remain far better than best open weights at least today. - compute is not just used for training, more and more is inf - even in training you don’t do one run, you do many. Final run is a small portion of total compute.

The world is extremely compute constrained currently, like extremely.

reply
This is evidenced by the prices on every large-model capable device/node increasing significantly over the last year. Demand is far exceeding supply.
reply
For inference (Anthropic and OAI are B2C on top of B2B)

To train much larger models. It is quite possible that 10T-100T models be on the horizon

reply
> what are we (in the US) even building these super massive data centers for?

Partially to make investors think it's worth giving US companies a lot of money. Also, I think distillation is a significant part of why Chinese models perform as well as they do. I think that's completely fair play (OpenAI/Anthropic/Google/Meta stole a lot of their training data). But I would expect if US model developers stopped right now Chinese development would slow down.

US companies are clearing the path, others follow in their wake.

reply
Today's data centers are being built for yesterday's inference need. There's a persistent cult belief that ai hasnt found a niche or that companies havent proven utility or use cases or whatever. the demand for ai (internal to hyperscaler, and external for everyone else) simply dwarfs what is available.
reply
Both Anthropic and OpenAI have been having major load issues though. Up until yesterday, OpenAI was serving at only 30t/s per default.
reply
You are in violent agreement with the comment you replied to.
reply