upvote
Whether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.

If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.

But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.

reply
By the way, this is the same argument that Michael Burry used to short Nvidia.

He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]

The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.

New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.

New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.

In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.

We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.

[0]https://inferencex.semianalysis.com/inference

reply
1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations.

2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation.

> If Amazon doesn't buy but Microsoft does

The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now.

I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.

reply
1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer.

2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens.

3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC.

Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.

reply
I wonder: in world where inference is cheap, how many engineering agents that use simulation as their feedback we will use?

In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system?

And maybe llm's could invent. So even more to simulate. And simulation is inherently compute-heavy.

So unless there are some other bottlenecks, we'll use a lot of simulation servers.

reply
True, I have to agree with you. The AI giants might be investing a huge amount of money in generation 1 technology. There might be a much better way to do it just around the corner. They might know this and thus the hurry to IPO.

A rough analogy would be if the first generation of ISP's spent billions on dial-up exchanges, when fibre could be invented next year.

reply
On the plus side, lots of cheap servers to swoop up :)
reply
But power hungry.

In that 5+ year timeline, the compute per watt could change by three orders of magnitude.

GPUs are to LLMs what CPUs are to gaming — not a good fit.

reply
A cursory estimate courtesy of ChatGPT suggests that there is a grand total of one order of magnitude or less of power efficiency improvement available compared to current Blackwell if the entire system’s power consumption outside the ALUs went all the way to zero.

If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, different ALU design, model architecture changes, etc.

reply
Look at their power supply, it’s not something you can run in a home lab. Unfortunately most of that will likely go to the bin eventually :(
reply
Exactly right, and nVidia is protecting their moat through business practices rather than genuine product innovation.
reply
deleted
reply
By the time these gigawatt datacenters are done being built the hardware will be so far behind state of the art they may be mostly useless.
reply