upvote
There's been a ton of optimizations already, it hasn't remotely reduced demand even temporarily. More efficiency just makes the compute have even higher ROI per $ and watt spent.
reply
With sufficient optimisation, there ought to be a tipping point beyond which local inference is good enough. And, sure, datacentre compute will still be needed for training but one of the biggest current uses will begin to taper off.

The question really is how soon we reach that tipping point, and whether it's before or after the current bubble runs out of steam for some other reason.

reply
>there ought to be a tipping point beyond which local inference is good enough

There's no such ought really. Even at current levels you'd need like a 100x gain from here to approach current top proprietary models (probably a lot more for say Mythos or Mythos 2), and it's not like they are stoppng to improve. This is before we even account that you'd just be running 1 agent then, and not a swarm like you'd be able to in the cloud or that you can do only so much compression before you are losing out

reply
> (probably a lot more for say Mythos or Mythos 2)

Not everyone needs that large of a model, though.

reply
It's not just inference, some things done in data centers like simulations, testing, are complementary to inference.

And in these types of hardware, the time between a successful prototype and a fully deployed product is pretty long. Maybe they count on that to know when to stop?

reply
Jevons Paradox shows that increasing efficiency can increase demand for a product by making it cost effective for more uses.
reply
deleted
reply