upvote
We could make LLM inference 100x cheaper to run at home efficiently, but that solution might need to be updated every 1-2 years, whereas current GPU are useful for various others tasks and last longer
reply
"Never" is a long time. Just think about how much ram we had 10 or 20 years ago. 1.5TB isn't a lot really.
reply
It doesn't matter if you have the RAM, running a 1.5TB model for a single context stream is fundamentally inefficient.
reply
The typical ram has surprisingly not increased very much in 10 years.

> April 2016, 8 GB was standard across the 13-inch MacBook Air range

... Now it's 16.

Rich nerds will have quite a bit more. But I suspect the standard of model rich nerds want to use will have gone up somewhat too.

reply
‘Never’ is a big word in the computing world. 10 years from now a model this size will probably run on a high-end phone.

Of course, by then we’ll want to run something commensurately larger.

reply