upvote
There’s a lot about this that would make sense.

But, not really at the current technologies. Kimi and GLM are fucking awesome, but I don’t have 3TB of VRAM to run them, and I don’t expect to even when ram prices drop.

So now you’re back to the scaling issue before talking about power and compute distribution.

reply
Do you think you'll realistically need 3 TB RAM to run a sufficiently good model 1 to 2 years from now? I certainly don't. Considering what can already be done with 128 GB of relatively slow unified memory, imagine if efficiency improvements continue apace, the memory becomes 128 GB of HBM, the flash device becomes capable of sequential throughput matching today's DDR5, and such a system was affordable as a routine purchase for the average person.
reply
Why do you think the public wants to self host models over using a cheaper solution hosted in the cloud?
reply
Almost all the non-tech people I know acknowledge concerns about privacy but have no realistic alternative available to them. I think they won't want to self host until and unless it's easy to deploy, there's nothing to administrate, and the hardware cost is in the same ballpark as a new laptop was prior to the apocalypse.
reply
Subscriptions are almost never actually cheaper.
reply
It is cheaper to subscribe to AI for $500 a year for the rest of your life than to buy a machine capable of running the a current frontier model with no subagents.
reply