upvote
Yeah, local LLMs are probably an order of magnitude less efficient at a fixed level of "intelligence" if not more.
reply
Inefficient per watt, yes, but local inference capacity is greatly underutilized in aggregate. If a model can run on a machine that already exists, that's a bunch of additional chips that don't need to be built.
reply