upvote
I think the question is even a bit more nuanced than that. Even if frontier models can maintain a big gap that gap has to actually _matter_. If a local model satisfies my everyday use cases adequately then I may not really care that a frontier model is 5, 10, 100x better at ultra high order reasoning tasks.

I think that reality is probably not all that far off for a huge swath of use cases.

reply
This is exactly the mainframe vs PC dynamic.
reply
100%, I thought about writing that out explicitly. I really feel like we are extremely close to reaching that kind of breaking point for most folks LLM use cases.
reply
Hell, Bonsai Labs 27B parameter model can run on phones with their ternary implementation which is quite efficient. Scale that up to frontier model parameters and it's quite likely we can run them on current laptops.
reply
Came here to say that, my bet is that in 3-4 years you'll be able to run Fable-level of intelligence models on your laptop or maybe even on you phone
reply
But isn't there the raw intelligence of a smart model and then the practical intelligence fuelled by how many parameters it has? You probably will barely be able to fit a 70 billion parameter model on a phone in 3-4 years let alone a 2+ trillion parameter model... so it depends on what you call intelligence
reply
I'm not willing to believe in phone-based frontier models anytime soon. Though, Gemma 4 12B is a beast that runs comfortably on the current top of the line phones (or would run fine if allowed to run, I think there's some kind of 6GB limit on iOS, and 12B is ~7GB). I'll believe in three years we'll be able to run ~30B models on the best phones. That's 16GB in a 4-bit quantization, and I believe ~30B models will be competitive with 120B models of today, based on the curve we've been on. Qwen 27B and Gemma 4 31B are competitive with much larger models of a couple years ago.
reply