So yes, I think your scenario is likely to eventually happen, but there will be a much more powerful, capable frontier model then.
It's almost a given that whatever is frontier intelligence today will run on a potato in a few years.
I'm pointing out that there's no known information theoretic constraint about the impossibility of frontier AI models being improved to fit/run on a small GPU.
Please do not make up plausible sounding science facts.
There are constraints of course- training takes way longer.
People will claim to have “enough” even though they already have the equivalent of last years capabilities locally.
the biggest winner in that scenario would be ai providers, who suddenly have a capable model that they can serve much more efficiently. and the incumbents have a whole lot of compute. wouldn't anthropic and openAI just start offering that open weights model at prices that nobody else could compete with?
Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre.
For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length.
Like even if you run it in a datacenter in this scenario, you could do it on a cheap GPU instance in Azure, you still wouldnt need OpenAI or Anthropic specific clouds.
>uses 300W of power to do so.
There are plenty of people with phat electricity pipes in their on prem server rooms that have been vacated for cloud. Companies who want the benefits of AI but dont want the risk of sending their data to foreign API endpoints.
RTX 5070 prices go up ~N times. Nvidia makes more money because it's easier to make these things than it's to make a GB300.
If it does happen then NVidia will sell a lot of 5070s though!