This claim isn't really outlandish in any way. It's not hard to imagine:
- Future models being able to handle current frontier models' workflows with much higher efficiency.
- Future consumer devices like phones having 2-4x the RAM onboard along with GPU/NPU performance greatly increased in 10-15 years.
Performance, storage, etc is definitely getting better, but it's a different scale of improvement
It could be that the company valuations crash tomorrow, and (almost) only performance gains achievable on hobbyist-level hardware come to fruition from there on out.
Or it could be that in the future, we have a custom "model FPGA" à la Taalas [0] in every home, and that it turns out we can still massively boost inference efficiency due to novel discoveries like TurboQuant [1] or a somehow-improved quantization method [2] again and again ten times over.
Point is, Moore's law in this context shouldn't be applied to just hardware spec sheets alone, but more the total number of "parameters potentially improving", IMO.
[1] https://research.google/blog/turboquant-redefining-ai-effici...