I feel like cost competitiveness of local has been going backwards if anything, compared to API providers. I assume it's because API providers can reuse hardware more and have other efficiencies (batching?). Do you see any reason this might change?
Because everything is converging on a backlog of huge efficiency gains established in research, waiting to be combined. Looped transformers, a whole host of diffusion techniques and new quantization techniques, maturation of ternary distillation and new ways to separate logic from stuff that can be looked up. It would surprise me if most frontier models were actually even that big at that point in terms of active params. I highly doubt it.
But he gave a reason for that. "Before that happens the frontier will move". Why do you think it will happen anyway? Do you think the frontier will not move fast enough that local models are unable to catch up, or do you think people will prefer local models at a point. Or something else?