I think most important thing is that Nvidia doesnt want to give 100% of market to frontier AI labs either.
It's way too easy for 1T+ frontier labs to ditch Nvidia. So Nvidia will also put effort to make sure there are open weights models and local hardware available.
There is zero chance that an LLM approximating a modern frontier model is going to be running on a phone in the next decade. Even if you grant that you could stack enough DRAM dies on top of each other in the package, that would be a three order of magnitude improvement in power efficiency just for the compute.