At that point just rent proper GPUs in the cloud, you'd have way more power and pay only what you use for.
I have a bit of paranoia/anxiety about AI, but it's not what most people are concerned with. I understand the limits of these tools very well, and still find them extremely useful. What concerns me is that it's going to become difficult to impossible in the future to run local models which have near-SOTA capabilities in a way in which you can exercise full control of the model. I see the writing on the wall, and its more than worth it for me to invest early to ensure my own capabilities. I am very much not a fan of our "you'll own nothing and be happy" directionality for the world, and I am (at least currently) privileged to have the means to slow that decline for my own self.
Are your referring to Taalas / chatjimmy ?
I definitely think we'll see an ASIC-like approach in the future, especially for embedded small models where it may require minimal silicon area and can result in near-realtime performance. But at the frontier, I don't think this is a solved problem and will continue to have model weight churn that will advantage more flexible general-purpose hardware.