But, not really at the current technologies. Kimi and GLM are fucking awesome, but I don’t have 3TB of VRAM to run them, and I don’t expect to even when ram prices drop.
So now you’re back to the scaling issue before talking about power and compute distribution.