you cant. The best you can do is Qwen 3.6 27b with a 24gig ( or cumaltive gpus ) to get to 24gb vram. ala 3090, mac with 36gb ram, amd cards, halo strix amd, dgx spark etc. Lots of youtube videos out there.
I haven't seen either of these running outside their creator's services yet, but typically you can watch services like openrouter or nano-gpt for it to show up at a (usually small) discount.
Well, if you're happy with around (as in within an order of magnitude or two of) 0.1 tokens per second... I believe that's around what people are getting when loading MoE weights from NVMe.
You've got to consider how much power that uses though. Depending where you live, some of these providers can serve it for less than you pay for power.