While it's great to see tok/s go up as high as possible, I think it's important to consider the actual usability of these models when you quantize down to something like 2-bit. From what we've tested it seems like going below 4-bit quickly leads to serious issues with thinking, tool calls, and overall model coherence.
Our plan to enable running bigger models on less GPU memory in a way that'll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.