upvote
Maybe, but those higher quants would need a large GPU accessible memory space - and the obvious candidates such as DGX Spark and Strix Halo don't have the bandwidth to run 27B at high quality quickly.

With Flash Next you only have ~6B active parameters so you can toss experts up into VRAM and/or run them on a CPU if you have enough RAM and bandwidth.

reply
The benchmark indicates the IQ3_XXS quant beats 27B. I've switched to that now and am ditching 27B. Genuinely better results so far.
reply
I'm not going to knock off my 27B-Q_6 for this. Good to to experiment though.
reply
from my experience dense models like 27b suffer less from quantization compared to large MoEs
reply