upvote
Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!
reply
Sometimes you want a decent model running in the background that doesn't take up all the VRAM.
reply
Or maybe even to run parallel threads of the same model!
reply
They mention 5090 with regards to speed, Q6 will not have that speed?

And speed matters a lot for many use cases

reply
150 tokens per second on a ternary model implies that it’s GPU bound, I’d bet a Q6 model is even faster because it’s existed longer and seen more optimization. You’d have to be insane to not run an NVFP4 quant over a ternary quant on Blackwell if they both fit.
reply
With Ninfer and Qwen 3.8 27b it uses a groupwise int mixed quant, and it gets 160 tokens/sec. The mixed quant is between 4 and 6 bits.
reply
it is like "my fridge is 2mkm (millikilometer) from my desk" m=0.001 h=3600 it should be just Ws or just J
reply
What’s wrong with milliwatt hours?
reply