upvote
Qwen3.6 is very token inefficient with it's thinking. Quantized versions often get into loops.

Glimmer is trained with 4 effort levels, not just thinking on/off. Maybe it's more token efficient in general. There's official 4 bit quantizations with reported 1% loss across 15 benchmarks -- so quants probably work good.

IMO that alone is worth trying for, even if they're otherwise equal.

reply
It takes 10 minutes to download and try, any model is worth at least that. In my experience benchmarks are generally dogshit
reply