upvote
I'd be skeptical w.r.t. "by all metrics".

Qwen3.6 is a definitive, significant downgrade from Qwen3.5 for creative writing and prose for example. Yes, it's better at agentic and coding, but it regresses in many non-coding areas compared to Qwen3.5.

Of course, I do expect the 3.8 ones to perform better for agentic coding.

reply
One thing I would caution is staying out of the prediction market like this.

Tech tends to get boring when you judge current products against the hypothetical capabilities of unannounced products that may never ship. It's like comparing Nikon cameras against Canon camera rumours, or comparing iPhones against unannounced and therefore largely imaginary Samsungs.

- If they do a Qwen 3.8 35B A3B (and I hope they do because I love the 3.6 version)

- and if it beats 3.6 27B by all metrics

… then the local open weights world will be a better place.

But they have said nothing about it and they dropped several weight classes for 3.6, so who is to say they won't drop the 35B? And even if they don't, this is a tall order; why would the MoE tradeoffs no longer be apparent? (Again, I really like both the Qwen and Gemma MoEs)

FWIW I am enjoying testing Muse Glimmer — it's really quite impressive on chat, has nice terse and even amusing thinking traces, a bit of brass to it, and I'm hoping it will be good on agentic stuff.

reply