It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.
Yeah, this is what I'm holding out for, the NVFP4 variant of 3.5 122B is blazing fast with reasonable quality and even with max context fits perfectly within 96GB.
Edit: as a concrete example, I'm working on a "optimization framework via agent harness" right now, Qwen3.6-27B-NVFP4 is often unable to actually complete the optimization within 100 turns, while Qwen3.5-122B-A10B-NVFP4 has no issues finishing within ~50 turns or so.
But why are you using Qwen3.6-27B-NVFP4 compared to the FP8 or full version? In my experience the Q8 of 27B is on par sometimes better than 122B. I am experiemnting witb higher quants for 122B to fit on my Strix Halo, but still, the difference honestly for my workflow is not that much. I just wish they released 3.6-122B version.
The fantasy is a 100B or 80B model, but MoE and highly tuned for coding.
I wholly disagree. Rather than going the "everything is a claude code skill" route, I've been hacking together purpose-built harnesses for all sorts of tasks, and in that environment a wee little baby model can do some really useful things. You end up burning lots of tokens making the thing, but then all that investment comes back when the resulting tool works perfectly fine on a dinky little model that fits on my 3060 Ti.
Europe will definitely be interested in democratizing these things if China starts losing interests; from there, there'll be more countries looking to keep their citizens entrained in their own Country's infrastructure.
It'll especially be true if the memory cartel keeps prices high and NVIDIA tries to gouge higher memory models.
It's an arms race everyone can join because PC hardware was mostly democratized in the last decade.
I dont see most model building as anything more than a pig at a slop troth, despite the level of sophistication; they're still rarely pruning the input beyond random sampling.