I wonder if I could get this running through vLLM on 6x Nvidia L4 - the 3.6 worked great on 4 cards but sadly TP6 just isn’t a thing and I don’t have 8 cards available, maybe it’s gonna be okay with like TP2 and MTP. I have no idea at this time, probably need to test out what even might be possible.
This "next" release adds a new concept, first public release with n-grams, I think. And it's in a MoE size that is likely to be very fast and cheap to serve (faster than 27b for sure). It's also well suited for inference on alternative compute (i.e. sparks, macs, etc) so it's relevant to local users.
:)
It is very relevant and for a certain group of us, far more impactful to our work the next month(s) than any blog post could be.
> trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board
Qwen's advances do (currently) have merit.