upvote
Oh!

I think I missed that and assumed it wasn't because it performed similarly to the dense models. Interesting!

reply
As an aside, the bigger S 2.1 runs faster than the dense Qwen 3.6 and Gemma 4 models on the same hardware (assuming the same hardware is big enough to run it) and definitely feels smarter, and more capable of long tasks, but it doesn't feel as heavily tuned for code as Qwen 3.6.
reply