upvote
The larger 3.6 35B is actually a mixture of experts (MOE). This means a small proportion of those B's are actually active. It's fast and suitable for agentic tasks but nowhere near good as the dense 27B model, which has all of its parameters loaded.
reply
I have found 27 to just be so much more coherent than 35:

https://humanparadox.org/local-vs-frontier-benchmarks-for-my...

It can complete multi-step tasks much better, and has a bit more curiosity.

reply