Both MoEs and dense models are always getting better, so I don't think this comparison is meaningful across generations. But still for a first approximation, this tends to hold (you wouldn't expect a lot from a 10B model in coding yet).
Don't use ollama. The entire project is just a series of stupid decisions like this.
Also try fine-tuning small models and see how they perform.
I think I missed that and assumed it wasn't because it performed similarly to the dense models. Interesting!