Also consider that for things like cyber work, the frontier models give you nerfed results and poor performance. Whereas the open weight isn't nerfed, and I routinely get 200t/s with my subscription. Finally, DS4.1 Flash is natively multimodal, while Haiku isn't.
They're all perfectly fine models, you should use any one of them you want. But DS4.1 Flash can do more for less. (That said: GLM-5.3-Flash is even better and cheaper...)
oai and anthropic also subsidize the hell out of their subs compared to what you will find in smaller competitors, a $20 codex sub gets you like $100-150 usage/wk which goes way further than 2x opencode go (which would only net out to $120 of deepseek 4.1 usage a month, on top of being low performing quantized trash).
For $20/month subscription, Charm Hyper gives you $12.50 per day, for a total of $350 per month. Like I've mentioned in other comments, OpenCode Go performance and rates are terrible now, there are several better options.
Quantization is not trash, there's a year of evidence that shows Q4 provides ~4% degradation and Q8 provides ~1% degradation, and you don't need that severely quantized to gain benefits in inference performance.
Are they. Luna uses way less tokens for identical tasks so its a bit of an apples to oranges comparison.