upvote
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.

DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.

reply
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
reply
Hopefully some faster providers will start offering mimo-v2.6-pro because it's cheaper and benchmarks better than Deepseek
reply
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.

MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.

I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.

I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.

Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.

reply
I read they identified a training bug and were going to push out an updated release to fix the looping. I really like it overall.

GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.

reply
I did like GLM 5.3 Flash but it's just way too often I'd run it on some long running task and come back to it repeating the same tokens or tool calls endlessly, just doing nothing. It wasn't unusable, but I couldn't trust it. That's really frustrating and I think new models have to do better not just on benchmark scores but general reliability and user experience as well.

At some point Anthropic and OpenAI models definitely could fall into similar traps so I do think it is a solvable problem and likely not a reflection of the models themselves being bad. In this case it may indeed be a training bug of some kind, but I also suspect mitigations on the inference side are possibly lacking or not effective enough for the open models and their runtimes.

reply
The latest Haiku release is almost a year old. Clearly they don't care about the small-but-capable part of the market at all.
reply
From TFA:

>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.

reply
This will be interesting. While no one cared about small models in the last few months except for the OSS community, there is a silent small model revolution with gpt luna and jev. Headless/background llm routines are cost-feasible, which will of course lead to exponential usage and cost.

My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.

Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?

https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons...

reply
> but openai went kamikaze

I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.

reply
Of course, my opinion is based on my personal experience + openrouter data that shows stickiness and low terra adoption; with openai confirming by making sol terra, astra sol.

I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.

reply
> This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

It's been out for an hour and you've already concluded this?

reply