upvote
I'm using DSv4.1 in OpenChamber (eg OpenCode) using the Superpowers skills and a lot of custom AGENTS.md instructions to iron out the kinks and I genuinely cannot see a difference between it and Opus and I've been building native iOS and AppleTV apps, Go servers, Typescript, Cloudflare workers, Svelte/Astro, etc.

It's a super capable model all around from my experience.

reply
Sol 6.1 on medium seems to benchmark better for less money than deepseek 4.1 flash: https://artificialanalysis.ai/#intelligence-comparison-tabs
reply
How are you getting down to $1-$2 per day running "all day long"? I've been using GLM 5.3 Flash and I also spend $1-$2 per day, but my use is pretty modest, I think. DS 4.1 Flash is priced similarly to GLM 5.3 Flash; can't imagine DS is significantly more token-efficient.
reply
Luna is 1 point being on AA's index at 1/4 the cost, yes it "doesn't score as well" but paying 4x for 1 point is crazy if you're going off benchmarks.

AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn't ideal but what can ya do) and a 4 point intelligence gap.

Why do people like to think open models are more competitive than they are?

reply
Because it's just Opus 5.5 and Haiku, and on OpenAI side Luna, that changed the calculus. Also, I find all of them, including GLM5.3 and DS4.1, to be 100% en par with American's frontiers model for CRUDS (90% of enterprise programming).

Tangentially, all of them would have broken quite badly custom ERPs from my own experience.

reply
DeepSeek's own paper advises against using Max, showing that it normally doesn't perform that much better. I am not using it on Max, so that's not a useful benchmark for me. I have seen other benchmarks where Flash does significantly (30%) better than Luna.
reply
i mean we use max benchmarks because they have the best coverage, which sucks because very few people use max day to day, but it's what we have. the performance curve is generally pretty similar across models and effort levels, weird outliers are pretty weird. max is generally a big cost bump from most providers (less so from OpenAI)
reply
It is super bad on a bit more complex workflows and starts repeating same errors with the same tool until the cycle breaker hits.

6 is worse than 5.6 here.

But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre/bad job on reports.

reply
You can't use a flash version and complain it doesn't think well you would use the pro version.
reply
While I agree, this particular discussion chain really frustrates me.

1. "This Flash model is really smart. Here is an article to discuss how smart it is. Why aren't people freaking out about how smart this Flash model is?"

2. "I tried using it for a smart thing. It doesn't work so well for it."

3. "You should know better than to use Flash for smart things. It's not meant for smart things."

reply
It's important to move the discussion past this model is smarter. The flash version is great at implementing a plan but because it makes trade-offs for increased speed it doesn't take the time to think through everything. Some models are great at making a plan but can't follow through.
reply
DeepSeek themselves said 4.1 Flash is better than their current Pro version.
reply