upvote
Beyond benchmarks, does it in day-to-day? Have always struggled to get competitive performance out of any Deepseek release going back to V3 vs Z.AI and Moonshot models. Maybe I really suck at whatever is needed to make DS models fly, but even tailoring my suite hasn’t gotten me far when I tried with V4 Pro. Happy for anyone who is able to leverage their models well, wish I’d be able to crack how to leverage them.

Will say their research is some of the best reads in the industry and I could not care less about their model release cadence as long as papers keep coming.

reply
Yes. V4.1 flash performs really well. I don't know what to say, maybe you don't believe, but my team has been using it mainly for almost a month now for programming. It is as bad and annoying as any of the US sota, but costs pennies. And if you look close enough you find providers that can push it 300-500 tokens per second...

Maybe it is due to us being all very experienced devs. And can steer the model. But my daily routine is just to have 8-9 Zellij tabs open, DeepSeek in omp in each, and grind research and code day and night. Really nice model...

reply
Is GPT Luna 6 dethroning Deepseek V4.1 Flash? It's price seem to be undercutting flash at a relatively similar capability.
reply
It is not super good in long horizon tasks and worse than 5.6 in our evals. It really failed the agentic evals where DeepSeek, Kimi and Opus are the winners.

It is great on creating summaries and content.

reply
Yeah, very similar benchmarks at 1/4 the price. I've been pretty happy with Luna 6 though I still think the gap between small and frontier models is larger than many people want to admit.
reply