upvote
The argument that most people are making isn't that dsv4.1f is better than frontier, but that it's good enough for most tasks, faster, and way cheaper.

> if one looks at the CoT, it's evident that it's way way stupider than frontier models

Frontier models don't show the full CoT

reply
It's better to look at the results than the CoT. As far as I know, the CoT is censored for US frontier models - it certainly was for Gemini last time I tried. When you get a condensed summary of the CoT omitting all the false leads, incoherent digressions and backtracking, of course it's going to look smarter.

> I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use

You've measured something, but I'm not convinced you've measured what matters, because that's a lot harder than people give it credit for.

reply
DS4F consistently gives poorer results than Opus 5 (speaking of previous generations) and the CoT shows why.

> You've measured something, but I'm not convinced you've measured what matters, because that's a lot harder than people give it credit for.

"What matters" is what matters to you, right? Who, by the way, don't know what "something" is.

Anyway, if you're so sure that DS performs as good as other frontier models, you're entitled to your opinion. For me that's just having low standards.

reply
An older model that didn't yet have obfuscated thoughts? I'm skeptical. I think the non-obfuscated CoT traces I've seen all look similar.

That's right, what matters to me is what matters to me, and the something you've measured I don't know - but that's not a point in your favor.

The worst sin a model can commit in my opinion, is to give an excellent dazzling response to a slightly different assignment than the one you gave it. DeepSeek seems really good at NOT doing this.

But if you ask the model what it expects to be asked, of course you won't have that problem. It could of course be that DS commits this sin, but just happens to expect the tasks I give it.

But I rather think that it's Claude which is good at expecting your tasks - because I have seen all your "high standards" models commit this sin.

reply
The COT isn't an end all be all. Research has shown that the COT isn't necessarily what the model is actually thinking.
reply