upvote
DS4F consistently gives poorer results than Opus 5 (speaking of previous generations) and the CoT shows why.

> You've measured something, but I'm not convinced you've measured what matters, because that's a lot harder than people give it credit for.

"What matters" is what matters to you, right? Who, by the way, don't know what "something" is.

Anyway, if you're so sure that DS performs as good as other frontier models, you're entitled to your opinion. For me that's just having low standards.

reply
An older model that didn't yet have obfuscated thoughts? I'm skeptical. I think the non-obfuscated CoT traces I've seen all look similar.

That's right, what matters to me is what matters to me, and the something you've measured I don't know - but that's not a point in your favor.

The worst sin a model can commit in my opinion, is to give an excellent dazzling response to a slightly different assignment than the one you gave it. DeepSeek seems really good at NOT doing this.

But if you ask the model what it expects to be asked, of course you won't have that problem. It could of course be that DS commits this sin, but just happens to expect the tasks I give it.

But I rather think that it's Claude which is good at expecting your tasks - because I have seen all your "high standards" models commit this sin.

reply