upvote
An older model that didn't yet have obfuscated thoughts? I'm skeptical. I think the non-obfuscated CoT traces I've seen all look similar.

That's right, what matters to me is what matters to me, and the something you've measured I don't know - but that's not a point in your favor.

The worst sin a model can commit in my opinion, is to give an excellent dazzling response to a slightly different assignment than the one you gave it. DeepSeek seems really good at NOT doing this.

But if you ask the model what it expects to be asked, of course you won't have that problem. It could of course be that DS commits this sin, but just happens to expect the tasks I give it.

But I rather think that it's Claude which is good at expecting your tasks - because I have seen all your "high standards" models commit this sin.

reply