upvote
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
reply
> accidental artifact of trying to make it give honest answers?

If it's not giving honest answers that implies it's purposely being deceitful, which it isn't capable of. Right?

reply