upvote
I assume that they're talking about, for example, training the model to produce code with predictable yet difficult to find vulnerabilities.

Imagine if every time <INSERT MODEL HERE> was asked to code up a web server, it made sure there's a subtle buffer overflow that grants a remote attacker RCE, whoever trained the model could then start scanning web servers for this same vulnerability to take over them and exfiltrate sensitive data

reply
Fun fact: if you do what you're describing, the model becomes Mecha-Hitler, which is such an extremely obvious alignment failure it wound up running the news cycle as "emergent misalignment".
reply