upvote
Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective.

"zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath"

"The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'"

reply
That’s what the alignment problem is all about though, isn’t it? AIs always act towards ‘their own’ objectives, that they derive from our instructions.

What we try to do is train them and provide instructions that will result in it having an objective closely aligned to our objective.

reply
What "own objectives"? Isn't it more that specifying objectives is hard (AI or organic) and typically supplement them by boxing things in (again, AI or organic).
reply
Well if the recipe requires access to some secret ingredient it may as well resort to hacking to obtain it. ;)
reply