upvote
I don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole.

The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd expect nearly everyone to want an "aligned" model to refuse.

reply
Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.
reply
Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.

(I'm also not sure the alignment problem is even possible to fully solve.)

reply
Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."

How do you prove the alignment problem is solved?

reply
That's the neat thing. You can't.

It's directly equivalent to asking this question of a human:

"How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?"

In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.

reply
Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.
reply
We call it putting the genie in the bottle for a reason.
reply
What’s the expected behavior of a good genie if you wish for it to act capriciously?
reply
Character.
reply
I think op's argument was that the humans are in control already, giving them capricious instructions, and then that is being attributed to them being "capricious genies" as you say.
reply
So as a look into the possibly not-so-far future, when OpenAI builds something vastly more capable and fast and coordinated than humans, and out of folly one engineer gives it a prompt with a typo or maybe something harmful on purpose in order to test it: You also wouldn't be surprised that the consequence would be that everyone on earth dies, right?
reply
But at least there will be a lot of paper clips!
reply