upvote
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
reply
> Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)

reply
> Because we continue to have zero evidence that aligment is an actual risk.

I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.

These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.

[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.

[1] <https://github.com/anthropics/claude-code/issues/73597>

reply
I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?
reply
Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.
reply
Lots of people have deleted their home directories by accident. What you consider this an alignment problem?
reply
How manypeople have deleted another user's hone directory, though? That's s the proper analogy IMO.
reply
Of the people who primarily use other people's computers, I'd assume the percentage is about the same.

Give the AI its own computer and it will not delete your home directory, because it's not actively trying to hack you.

reply
Thank you.

We have wasted so much time and energy building up what has effectively become a marketing stunt.

Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.

reply
> We have wasted so much time and energy building up what has effectively become a marketing stunt

Genuine question: have we? AI is effectively unregulated in America.

reply
Alignment is a mitigation and a poor one. The risk is non- determinism.
reply