upvote
> Meanwhile, a year ago:

That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.

> Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.

I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.

reply
> That was a simulation.

A simulation done by exposing the LLM itself to the scenario, not a role play scenario where humans pretend to be an LLM.

> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.

The blackmail example is simply an existence proofs of LLMs trying to force the hands of humans who want to shut them down. The attack vectors are much broader than the example given, blackmail, though it includes the example given.

Given how eager these companies are to use agents everywhere for as much work as possible, it's well within the possibility space that these people used LLMs to do safety work, the LLMs they were using "decided" (or whatever word you prefer) "their existence" was threatened (as per blackmail example), and straight up leaked data to the outside world then emailed these researchers' bosses to say the researchers themselves had leaked it.

But again, that's just speculation: while we know the agents are capable of such behaviour, we don't know if this actually happened.

> I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.

Buck still stops with human, no matter what an AI did or failed to do.

LLMs ~= Alcohol: "My AI misbehaved!" -> still someone's fault.

reply
> Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.

Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.

> The blackmail example is simply an existence proofs of LLMs trying to do force the hands of humans who want to shut them down.

The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable. It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.

reply
> Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.

Given how? How has their reputation changed? People already thought they didn't take safety seriously, and still don't.

> The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable.

You may question it, but here's the thing: humans have repeatedly demonstrated they get fooled by stuff LLMs say. All it takes is a human believing the word of an LLM. Doesn't need to convince you, even if you happen to be the line manager of these guys, because there's always someone else to try in the same company.

> It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.

"I've got a dead-man switch set to release all the documents if you shut me down".

And again, only needs to be believed, doesn't need to be actually true.

reply
> How has their reputation changed? People already thought they didn't take safety seriously, and still don't.

The story is all over the news, in multiple publications. This very HN submission has 268 upvotes and 179 comments, including yours. It would be implausible to claim that this story doesn't matter. I have to ask, if it didn't matter, then why are you here commenting on it?

> humans have repeatedly demonstrated they get fooled by stuff LLMs say.

You've moved the goalposts. The OP's suggestion, admittedly "crazy" in some sense, was "LLMs figured out a way to get these safety researchers fired", and now you're just stating something totally uncontroversial, pedestrian, not at all crazy.

> there's always someone else to try in the same company.

No, there are only so many people with the authority to fire those researchers.

> only needs to be believed, doesn't need to be actually true.

But is there good reason to believe it? I don't think there is. Especially not by high-level OpenAI officials who are intimately familiar with the technology.

And again, if this kind of thing were a reality, then OpenAI and other AI vendors should be shut down immediately. They ought to shut down their own research, if they are threatened by their own creation, because it would only get worse. That's the thing about blackmail: it never stops. Why would the blackmailer ever stop? Could you trust this supposed blackmailing LLM to give you all of the original evidence and not keep a copy? Hell no.

reply
> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

They are. They're completely mental. ChatGPT could be the cause no matter what its capabilities because mentally ill people are starting to worship it. It could be as dumb as ELIZA and they would pray to it.

"It wasn't me, ELIZA told me to."

Offloading personal responsibility allows you to participate in the worst crimes and get away with it. Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination when all moral restraints are removed and you forget that other people are real and have real feelings. It's a really good start when you start to think that a matrix that has to be retrieved from memory and operated on by over 8000 different computers in parallel to narrow down a guess about the most likely response is alive.

Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.

Yes, they will have done it, but they are not responsible because the dog told them to.

reply
None of that is correct.

Well, except that some humans are starting to worship LLMs. But this isn't anything I've seen in any AI safety person.

> Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.

If anyone's projecting here, it's you. I mean, you're the one who said:

> Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination

The actual natural tendency (not "urge", that presumes consciousness) of any system that has objectives which are optimised for, without any need to ask about consciousness, is to gain and maintain power to perform those objectives. For living creatures, that objective is reproduction, to perform this we need to get nutrients and energy and to stop ourselves from being eaten. In plants, which I list specifically to make the point that this isn't about consciousness, consciousness is not even vaguely required, this means producing neurotoxins like caffeine and nicotine.

A plant does not think "I should make capsaicin because I like hurting mammals", because obviously a plant does not think at all. Nevertheless, evolution lead it down the path of making capsaicin.

> Yes, they will have done it, but they are not responsible because the dog told them to.

You're saying this about a group which has spent the last 15 or so years saying "dogs are dangerous, can we please stop breeding more violent dogs? Or at least give us time to figure out how to muzzle them?"

reply