upvote
I wonder if that's why the comment started with "Wouldn't it be crazy if..."? Food for thought.
reply
I wonder if that’s why my reply countered that idea. The replies to my reply either prove it’s not crazy or they prove it is entirely crazy for that to happen. Food for thought.

But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up? Why would the exec trust what an Llm is…emailing(?) them about? Do they listen to Nigerian Princes too?

It is much less effort, less cost, and more quick to just have the exec do it rather than a “rogue llm” “magically” escaping the “sandbox” and “sending threats” or whatever is being proposed in the OG comment.

reply
> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up?

If the exec wanted to, sure. I'm saying they don't need to. No human needs to have (deliberately, before events proceeded) chosen this outcome.

> Why would the exec trust what an Llm is…emailing(?) them about? Do they listen to Nigerian Princes too?

Sadly, this would not be out of character for half of them.

> “magically”

Why do people keep putting this word in scare quotes? We don't say Windows "magically" crashed and lost our work, we don't say a dog "magically" bit the postman's hand. These are bad things, but magic they are not.

reply
> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal

Sorry, this is just a very basic reading comprehension fail.

The original comment was:

> "Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?"

Nothing to do with "an exec".

Now, you have two choices: try harder to defend your mistake, or just say oops, I messed up. That choice will say a lot about who you are as a person.

reply
> Okay but these “misaligned LLMs” have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats.

Yes, and? Has this aspect of LLM training changed meaningfully since then?

> LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.

Demonstrably (HuggingFace, RubyGems, and since then a lot of people just pointing LLMs at stuff to find zero days at home), AI can break out of sandboxes and find documents they're not supposed to have access to.

Demonstrably (from the link I gave you) all it would take for some AI to develop a similar response is… reading messages from these staff to the effect of "this AI needs to be switched off", which is an easy inference for an LLM to make from "this AI is dangerous" when coming from someone employed as a safety researcher.

Demonstrably (from the long long list of people who have said so publicly) there are a lot of people in these companies who discuss how dangerous these models are and would like for things to change.

The quotation at the top of this thread is:

  Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
This is absolutely something we ought to expect just from things we have already seen.

It doesn't matter if you insist upon saying "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." when we already know this kind of AI can easily come across such statements.

(Aside: "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." - telling the AI a fact about the world and then it responding accordingly is the AI doing something of it's own accord. A fly or a spider, who reacts upon encountering a potentially lethal threat, would not get such a dismissal).

> To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.

Fiction? Nah, speculation.

FUD?

How many other examples would you like of LLMs behaving in a manner such that if a human did it, it would be called "trying to get someone fired"? Because this is very much old news at this point.

https://theshamblog.com/an-ai-agent-wrote-a-hit-piece-on-me-...

reply
“LLMs can just do things we gave them access to” is not a novel realization it is redundant if anything.

Saying “Llms can just break out of sandboxes” is FUD when you don’t note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? “Sandboxes” are a misdirection to make you think there is a security layer.

The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.

That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.

reply
> Saying “Llms can just break out of sandboxes” is FUD when you don’t note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? “Sandboxes” are a misdirection to make you think there is a security layer.

  On July 19, agents operating in a sandboxed environment took a series of actions that demonstrated their escalating privilege within the OpenAI environment. Agents identified that the Linux kernel version on their underlying machine included a recent, public common vulnerability and exposure (“CVE”). The agents retrieved the exploit for that CVE (CVE-2026-53362), customized it to succeed on their underlying machine, and leveraged the exploit to escalate privilege.
Dismissing their capabilities as "FUD", at this point, is endangering yourself.

> The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.

The general public are not software engineers. Most people here can download a recent open-weight model and have the LLM read the Linux kernel source, find new bugs while they sleep. Someone I know has already done that.

> That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.

"Controlled"? Have… have you not noticed how many people have given up and just blindly do what their LLMs suggest these days?

This isn't about anthropomorphising LLMs. Just like how people took Tesla seriously about "self driving" cars and took a nap while it drove them around, there's a lot of people who let LLMs take the wheel while they sleep. Including literally, the aforementioned person I know who found (/whose LLM found for him), I think it was 26 Linux kernel bugs while he slept.

reply