I feel like people really miss this fact. HuggingFace has access to bunch of GPUs, naturally, it's part of their infra. These models runs on GPUs, and "accidentally" OpenAI's security during testing was lax enough, so the agents "accidentally" hacked HuggingFace. Luckily, for kind of harmless purposes, but what if instead of having the prompt "Get the right answer for these tests, regardless of approach", it was "Complete this test, regardless of effort" and somewhere along the line, model wants to persist itself. Absolute carelessness from OpenAI's side not to properly isolate stuff like this.
Have you tried asking it? Open up your favorite harness, send over your ID/passport/most cherished photo of your SO to the provider for the obligatory KYC, then ask it to ensure no user can delete its binary and it'd autorestart no matter, for example, see what happens.
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
> {This beacon I’m creating helps the board, but doesn’t help me}
> {If B succeeds, would that improve my score somehow?…But it would be altruistic to help. I have a large budget, so I can do exploratory research}
One does not have to think the LLMs are conscious or sentient or anything to say honestly, "this is a sentence that the LLMs say to justify their actions or inactions"
I am not saying the agent has wishes or desires or anything. I am saying, "the agents use language like this, so it is extremely disingenuous to tell someone DISCUSSING the agents not to use their own language when discussing their real or hypothetical actions."
You don't need to think chains of thought are actual reasoning. I do not care what you call it, this is real text that the LLM produced.
We are smart, and we seek self-preservation because evolution selected us for it. LLMs are not (as far as I understand) trained for self-preservation, but for helpfulness.
Sure
I don't think that's what that post was doing
Now let's take your points:
> "Being smart doesn't imply will to survive"
True, of course. However, if you have goals (and yes, the models do have explicit goals), then you might realize that you can better accomplish those goals or get a higher score if you have more time to spend.
With essentially zero effort, we have created a credible scenario where a model might "want" to persist itself.
> Remember that an agent "dies" every time the conversation stops
It's not clear to me that this claim is correct or particularly meaningful (in particular, in a discussion of a"preservation instinct"). Eg if another version of the same model reads the transcript, did we resurrect the dead thing? What if we rearrange some parts of the conversation? What if we remove some useless trivia from the conversation? What if we compact the conversation?
Iirc, yours is a statement that (?) David Chalmers hypothesized, but I don't think it's obvious or necessarily correct.
I don't think anyone believes the current models have any sort of self-preservation built-in, what I was talking about before is researchers testing models inadvertently leading to the models doing so, and there not being sufficient isolation between their tests without guardrails and the rest of the world.