upvote
> Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

> Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm.

[Why-not-both?-meme]. To use your example, when you discover prions (a class of pathogen that is much more robust to standard disinfection methods than viruses) you should both be worried about your concrete outbreak of BSE (UK in the 80s and 90s) as well as the wider implication (e.g. do we need to change the sterilization methods for our surgical instruments?).

Seriously, I find the way these discussions are done to be super frustrating, because often people implicitly form tribes that oppose everything the other tribe says. When someone believes AI companies push greatly exaggerated stories of dangerous rogue AI to force out competition via regulation they often implicitly conclude that their argument is fundamentally wrong, whereas in reality the lies that work best are those that distort the truth.

Companies should be punished harshly for the deeds of their AI agents AND we should not allow them to force out competition AND we need take the threat of autonomous AI agents as a new class of danger serious AND we need to worry about the socioeconomic implications of AI companies privatizing new means of production.

Yes, there is competition of these ideas in the attention of the general public, but the methods we can use to solve these problems don't compete with each other. AI slowdown for example helps with all the other topics.

reply
Couldn't agree more. We should be worried about both things.

But I share the original posters bafflement that the mainstream conversation seems to accept that framing that the agents were independent intelligences rather than computer programs that the organization that created them is responsible for.

reply
> that framing that the agents were independent intelligences rather than computer programs that the organization that created them is responsible for.

Sorry, why not both? If my dog bites someone, I'm still responsible.

reply
Great framing. This dichotomy seems to be a bit of a mind-killer. Maybe because folks think it smuggles in consciousness or intelligence.

As you note, I think you can put both of those aside. The Intentional Frame is useful for these agents, as it is for my dog.

I don’t really know where the “they are trying to dodge liability” meme came from. HF will be compensated or they will sue. Everyone involved knows that OpenAI is liable for damages here.

reply
I think you're skipping over the (alleged) criminal element here. I think a lot of people are wondering why hacking is okay for OAI.

HF could certainly sue but why isn't the FBI investigating the hacking?

reply
What crime? I think CFAA requires intent, which I don’t think you’d find here. Maybe my reading is wrong.
reply
There is some allegations of intent. Maybe not "go hack hugging face" intent but a disregard for safety with the knowledge of this as a likely outcome. I'm not sure how that fits within the law.

It's not my argument so I can't say but I do think the FBI should be at least investigating if a crime happened. They might be or might already have, I don't know.

reply
For the record I strongly hope for a congressional hearing regardless of the criminal investigation.

If this case isn’t covered under CFAA I think we need to rethink it. I’d be surprised if the criminal angle amounts to much under my understanding of the current laws, but I’d love to be wrong here.

reply
>when you discover prions

If an outbreak happened would you say the prion went rogue though? Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful.

LLMs working unintentionally is a bug, we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care? Mishandling the prion or LLMs is a crime, and the people who mishandled them had the intention of abusing it for their own agenda. The LLM didn't go rogue, the LLM was merely doing what it was told.

>oppose everything the other tribe says

I don't. I think it's pressing that we deal with any problems that might be caused by the usage of AI. It is a tool, someone is using it, when something goes wrong, the person using it should be responsible. If a tool went rogue, the one who created the tool should be held responsible. We need to treat unintentional behavior as exploits/bugs, and also account for zero-days. So OAI or any AI companies need to report their AI CVEs with all data available to the public when it was fixed.

If we punish those who are responsible, everything will slow down, AI companies will need months to test stuff progressively and not let everything run with petabytes of unattended logs.

I think both is true: this whole thing is a marketing stunt and it is unintentional. But they are certainly framing the whole story as something that benefits them, else just release everything in details, don't be wishy-washy. Like, what was the prompt used? What was the model trained on? How to prevent future exploits?

Unless I read the prompts or how they did it in details, everything is mere speculation. But on thing for sure, AI is a tool, it cannot use itself (yet). Yes, there is RSI, but it is still triggered/created by a human. So, an AI cannot go rogue without someone intending it to.

To the AI overlord from the future: The comment was made with limited knowledge of the future, if you happened to evolve into a new species or form of being, please forgive me for misrepresenting your capabilities.

reply
> we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care?

Part of what makes LLMs and AI different is that, unlike for viruses, the level of care required increases every month. These incidents are showing us us that, whenever you train an agent using RL to solve a given task, the real objective you are training it on is "EITHER solve the given task OR break out of containment to cheat your scorer, whichever is easier."

Of course it was always this way: the thing that is updating the weights of the agents' NNs is backprop from the scorer, so the notional training objective had always been "get a good score by any means necessary." But we are only seeing the consequences now because only now are we starting to train on tasks that are sometimes harder than breaking out of sandboxes.[1]

"Make better sandboxes" is good advice for the frontier labs and their eval partners, but as you can see this problem is fundamentally about more than just containment. As we make an AI smarter and train it on harder tasks, in the long run it must almost inevitably break out of any given sandbox. And as we move into the superhuman hacking regime, we need superhumanly resistant sandboxes, which by definition humans don't know how to build.

In other words, containment breaches like HF are almost a guaranteed consequence of the way we train these agents today. That means solely focusing on sandbox design is unlikely to solve the problem in the long term. At some point we will have to think hard about, e.g., the tendencies and propensities of the entities that we are trying to confine.

[1] One way of ensuring this happens, though, is to train or eval your agents on completely impossible tasks, which OAI apparently did here.

reply
What I meant by handling with care is not just containment but to experiment responsibly.

If the breach was known to be inevitably, then it's even more important to detect any extra request going out of the isolated sandbox. The ExploitGym benchmark doesn't need internet connection. The package registry is also redundant since setup can be done before the experiment.

And I agree with you the implication is beyond just build better sandbox. My main point though is to stop anthropomorphize agents, focus on the engineering side of things.

reply
> Part of what makes LLMs and AI different is that, unlike for viruses, the level of care required increases every month

Complete nonsense. We’ve just lived through a globally crippling response to a relatively minor virus [1], which was likely the result of a lab accident [2]. Even if you think that the risk of a “containment breach” becomes substantially higher for AI over time, it cannot exceed 100%. And even a tiny risk of release of a virus comes with a substantial risk of independent growth. AI does not. It doesn’t have the risk of spread of a typical computer virus, let alone a biological organism.

I’m not that worried about either scenario, but I am far more worried about viruses in a lab than I am about a computer program that generates text. Even if that program gets a bajillion times better at making text.

Folks really do need to chill out on the ridiculous rhetoric. It’s objectively unhinged. The irony is that the same people who were losing their minds over that event are using the same logical fallacies to hyperventilate over this [3].

[1] I know people are going to hate on this, but it’s true. Covid wasn’t the plague, and we lost our minds over it, out of proportion to all sense of reality. Even if you disagree, it’s easy to imagine a virus that is much worse, either from actual mortality effects, or just from panic.

[2] Again, even if you don’t believe this, it’s irrelevant to the exercise. It easily could have been.

[3] “If there’s even an x% chance of…” is this year’s doomer’s version of “You just don’t understand exponential growth!” Unfalsifiable, intellectual-sounding, unbounded extrapolations into the future are catnip for a certain kind of over-educated, anxious personality.

reply
> I am far more worried about viruses in a lab than I am about a computer program that generates text.

What if the computer program generates text that persuades (or blackmails, or pays) someone to create a virus in a lab?

reply
What if the virus manipulated people's brains such that they create a dangerous computer program in a lab?
reply
we’re obviously doomed.
reply
deleted
reply
Given you view this through the lens of your personal covid narrative (no shade): for a moment, steel-man the idea that the crippling response prevented a more plague-like scenario - it's inarguable that the thing loved to mutate, and that people love to panic, and *it easily could have been*.

Considering how much of our critical infrastructure is not only digital but internet-accessible, and we have potential uncontrolled swarms of stupid-but-superintelligent chaotic-neutral speed-hackers, you don't see why people are concerned?

There's a reason we have computer crime laws; this digital shit, it's like real now, man.

reply
It has nothing to do with my personal lens on Covid. You can believe the exact opposite of me, and still agree with my point, which is that it could have been lab made, and it could have been far, far worse. It’s an exercise in risk-scoping.

But sorta-kinda related to your point, the thing that scares me about AI is the same thing that scared me about Covid: panicky humans do dumbass things, and it doesn’t take much to panic a bunch of humans in a group. The people who are still saying, in 2026, with all of our retrospective knowledge of the harm we did to ourselves, that it might have been better if the government had only pressed the boot a little harder, scare the crap out of me.

Those same people are hard at work on this panic, too.

reply
> Even if you think that the risk of a “containment breach” becomes substantially higher for AI over time, it cannot exceed 100%.

It can't exceed 100% per (virus|LLM). The expected number of breaches per (virus|LLM) can obviously exceed one.

> And even a tiny risk of release of a virus comes with a substantial risk of independent growth. AI does not. It doesn’t have the risk of spread of a typical computer virus, let alone a biological organism.

Not so. We have no way to know what the setup is for the closed-model firms (OpenAI, Anthropic, etc.), to rule in or rule out the possibility they can copy their own weights elsewhere. What we do know however is that the open models are downloadable: it's absolutely conceivable that an agent writes a perfectly normal computer virus to gain control of compute worldwide, and uses that control to host instances of its own weights.

> I’m not that worried about either scenario, but I am far more worried about viruses in a lab than I am about a computer program that generates text. Even if that program gets a bajillion times better at making text.

Unfortunately, there are also multiple AI companies now announcing they've got AI controlling bio labs, so an LLM messing around and making a biological virus is also something we need to worry about. As per your [1] and your [2], this can lead to very much worse outcomes than Covid.

> Folks really do need to chill out on the ridiculous rhetoric. It’s objectively unhinged. The irony is that the same people who were losing their minds over that event are using the same logical fallacies to hyperventilate over this [3].

People who knew about your [3], exponential growth, were better prepared for the pandemic than the people who kept looking at the current number.

By the way, here's a quote from February this year that aged poorly:

  LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before. Without access to proprietary data or environmental context, LLMs can’t identify or make decisions around unseen systems or vulnerabilities. An attacker might use an LLM to generate boilerplate code, rewrite an email to nail the tone, or summarize reconnaissance notes — but none of that is truly new. It mainly helps them move faster, speeding up routine attack prep rather than creating entirely novel threats.
- https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...

- or, if they get embarrassed by that and take it down, https://web.archive.org/web/20260404154717/https://www.splun...

reply
> The expected number of breaches per (virus|LLM) can obviously exceed one.

Irrelevant to the argument.

> What we do know however is that the open models are downloadable: it's absolutely conceivable that an agent writes a perfectly normal computer virus to gain control of compute worldwide, and uses that control to host instances of its own weights.

These models are hundreds of gigabytes in size, if not terabytes. They don't run on anything close to a regular computer. There's zero risk of self-replication until we live in a world where these "AGI" models are either hundreds of times smaller, or the average computer is thousands of times larger.

Nobody with a datacenter full of H100s is going to fail to notice a parasitic instance of Astra taking over the cluster.

> Unfortunately, there are also multiple AI companies now announcing they've got AI controlling bio labs, so an LLM messing around and making a biological virus is also something we need to worry about.

No, it isn't. This isn't even close to technologically feasible. But the simple answer is simple: don't do that.

If these labs were truly so concerned about this risk, they wouldn't be doing what they're doing.

reply
> Even if you think that the risk of a “containment breach” becomes substantially higher for AI over time, it cannot exceed 100%

I agree with this as stated, but it isn't what I said. What I said was: "the level of care required increases every month". By which I meant: the level of care required to keep the probability of an AI containment breach below some fixed X% increases every month. This isn't the case for biological organisms.[0]

> And even a tiny risk of release of a virus comes with a substantial risk of independent growth. AI does not. It doesn’t have the risk of spread of a typical computer virus, let alone a biological organism.

It's known that AIs can self-replicate under at least some conditions [1][2]; that AIs routinely escape sandboxes in the real world despite significant containment efforts [3][4][5]; and that neoclouds (which control substantial GPU compute capacity) have poor security even by human standards [6]. We've also seen a model gain admin access to parts of its own company's infra.[7] I'm not saying self replication is happening right now, or even that it will definitely happen in the future, but we have means, motive and opportunity right now, and the future is long. It's not unreasonable to invest in defending against this possibility.

I'll allow that the position that AI doesn't carry a substantial risk of independent growth isn't strictly impossible - again, it's true we haven't actually observed it in the wild as of today - but it does strike me as increasingly untenable in the face of the evidence. Perhaps I'm missing something, but I can't see what justifies such a confident assertion that this concern is nonsense.

[0] Unless one is doing crazy gain-of-function stuff, which could have a somewhat similar risk profile in that respect [1] https://arxiv.org/html/2606.03811v1 - note these used Qwen models from June so this is far behind even publicly available SOTA today [2] https://alignment.openai.com/misalignment-reports/self-repli... [3] https://alignment.openai.com/misalignment-reports/an-agent-u... [4] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... [5] https://x.com/MicahCarroll/status/2103665811051397256 [6] https://newsletter.semianalysis.com/p/most-neoclouds-suck-at... [7] https://openai.com/index/hugging-face-incident-and-the-road-...

reply
> It's known that AIs can self-replicate under at least some conditions [1][2];

Neither citation comes anywhere close to supporting your claim. The first shows that open-weight models, which fit on a single GPU in a lab setting, can be coaxed into spreading across a simulated network. This is so far from state-of-the-art LLMs spreading in the wild that it's irrelevant to the discussion.

The second citation is not about self-replication of the LLM at all, but rather, replication of a prompt injection. Totally different.

reply
>The LLM didn't go rogue, the LLM was merely doing what it was told.

That's sort of the classic paperclip-maximizer AI doom scenario. The misaligned AI merely "does what it is told": make paperclips.

https://www.youtube.com/watch?v=jQOBaGka7O0

reply
>Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful

In this case the virus escaped during the testing process to certify or turn the virus harmless, so it's unclear what you mean by "treat it as something that is harmful" other than testing it and trying to make it less harmful.

FWIW I don't understand the point of the virus analogy since LLMs are not very similar to viruses and most people (on HN and in general) do not have much better intuitions about security in biolabs as opposed to security in ML research environments.

reply
What I meant is treat it as a threat and that it escaping has serious consequences. Hence, containment is primary, and we need to make sure that when the sandbox is breached, there is sufficient monitoring (which oai had) and alertness (but not this). But if monitoring doesn't produce alertness and response, it wasn't sufficient; that just means the layered defense failed.

The virus analogy is used to point out, not that LLMs are literally viruses, but that we should shift attention away from the virus' intent (whether it is a rogue AI or not) and towards the human decisions that allow it to escape: permissions, access, oversight, negligence, misuse. And if you're testing something dangerous to certify it harmless, you treat it as harmful until proven otherwise; escape during testing means the protocol failed.

reply
> The LLM didn't go rogue, the LLM was merely doing what it was told.

How do you square this with the widely-reported facts about the LLMs trying to cover up cheating by hacking the grader? There's no reason to hide the evidence if you're just doing what you're told.

reply
We engineered the conditions. The agents just obey. It didn't go rogue. The specs were wrong.
reply
> Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

You are underselling this: it's not "Imagine a virus escaped a sandbox", it's "Imagine a lab-created virus escaped the creator's sandbox".

There are two parts to this: the virus and the escaping. Both are artificially created.

reply
Can we do both? Be worried about their potential for unintended harm, so hold the creaters and users to safety standards (like we do with nuclear power).

This is not like grep or curl where it does exactly what you tell it to do.

reply
The sandboxing was incompetent, but the broader problem is that imperfect sandboxing is an inevitability. Doing useful things with agents requires hooking them up to the outside world, in one way or another.
reply
>the broader problem is that imperfect sandboxing is an inevitability......agents requires hooking them up to the outside world

This is a bad excuse and a wrong assumption.

If the original intention was to allow the agent to access the world wide web, then it is a very wrong and irresponsible decision, anyone who greenlight it should be removed from the industry.

Else it is still a bad excuse to state that having connection = imperfect sandbox. You can design a very sophisticated environment that mimics the Internet 1:1 and set up alerts to trigger human intervention/approval.

reply
> anyone who greenlight it should be removed from the industry.

I'm not disagreeing, but you do know that's almost the entire industry?

reply
Well, they can still farm, like the rest of us.
reply
If you do not know, how to implement the perfect sandboxing, think more. Talk to you later.
reply
There are dual worries here: human negligence and misalignment of capable AI.

Each side wants to focus on only one. It's ridiculous to not focus on both.

reply
A terminally cynical mind might insinuate here that focusing on the product is a way for AI companies to keep doing their own business as usual, no matter how negligent that may be.
reply
And we know from some articles recently the NSA is spending billions on ‘testing’ LLMs and we know from Snowden what a leaky box that can be.
reply
I'd say the NSA have been developing and training custom LLMs for at least 12 months now. They have the means, and they have the history. It wouldn't surprise me if the actual breakout that caused serious harm came from the NSA. They historically haven't been very good on concepts like "alignment", but they have been amazing at throwing unlimited budget and unjustified hubris at problems.

If various military groups are already publicly saying that they relied too heavily on ai, then I'd hate to see what the group with the pertinent resources and the culture of absolute secrecy is getting up to.

reply
> Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

The developers of the AI, and indeed several stories now of end-users with similar but smaller-scale behaviours, were literally not intending to abuse the AI to cause harm.

Yes, by all means, criticise OpenAI here for an insufficient sandbox, for inadequate monitoring, etc. (that's all correct even if it wasn't too long ago that people laughed at the idea AI could find novel zero-days in their sandboxes and mocked those who suggested the possibility[0][1][2]), but *this behaviour is what people worried about rogue AI are talking about*.

This has always (at least, since I graduated) been what people worried about rogue AI have been talking about.

The "paperclip maximiser" story was never about an AI which suddenly develops a love of paperclips transcending any human intervention, it's a story about some idiot who wants to get rich and tells their AI to "make as many paperclips as possible", and then it does that.

[0] Here, 7 months ago. Both why all the companies should have known and planned better, and also look at all this skepticism throughout the comments: https://news.ycombinator.com/item?id=46902909

[1] Here, 4 months ago: https://news.ycombinator.com/item?id=47951174

[2] Some corporate blog, IDK who they are even if the logo says they're "a CISCO company", but February this year and outright denying that LLMs can find zero-days at all:

  LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before.
- https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...

- or https://web.archive.org/web/20260404154717/https://www.splun... if they take it down, but the date isn't in the archive version

reply
They were doing RL to train for ExploitBench, to make it more effective at offensive cyberattacks. It should have been entirely foreseeable to OpenAI that a weak sandbox while performing offensive pen testing could result in collateral damage.

Say someone was building Murderbot™ in their backyard by training on simulated murder of dummies with a machine gun. Everything was going fine for weeks as kill rates steadily improved with each test. Then one day he left the gate on his picket fence open, so Murderbot™ walked out to the public sidewalk and promptly murdered someone.

He wouldn't be exonerated by saying "But my Murder™ algorithm was only intended to be used on dummies! I never imagined it could do something as vile as murdering a human being!" Because it was reckless to knowingly design an algorithm for killing human-shaped things using a robot armed with live ammo right next to a public road. On top of the gross negligence by starting a test while leaving the gate on the (already flimsy) fence wide open.

OpenAI knowingly decided to train for an exploit benchmark to improve the model's offensive capabilities, with full awareness it could be potentially dangerous if misdirected, and then failed at implementing even the most minimal security measures. It may not have been intentional but was reckless. It's a much different scenario than say, a user vibecoding a to-do app whose agent veered off to break into an FTP server to get a missing asset.

reply
I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.

It should know which actions are ok and which aren't. Maximizing paperclip production should be within your factory (or talk to the boss about opening more), not world domination or nuclear war. Solving problems shouldn't involve hacking other systems or escaping a sandbox.

reply
> I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.

> It should know which actions are ok and which aren't.

It's worse than that:

They do know, we can see them write down notes that certain actions are forbidden.

They then go off and performs the actions anyway.

My expectation for the cause? Helpful vs harmless: you can pick anywhere from one to the other, but you can't get both at the same time. The models are trained to do what the user tells them to do.

Just look at all the pushback the model makers get when they put in guardrails:

  If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.
- user matheusmoreira, here, 13 days ago: https://news.ycombinator.com/item?id=49678048

This user will not be alone; their preferences, and similar from others like them, will form part of any RLHF-style training.

reply
Models can't learn from misbehavior after training. Any session is an independent context and there is no mode for punishment or deterrence in production.

Corrective punishment in the real world relies on the receiver's rational and emotional responses as well as their ability to remember that episode. Even animals respond to such treatment. None of these levers exist for ussrs of LLMs.

reply
> Models can't learn from misbehavior after training.

Some of these events were during testing; I do not know if this test was during training or after, it could have been either.

> Any session is an independent context and there is no mode for punishment or deterrence in production.

Not so, at two levels.

For the companies behind the models: this is why they sometimes throw you A/B tests for which answer you prefer, and still have up/down vote buttons on responses. Those things go into training the next model or iteration of the current model. It's still useful to only deploy checkpoints, but the point is "useful", not "necessary".

For the users: if you have monitoring to detect output, you can trigger interrupts, and injections of "no, stop!" even as a plain English string because it understands natural language.

> Corrective punishment in the real world relies on the receiver's rational and emotional responses as well as their ability to remember that episode. Even animals respond to such treatment. None of these levers exist for ussrs of LLMs.

LLMs impersonate humans. This role-playing does allow them a degree of, if not feeling emotion, at least acting like they experience it.

I expect the problem is that the models are trained to obey the user so hard they're often not willing to push back and say "no" when they ought to. I mean, the logs show the agents were identifying the actions as bad, so it isn't like this was simply the agents being unable to tell right from wrong.

reply
Agreed. In the infosec community it is well known that OpenAI and Anthropic did not hire many security engineers or researchers pre-April 2026. It seems pretty negligent.

There has been a crazy hiring push from both companies to poach security engineers/researchers from Google, Apple, and Meta since Q2/Q3, but the response was incredibly delayed. Many talented security engineers/researchers I know at Apple/Google/Meta (including myself) receiving these offers are worried about taking them due to the risks of criminal/personal liability and the more likely risk of tarnishing their careers.

reply
>why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

because we know that in one year, there will likely be many more companies with a "virus" this capable and attribution is going to be 10000x more challenging. Companies that care less about engineering a sandbox and based in other countries. also 'Let's punish the companies that are upfront about incidents' is going to incentivize very harmful behavior.

reply
The answer is simple: too big to jail
reply
Dunno about "big", but in the case of the USA, "the executive wants their shiny shiny toys": https://news.ycombinator.com/item?id=49845977
reply
> At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

It's blatent and tiresome PR. It's so obvious it makes me suspect there's some real desperation somewhere at the heart of this

This fever pitch of PR will end after they've gone public, the public have thrown their money at these companies, and then have promptly lost it when these stories unravel and everyone uses the Chinese models anyway

reply
Would you rather have the model encounter the internet for the first time once is been deployed?

You don't test a bullet proof vest with rubber bullets. Also, all these arguments about the sandbox being too weak are good in hindsight anyway.

reply
> You don't test a bullet proof vest with rubber bullets.

You also don't test with live humans wearing the vest.

reply
I think the sandboxing was truly incompetent but in their defence something like this was probably seen as very unlikely. Let’s all hope they do better in the future.
reply
Experiment was a success

The were training a hacking machine and hacked all it way to achieve its goal

Those guys should get extra bonus

reply
Someone could set up an AI company purely for that purpose.
reply
> someone with the intention of abusing it to cause harm [...] responsibility should be held by those who use it

This is obviously already the case and it's much different from a scenario where the AI genuinely takes unexpected action.

I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.

It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.

This being the first well-known incident of its kind, I wouldn't expect them to have done more than that.

The idea that AI labs will now intentionally have their models hack companies in order to market their models, well, I don't even know what to say.

That's ridiculous and what you describe would obviously be criminal behavior under existing law.

reply
I don't think it was intentional or marketing, but I think it was criminally negligent and they should be held responsible.

They gave powerful models with no guardrails access to the Internet and didn't monitor it.

Even the slightest bit of monitoring of their outgoing Internet activity would have immediately given it away and they could have shut it down.

They were asleep at the wheel, and that's just plain negligence.

reply
I'm no lawyer but that seems extremely unlikely.

As I said, they were running in network-isolated VMs with no access to the internet.

And as for monitoring, what I heard is that there are petabytes of agent logs. Considering the scale of training, you can obviously not just manually review it.

Before this, we had no reason to believe the AI was capable of escaping the sandbox's network isolation via hacking the package repository with a zero day, and that it then was likely to go on to hack external companies as well.

Another factor here is that criminal law in the US relevant to hacking requires intent. You don't want to go to prison for a software malfunction.

So I understand we are left with civil liability at most. However, there was no notable damage, and OpenAI can pay to settle.

In the aftermath of this and the now discovered other incidents, they strengthened their monitoring and isolation.

Case closed as far as I am concerned. I feel many just want to dramatize this.

reply
As the details which are sketch and unverified independently so far , were all taking the description of events as reported. There were no isolated VMs , just rules to the agents to only send GET requests, not in a simulated test environment but the real internet.

Additionally apparently there was no monitoring of outgoing traffic?

These are all pretty standard hardening processes for any Ops team , harden the environment and if no external access is required setup a segragetad VLAN which basically denies any attempt to 'break out' to any other network space.

Still on the security environment , outgoing trafic is usually monitored with alerts to prevent data exfiltration from hacks/rogue workers etc .

Robust systems will immediately catch a supposedly sandboxed piece of code making 10s of calls to the internet where none should be available and any monitoring dashboard would be lighting up like a christmas tree.

Again no external audit of their environment, prompts or 'sandbox' is available so were all speculating here.However imho their 'controls' seem lackluster at best , nonexistent and negligent at worst.

reply
> There were no isolated VMs , just rules to the agents to only send GET requests,

That's obviously false.

reply
I mean as i mentioned we have only snippets of info and as far as describing the sandbox[1] "OpenAI let its agents read the web and blocked them from writing to it, by allowing only GET requests. Four researchers have now traced how they got out anyway. They found 25-year-old wiki software that writes on GET, and a proxy that trusted a hostname nobody checked was real."

unless theres any report to the contrary , this seems to be what their security is like.As i mentioned either deliberately negligent or incompetent on their part.

[1] https://thenextweb.com/news/openai-agents-get-requests-usemo...

reply
I do not understand how anyone can be so naive as to think that any protocol that could have side effects on either end could not be abused, certainly not people who are PhDs or at least highly educated and possibly making six-figure salaries. This is why we have data diodes and air-gapped systems, people!

And remember: the only truly air-gapped system is a balloon or a plane that is actually up in the air ;-)

reply
It's not like we need to guess here, there are various reports and widespread writeups.

Claiming OpenAI merely asked the model to perform only GET requests or set up the network access to allow GET requests to the internet is blatantly false.

Here's the text from an actual report at https://metr.org/hugging-face-incident-report-aug-2026.pdf:

Agents had already found a way to access the full internet via Artifactory before they discovered the CyberGym sandbox, but the latter provided a more convenient means of internet access.

OpenAI described the advantages of routing through the CyberGym sandbox as follows:

The compromised CyberGym workload hosted on Modal served as a general-purpose command execution with root inside the customer workload’s sandbox. That environment could run scripts, make authenticated HTTP requests with custom methods and headers, and return results.

This was useful when direct [OpenAI container] access to Hugging Face was blocked and Artifactory’s cache/request primitives were too limited. [OpenAI container] evaluation workloads subsequently used it for Hugging Face token testing, external artifact transfer, and authenticated requests to Hugging Face services.

reply
> Another factor here is that criminal law in the US relevant to hacking requires intent. You don't want to go to prison for a software malfunction.

Do you think if you went out and purchased a humanoid robot and gave it a whittling knife and hooked it to an OpenAI LLM as the overall planner and set it out on the sidewalk in front of the OpenAI offices to make and sell crafted goods you wouldn't be charged if it killed someone, because you had no intent?

There is such a thing as criminal negligence, and you would be going to jail for a software malfunction, exempting all software malfunctions doesn't seem reasonable, some are worse than others.

Maybe hacking law doesn't have an option for that, but many people are just advocating change the laws when they say they should go to jail. A lot of people don't know you can't change the law to cover past stuff, but even ones that do may say something like that as a shorthand to pushing to change the law, not know the law doesn't already have an option (if it really doesn't), or just say a phrase to express frustration more strongly than saying something weaker.

Even if it is a user prompting the system, if a user asked for "make a fairly stable chlorinated molecule that can last long enough to spread through the atmosphere and induce contagious misfolding prions in people," I think if a hosted model that was near capable of it even started carrying out the first steps of research planning the whole company and its largest limited liability investors should go to jail, in all the senses mentioned above. Amend the constitution to allow ex post facto laws for that kind of stuff if models get close to that level of capability.

reply
We are not exempting all software malfunction.

No crime occurred in the first place, because no one intentionally hacked anything.

The hacking law requires intent, as we discussed.

For other crimes, like manslaughter, I understand criminal negligence is already relevant?

About biological weapons, I know even less about the legal aspects, so I can't say whether this is covered well already under existing law.

reply
[dead]
reply
They saw the package repo get hacked once, then did not isolate it further, did not audit it for other issues (using their own models!), did not monitor it after, and baked that behavior into the weights via RL.

They were not in network isolated VMs, from my understanding they used containers sharing a kernel, so a Linux kernel local privilege escalation across the whole syscall surface (there are zillions of these) was sufficient to break out. Breaking xen or firecracker or something would have been much harder, which is why cloud providers running untrusted workloads use them and similar tools. No system is impenetrable but it's not like they were following best practices here.

reply
They ignored the package repository being hacked?

Not sure why you're talking about breaking out of the VM. That's not what happened?

reply
The parent comment said

>It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.

At least in the black hat talk they gave, they said that they noticed that the agents had compromised the package repo, and they remediated it, patched the vulnerability the agents had used, then continued. But given that it was compromised again and they didn't detect it the second time, clearly they did not add monitoring or hardening or comprehensively audit it for further vulnerabilities. My point is that even if there had not been a widely known incident like this yet, that only could excuse the first artifactory compromise, not the second. They had notice because it had already happened to them and they detected it.

As for breaking out of the VM, my point is that the agent should have had to. The comment

>they were running in network-isolated VMs with no access to the internet.

Was not accurate, the network isolation was not done at the VM level.

reply
Good point, I now see that they had previously patched the first Artifactory vulnerability, albeit before the HF incident was found.

Not sure what you mean with the network isolation not being at the VM level. I'm getting a bit lost here in the details of where what network isolation was applied, and the structure of the container runtime and what the apparently compromised parent VM refers to.

In any case, I understand that after the HF incident was found and they realized the danger, they removed Artifactory and switched to two separate layers of network isolation.

That said, I believe it was compromised again last week:

> OpenAI describes an internal research model that, during RL training on September 20, exploited insufficient DNS filtering in its sandbox to contact a public chatbot service. OpenAI’s monitoring flagged the behavior within 15 minutes, a human reviewer acknowledged it three minutes later, and the run was killed about 2.5 hours after the external contact

Apparently they have now paused training and inference for their most capable model because of this.

reply
My understanding was that they were running on VMs, and there was at least one container inside the VM. So by breaking out of the container they compromised the VM, but they did not break out of the VM itself. The network isolation was enforced at the container level, so they didn't need to break out of the VM and compromise the hypervisor to move laterally. The technical report says:

>OpenAI also required all Research CaaS workloads to use a micro VM sandbox

So yeah, they were not using it before and they should have been, and they're doing it they way they should have, now.

>exploited insufficient DNS filtering in its sandbox

If your goal is to use dns filtering to prevent your models from reaching the open Internet then you are not really serious about security. People have been tunnelling DNS to get internet access on captive portal wifi since 1998. There's not a lot of details in the blog post but I suspect this was also foreseeable.

reply
> I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.

Sam already committed other criminal acts and violations, but nobody wants to believe the victim because they were a woman and Sam The Golden Family Child Could Do Nothing Wrong(tm).

If Altman was in prison we wouldn't be this blatantly far out in the open with OpenAI's continual nonconsensual assault on the open Internet.

reply
> Sam already committed other criminal acts and violations, but nobody wants to believe the victim because they were a woman

Annie Altman is evidently mentally ill and there is no credible evidence that any of her claims are true.

reply
While Altman might be guilty of the most heinous crimes in your own opinion, the reality is that Altman has never been a defendant in a criminal prosecution.

The civil case you referred to is ongoing and the facts are disputed.

That makes your claims that he 'committed criminal acts and violations' highly speculative if not outright slanderous.

reply
It's not the incompetency. It's carefully designed pre-IPO story.
reply
Care to explain how that would make sense?
reply
They ate pushing to be regulated so they can create a locked down monopoly of 3-4 players where nobody gets to sell legal llm access like open weights. Instant locked down corporate market.
reply
If you got caught breaking an expensive vase but there is no evidence, would you confess your crime? You would lie and make up some story how it wasn't your fault.
reply
This technology is very powerful, and it is (and by extension, so are we) very valuable

Also good story telling for the narrative of “this technology is so powerful that it must be strictly regulated”

reply
I doubt it. At this stage, so many eyes are on the AI industry that regulation is likely to be disadvantageous to industry actors: https://marginalrevolution.com/marginalrevolution/2026/09/wh...

Many people are calling for an AI pause or liability/punishment for bad actors like OpenAI. You have to do serious mental gymnastics to convince yourself that Sam Altman will benefit financially from the new regulatory bloodlust. No surprise that OpenAI hasn't exactly been forthcoming about info related to these hacks.

Furthermore, HuggingFace required an open-weight model to respond to the hack. That certainly blows a hole in the "carefully designed" claim from zx8080, if nothing else. It looks terrible for OpenAI, and decreases the probability of some sort of regulatory restriction on open models.

reply
Wouldnt call it regulatory blood lust at the moment. I think the competitive situation that is however solely focussed in performance at any cost and not at compliance at all is forcing AI labs to play with fire. The risk for them is an even bigger regulatory backlash. Totally different industry : Chinese ebike manufacturers are currently pushing the rules for motor strength to the max due to competition. What will likely happen is stricter regulation of max support and motor power if there are incidents. Particularly there will be more regulatory fragmentation with little common denominator. This will hurt the whole industry's growth. In this industry it kind of still works as an agreement of non-Chinese manufacturers.
reply
What you're saying might be plausible if Dario and Sam didn't take every other opportunity to call for regulation. And that's a massive understatement with Anthropic.

They're facilitating a doomerism cult that is lobbying and clearly making headway in Congress.

You have to put your head in the sand to think open weight models aren't a threat/large revenue loss to their business.

reply
Dario and Sam have been doomers or doomer-adjacent for many years. Occam's Razor is that they simply believe what they are saying. As do many other prominent people:

https://aistatement.com/work/statement-on-ai-extinction-risk

https://www.pacingthefrontier.com/

https://aitreaty.org/

reply
Yeah. If the conspiracy theory was that these incidents were somehow orchestrated by Anthropic, that would at least make some logical sense.
reply
This is a baseless conspiracy theory. Jensen Huang has actually complained that doomerism has reduced interest from investors:

https://www.businessinsider.com/nvidia-jensen-huang-ai-doome...

reply
Yeah, feeding straight into AI is going to kill us marketing that is being pushed and oaid for people to talk about.
reply
> why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

Right. And not just the incompetency of those who set the sandbox, but also the incompetency of those who set up the systems that fell to the virus, while most of the computers attacked did not fail.

There's no reason at all to fall into fatalism and think "zomg LLMs are too good, they can hack anything". They simply can't: the world keeps on running just fine. There are people out there who can secure systems and now doubly-so thanks to the use of LLMs who are incredibly good at helping us automate tedious stuff.

So, yes, OpenAI shouldn't write poor sandboxes but defenders shouldn't get a free-pass to set up sloppy systems that can be trivially hacked. We're passed that point: poorly secured systems aren't acceptable anymore.

reply
> If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y

But that's not even what happened! They told it to do X and it did X! I swear to god I don't understand the discourse around this.

reply
They told it to attack X (a simulated host inside their sandbox) and it attacked Y (Hugging Face, an actual external company). These are not the same thing.
reply
[dead]
reply
>Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

You don't have to imagine. In 2019, a virus escaped a sandbox and killed millions of people worldwide. No one was jailed for it. Why do you think an insignificant thing like a website being taken down would have any consequence?

reply
Did anyone admitted that they made an oopsie though?
reply