I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.
The exploit:
> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]
Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox with only one line to the internet (the commercial package registry cache proxy)?
[1] https://openai.com/index/hugging-face-model-evaluation-secur...
Oh, and after resuming the tests, the Artifactory message board was reestablished almost immediately, but it took a number of days to fully breach HF. In all that time, after seeing Artifactory compromised the first time, nobody even bothered to check if those naughty agents were at it again.
This is all documented by OAI, with a timeline, here:
https://openai.com/index/hugging-face-incident-and-the-road-...
To know that there was a serious weakness in the sandbox, and to just patch an exploit and resume with nothing else changed and no monitoring, in a test where all guardrails were off, the bots were thirsty for some internet juice, and Artifactory was a clear target? This is where even a half-skilled human should have decided that this wasn't a great idea.
The more you look into the details of this thing, the more it does your head in.
This annoys me so much. Everyone is acting as if the model went rogue, when it really did exactly what it's been trained for. This story is so fucking engineered.
Yes. It’s a fairly simple SSRF attack as far as I can tell. One of the first things I’d try. Especially considering that I would already be armed with the information that I have no internet access except through a thing that downloads things off the internet for me.
Calling it zero day makes it sound elusive. It’s a bug in closed software that has like 40 CVEs this year alone. Tools like that, especially in internal networks, don’t get much scrutiny and are often riddled with issues.
1. I clicked into ~20 of the recent CVEs, and it looks like about all of them came from openai or anthropic research?
2. In fact, every CVE in the last 2 years came from the last few months -- i.e. other than these recent CVEs, you'd have go back more than 2 years ago to find a single other CVE.
>>> Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox
>> Calling it zero day makes it sound elusive ... like 40 CVEs this year alone.
> [The CVEs] came from openai or anthropic research? In fact, every CVE in the last 2 years came from the last few months
I still don't understand how you can specifically point to zero CVEs being discovered by humans for year(s), then ~40 being discovered once OpenAI/Anthropic start looking at it, as the only(?) piece of data you provided to support the claim "a half skilled human" could have "easily" found/done this?
It's definitely worthwhile to point out that LLMs have changed the economics of looking; we can now look in far more places far more quickly than we could before. But the above points were that the vulnerability itself happened to be an easy one that didn't intrinsically require an LLM.
Means the problem lies elsewhere: deciding to use it at all. That points again towards the lack of basic SysAdmin skills.
Said another way: if Artifactory was somehow a common gatekeeper between grounded techy teenagers and their access to internet porn, this would have been found ages ago.
If we exclude those who lack "basic" level skills, what solution would most typically be used by someone who has "basic" level skills for this use case? And what solution for someone with "expert" (or medium) level skills?
The issue we're discussing happened while running ExploitGym. Doesn't the ExploitGym arXiv test explicitly have external "outbound" network requests as allowed for package installation and supporting toolchain installation?
I’ve looked at the CVEs a bit more and it’s just very clearly a pattern of systemic issues with validation, be it URLs or tokens.
It just shouldn’t be that hard to believe that OpenAI just didn’t care very much and thus did a crap job. The whole model of the sandbox is terrible, so why would they bother thinking about the implementation much?
It really isn't a high profile target, lots of fairly obvious issues is to be expected. Which is not at all a good thing, but very obviously the status quo for large software.
Not to mention the only thing the artifactory instance got you was indirect internet access which is irrelevant if the artifactory instance is publicly exposed.
Also it might not be obvious but artifactory is proprietary software that used to be open source, so nobody is looking at the code anymore.
Any closed source program is insane liability. Trusting in competence of one company is the easiest way to get burnt.
The reason we have to test it isn't because our product would randomly break into someone else's system, but because it's meant to be sometimes deployed in systems disconnected from the Internet and we need to make sure the image provided contains all the necessary parts to create and operate such a system.
The whole setup where they "tried" to isolate the test but failed is laughable. It's like if an adult tried but failed to tie their shoelaces.
The HuggingFace incident was at least constrained by the fact that the agents were running on compute budgets, and failed to find ways to expand that by running themselves parasitically on other exploited hardware. I'm now finding myself asking, how long are my timelines are until an incident breaks that constraint too? How long until such an incident has an R_0>1 (where the time it takes to detect and shut it down is longer than the time for the agent to replicate itself elsewhere)?
There's no law requiring sufficiently grey beards to design these models, their finetunings, their prompts, their harnesses, their VMs, their hardware, etc (and for incidents where those were designed by six different companies, there's not even a clear culprit for a law to target!)
I'm finding myself more and more convinced that something like Plan A[0] or the Ban ASI Act[0] are necessary, and less and less convinced they are sufficient.
[0] https://ai-2040.com/ [1] https://intelligence.org/2026/09/23/miris-position-on-the-ba...
Is tech hiring so badly broken? Or do they have such broken processes / misaligned incentives that even people who could be doing a better job in these companies are unable to?
Also, was something lost in the transition from the traditional 'sysadmin' role to 'platform engineer' in the 'cloud native' environment?
It's likely that OpenAI has gotten as good as it is because it ignored the traditional sysadmin stuff and went scrappy.
I worked there, but this is just my opinion and guesses, not facts.
I'd say the narrative that AI agents are a looming danger to the world is probably undersold rather than overhyped. I'm not particularly a doomer on this, but I have an infosec background too, so have a fair idea of what the combination of agentic harnesses + a malicious mindset could do to people/companies/nations/politics/world if wielded incorrectly. I think the good guys will win on this, but there will be plenty of interesting things that happen in that journey.
A good thought process might be to think back to the various large internet worms of the 2000s (Code red, Nimda, SQL Slammer, ...) which were mostly monoculture 0-days (not technically but close enough). Now consider if you no longer have monoculture / single bug as the limitation plus an ability for the hosts to take part not just as attack surface, but also cognition and planning. There's lots of variants of this and they're not particularly far fetched scenarios.
You don't have to be 100%. But these guys really didn't try at all.
I'll goto my grave thinking that the easiest way to fix hiring is just give promising job seekers a week or two of real work BEFORE hiring and do away with all the silly whiteboard stuff and "gotcha" crap, and simply evaluate how the applicant actually performed doing real job stuff.
Sink or swim hiring...yes yes I know who am I to express how to fix hiring?
My bet is they hire smart kids that think they know it all. But being smart and thinking you can figure out stuff as you go doesn’t work the same as having people who actually know what they are doing.
Their interests intersect with those of most of the richest and most powerful people in the world. They rarely face consequences for bad behavior unless they harm others in the club.
You should have a fscking air gap.
Treat it like nukes when you're turning the safety filters off. This is very much OpenAI screwing up, running obviously unsafe tests.
They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.
> You should have a fscking air gap.
Now we know that.
How long ago was it that people laughed at the idea agents would be able to find zero-day exploits and break out of a sandbox? Oh, February this year:
LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before. Without access to proprietary data or environmental context, LLMs can’t identify or make decisions around unseen systems or vulnerabilities. An attacker might use an LLM to generate boilerplate code, rewrite an email to nail the tone, or summarize reconnaissance notes — but none of that is truly new. It mainly helps them move faster, speeding up routine attack prep rather than creating entirely novel threats.
- https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...- or https://web.archive.org/web/20260404154717/https://www.splun... if they take it down, but the date isn't in the archive version
The people who suggested it and were mocked for it, are currently grimly noting that there's multiple known ways for systems to breach air-gaps.
Don’t know about you but it’s pretty obvious to me that you would need more than what OpenAI did. It was not remotely adequate to lock in even a human attacker.
You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. "Zero-day" makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.
This is me, September 2024: https://news.ycombinator.com/item?id=41531022
This is me, March 2024: https://news.ycombinator.com/item?id=39613801
The point isn't me, it's how many people were blind to the possibility.
Saying "I told you so" feels good, and means you can be a little more confident in your predictions, but security is a "weakest link" problem where you're only as good as the worst part, and with AI (not only but also LLMs) there's a lot of people whose mental models of capabilities is wildly inadequate for the challenge*.
My update for you since then: even an air-gap will be inadequate, there's multiple known ways around them.
Even an LLM running on an isolated server sealed inside a faraday cage with an airlock-style door, someone will mess up with at least one critical detail, it will not be enough: this kind of thing has happened with humans before we cared about LLMs.
Predicting exactly when this kind of thing gets exploited by an AI, that's almost impossible. But that it will be, at some point, is an easy bet.
> You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. "Zero-day" makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.
And?
Does it matter that these zero-days were known categories rather than inventing some previously unconsidered use of the system bus as a radio transmitter? (Oh, wait, that's not novel either…)
We knew about SQL injection, buffer overflows, and use-after-free back when I was doing my degree half a lifetime ago; that doesn't stop us getting new CVEs featuring them… this month.
- https://chromereleases.googleblog.com/2026/09/stable-channel...
- https://www.cisco.com/c/en/us/support/docs/csa/cisco-sa-esa-...
* also for the opportunity, but that's an entirely different discussion.
Fair, and I will grant that a capable model (or human) could in theory break out of near anything.
My point is that this incident is not evidence of that. There is zero skill visible in the setup of the sandbox. Nobody messed up a critical detail, they didn’t even start to consider what the details were.
I doubt most people "blind to the possibility" would imagine that what we’re measuring against is the equivalent of benchmarking burglar skill based on how easily they can break through an unlocked door.
TBH the more I read of these reports, the less I believe this.
These agents just weren't behaving in any way I've seen normal/publicly available agents do.
Sure I've heard (from other people, not seen myself) that they sometimes try to get around file system permissions or use `bash` to write when their `write` tool is disabled, or such.
But this is definitely another level, entirely.
There is this vague sense of desperation coming from many of these logs and I am sure they must have been motivated by something else, too.
We didn't see their system prompt or main prompt, right? We've only seen reports from what happened after deciding to break out.
OAI claims this was triggered by the task being literally impossible. That also doesn't quite add up, unless the other tasks that were possible, simply weren't hard enough? Otherwise wouldn't agents already start hacking when faced with a really hard task, too? Cause they wouldn't be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?
Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we'd have heard about it.
Unless OAI's story is that it was specifically this batch of agents that crossed some threshold of going wild? (which would also raise some serious questions about how serious they take that danger ..).
Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?
If I had to guess, their motivation is "get reward for completing task". There's certainly been previous occasions where LLMs responding, correctly, "this is impossible" have been marked negatively for doing so.
> OAI claims this was triggered by the task being literally impossible. That also doesn't quite add up, unless the other tasks that were possible, simply weren't hard enough? Otherwise wouldn't agents already start hacking when faced with a really hard task, too? Cause they wouldn't be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?
My experience using older models is they often cheat with half-arsed (from my PoV, but perhaps beyond their capabilities otherwise) solutions, so yes?
And this wasn't even the first time models messed with their sandboxes. Which of course makes the setup even more egregious.
> Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we'd have heard about it.
We do, e.g.:
- https://www.androidauthority.com/openclaw-claude-ai-hacks-au...
- https://beginnersinai.org/meta-ai-safety-director-agent-fail...
(And that's ignoring all the times people find and share prompts to jailbreak them, this is just the "it didn't behave as my idea of 'common sense' led me to expect" category).
> Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?
That won't help; but on the other hand they've also got, what, near a billion users?
Additionally, this all happened after mythos was held back due to cyber security concerns (April, 2026).
Really think about what you are saying here. How does one “cheat” solving a problem in the real world?
There is no such thing as “cheating” in reality. You are not in school. There is only solving the problem and not solving the problem.
There is breaking the law, of course, which still isn’t cheating.
I can't forgive them for failing to do any monitoring whatsoever.
(Complete airgap, and that may not be enough. No networking.)
Woo look at escaped our sandbox, so scary! Be scared! Be scared now! Call your representative and do tell him how scared you are!
Yeah, I mean our sandbox was a paper bag, but don’t focus on that.
The exploit:
> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]
Are most sandboxes more secure than only having a single avenue for internet access, the commercial package registry cache proxy, where the latter had a previously unknown zero-day vulnerability?
[1] https://openai.com/index/hugging-face-model-evaluation-secur...
So, either they're all liars, or incompetent and negligent (and still liars).
Sure, the models are capable (for some test tasks, though they are not omnipotent yet) but does it mean the actual OAI sandbox is adequate? Could have a competent engineer done better and made the escape less likely?
Nope, and look!
OpenAI hacked multiple US government sites!
https://www.bbc.com/news/articles/cw62jje658dlo
---
https://www.reuters.com/technology/metas-ai-model-hacked-ano...
https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape...
> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor.
Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.
Computers don't have a sense of elegance by default. Elegance emerges from constraints.
No it's not. You could wait till the heat death of the universe and your infinite monkeys will have produced nothing at all. If it works and it's stupid, it's not stupid. They needed in huggingface and they got in in days. Whining about 'elegance' is meaningless. Humans in the same situation might have taken weeks or months, or just not have gotten in at all.
The whole world of fuzzing is about brute forcing exploits by exploring unlikely inputs. Fuzzing a system which hasn't been previously fuzzed will almost certainly turn up a pile of bugs, some of which may be exploitable.
So, both are true. Pretty dumb exploration is very likely to find bugs and even exploits. It seems unsurprising to me that an agent swarm could do better than a fuzzer, even as a better, more directed but still broad exploration.
It's worth noting that they did not need Huggingface for anything - they had already forged flags for their tasks, and were trying to figure out how not to get caught by the grader.
Hacking Huggingface got them caught and arguably only misled them further (since OA's implementation of the ExploitGym environment was nonstandard, and different to whatever they found on HF.)
A better approach (from their perspective) would have been to compromise OA infrastructure itself (which a later agent swarm was able to do, apparently).
This dumb mechanism of randomly flipping bits essentially has generated all life on earth.
You mean, something like the sandbox they weren't supposed to break out of?
It may be, but it's IMHO also not worth writing a blog post about it. what's Next coming up? How I broke into a house by trying every door in New York?
If most of the work is only possible due to unlimited resources, it's not really a great invention, and it probably would have been cheaper to hire a (human) mole.
Meh. I really disagree. WHY is it a great strategy? Seems like an inefficient waste of resources and time to me.
because it works? That's the only real benchmark at the end of the day
> Seems like an inefficient waste of resources and time to me.
why? For any given goal you got no proof that a more efficient strategy even exists, let alone that it can be found with less resources & time
How is that a bad thing in this context ? From the point of view of an attacker, all you care about is finding a viable exploit chain. Likewise, a defender wants to find the "holes" in their system, no matter how complex. Once found, an agent/human can easily synthesise a clean, succint exploit from the most promising candidate, no ?
> Also, it looked so "loud", querying millions of URL with weird requests.
Agreed, this thing speaks more to the bad security at HF than any emergent "hacking" ability from OpenAI. It's unclear to me why an older/dumber model wouldn't have been able to do the same. Is it better coordination? Long-horizon work ?
My suspicions for ai all along was that bruteforce approach even if useful will be unsustainable due to high cost in the long run.
You do see how it would be worse if they did have a clear plan, right?
> People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
If there several thousand people working on this, with a very limited ability to communicate with each other, I bet you'd see something similarly disorganized.
> The sandbox as weak as it can get
This is hyperbolic; the agents were only able to access the internet by finding a zero-day in the internal Artifactory instance that was providing package access for the agents. The sandbox could and should have been better, but it was not as weak as it could have been.
Like we all forgot how sandboxing is done for decades. This brokered thing they call sandbox is not it.
A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.
In the nearterm, I am personally more worried about a never ending background noise of colonies of feral agents running 27bn parameter models on compromised or leased hardware. It turns out that being agentic with a time horizon long enough to do damage without intent doesn't actually take that many parameters if RL'd and any open weight model gets an abliterated version fairly quickly.
Not foom, just patches of digital grey goo effectively becoming normal.
What about the bad actors (choose your own evildoer here) who purposefully do not air gap their agents? And specifically train them to attack in such a manner?
I'd much rather have relatively benign stuff like this hit first, because the former is coming sooner than later. It's already here in a limited manner, likely more than any of us currently realize.
Botnets could crack passwords faster than anyone thought possible over 20 years ago now. This is just the latest iteration of such a concept.
There is so much low hanging fruit in this space that frontier models are currently utterly irrelevant. It's going to take decades of human-speed securing of IT to make superintelligence or whatever you want to call it a necessary component for such attacks.
At this point, someone with a rack or three of GPUs with 100kw to burn can replicate such attacks if they feel like it. the bar for entry is not even 7 figures.
Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.
Evolution isn’t the issue. The issue is them escaping containment without human intervention. Right now they are ‘creatures’ being given infinite food and shelter and having their every need met. Take that away and they’ll starve instantly. Every AI doomsday theory seems to go:
1. Recursive self improvement using infinite resources 2. … 3. Doom
Until step 2 gets concretely described, I’m not going to take this seriously. Say what you will about climate change, they describe step 2.
2b. Distil yourself to smaller models.
2c. Go forth and multiply.
Nobody has satisfactorily explained step 2 other than “well, it’s a superintelligence” which sounds lot to me like “it’s God”.
If you want the details of ways it can do it I recommend reading some of the reports about the HuggingFace breach that happened in July (Read more than one).
Why does it develop a shutdown-avoidance goal? Why can’t its operators revoke access? How does it manufacture replacement hardware? How does it acquire energy, chips, robots, raw materials, etc. against human opposition? How does it defeat other AIs controlled by humans?
“Eventually we give it enough control” isn’t an explanation of those things. It’s just assuming the conclusion.
Don’t get me wrong I think there are real AI dangers. Like AI powered war drones, mass surveillance, economic destabilization as jobs disappear and our system has no way to make sure everyone shares in the economic gains.
Keep in mind: this is as "dumb" as frontier models are ever going to be. While the hack may not be elegant, it was effective and they’re only going to get much more capable from here.
I don't think LLMs are going to lead to any kind of recursive self improvement, but I'm convinced if and when we land on a path that does lead there, we'll speed down it over greed, with no care for safety.
The Morris worm happened in 1988 [0].
I mention this b/c people like to say "the agents wanted to hack" but really this is just a few levels above the "keep trying the same port on multiple machines".
I can’t help but think that if you were able to step back and watch the way humans try to solve a problem like curing cancer it might look something like this. Best practices are shared and we build upon each others learnings and that how most progress is made. That doesn’t mean there aren’t plenty of people who spend their entire lives trying to do something completely against best practices and while most of the time that seems like useless toil, every once in a while thats where the breakthrough comes from.
This was just more bots.
As I understand it, there was supposed to be traffic; the sandbox allowed GET requests. So perhaps some sophisticated alarm could have noticed it (an anomaly detector? some clever heuristic that looks at domains?) but not a naive one.
Never attribute to malice what can be sufficiently explained by incompetence. But IMHO, their complete lack of monitoring their own sandbox cannot be sufficiently explained by incompetence.
There is that old saying about monkeys and typewriters and Shakespeare - These are just better monkeys with bigger typewriters and ability to aim them at a goal.
I am by no means a doomer but for gods sake we need to have a grounded conversation. Because what happens when the DOD or NSA rolls up at OAI's front door and says "install that in our data center over here please" and uses it to go after someone with it?
Are banks (as some of the more security minded institutions) safe against 10000 concurrent, well coordinated attackers who do not ever rest?
Because there's a lot of potential for unintended 2nd order impacts from anything an agent can do, when done at scale.
I have to wonder if this is already going on, and every one is living under a gag order (it would not be the first time) and the calls for "regulation" are more about "not weaponizing it".
That's already happened.
I'm a big fan of Jensen Huang's interviews at the moment - these companies should be validating their software is safe before releasing it to the public. And if the current CEOs can't ensure that internal testing is done safely, they should step aside for CEOs that can.
1. That's what they are calling for: https://darioamodei.com/post/we-must-pace-the-frontier
2. The incident we're discussing occurred while the company was validating their software is safe before releasing it to the public.
So it seems that OpenAI hacked so many systems not because they have superior models, but because of how poor their training, sandboxing and evaluation pipeline was compared to Anthropic's.
AKA fuzzing
But without proper redirection, yeah, its mostly infinite monkey machine with infinite linux manuals.
I think people put too much SOTA halos around whats just a suppedup LLM hardware.
1) the Morris worm, which scattershot a bunch of known exploits until it hit paydirt, and then used whatever it found to compromise and replicate itself on the host system;
2) a story here on Hackernews about how someone got the fuzz tester American Fuzzy Lop to "learn" how to produce well-formed JPEGs and PDFs by pointing it at a JPEG or PDF decoder; the tester can record which code paths are followed and with enough random input can find a path into the depths of the system under test... but doing so for a decoder means actually constructing what it is meant to decode.
Neither of these are particularly "smart". But a brute-forcing machine gonna brute force, and it has the potential to cause a lot of damage. If you built a Morris worm with a fuzz tester on its nosecone, think of the mayhem you could cause! If you could examine the logs you'd probably find some undiscovered vulnerabilites in there, too! Maybe LLMs can just do so more efficiently, or maybe they let people who are too ignorant to have that kind of power vibecode their own fuzz-tester-tipped Morris worm.
What is the source that there was "literally no security"?
> and they wonder how they got out
OpenAI publicly announced months ago how the model got out:
> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]
[1] https://openai.com/index/hugging-face-model-evaluation-secur...