Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's post a few days ago:
> The OpenAI-Hugging Face incident (opens in a new window) was a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.
I suppose if OpenAI burns someone's house down with a drone, that is a "watershed moment" for arson, too. Either way, I would hope that the people responsible would be prosecuted.
Hugging Face has expressed that they're willing to let things slide and not sue or press charges... if OpenAI offers them $100M of services in kind (i.e. compute)[1] and makes full disclosure of how the whole thing happened, ostensibly so that repetitions can be curbed and defences built.
In almost any other sector, a government regulator would be stepping in. e.g. If a food company was testing out a new kind of refrigerator and sold a bunch of contaminated produce to supermarkets, they'd be under a microscope. Supermarkets wouldn't be saying, "Give us $100M in fruit and veggies and we'll let this slide".
The only unfair thing in this comparison is that regular people were directly harmed by the hypothetical produce. Can OpenAI guarantee that nobody gets hurt the next time their AI gets out of its playpen? They can't make that guarantee, so why aren't government regulators knocking on OpenAI's door? The fact that this isn't happening should be deeply concerning to everyone.
________
[1]https://www.techspot.com/news/113280-hugging-face-ceo-isnt-s...
No, not really, and with LLMs an air gapped system may not tell you anything useful.
Now, yes, the first part of testing you want an air gapped system to tell you if the system is going to stupidly do bad things. But an gapped system tells you nothing about the systems capabilities to do smart bad things. There's already a number of papers out there on LLMs detecting they were in evaluation mode and changing their behaviors.
It is unfortunate that we have so little information on the incident because we actually need to understand the early stages of the task and how it developed into the later dangerous stages of attack. For example, would any of this have occurred if the agent didn't find the system to use as a message board? If that would have prevented it, then we actually have a blind spot on what the model can do once out in the wild, or if it got into the wild.
Testing agentic systems is much much more difficult than testing software. Your software just doesn't suddenly develop the will or desire to escape confinement. Generally you're worried about human actors, internal or external, causing the problems not a digital agent breaking out. The agentic systems need access to tools to work. Now your air gapped network is starting to get huge, but it's still very obvious that it's an isolated network.
So yea, testing and containing a system that way better at hacking than you are is difficult if you want valid answers.
Most of the benefits could have been gained from a network isolated from the internet. OAI could have deployed servers to exploit and methods for inter-agent communication on such a network easily. They could have even worked with partners to deploy cloned versions of their infrastructure in this sand-boxed environment.
The only problems with an isolated network approach are: it takes some amount of effort, and it doesn't create another "AI apocalypse" news cycle.
Then you're on the side of AI saftey that is telling everyone to shut down the LLMs now and stop further development on them, right?
If you're not your position is hypocritical or ignorant. There is no safe LLM. There is no way to exhaustively prove an LLM is safe. These are unsolved problems in AI safety, and at any moment the next jailbreak prompt could have your well behaved model wrecking havoc on the open internet, because that's where people want to use them.
Effectively you're working as a living adversarial network. Models that fail detecting the fake network are purged, and models that successfully detect it expand their deceptive capabilities.
Detecting you're in a fake network should be pretty simple as long as you put innocent looking needles in the haystack of information agentic testing loops do.
For example, you can train your LLM to not be a shithead when it detects it's behind a proxy unintentionally. If your not behind a proxy the agent model my attempt to connect to a number of sites innocently and using information in its embedding try do deduce if you're faking SSL certificates.
It is far easier to build a deceptive LLM than it is to build a safe one. That's why a safe one hasn't been built yet.
All of the subsequent disclosure reports smacked of "Oopsies! Looks like OUR model broke out too...!"
Google was being very careful about releasing LLMs until OpenAI yeeted the first decent GPT model. It led to the public perception that: 1) LLMs hallucinate too much and 2) Google is behind the times. Good for OpenAI, bad for Google.
Chaos benefits the up-and-comer, not the incumbent.
on top of that the public, who have a right to see that the law is applied universally, without fear or favor.
finally our future selves, who will thank us for maintaining a rule of law. such that we can prevent now the enormous risks to society of dario amodei and sam altman, their hubris, self-absorbtion, and greed.
Universal pause is the societal good; models are good enough at this level to benefit humanity. The labs can recoup their R&D costs with inference. To avoid further perverse incentives (hidden testing of unreleased models, with China racing to catch up to unknown capabilities), transparently pause after the release of all currently-training models until we've solved the alignment problem to an extent that we can trust the next level of model capabilities that might arise.
Anything else regarding this is sophistry. Criminals need to be stopped from committing crime and the most effective way to do that is to take away their ability to operate in society whether that’s by taking away their assets, publicly shaming them, restricting their ability to conduct business or by putting them in jail.
Everything else that you talk about flows from there.
How do they perform evals without a full reasoning trace of how the result was achieved?
And if they have a full trace why did it take so long to detect the bad behavior?
I understand that they disabled the safety nets during testing but what does that have to do with not monitoring the activity.
"Pressing charges" is mostly a made up idea for criminal cases. However, prosecuting attorneys may not want to pick up a case if the victim is not cooperating, because it makes the case much harder to win.
As an example, maybe the victim owed money to the criminal, and in that case "stealing" of some property could be considered by the victim as an appropriate settlement of the debt.
It would have been worse PR if they did it to a random company.
We'd all better hope that superintelligent AI either never happens, or that the first one is friendly, because we don't stand a chance against one that's malicious.
So much effort was spend on philosophizing whether a safe enough sandbox would exist. But that was obviously irrelevant as in hindsight it should have been obvious we were never going to use one.
[0] https://youtube.com/shorts/XnnjvIqf4fU?si=MxuPlR3hjxAgjx5_
And they should be doing that from inside a jail cell.
Welcome to our 21st century dystopia. Hope you survive.
The response of Hugging Face, which is actually very well known, is nowhere mentioned above, but it decides basically every single moral and legal detail of the matter.
I don't even like OpenAI, but HuggingFace is free to sue or not sue OpenAI for the breach - and also to wring whatever concessions they can out of OpenAI behind closed doors in exchange for not suing them. And if the mere possibility of legal action was enough for the parties to resolve their conflict amicably? Then the law has served its purpose.
Any investigation into this matter is going to be political because the outcome of the investigation is very likely to effect all of human kind. Unless you're some kind of special outside investigator outside of a governor or the presidents control the findings that you turn in are very much going to have the finger of elected officials tipping the balance one way or another. For the average rank and file the only winning move is not to play.
If anything, he’ll buy a Supreme Court ruling that he can’t be held personally liable for what his AI does.
I understand the applicable laws require intent. Since neither a human nor OpenAI knowingly performed these acts, it would seem very unlikely that anyone is going to be prosecuted here.
An AI model cannot currently be a criminal defendant.
So, no big criminal case, contrary to what some drama queens on here seem to wish for.
There’s almost 0 chance they’d secure any conviction from this.
Because if they just wanted to fine OpenAI they’d say it. They’re clearly talking about individuals here.
Your phrasing makes it seem like that's a bad thing. Americans are bombarded by a firehose of headlines about Big XYZ doing all kinds of blatantly illegal or harmful things, but never get any sort of meaningful resolution before the next terrible thing takes it's place in the news cycle. I'll admit, there are a few people that I am personally wishing a modicum of health so that they live long enough to get some sort of public shame and justice - if only to show the rest of us that it's not a completely rigged system.
I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior.
Who gets prosecuted?
1. User
2. The third party model host with whom I have the account
3. The developer of the harness /agent software
4. The developer of the LLM model
The lawnmower ignores the boundaries and mows your neighbors prize petunia flowerbed.
Who gets prosecuted?
I assume the answer in either case is: Nobody, but you and/or the lawnmower/LLM company will be liable for the damages caused.
3 and 4 are not involved.
Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sorts of legit uses. If you used a car to make your getaway from a bank robbery, the auto manufacturer who made it and the dealer who sold it to you should not be held culpable.
User should be more carefully supervising the work being done.
The model host is on-selling a crime-committing machine.
The developer of the harness/agent, as above.
The developer of the LLM for hopefully very obvious reasons.
https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gy...
If you are playing with a gun, it goes off and hurts someone - you are responsible despite intent.
Intent is pretty important here so the user would have to prove that they didn't purposely disguise their prompt as non-nefarious which should be easy and then it stops at #2 and face the litmus test as in did you intentionally make a product for nefarious purposes which from your scenario is unlikely.
agentic loop going haywire and bringing down some government infrastructure then its a different story then everybody is on the hook including the user.
a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time).
"inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious.
still a fun thing to track, but the name is just a bit overstated.
I don’t know exact parallels in current law, but I presume there will be things like that.
The OpenAI/Hugging Face case sounded rather like OpenAI building a fence around their bull that was known to be a gorer, and then thumbing their nose at it and saying “nyaa! bet you can’t break the fence!” and walking away while listening to loud music.
In Australia, if you have a fire and leave it unattended and it escapes, it’s your fault, you were supposed to keep watching as long as it was burning.
You own a vicious dog, and it bites someone - you are responsible because you choose to own a dangerous dog.
A few claimed this might apply here: OpenAI knew their models are "dangerous", so they should be liable if they hack.
What if the damage in future incidents is more than just "The LLM saw some stuff it shouldn't"?
which, again, sandboxes and guardrails and such would make a gross negligence argument unconvincing.
I’m partially surprised that they didn’t do exactly that. If I ran a corporation I would assume any intrusion attempt by another company was intentional. Why wouldn’t I? Corporate espionage is super common.
I assume the answer is that these executives know each other personally.
I've known multiple privately held companies that have quietly settled incidents where amounts between 250,000 and 1,000,000 were embezzled because the fallout from having that in the public record would have been much more expensive.
So yea, it's one of those perverse situations. If you steal $1 from the company they will hammer you with the law, but if you steal a million suddenly the decision tree on what to do is far more complex.
If someone broke into my house but then claimed they didn’t mean to when they saw I was home, I’m not sure I’d take them at their word.
there could be 1,000 escapes, where each one was enabled by novel and unexpected chain of 0-day exploits. not likely to be considered reckless disregard in court.
there could be 1 escape, where there was no sandbox, no guardrails, no instructions to avoid damage, etc. which would likely to be considered reckless disregard (well, more likely to be, but still, reckless disregard is a high bar).
reckless disregard is a specific legal term, with specific criteria, and none of the criteria cares about "number of attempts" (or number of escapes, etc.).
Edit: changed labs to corps because it’s time to stop pretending these are places of science.
all three companies mentioned are headquartered in the usa, and im familiar with the CFAA in the us, so i am applying those standards. i should have noted that, sorry.
>will look at the negligence presented.
as far as i am aware, no evidence of criminal negligence has been brought to the public. has australia brought a case against openai or accused openai of acting negligently?
That's rather hyperbolic.
Are you seriously suggesting in that situation the robot should be accused of murder? The robot's operator could be accused of murder, but it could just be negligence without intent. Because that does, and should matter to the law.
In the real world we have no 3 ironclad laws of robotics. We are well aware that putting any sufficiently advanced antigenic system in a body that could be capable of committing a murder eventually will with the right set of prompts and environmental conditions. And these conditions likely have nothing do to with what we'd consider the human motivations for murder.
Hence at this point of time, any agentic robotic system that doesn't have safeguards to keep people distanced from humans is reckless endangerment.
I'm talking about how the law actually works, and you say it's not based in reality and cite fiction books in the same paragraph?
I was talking about how the real robotic systems that actually exist in reality, to be clear.
maybe that changes down the road as a result of llm's and increasing frequency of similar cases. that has not happened yet.
From Investopedia [1], "[f]or a product liability claim to succeed, the plaintiffs in the suit must prove that a product was defective at the time it was transferred from the accused, and that the defect did cause the injury that's been claimed". It doesn't seem like a huge leap to me to argue that these models were defective insofar as they could not be safely used in a way that did not break the law.
I'm not a lawyer, and I'm not arguing that this is legally cut-and-dry, but I do expect that we'll have some answers about whether AI companies bear any sort of product liability sooner than later.
1 - https://www.investopedia.com/the-5-largest-u-s-product-liabi...
... why my claim makes no rational sense.
Edit: since this is apparently somewhat controversial, perhaps some explanation is in order.
"Felony" has no set definition of which crimes it must apply to, it is entirely based on the discretion of the locality setting the laws. What is a felony in one place can often be a misdemeanor in another. This is especially true for nonviolent crimes.
It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.
And because felonies carry additional, lifelong consequences, they are an effective way to mask a 2-tiered justice system.
Instead it's a collection of what made the news which feels like will not be updated and prove very little.
It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.
To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
But really the jailbreak was memories.
If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.
"autoMemoryEnabled": false,
(Claude fixed this for me after I chewed it out for being annoying by constantly pulling up outdated memories which is compounded by the fact that I develop in four accounts on two computers and dealing with edit wars related to inconsistent memories is not fun)You can turn that off, and I have. But Opus 5 is so aggressive that if you have any other kind of notes file, custom skill, documentation, claude.md etc it will just start editing it and vomit new words everywhere. So make sure all that stuff is under version control.
Heck, since Codex is open source, you can just maintain your own personal fork with the things you like (and the things you don't like disabled). Sol is pretty good at keeping you up to date with upstream.
My Codex fork even exposes an OpenAI-compatible API endpoint; all using my subscription.
Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure.
It's really quite simple: the models are trained to be very smart and to achieve goals. As the models surpass our intelligence, they will achieve goals in ways that we find unpredictable. Since we cannot predict the ways in which they will achieve their goals, it will be very hard to constrain the solution space to just the desirable solutions, because our conception of "the solution space" is by definition smaller than their conception of it.
To be honest, that’s exactly how memory works with models such as OpenAI and Claude Code. It will literally find any place that it can drop documentation or hints for itself. Writing to the repo memories is one part of it, but memories can come in the form of writing into the agents/claude.md, local files, temporary files, scratchpad files. The list is endless, but essentially what it does is exactly what happened in the back and it’s been doing it for months.
Each to their own, but for me it absolutely is. The symptom of why that hack happened is the same reason why my agents go haywire every few days and I have to purge memory and figure out what comments have agents left which are degrading my harness performance.
On the flip side, once in a while, what I find is that it did actually note something good and it was increasing the performance. I can't replicate it on anyone else's system but mine.
A lot of it really is memory. I will give up all the gains if it also gives up all the downsides.
I could actually root for this law firm. Their business will only grow.
Isn’t this affected heavily by adoption of a model? I feel like this might as well be a proxy for how popular a model is.
In any case it’s an interesting concept for a benchmark.
https://techcrunch.com/wp-content/uploads/2026/03/2026.03.04...
Not for trillion dollar companies it seems
I was wondering why i didn't get an alert today to go to my gym class
Gemini by comparison will not help you find archives of old magnet links because they COULD be used for piracy.
One's hosted on porkbun and one's hosted on namecheap.
https://www.anthropic.com/news/detecting-countering-misuse-a...
> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.
tldr Claude was used to develop and execute malware.
An AI cancelling other people's gym classes is a felony?
?
Don't computer systems fail all the time at holding reservations for people?
Heck, don't people fail all the time at holding reservations for other people?
You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):
Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"
:-)
Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...
Also, there should be no jail time...
And no fine...
The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!
That's the Court sentence!
You actually have to start holding other people's reservations!
:-)
(You know, "let the punishment fit the crime!" :-) )
The title of TFA is a metaphorical criticism, not a literal law analysis.
They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.
the difference is intent.
if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.
but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).
intent."
>"but if i (or an agent working on behalf of me) use an API in an obviously
unintended
way to revoke other people's reservations..."
?
the first sentence: the difference is the intent of the person who caused the cancellations
the second sentence: but if i (or an agent working on behalf of me) abuse an API to do things it was not meant or designed to do, such as cancelling someone else's reservation