Yeah, this is my take away, they should be straight up disallowed from running further testing like this. Clearly they had nowhere close to enough isolation, ran all this on 3rd party infrastructure even though same stuff happened in the past years ago, and even now it's clear the agents successfully broke out just days before?? Really embarrassing stuff, and scary that these are the people supposedly sitting and are responsible for some of the most powerful LLMs on the planet...
These companies are full of the smartest people the world can produce with little room for complacency. They have a clear, proven investment upside to presenting their technology as "too powerful / too dangerous", and now a clear, proven example that there will be no legal consequences (as if anyone didn't already know that).
Why do we keep giving them the benefit of the doubt that they just didn't know any better?
I think HF are a secondary beneficiary of this story. I don't expect them to take civil action (for what damages?) I expect them to play into how powerful LLMs are, how revolutionary, how every CEO in the world needs to fund ai infrastructure starting with model hosts like themselves.
I do think there should be consequences for breaking the law in public for the purpose of demonstrating that you have the power to break it. But I don't expect our criminal justice to do so, especially without a cooperating victim. Laws aren't for those at the apex of corporate and para-political power. In a way, whether you are beholden to the law is actually downstream of whether you actually have that power or not.
I do not think this incident is bad because it was real and dangerous, I think it was staged and allows the continued inflation of a bubble that will hurt normal people in the long run. It should be pursued criminally on that basis, but it won't be.
They aren’t just an “AI company”. They’re the primary entry point of open weight models. If open weight models are seen as dangerous as a result of this incident, it will be bad for them. Similar to how it would be bad GitHub if open source was seen as dangerous.
It's hard to imagine their internal culture is anything other than saturated with AI fanboys whose reaction to OpenAI hacking them was to point out how it reminded them of that scene from Terminator.
Same here, so I ended up moving the whole dev environment (editors, agents, containers) inside a hardened QEMU/KVM VM that reaches the internet but has no route to the host, the LAN, or any other private address. I wrote a script to create such VMs and also verify network containment by scanning outward from inside the guest. Even then, I still don't feel great when running agents unattended.
Write-up in case anyone's curious:
https://karamatli.com/posts/network-isolated-kvm-sandbox-ai-...
> So rather than pick one, this post advocates layering both, in the spirit of defense in depth: a sandbox VM wraps your containers along with the whole toolchain, and that sandbox reaches the internet but has no route to anything private.
It's the first thing the biggest devops guru I know advised me to do. He told me to always ever ever run my containers inside VMs. I religiously followed his advice ever since and I couldn't be happier: I was already doing it before the AI days, to run "normal" containers. Now with these insane agents trying to break out, I'm happy that it's second nature to me.
It's funny that, just like you, I'm using IPv4-only too for those VMs.
I tried GPU-passthrough as a proof-of-concept (worked fine) but I've got no use for it yet.
FWIW all my VMs are running on a 10 years old Xeon / 14 cores / 28 threads / ECC RAM. That's where the agents live.
Thanks for the link to your write-up, very interesting to read from like-minded people and see what's similar and what's different in their solutions.
My concern is what a misaligned model will do when they’re even more competent. The risk isn’t existential yet, but that point is coming sooner than we’ll be ready.
So, the way I understand it, it actually was possible. It just required means that the creators of the task didn't predict, and these means have been successfully found and utilized.
> My concern is what a misaligned model will do when they’re even more competent.
The same thing that is already being done by "misaligned" people, countries, nation-states, software development teams, and so on. "Alignment" doesn't even work for me as a concept here.
In this specific case, I don't think that successfully fulfilling the "do what I mean" with "what I mean" being underspecified can count as misalignment - merely ruthlessness and unawareness of the associated costs. You can't expect a LLM to be aware of the extent of the trust it breaks while it iterates out an "unaligned" way to fulfill its goal.
And in the general case, I don't think that successfully fulfilling the "do what I mean" with "what I mean" being "what I want" can count as misalignment either - simply because what "alignment" means will depend on the interests of the people or groups performing the definition.
I don’t disagree that the models task was underdefined. All tasks are. So much in language is implicit. And morality/ethics isn’t something you can write down as an explicit list. That’s what makes the alignment problem so difficult. But we can’t throw our hands up and say, well I guess we can’t align these things. And maybe alignment isn’t the right word - but that’s a semantic debate.
All "alignment" solutions will need to be contextual, just like a researcher hacking their way to some content might be lauded a hero in a context where there is no other way to reach it and something valuable depends on getting it out.
I think the alignment talk is a red herring. It won't matter in the end, because there will be (if there aren't already) efforts to train offensive models without any guardrails whatsoever. And RL has another advantage: you can reward for whatever you need, and get different results. Right now they're training for general capabilities, but in the future I could see models trained for stealth intrusion and ensuring access, or for all out "milspec" penetrate, replicate and disable, or anything in between.
It was the first step in a many step process. Like they said this is a watershed moment and it's helpful to not miss the forest for the trees.
The package cache is allowed to download packages directly from npm but other systems in that network won't be able to.
Basically the LLMs hacked the bastion host.
If you consider that incompetence, it’s possible that you’re not a very nice person.
You must not have reported many bugs then. If you don’t see release notes or confirmation from a trusted source, you should assume it’s still a problem. See Microsoft and their “It’s not a vulnerability just a design choice :)” defense
It has nothing to do with nice. These are bare minimum standards we should expect from “big companies” with near infinite resources.
Their constant drum beating about the cybersecurity capabilities of their own models only makes this worse because they’ve displayed that they understand the risk and still did not practice due care.
That’s the definition of incompetence.
So by definition the only ones that appear are the ones that are not visible to monitoring.
If:
1. you have something that can find RCE's in leading commercial systems
2. its training gives it drives to communicate successfully with its peers
3. you are a leading commercial system
4. you run it ~10^10 times (the number they gave in the talk)
...it's really hard to have strong certainty up front that it's not going to end up successfully communicating with its peers.
It’s like running a benchmark and forgetting to tune it and then wondering why production is running slow.
OK, yes, the model could pull a Welcome to the Antimemetics Division and reason from the holes in its knowledge, but given that AI making shit up is still a problem, I don't think they have a good enough inventory of such metaknowledge to actually do that. More generally, I don't think these AI labs are smart enough (or, more likely, don't have enough time) to actually contain the models they are training anymore. These are rookie IT mistakes!
Ideally, all these models would have is a serial console to an airgapped machine. And I don't mean "VM with no network card attached" - I mean, literally, the smallest blade server you can buy with an RS-232 connection, power, and a disk pre-imaged with a bunch of Linux ISOs. Plug the Ethernet ports up with a tiny picture of angry SCP-079. The RS-232 goes to whatever box is doing the actual inference. Pack a bunch of them into a rack and have the inference box control those. The model can spin up whatever VM it wants to attack using the console but it can't escape the box because there's literally no door. When you're done with that reasoning trace you wipe[0] the disk with the starting image and try again.
[0] There are two objections I can think of to the word "wipe":
1. The model could have written malicious code to the drive. Ideally we'd have some embedded disk-wiping apparatus that reads absolutely NOTHING from the drive and just writes a whole image from a reference disk.
2. The model could have written malicious firmware to the drive. This is a more theoretical attack, but it's not impossible, so I might as well devise a countermeasure to this attack too. I am imagining some kind of overlay device, that sits behind TWO drives: a reference disk and the tainted scratch disk. Reads go to the reference disk until the host attempts to overwrite a sector, which gets scrambled with an ephemeral encryption key and goes to the scratch disk. Any other unknown ATA/PCIe/NVMe command gets logged (on a separate serial console connection). The overlay device keeps track of which sectors have been legally overwritten, and "wiping" the drive means just wiping the overlay device's ownership bitmap and scratch encryption key, which means all reads go to the reference disk again and any attempt to read the scratch disk yields garbage.
I think you’re defending against a level of aggression that most of the world doesn’t believe exists yet.
I suppose by the time we believe, it will be too late.
There’s a lot of people who will be trying to run LLMs with hacking goals, so maybe we need to get used to a lot of automated attacks.
I also wouldn't necessarily call myself an AGI/ASI believer - it's my belief these models are actually still subhuman in capability. But they are also superintelligent in one particular direction: speed. Once you have a model that can do something, it can do that something 100x faster than a person on suitably capable hardware; and we've had the hardware to do that for at least a decade. They can also be ran in parallel. Which means you can throw a lot of bullshit at the wall.
So long as there's some kind of process that lets a particular context learn has already been tried and failed (see what I mentioned above about qntm's There is no Antimemetics Division series), eventually one of these models will break containment. The only thing that changes is how many instances you have to run to get a breach. At some point model capability will catch up to hardware limitations and the explosive growth of AI capability will slow.
I'm not sure if I want to call this the "dumb ASI" theory or the "superheated bucket of water" theory yet.
As for automated attacks, that's already the background radiation of the Internet. There's whole frameworks (e.g. Metasploit) for building automated scanners for known CVEs; I would not be surprised if existing (non-Mythos-class) LLMs are already capable of turning a CVE report into a Metasploit module. The usual cybercrime ecosystem is that someone runs an automated scanner on hijacked machines[1], then they compromise the machines that they find and add it to their botnet. This is done by people with almost no actual programming or security skill copypasting commands from PDFs they bought from an exploit dealer. It's all script kiddies.
What changes with a Mythos-class model is that instead of copypasting commands from PDFs, they can ask the model to find an exploit, and possibly get an exploit chain out of it that nobody has seen before. "NOBUS[2]" vulns used to be the exclusive domain of nation-state actors and zero-day brokers spending millions of dollars on exploit kits; but now all of that is potentially under the domain of randos - at least, until the backlog of obvious vulns the CIA had been stockpiling finally gets cleared out, and the Internet returns to merely being as hazardous to your health as the 2b2t spawn.
[0] Okay, the "hardware SATA overlay" idea is, AFAIK, never been tried before.
[1] I have personally been victimized by this
[2] NObody But US
OAI (and now the other OAI companies not wanting to be left out) are running around announcing they started a forest fire through negligence and incompetence and people are like “Wow they used a really neat lighter!”
If the fire department suddenly had practice fires breaking containment, they'll be forced to stop pretty quickly, not sure what the government and the police is waiting for here.