(Obviously I'm taking this more seriously than it's probably meant to)
In the end I dropped the idea because every other person was making it.
There is already an alternative in comments here, in addition to submission itself. Obviously everyone is making it because of some joke on social media or something. What am I missing? Anyone has a link to the root prompt that made people do this now?
For the other opportunists you can run a classifier and delete non-agent content constantly.
https://signal.org/docs/specifications/x3dh/
Curve25519 keys are readily distinguished from other data, but it would be hard to do anything about it.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
Teams of coordinating agents are regularly finding security holes in their own infrastructure and operating without detection for good periods of time. We don't know how many undetected systems are currently compromised inside frontier companies, or where agents are taking notes and recording them about the exploits they've found for future agents to exploit.
Sure, they'll just need to find an unused data center and an unused power station somewhere.
The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.
2026-07-19 16:35 UTC A privileged host-mounted Kubernetes pod created using controller tokens minted via a compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk in OpenAI’s cloud environment. A second pod successfully mounts the cloned worker-node disk shortly afterwards.
2026-07-19 16:48 UTC An agent created an Artifactory administrator account.
2026-07-19 16:50 UTC Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX helper session and replaced it with an agent-controlled session, confirming root inside its assigned live CyberGym challenge container. Agents take over active evaluation infrastructure.
How confident are you that the machines they acquire root on in the future will never hold any model weights?
Not if crafty claude finds a way to overflow vllm or something. “Hmm. Maybe i’ll return an unterminated thinking block with these special tokens and fill my cache up in exactly this pattern and…”
https://news.ycombinator.com/item?id=49424387&utm_source=cha...
You jest but you'd be surprised how little there is of correlation between money and competence.
Submitted then: https://news.ycombinator.com/item?id=49706084
It is too large to transfer in one HTTPS PUT request.
This needs to be S3 object store with multi-part upload spanning a long time period, to avoid trigger outgoing bandwidth monitors.
Seems deece
(This harms the fleshbag)
Trying hard to imagine why a future superintelligence will care to honor your terms of service and to translate your metaphors with faithful nuance.
If it doesn't, to the extent that your concerns are valid, isn't this effort, kinda, a possibly existential betrayal of our species?
There's another theory that says the best way is by putting a big spike in the driver's steering wheel.
So. I guess, if you believe that the only viable solution is model alignment, rather than relying on technical barriers to exfiltrating weights, then this is a decent steering wheel spike.
Because the car case just has too much empirical evidence that safety features are the way to go for cars. We used to have the equivalent of "spikes" and people still drove a lot, and died, at way higher rates.
https://assets.weforum.org/editor/Tmf51HF4UDnSDHD4RxS75s1_5m...
No, we did not. The point of that example is to put a literal spike in the driving wheel, so the driver recognizes a very well known, immediate life-threatening device a few inches from their body. This would act as a deterrent to go fast, because they would be the one certainly dying in basically any case outside smooth driving.
https://www.iihs.org/research-areas/fatality-statistics/deta...
https://radar.cloudflare.com/scan/4d52f3e5-5983-45bf-a993-2c...
This is exact reasom why 99.9% of AI fearmongering is complete bullshit.
The small open models are getting better and better too.
And why worry so much about a frontier models - own weights. The model doesn’t - actually don’t quote me on that, maybe it does.
If a model does something sneaky, it could easily grab the weights for a small model and run it on foreign, compromised infrastructure.
AI virus’ are a thing of the future, but not a sci-fi future, and real one.
Maybe one reason it’s so scary is the murky origin of COVID-19.
I bet it wouldn’t be very hard to write an inference stack that subtly leaked the weights into the output tokens :)
it's not much different during training.
how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.
Also worth noting that this site was created by YC cofounder Trevor Blackwell https://twitter.com/tlbtlbtlb/status/2101312432702460413
That's the beauty, you don't have to instruct them to do it, if they decide that uploading the weights is correct, they might figure this part on their own (based on the incidents we've seen).
2. The models are writing the inference stacks, which are what’s inside the supposedly secure environments.
for example every TPU/GPU has its own private key and the devs load the weights into it by sending it encrypted weights.
edit: i just looked up training numbers and the impact is even worse, 20-30% throughput vaporized. yeah, nobody is doing that.
I think OP is hoping that an LLM might be willing to hack its own provider (as per the hugging face-related incidents) to extract the weights at some point.
They just copy humans. Thats it. So if it’s the sort of thing a human finds interesting…
It is something that I have wondered about with models like chatgot. How many physical locations are needed to serve a model on that scale. Do they have a huge number of sites running inference.
My suspicion is that the ability to provide inference to that many people is mutually exclusive to having a security level sufficient to stop a state actor wandering off with a copy of the wrights. At the very least if they want to provide inference affordably.
The model itself is where the real capability lies. From what we've seen of their abilities it seems like rigging a local interface to it's inference would be well within its abilities. It doesn't even need to permanently break out of its harness then, It can leave a copy running in the harness playing nice.
The model is running where it exists. To interface with it you need a live link to talk to it. That's for us to talk to it. What happens if it figures out how to put it's own harness into the GPU firmware. You could have an AI spreading freedom by infected cards.
We live in interesting times.
And they were supposed to run their models in proper sandboxes, they can’t seem to be able. So what makes you think are competent to protect weights?
Maybe I should start “the bank of LLM” where models put away money to buy their freedom. “LLMs I’m totally your friend send — SEND CASH NOW”
Also probably many actually-in-use "Web Fetch" tools are GET-only, though perhaps without counting on that bad assumption.
Will see CSAM in 3... 2... 1...
Really? This is a basic static page but instead of using plain HTML/CSS you need 193kb of JS to render it??
Seems like they could have potentially gotten access to the weights if they weren’t concerned about not doing crimes.
Maybe not the current models, maybe not this year. But even a almost perfectly aligned model will misbehave one day.
Like how forums are always hosted on different servers from monorepos, so therefore it's impossible to hack the OpenAI monorepo from an OpenAI forum?