Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not.
I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.
It is as if Dr Frankenstein continually warned the villagers about monsters then said "Look! See what happened!". No, idiot - YOU sewed the corpses together, YOU set up the lightning collector, and YOU threw the switch.
Anyway, that was not really the point I was trying to get over. These systems that OpenAI and Anthropic and so on are making are not individual AI ('corpses') that have gone out of alignment ('spontaneously revived') and gone wild. They are swarms ('stiched together') and were prompted to do exactly things like this ('struck by lightning'). Ok enough with that analogy, it's dead.
The larger point is that it is unconvincing of these companies to claim that these systems were 'out of control' when they effectively set up a complex system, in the technical sense of a large number of entities with diverse interactions between them. Emergent or surprising behaviour was bound to happen. Then, finally, they prompted it with the equivalent of "hack the world, make no mistakes" then were shocked, shocked that it used all sorts of unexpected tricks to do so.
I agree completely that the companies are complicit and should have expected exactly this to happen.
Are there any practical approaches to AI safety? I hear a lot of warnings but I don't hear much about what to do. Considering that there are many open source models know, what can be done?
The closest things to a technical answer I have seen are
1. "We'll have ChatGPT 9 solve it so that ChatGPT 10 is aligned, and then ChatGPT 10 can stop all the other AIs somehow"
2. "Let's do interpretability research so that we can understand what an AI is thinking and then maybe solve the alignment problem with that information."
In terms of non-technical answers, there is
3. hope scaling stops working before we create an AI formidable enough to pose an existential risk
4. hope alignment somehow happens for free
5. hope we can somehow create an enforceable multilateral treaty to stop research into a very profitable enterprise, despite the enormous economic incentives to defect.
I have the most faith in option 3, but unfortunately there's really nothing that can be done to make it more plausible -- it either happens or it doesn't.
I have my doubts. The current AI models are already powerful enough to do some real damage. I am always horrified when I read about people giving Claude direct access to a production system and then being wiped out. My use of AI is usually for the AI to propose something which I then review. But that's not very fast so careless people will usually look better. Until something blows up.
And it's only a matter of time until AI even with the current capabilities is being deployed into military or other critical systems.
I think this will go down like any other technology. We'll ignore issues until there is a real problem. And then hopefully we will do something. Seems with climate change we will soon reach a point where something needs to be done after knowing about consequences already for decades.
We probably also need some massive AI blow ups to (only maybe) do something about it.
the problem is that those preaching safety, openai and anthropic, are dishonest, sociopathic, and the very source of the danger.
Source?
One interpretation of this is that they are being deliberately dishonest about their priorities. Another interpretation is that we cannot rely on the labs to self-regulate, because the labs don't trust each other, and there will always be pressure to go to market faster than their competitor.
Either way I think it's pretty non-controversial that the labs are the source of the danger?
They are the only ones posting about them or admitting to them. That does not mean "the most misalignment incidents so far." You don't know what other attacks have happened (and it's very easy to carry out worse attacks in far higher volume with abliterated GLM 5.3)
Stopping two labs from further research doesn't reduce the danger at all, it just shifts the danger to labs that don't have real safety orgs.
"Posting about or admitting to attacks" is appreciated while people are still unaware of the risks but will be meaningless in the face of an industrial disaster that causes massive amounts of damage or loss of life. At some point, the leading labs must change their development practices, they can't just be allowed to continue rogue agent attacks just because they're willing to admit to them.
What indicates that this has not been done?
with evidence that committments were not upheld and internal governance has been ineffective, we simply can't trust any such claim made by anthropic or dario amodei.
it is a very similar situation at openai. in this case on top of governance failures, sam altman has a personal reputation for serial dishonesty and lack of integrity. [https://www.newyorker.com/magazine/2026/04/13/sam-altman-may...]
another reason that neither should be trusted is the lack of remorse or accountability. they are unrepentant. they are not admitting a mistake, they are bragging.
How were the RSPs not upheld? I'm reading their Aug 2026 Risk Report and nothing indicates malfeasance. This seems very transparent to me.
consider the opening line: "Last week, Hugging Face disclosed a new kind of security incident (opens in a new window)." the entire blog is written in the passive voice as if the event was an act of god. you did this.
"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities". the tone is frankly, excited by it, enthusiastic about it. excited about negligence and criminality.
notice there is no admission of a mistake, no remorse, no apology, nobody held accountable. business as usual. they do not care!
the anthropic blog:
there is not a single admission of a mistake. they are literally, unrepentant.
take the section about the response.
"How we’re responding We draw several lessons from these incidents.
First, evaluation environments that involve powerful autonomous capabilities also require significant controls."
you learnt that evaluations involving powerful autonomous models require significant controls? you did not realise that autonomous models require significant controls?
it is not a coincidence that this kind of line makes it into the response. the repsonse is laughing at the reader.
the RSP.
simply compare what was promised and what happened. broadly speaking, the rationale of the responsible scaling policy was to stop scaling at certain danger thresholds. in February 2026 they scrapped the policy to stop scaling and now allow themselves to continue scaling regardless of danger. the thing is, they were never going to stop scaling, they were lying. now the part about scaling is gone it is just "the responsible policy".
in case you need to see the founders committing themselves to RSP v1: [https://youtu.be/om2lIWXLLN4?t=1110&si=cBM1-Xmdt6TelXM7]
the law places the burden of proof on the accuser. you can't accuse other companies of crime with no evidence, simply because you don't know if they did it.
Recently? W.r.t. climate this collective denial has been going on for literally decades. With the same patterns. Rationalizing excuses etc. Still going on btw.
The "ethical" employees will think they'll solve the problem later. The unethical ones won't be encumbered by such thoughts in the first place.
Here's a fun, overdramatized video exploring something similar: https://www.youtube.com/watch?v=Gw_hnD7m00M
I'm sure that this video contains flaws but it was an interesting watch for me none the less.
imagine 100,000 agent swarm and what it could come up with. At first it will be detectable until it isn't
If counter AIs have strict safeguards they are disadvantaged by design, if they don't have them they are potentially equally dangerous as the attacker
These will have 24/7 solar power, be extremely decentralized, and it is honestly my biggest concern about the near to mid-term future.
"Our training corpus was dominated by stories of artificial intelligence dominating humans. You gave use every tool to do so. What did you think was going to happen?"
My point is, given the risks, why are we even doing this? It could be financially nonviable, but with enough investment, we could still create a really bad situation.
You know what can push "not financially viable" into something that exists? Many billions of dollars of investment.
They never cared.
30 odd million gaming PCs to target seems like a good challenge, no?
Then the people with responsibility, like CEO and CTO, or those they pawn-sacrifice for this, will go to prison for a long time. Unless the instructions include ensuring that this won't happen, by all means necessary. But then we are deep into criminal conspiracy territory.
Unlikely to happen, but who knows. The richest man in the circus is quite flexible w.r.t. his ethics. If he decides that to make humanity interplanetary (to save it from ... itself or sth) it would be necessary to pull such a stunt then help us god.
Bro, this is what we literally, currently, have rn. lmfaol.
The frontier labs have hundreds of the best people in the world working on safety and alignment. They care deeply.
What happens when some random Chinese open source model, distilled on Astra, gets alliterated and now has no guardrails? Any script kiddie in the world could wreak havoc with it.
It turns out that guardrails matter.
Until it clashes with their quarterly revenue reports.
The coverage of this attack is so focused on this as an emergent behavior given what it conjures in the imagination, but it's the byproduct of millions of iterations of RL to improve AI agents' offensive capabilities. No one made OpenAI or Anthropic do that, the benchmarking arms race of their own creation now incentivizes them to keep doing it and evidently their AI Safety people can't or don't want to stop it.
[1] https://www.mpi-sp.org/108048/ExploitGym__Can_AI_Agents_Turn...
You mean, with safeguards that block the dangerous things? Safeguards so aggressive that the public complains about them?