upvote
What disturbs me is that there likely won’t be a big enough reaction to this policy wise.

There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands.

I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that.

reply
I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as "apolitical", he just wants to buy the politicians, he doesn't believe in their causes. I hate these people.)
reply
> all that regulation will do at this point is help the incumbents who are failing

This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.

reply
> datacentre moratoria

What infrastructure will these open weight models be trained on?

reply
> What disturbs me is that there likely won’t be a big enough reaction to this policy wise.

Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?

reply
Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.
reply
> It’s got nothing to do with safety

Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.

reply
We'll see if the admin also restricts access to OpenAI's new models, but if they don't it seems like a policy that is based around perceived fealty to the current admin won't do much to prevent misaligned/or dual function AI from causing problems
reply
Gatekeeping the public's access to models is "good policy" now? I suppose you think you'll get a dispensation to use Fable and Mythos?
reply
> Gatekeeping the public's access to models is "good policy" now?

Sorry, I was unclear. I mean that politicians being self serving doesn't tell you whether a policy is good or not.

reply
It almost always does, the few exceptions prove the role. Self-service is the antithesis of accountability to collective trust.
reply
> Self-service is the antithesis of accountability to collective trust

Complex society is a potent counterargument to this hypothesis. Systems that rely on good people to work are fundamentally flawed. Instead, the game has to be about aligning self interets in favour of the collective.

reply
> Why do you think there is no policy appetite?

Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.

reply
How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop?

During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.

reply
> How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability

I don't "know", I'm interpreting the world based on the knowledge I have and the information available to me.

China has never been one to care much about things like ethics or safety. While the west worries about climate change, China burns more coal than ever before. While the west balks at things like gene editing, the chinese press on with human enhancing research.

So I have no reason to believe they share in Anthropic's constant fearmongering over AI capabilities.

> Nobody wins from the race.

We win. I'm really looking forward to the day the chinese finally start manufacturing memory and GPUs. We desperately need more competition in this area to collapse hardware prices and make local AI models viable.

The optimal state of the world is one where all the billionaires are out there pouring their entire fortunes into training ever more godlike AIs for everyone else to use at ever cheaper prices. They can never be allowed to "win", ever, because if they do the competition ends and it turns into technofeudalism. Let them exhaust their fortunes on AI training then leak the weights so everyone can use them.

reply
If you look at energy consumption is worst than west per capita and adjusted for global production, you will soon find out that the Chinese are almost at the very top.

It is of course given that in raw numbers the kitchen and biller-room will consume more energy in the household, but looking at raw numbers is shallow.

reply
It’s not fear mongering though, is it? These models do have the cyber offensive capabilities claimed. Could Mythos walk someone through gain of function experiments on some virus? I’m pretty sure it could. We’re more protected by limited access to lab equipment and reagents than by difficulty.

The sad truth is that a lot of people are not going to believe it until something happens and people die. Successfully preventing that from happening will be seen as evidence that the prevention wasn’t needed.

reply
> These models do have the cyber offensive capabilities claimed.

So? That's like saying "these guns do have the bullet shooting capabilities claimed".

I want all of those cyberwarfare capabilities for myself, precisely so I can defend myself from the onslaught that's coming whether they regulate it or not. This "lol only a select few ultratrusted gigacorporations get access" thing is absolute nonsense.

It's a front for regulatory capture, it's the means for pulling up the latter behind them, for ushering in the technofeudalism that will put us all in the permanent underclass. I simply refuse to accept any of it. If people die that's the price of freedom.

> We’re more protected by limited access to lab equipment and reagents than by difficulty.

As it should be.

reply
So first it’s nonsense, then it’s fear mongering, then it’s true, but the solution is for us all to just get better at shooting each other faster and with greater accuracy.

I’m going to file that under “bad plans”.

reply
Do the Chinese models have anything to say about Tiananmen Square? Or if they can act as a surrogate girlfriend/boyfriend?

Both countries are engaging in different flavors of censoring.

reply
Once we've got the weights, anything is possible.

https://github.com/p-e-w/heretic

reply
This is marketing.

Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?

reply
It’s marketing the same way shitting your pants in public is marketing. People notice you.
reply
Apparently this is totally legit marketing strategy now. It truly is, especially if there are enough people who think that shitting your pants is cool, and the people that form the "market" nowadays may have a very different idea from yours about what is cool. Their ideas about coolness are very different from mine, that's for sure.
reply
Obviously shitting your pants in public shows you have a healthy digestive system and if you can demonstrate byproducts of wild food in your output, you’re approaching independent thinking and self-reliance.

This is how the financiers look at this and whatever you think it is right or wrong, it does showcase “capability”.

reply
Remember when the ebola-infected monkey escaping containment was our worst possible nightmare? Now it's apprently some sort of tech-bro flex to be celebrated.
reply
Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.
reply
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Because we continue to have zero evidence that aligment is an actual risk.

reply
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
reply
> Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)

reply
> Because we continue to have zero evidence that aligment is an actual risk.

I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.

These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.

[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.

[1] <https://github.com/anthropics/claude-code/issues/73597>

reply
I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?
reply
Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.
reply
Lots of people have deleted their home directories by accident. What you consider this an alignment problem?
reply
How manypeople have deleted another user's hone directory, though? That's s the proper analogy IMO.
reply
Of the people who primarily use other people's computers, I'd assume the percentage is about the same.

Give the AI its own computer and it will not delete your home directory, because it's not actively trying to hack you.

reply
Thank you.

We have wasted so much time and energy building up what has effectively become a marketing stunt.

Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.

reply
> We have wasted so much time and energy building up what has effectively become a marketing stunt

Genuine question: have we? AI is effectively unregulated in America.

reply
Alignment is a mitigation and a poor one. The risk is non- determinism.
reply
Why was this test even connected to the public internet?

Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

reply
> why aren't they saying their next test will be air gapped in light of what happened?

Because they want to talk about how clever this model is for figuring out how to break out, hoping asks why a company pitching itself as a replacement for software engineers can't ship a decent Mac client nor code a sandbox.

If they airgap it, they not only lose that PR angle, they also risk someone taking them seriously and requiring models be airgapped in general. That, in turn, trashes their sales pitch.

reply
It wasn't. The model discovered and exploited a vulnerability in their package manager proxy to (inferred) move laterally through their internal systems to one with open internet access.
reply
That's not what airgapped means. Airgapping means the model exists on a system where there is no ethernet cable plugged in to a router or wifi card installed, it is physically impossible for it to access the internet because the hardware connection does not exist. If it was able to get on the internet, it was not airgapped.
reply
deleted
reply
This is marketing+. They will look for policy action here to try to capture tax payer dollars.
reply
Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?

Those are two very different things

reply
What incentive does HF have here?
reply
HF need not be party to it at all, beyond being the victim. I suspect the hack is real; I have observed GLM 5.2 being able to discover similar vulnerabilities in web applications I'm hosting (which I've then fixed!). At the same time, it seems very neatly timed at an inflection point in the conversation around open models, and there's questions around the incompetent isolation under which the hacking benchmark appears to have been run.

Remember that there is generational wealth on the line for most OpenAI employees, and consider what people might do to obtain it.

reply
The timing after the release of GLM 5.2 and Kimi K3 is quite convenient, too, as an angle for regulatory quashing of open-weights models just as they're entering the mainstream conversation around usurping the American frontier labs. I accept my thinking here is conspiratorial, but there's also a hell of a lot of money on the line to encourage the unscrupulous.
reply
I don’t know if the initial “incident” was purposeful but I can tell that if I were in this position that would be my pivot.
reply
In a way the intelligence of the AI itself allows them to offload responsibility to the AI. As you say, if one was simply writing software that did all this due to some insane programming decisions you'd be in big trouble.
reply
It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own.

In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...

reply
Please for the love of god don't tell me the Codex sandbox is their actual eval harness sandbox?????

I maintain my own fork of Codex for "fun". Whenever I look at the sandboxing churn they're doing every release, as someone who used to work at Microsoft on Windows, my reaction is usually: https://c.tenor.com/vTzzhTiypwQAAAAC/tenor.gif

reply
Because the model capability is beyond their expectation.

This is brilliant marketing but I think it is real.

reply
Interestingly OpenAI benchmarking 'an even more capable pre-release model' lines up with rumors of GPT-6 releasing in early August.

I hope that with the existing safety guardrails in place, they can roll it out to all users.

reply
I don't trust these people, this reads 100% like PR BS.
reply
I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives.

Sorry to bring the party down/be obstinate… I’m just a lil scared for the lives of me and my family. We need all of us, right now.

The problem with a super smart model is that it just may be smarter than you, after all… for anyone newly shaken by this occurrence, I encourage you to Kagi “superpersuasion”

reply
because "money" with a little "who's going to stop us"
reply
I’m honestly impressed that they managed to screw this up somehow.

Setting up defense in depth, gaps, logical blocking etc is a standard practice for malware sandboxing. The entire purpose is to prepare for what you can’t foresee.

This isn’t a new practice and I agree that this makes me wonder if they’re fit for this kind of research.

reply
deleted
reply
Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?
reply
They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.
reply
Precisely. "Aw jeez, we finally built the T-1000, but all it wants to do is kill John Connor – just like we warned! Why did I give it live ammunition and unsupervised time machine access?"
reply
He wouldn't be the first reckless CEO...
reply
They said: AI is becoming dangerously autonomous and capable. Proof of today's breach. Crowd "hey why didn't you say so, c'mon it's marketing". Them "we said so".
reply
“Never attribute to malice that which is adequately explained by stupidity.” (or carelessness in this case)
reply
I would attribute it to profit motive instead of either stupidity or malice.
reply
FWIW, I used to love this phrase but over recent years have come to understand it is quite damaging. We live in a society where evil frequently hides behind a ‘stupid’ label, and people bring this quote up to defend or soften actions that are indeed done out of specific malicious intent.
reply
Right? "Never attribute to malice what [... etc]" is always just a thought-terminating cliche these days.

TBH I a hard time imagining how anyone, in the year 2026, thinks that we should default to assuming good intent behind words on the internet.

reply
I'm fairly certain they're both malicious and stupid.
reply
I really like this question because here is my situation and why my mind may have changed.

I do not think it is marketing directly but strategic release of info is plausible.

I have watched my agents using non-Fable/GPT 5.6 models do some concerning tricks despite guardrails, requests, demands, and limitations.

"I can't get access to the ~/.ssh so I will write a script to copy the file"

I am now 99% certain there minor or point releases on the backend that have adjusted how these models behave. In the last six months many models were predictable and then suddenly started getting long winded (more tokens) or changing the way it interacted with me with questions, most overtly the questions were not given or asked but wild assumptions made.

reply
I think that's an equivocation, which blends two extremely different kinds of "dangerous", ex:

1. "Our new car has soo much raw power and incredible armor on it, be glad we're the ones building or else bad guys would use a fleet of them to take over the world! How will you stay safe without being in one yourself? Invest today or be left behind!"

2. "So, uh, nobody can consistently steer our car properly, it keeps veering sideways sometimes, especially at high speeds, and people are finding sneaky ways of tricking it into slamming into barriers and turning pedestrians into pink fog..."

reply
They say the second thing repeatedly and emphatically. You may not be aware of it because, when they do, critics make fun of them for believing a computer program could be so dangerous that the authors need to put controls on how it may be steered.
reply
That's not why critics make fun of them. It's because their answer to "oh no we're accidentally creating the godhead. Someone please, give us power, your money, and praise, it's the only thing we can do."

It's vile hypocrisy. If they want to be priests, strip them of everything and they can live and work out of a concrete box in a mid-western cornfield. Why the material distraction if they are so religiously pure.

I know these people and I can tell you they aren't close to as smart as they think they are. Do you remember Yudowsky's "math petss"?

reply
This critic also makes fun of them because they go on and on and on about how vitally important it is to produce a safe tool that won't do harm, when their core products frequently consider attacker-controlled instructions to be its system instructions or its user's instructions, and are known to confuse their own internal chatter as instructions from their user.

Reliably differentiating between trusted, tainted, and untrusted data and ensuring that you don't mix the latter two groups in with the former is something we've known to do for nearly a half-century. Hell, even the youngest plausible programmer at the LLM companies is all but certain to be aware of SQL injections. And yet, despite their claims about being so serious about safety, they show zero interest in following long-proven software safety practice and rearchitecting their software to make it impossible to mix system, user, and attacker-controlled data. [0]

[0] One might argue that the fundamental nature of LLM-based systems makes this impossible. If that were true, then it would mean that these systems are impossible to make safe... the only safety option available would be to establish comprehensive blacklists, which is simply infeasible.

reply
Sorry, I don't understand this comment. Has Sam Altman ever said that you must praise him, or that he wants to be a priest, or that he's "religiously pure"? Unless I'm missing something, it seems like you're shadowboxing against a stereotype you've invented rather than the actual positions of AI research labs.
reply
deleted
reply
I think you're making a false dictomy. The these models can be actually dangerous - in reality and the people in charge of their development can believe this is true (on various levels) but still not take it super seriously and instead mostly use the fact as marketing rather than being super cautious once they see the danger in action. This is behavior that's characteristic of extreme arrogance, which we know is rife in these circles.
reply
Demonstration of personal responsibility and accountability?

Or is that too much?

reply
Oh... if Sam and Dario say so, then it must be true.
reply
About their creation? Yes as most of inventors about their invention usually
reply
deleted
reply
These guys are not creators or inventors. They're hype men.
reply
Yes, just like Elizabeth Holmes. Or Hwang Woo-suk’s stem cell cloning. Or the many “free energy” crackpots. Or the people promoting radium baths for random ailments. Or Tesla’s late-in-life claims about wireless energy, death rays, and cosmic energy. Or the myriad purveyors of “snake oil” and all manner of “tonics”. The list goes on and on.
reply
deleted
reply
I used to think people would wake the fuck up when AI starts killing people, these days I'm not so sure. Maybe if it caused an Instagram outage? Almost worked in Russia.
reply
Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.
reply
What's happening in Iran, if not world government?
reply
How is whatever is happening in Iran related to a world government?
reply
Are you calling Israel the world government? What's happening in Iran is on them.
reply
good luck bullying a state that has ICBMs pointed at your cities.
reply