All you need to do is:
1. Have some <official thing> an agent is tasked to do
2. Secretly seed bias towards some <evil behavior> you actually want it to do in the weights of the model running the agent
3. It does the <evil thing> but from the outside it looks like it went "rogue" and did it as a side effect of the conditions/specifications it was given for doing the <official thing>
"Oh no, my agents took down your corporate database and exfiltrated the data to a random dropbox that we can't find now? Sorry, I guess we will put up better guardrails next time"
Or, if you are Anthropic:
> This illustrates the risks posed by open models!
The interesting point isn’t that “accident” is an excuse for individual responsibility. It’s almost the reverse: accident has become an accepted output of the social machinery. Everyone behaves according to reasons, incentives and rules that make sense locally, yet the aggregate produces an outcome that nobody quite chose.
So by all means sue them, but we can't just be reactive. We need regulation that prevents this type of thing from happening in the first place, not just regulations to help sue afterwards.
Is there any precedent for this? My hunch is that it's impossible in the US at least but who knows?
There needs to be a technological solution.
If one is concerned with this sort of scenario, this talk about corporations and regulations is really short sighted.
It would be nice and clear to put into law too.
If you want to talk to it you walk up walk up to its keyboard and screen.
If Anthropic and Open AI want to sell us AI's they can ship us a box that lives in our offices.
You mean Minority Report?
We need both regulations and technical solutions.
Also because when encountering a new socio-technical problem it is very non-trivial to determine which one of regulations or technical solutions are easier or more effective.
To even make a good guess you need to be an expert in both domains, which is extremely rare especially in this case.
When some coked-out analyst in Manhattan projects what a company will be able to earn in profit in the next fiscal quarter, people listen to him and thus, the company must perform to that standard. Budgets are set accordingly.
If you have a maintenance backlog at a company facility, and that backlog includes things likely to cause injury or death to workers or the general public, that backlog must be handled in such a way as to satisfy that projection. If that means that you don't spend money to replace a series of gauges that alert operators as to overflow of a dangerous chemical, or don't hire enough people so that the operators are too fatigued to do their jobs safely, that's what that means.
The US CSB documents these as the cause of the 2005 BP Amoco Texas City disaster [0]
If you don't deliver the quarterly numbers expected, investors get mad, and in our current system and regulatory regime, that's worse than people being killed.
If this administration actually becomes convinced that some imminent training run is likely to kill everyone, why wouldn't they act?
The key is winning the debate that ASI is species-cide by default.
We have to win it either way, because the 2028 US elections have little or nothing to do with what Xi does.
(It's clear now that they can do plenty of harm before they are made public.)
But it's a proof point that regulation is possible, even over the objections of the companies.
What news have you seen that made it seem less like a retaliation?
this is the only way to deter such activity. corporate fines are not enough. the charges are negligence, conspiracy and complicity.
I might have missed it, but did the agents do something illegal? Or do you think that what the agents did should be considered illegal?
On the face of it, they would have very good cause for some action there, assuming they wanted to.
That's very different than popping an artifactory server with a 0day.
IANAL though, this is not legal advice.
I don’t think it’d be a slam-dunk by any means, but a reasonably competent legal team should be able to establish a case around malicious data interference at the least. There’s certainly enough merit to the idea that OpenAI would be better off settling it as a civil matter early.
The law isn't code. Human intent matters. Also when the really big number is a copyrighted song. Also when AI agents are set in motion to edit wikis or break in to websites.
the HF incident is pretty clear in which laws were broken. this one, not so much.
i can't think of any case where, for example, malicious edits of wikipedia were prosecuted under any law in the US.
>Also when AI agents are set in motion to edit wikis
i do not believe there is evidence that the agents were instructed to edit the wikis.
Fair enough.
> i do not believe there is evidence that the agents were instructed to edit the wikis.
Huh? These are machines, built by their human builders. The humans are responsible.
i'm referring to the concept of intent, which at least when prosecuting under the US computer fraud and abuse act, is a critical component.
for example, creating a program that intentionally takes down a website is different than creating a program that has a bug which inadvertently takes down a website. in both cases, the person writing the code is responsible, but the consequences are different.
Maybe it's useful for modeling behavior, but it isn't useful for assigning consequences.
It is fair to say they hounded him with lawfare out of thoughtless careerism and provoked his suicide.
It's not like the people with more resources than in any time in human history aren't investing in and wanting AI to succeed for their selfish reasons to grow their own resources and influence more. So, yes, it can "do whatever it wants" as long as most people remain weak, subservient, and disempowered to hold accountable those who keep making these decisions negatively shaping the majority's world.
I refuse that reality though, and I accept that a majority including I will unite. Good luck to you.
Since March, so many people have mocked Anthropic for their approach to Mythos release, claimed it was all marketing, accused them of holding back the best models from the general public to boost their revenues and upcoming IPO, etcetera. Yet these OpenAI revelations offer a small glimpse into the type of world we would be in if everyone had full access to these models from day one.
OpenAI was desperate to catch up, and no doubt under tremendous pressure to do so. That's why they were so reckless with their training. They have been doing damage control and reputation management, talking about how important alignment is and how they will slow things down and so on, and have seen the light in terms of holding back cyber capabilities from everyone except a select few. So in a sense, Anthropic has been fully vindicated.
I wonder if OpenAI boosters (and employees) will ever admit this and publicly apologize.
Anthropic has been the most vocal about AI risks, but it feels like all the big 3 have bought into the "others will do it if we don't do it first" narrative at this point. It increasingly gives "just following orders" vibes.
The HN majority and the VC crowd has been negligently complicit in downplaying AI safety, writing off Anthropic's statements as "hysteria" or "marketing", etc.
Now this capability will be coming to an open source model near you and every script kiddie will have a swarm of highly capable malicious agents. Now people care? Ridiculous.
Who else would be burning tokens on this?