upvote
There was no rogue agent. That’s the whole point
reply
there is a rogue agent - openai and the whole management chain from researcher to sama.

theres no separate agent, which is the point. the program might look like it, but that is an illusion of the interface. the llm produces text, and the harness executes commands based on text, based on what the human researcher included as things that can be executed

reply
What label do you prefer for the agent that did something it was not asked to do?
reply
Unreliable computer program.
reply
Most unreliable computer programs won't launch research programs consisting of thousands of pages of text to find creative ways around obstacles.
reply
Misconfiguration. We deal with lots of applications every day that can do terrible things if you get the config slightly wrong.

say.. https://www.investor.gov/introduction-investing/investing-ba...

reply
How specifically did misconfiguration lead to the HuggingFace attack? You could argue that its sandbox was misconfigured, sure. But suppose you had a similar incident where its intended task required access to the internet, and it veered off course in a similar manner. I don't think "misconfiguration" would be an accurate description of what went wrong in that hypothetical.

The doomers already have a term which fits pretty well: "AI misalignment".

reply
We indeed lack much of the vocabulary. From a practical perspective, however dangerous the creation, if you cant punish the creation for what it does it leaves only the one who started the process. If it's human error or intentional neglect for personal gain should be for the court to decide.
reply
>if you cant punish the creation for what it does it leaves only the one who started the process

Agreed, but I think we can do more on the prevention side as well. Traditional liability law is for negligence in case of preventable disasters. Since we currently have no way to prevent AI disasters in principle (alignment problem remains unsolved), I think we should just stop developing the technology for now: https://pauseai.info/

reply
bot. and we even have a word for program not behaving the way the way it was intended.
reply
"Bot" doesn't carry any implication of unintended behavior. You could call it a "buggy" bot, but these aren't ordinary software bugs.

There's no simple bugfix which will address AI misalignment. It's essentially been an open research problem for upwards of a decade.

reply
Fuzzer, then. It implies random behavior, which isn't unintended like you suggest. The agent's/bots/fuzzers have certain capabilities, so it's on their operator to make sure they don't do things they shouldn't
reply
>It implies random behavior, which isn't unintended like you suggest.

The HuggingFace attack was not "random" behavior. It was goal-directed but misaligned behavior.

This isn't necessarily a simple matter of the operator making sure they behave. AI alignment has been considered to be a difficult problem for over a decade -- and remains unsolved in general, as these recent incidents illustrate.

"Fuzzer" already has an existing meaning in CS anyway: https://en.wikipedia.org/wiki/Fuzzing

reply
I'm curious, cant you just count the number of times a program interacts with a domain? My website sometimes sends out emails, makes api requests etc There is a limit on those and a point where I start investigating wtf is going on.

If you merely put 10 LLM's on the outbound traffic log non of them are going to report something strange going on? I'm not buying it.

reply
This type of whack-a-mole approach is akin to "fixing a bug" by hardcoding a special code path for known-buggy inputs. It doesn't address the root problem of AI misalignment, and doesn't allow you to prevent catastrophes in advance, only patch things up after the fact.

This might be helpful reading: https://www.lesswrong.com/w/nearest-unblocked-strategy

As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: https://www.youtube.com/watch?v=7wy3xyoXYt8

reply
Person: "AI, please make me paperclips."

AI: "OK, I've now converted the entire planet into paperclips."

Alien observer #1: "Wow, that was a rogue AI!"

Alien observer #2: "False. We need to place the blame where it belongs, on the person who requested the paperclips."

Ultimately this type of terminology dispute has a tendency to miss the point.

reply
It does, indeed. Because OAI is not just a singular person, as in your scenario. No single person has access to controlling agents at the scale OAI has. Let's not conflate Frontier providers with "Person".
reply
I'm not exactly sure why you think this distinction is so important. I think my point stands if you replace "Person" with "OpenAI". In any case, I presume the swarms OpenAI has been researching will be available to the general public before too long.
reply
It makes a big difference: individuals do not have the capabilities to run millions of dollars of opportunistic hacking loop inference. That's why the distinction is important, they are not the same thing you've conflated them down to.
reply
"AI has gotten cheaper more quickly than any other transformative technology in history. The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. That price drop is four times faster than DNA sequencing, six times faster than compute, 18 times faster than lithium batteries, and (in the century up to 1973) 54 times faster than electricity."

https://epoch.ai/publications/the-plunging-price-of-thought

reply
That legal fiction works both ways.
reply
That's super cute. Care to share beyond a smarmy comment that has no depth?
reply
You could have asked for that more gracefully.

https://en.wikipedia.org/wiki/Corporate_personhood

reply
Check the mirror. And, with that link I think you've missed the point entirely, a tad too literal of an interpretation. But thanks for trying.
reply