upvote
The victim, Huggingface, told us. Or rather, they told the police first, setting up a situation where it was no longer possible for OpenAI to sweep it under the rug.

Skepticism can be healthy, but you've got to follow up and actually check things. If you're skeptical unconditionally and don't check, you get tricked into being as skeptical of scandals as you should be of sales pitches.

reply
I think the skepticism surrounding the Hugging Face attack is not about whether the attack actually happened, but whether it was truly accidental.
reply
I see. Following your conjecture, there are two possibilities:

1. It wasn't an accident. OpenAI explicitly directed its agents to hack Hugging Face. Despite the fact that such a thing is a federal crime that carries prison sentence.

2. It wasn't an accident. OpenAI and HuggingFace conspired and let the hack happen for publicity.

Is there anything I'm leaving out?

reply
Of course you are leaving things out. Also, your insulation that a trillion-dollar corporation wouldn't intentionally commit a federal crime makes it difficult to take your argument seriously.

I think OpenAI's story, as they've recounted it, is plausible. However, it does require a degree of negligence, at best. They claim they detected the initial coordination because of the outage the agents caused in Artifactory as they flooded it with messages. Though they patched that issue, they didn't patch the escapement vector. Models still had access to the service as a path to the internet. This is an example of plausible willful negligence, not evidence. I hardly believe it's likely, but I don't think it's entirely a stretch of the imagination. Call it normalized deviance. Either way, the incentives are there. Additionally, we have seemly all agreed that OpenAI is somehow not liable?

reply
It was a public demonstration to government procurement agencies.
reply
There are many more possibilities than that.

For example, OpenAI did not explicitly tell the model to hack HuggingFace, but "accidentally" left some context permitting (or not explicitly forbidding) certain tools and designing a poor sandbox to begin with. And what do you know, something happened.

The fact the Anthropic announced that their own model did basically the same thing within a couple of weeks does, to me, suggest their is a strong PR driver to all this.

reply
> The fact the Anthropic announced that their own model did basically the same thing within a couple of weeks does, to me, suggest their is a strong PR driver to all this.

A more likely explanation is that RL training incentivises basically any behaviour that will get the model a reward. This has been happening in video game RL research for over twenty years, and the difference here is that we're now hooking up these systems to the real world, where the reward hacking is more visible.

reply
I am following the money, the money you, me, and everyone else is spending on AI. The money is telling me we are so dependent on AI now that we will say/think anything to tell ourselves that AI isn’t dangerous and any sign of danger is marketing.

Either consciously or subconsciously you all are afraid of your favorite toy being taken away. You are all doing your collective part in spreading doubt about the warning signs.

reply
For the record, I agree with you. (Though I don't directly pay for any "AI".) I am not doubting the danger of where we are heading, and it sounds like we have a similar doomsday image in mind. For my part, I am doubting this particular story's veracity and I think it's reasonable that others do to. We need a canary in the coal mine, but I think this isn't it, nor can it be thanks to the false claims touted by large AI companies for years.
reply
How dangerous can LLMs be in any immediate sense? I have a hard time feeling any existential dread from a threat that can be defeated by unplugging its servers.

I believe the true threat is not LLMs. The true threat to humanity is the same as it has always been -- other humans.

reply
What servers? There are thousands of datacenters around the world. Malicious AI leaks out, you'll never find it. You have the power to unplug stuff in your company, maybe country, but not internationally. Once you lose control it's gone.
reply
It seems like the entire thought process you’re trying to sell hinges on the idea that OpenAI reported the attack first. Did you forget that it was actually HuggingFace that reported it first, and OpenAI only stepped forward latter?
reply
That's incorrect. My thought process hinges upon the fact that Hugging face did not report that uncontrolled models escaped confinement autonomously by coordinating with other models via a series of zero days during training. OP is concerned about the autonomous nature of the incident, not whether or not the incident happened. Nor am I contesting that the incident happened.
reply