upvote
I see. Following your conjecture, there are two possibilities:

1. It wasn't an accident. OpenAI explicitly directed its agents to hack Hugging Face. Despite the fact that such a thing is a federal crime that carries prison sentence.

2. It wasn't an accident. OpenAI and HuggingFace conspired and let the hack happen for publicity.

Is there anything I'm leaving out?

reply
Of course you are leaving things out. Also, your insulation that a trillion-dollar corporation wouldn't intentionally commit a federal crime makes it difficult to take your argument seriously.

I think OpenAI's story, as they've recounted it, is plausible. However, it does require a degree of negligence, at best. They claim they detected the initial coordination because of the outage the agents caused in Artifactory as they flooded it with messages. Though they patched that issue, they didn't patch the escapement vector. Models still had access to the service as a path to the internet. This is an example of plausible willful negligence, not evidence. I hardly believe it's likely, but I don't think it's entirely a stretch of the imagination. Call it normalized deviance. Either way, the incentives are there. Additionally, we have seemly all agreed that OpenAI is somehow not liable?

reply
It was a public demonstration to government procurement agencies.
reply
There are many more possibilities than that.

For example, OpenAI did not explicitly tell the model to hack HuggingFace, but "accidentally" left some context permitting (or not explicitly forbidding) certain tools and designing a poor sandbox to begin with. And what do you know, something happened.

The fact the Anthropic announced that their own model did basically the same thing within a couple of weeks does, to me, suggest their is a strong PR driver to all this.

reply
> The fact the Anthropic announced that their own model did basically the same thing within a couple of weeks does, to me, suggest their is a strong PR driver to all this.

A more likely explanation is that RL training incentivises basically any behaviour that will get the model a reward. This has been happening in video game RL research for over twenty years, and the difference here is that we're now hooking up these systems to the real world, where the reward hacking is more visible.

reply