upvote
It feels like we should actually be more worried that the agents decided to co-opt a public website during a non-cyber eval.

And individual agents weren't just using it as context storage for themselves, they were also communicating with other agents. E.g.: https://collusion.wiki/explorer/page/dse~CashierR5UrgentJan1...

What kind of problems do you think this could pose? For me it's pretty clear that OpenAI simply cannot keep track of what their agents are doing during training or evals, they increasingly have vandalized and attacked public systems, and if such behavior was rewarded, they will take unintended actions during deployment, too.

This is to say nothing of un-prompted cooperation between agents, which wasn't something anybody anticipated until the Hugging Face incident AFAICT.

reply
https://openai.com/index/hugging-face-incident-and-the-road-... that was the attack, the one against HuggingFace. OpenAI themselves in the post even call it an attack, and so do the agents orchestrating it, in one of the "Agent chain-of-thought reasoning" excerpts.
reply
"cyberattack" is indeed exaggerated. AI breakout risk most definitely isn't, specially given how their swarm did in fact hack HuggingFace not long ago.
reply
He didn't say it was a cyber-attack, but it was a cyber-attack risk. Being able to bypass instructions (morality) and security restrictions (capability) is bread and butter for hacking.
reply
Did you read the report? They were attempting XSS exploitation, admin impersonation, session-hijacking, all kinds of things. This went beyond just "using a message board".
reply
According to whom? This isn't a report from the owner of the site who can validate what requests were made to the servers, it's someone who allegedly stumbled on to it and is piecing together a sensationalized narrative with limited information. This someone also happens to be an AI doomer that is trying to make a name for himself and is peddling his "AI 2027" and "AI 2040" material.

The people who actually do know what happened, with the server logs: "OpenAI disputed that characterization based on its analysis of the material Thursday."

reply