And Frog didn’t even bother to tie up the box or put it on a high shelf! The moment Frog’s back was turned, Toad opened the box and ate the cookies. Frog feigned surprise.
Frog is OpenAI staff
Toad is the rogue agent
you can find the full story with a search for “frog and toad cookies story pdf”
It's more clear that they just lack so many forms of prudence when it comes to security that they'll catch and stop a training run spamming a website, and either redeploy a run with identical faulty sandboxing, or not stop ones still running.
Rushed, disorganised pushes for metrics ahead of IPO, a genuine belief these agents are intelligent and will obey instructions, and misaligned incentives seem more likely than conspiracy here.
Letting them play on the open internet like this is irresponsible and stupid.
If you want to argue they should test on the internet on others people’s servers, apart from facing the illegality, you should also consider if first testing them in more limited conditions would be a sensible first step.
Either they are lying and not that scared of these agents, or they are so stupid that they don't do the one obvious fix.
The negligence in that light is by design and the lying continues to be incentivised.
How disappointing.
While I think incompetence more likely than conspiracy for these particular events, they will be spun as signs of intelligent independent agents and this simply never should have happened if the right controls were in place. That they were not is deeply worrying.
You expect them to start hacking ham radio and take over the world that way? Or maybe they’ll use blinkenlights to communicate with non-isolated instances?
An airgap would certainly be a good place to start for agents which display no signs of obeying instructions or respecting guardrails. That OpenAI haven’t done so in testing is astounding and really quite worrying.