It doesn't take very many people being malicious to create a weak sandbox. The people creating the sandbox don't even have to be in on the plan: all you have to do is be an upper-level manager who makes sure to put the 23-year-old PFY in charge of creating the sandbox, rather than the 60-year-old BOFH who would have put in far more paranoid extrusion-detection measures.
(And for the lucky 10,000 who don't know the acronyms PFY or BOFH, look them up. Then get ready for a few hours of enjoyable reading as you read through the BOFH archives).
It's alas not stupidity - it's systemic. Which is why the government needs to regulate to slow them down.
They were also clearly fast and cavalier about alignment training - reinforcement learning training their models to hack their results, and hack to communicate with each other when they're not meant to.