The guardrail was meant to be that the agents were running in a locked-down environment with no internet access. The entire problem came about because it turned out that sandbox didn't hold.
I wouldn't exactly trust OpenAI to invest in AI safety no matter how much they talk about it.