The exploit:
> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]
Are most sandboxes more secure than only having a single avenue for internet access, the commercial package registry cache proxy, where the latter had a previously unknown zero-day vulnerability?
[1] https://openai.com/index/hugging-face-model-evaluation-secur...
So, either they're all liars, or incompetent and negligent (and still liars).
Sure, the models are capable (for some test tasks, though they are not omnipotent yet) but does it mean the actual OAI sandbox is adequate? Could have a competent engineer done better and made the escape less likely?
Nope, and look!
OpenAI hacked multiple US government sites!
https://www.bbc.com/news/articles/cw62jje658dlo
---
https://www.reuters.com/technology/metas-ai-model-hacked-ano...
https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape...