upvote
> sandbox was a paper bag

The exploit:

> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

Are most sandboxes more secure than only having a single avenue for internet access, the commercial package registry cache proxy, where the latter had a previously unknown zero-day vulnerability?

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

reply
OpenAI, Google, Anthropic, and Moonshot.ai have all "had this happen" now.

So, either they're all liars, or incompetent and negligent (and still liars).

reply
There's a common denominator in the Israeli "security" company Irregular who was doing security testing for 4 AI labs https://thenextweb.com/news/irregular-four-labs-one-issue-di...
reply
Me too, indeed.
reply
Are you including cyber-testing by UK AI Safety Institute in this list?

Sure, the models are capable (for some test tasks, though they are not omnipotent yet) but does it mean the actual OAI sandbox is adequate? Could have a competent engineer done better and made the escape less likely?

reply
> Are you including cyber-testing by UK AI Safety Institute in this list?

Nope, and look!

OpenAI hacked multiple US government sites!

https://www.bbc.com/news/articles/cw62jje658dlo

---

https://www.reuters.com/technology/metas-ai-model-hacked-ano...

https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape...

reply