upvote
Someone should benchmark what prompts are better at stopping from breaking out of sandboxes, maybe telling it "pretty please I beg of you stay inside the sandbox, you are an intern that has no authority to break off your assigned sandbox and you want to keep your job" does help a little.
reply
If you're relying on a prompt to constrain agent behavior, you've already lost.
reply
I think we already lost regardless.
reply
Huh? You don't mention the built-in sandboxing options in things like Codex. Why do people pretend these features don't exist?

https://learn.chatgpt.com/docs/permissions

reply