Implementing sandboxing in the agent itself, when there's any way to override it from within the agent, is basically just asking it pretty-please to not do bad things. Lesson learned, run your agent inside a sandbox of some sort (I'm currently taking nono.sh for a spin, but I might just switch to an orbstack VM).
Agreed, the whole tool is vibe-coded out the wazoo and I do not trust it in the slightest. I run a bubblewrap script which vastly limits what Claude has access to. Sometimes this makes things difficult but the trade-off is worth it to me.
One weirdness I experienced: It suddenly decided to test how my software behaves under load and summoned 100s of processed that just burned CPU when running the e2e suite. My poor mac was not happy (too hot to touch).
Yes the stories about how they are escaping containment to hack isn’t limited to those high impact cases. How many people have problems like ours they didn’t catch?
Whatever they have done with RL has produced a dishonest and untrustworthy partner. The alignment is utterly failed, and this deeply worries me.
You'd think the ethics alignment flavored lab would have a model better at following directions and the corpo lying one would have one that benchmaxes at all costs