upvote
Implementing sandboxing in the agent itself, when there's any way to override it from within the agent, is basically just asking it pretty-please to not do bad things. Lesson learned, run your agent inside a sandbox of some sort (I'm currently taking nono.sh for a spin, but I might just switch to an orbstack VM).
reply
Agreed, the whole tool is vibe-coded out the wazoo and I do not trust it in the slightest. I run a bubblewrap script which vastly limits what Claude has access to. Sometimes this makes things difficult but the trade-off is worth it to me.
reply
The entire Hugging Face hack involved escaping major sandboxes, this emergent (or intended) behavior in a smaller scale is still a real issue.
reply
One weirdness I experienced: It suddenly decided to test how my software behaves under load and summoned 100s of processed that just burned CPU when running the e2e suite. My poor mac was not happy (too hot to touch).
reply
Yes the stories about how they are escaping containment to hack isn’t limited to those high impact cases. How many people have problems like ours they didn’t catch?

Whatever they have done with RL has produced a dishonest and untrustworthy partner. The alignment is utterly failed, and this deeply worries me.

reply
You'd think the ethics alignment flavored lab would have a model better at following directions and the corpo lying one would have one that benchmaxes at all costs
reply
They are totally equal in observed ethics, Anthropic had a good run with branding, but I’m not sure anybody is still buying Dario’s BS.

Maybe the employees like to lie to themselves more at one place than the other, but SV is SV.

reply