upvote
The defense has to work 100%, the offense just needs once.
reply
We'd basically need frontier models to be superhuman hackers before this would be a risk. Do we have any evidence of this? Are they gaining access to systems they shouldn't have access to?

Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.

reply
Ok, open AI had terrible sandboxing... what about huggingface?
reply
Distillation is a thing.
reply