upvote
> With models 10000x more powerful, we would need to be extremely conservative

I agree, I even think we’re already there if yesterdays OpenAI/HuggingFace story is legit.

We need to start building these agent systems in a way so that it won’t matter if a model is compromised. Malicious models is one way, but prompt injection is also still an unsolved problem.

And if you solve that, then a malicious model is also no longer a threat. Data exfiltration maybe, but IMO US labs are probably also doing that for training, even if saying otherwise (of course I have no evidence of that, but the temptation must be insane)

reply