upvote
CSAM, and other harms, are typically detected using a set of specially trained, faster and cheaper models (and out of band matching techniques) that run before and after the main model.

Any mention in the system prompt is mostly defense in depth, and to make refusals more graceful.

reply
Also, the system prompt, or even something reinforced on every message, is nowhere near as strong as its internal training or as an external safeguard.

If the prompt were the only protection, it would be extremely easy to produce illegal content after a long session.

reply
I don’t think so. If you start a new Claude Code session without a system prompt, it doesn’t even know what model it is and hallucinates being some old variant of Sonnet.
reply
How do you start a session without a system prompt if you use ACP in Zed for example?
reply
The system prompt is (and cannot be) the only guardrail against things like that, because any system prompt is little more than a good suggestion.
reply
I wouldn't put auch limitations in the system prompt. A mix of fine-tuning and out-of-band detection appears to be a better fit.
reply
at least according to their documentation they do not

afaiu they have other systems for denying and re-routing requests

reply
They use non-LLM gates for this.

Otherwise DANmode and similar jailbreaks would still be as easily accessible as they were at the beginning.

reply