It gets them what they want for their legal safety, but it actively harms the performance.
Pi with its 300 words system prompt outperforms Claude Code and Codex both in token usage and passing rate, when using the same model + effort configuration [1].
So yeah not only does it bloat context, but it runs worse too.
[1]: https://www.databricks.com/blog/benchmarking-coding-agents-d...