If these companies will steal from deep pockets like Disney or Sony (some of the most infamously litigious copyright trolls to ever exist), they won't think twice of stealing every bit of code you upload to them.
If your code passes through an AI company's servers, you can assume you just gave it to them. In turn, when your competitor tries to copy that new feature you just added, the AI is now trained in exactly how to copy you and eliminate your competitive edge. Unlike your employees, the AI isn't bound by the same rules and even if it were and violated them, your company probably doesn't have enough money to prove it in court (and that's if we somehow reverse some of the stupid "AI is the most transformative use of copyright I've ever seen" judges who have drunk the coolaid).
Most companies could build the compute to run GLM or Kimi models for way less than the potential loss due to IP theft from using third-party systems.
A simple loophole, use the code to create an RLVR environment where the resultant code is the end goal / max reward. Technically the customer data is never trained upon, but effectively you’re using it. Even better, use the code as a seed to generate synthetic data similar to it and use that synthetic data as rewards in an RLVR model.
Unless you can host the ChatGPT model on your own servers, which I know some enterprises are doing, I don’t think there’s any hope of protecting your data / competitive advantage from these frontier companies. Better to be paranoid, than be commodified by these companies.
> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training. The OpenAI internal model used for this result was developed through large-scale reinforcement learning on top of a previously pretrained model. Our proofs also differ significantly. In the Euler case, Alpöge and Buckmaster proved a result with external forcing, while OpenAI’s system proved a result without external forcing.