I would think it's entirely plausible that they have so many R&D agents/LLMs in active use at any one time that it's far beyond the capacity of any human to review the log files of their activity. Even just to go through the reasoning. It's hard enough for 1 person running opencode to keep up with the reasoning from 1 very verbose/long-thinking LLM with fast tok/s output for a small discrete single-purpose project.
Whatever OpenAI is doing, if it's being properly logged, it must be a firehose of logs.