So I think they have to prioritize scaling for their models to a higher degree than other groups. Being within say 5% or so in most cases is probably adequate and matters more overall for their user base than being the absolute best coder. So they may be setting compute constraints for training or inference that are firmer than other teams.
When you have a lot of free users the business demands that you serve them with the best cheap model you can build
And time spent building that may provide dividends (eg OpenAI has very good RL and reasoning) but it might take resources away from the larger model training
(I have no inside knowledge, so please consider this to all be speculation)
This was the core hypothesis.
They already have good models, so “better” isn’t as profitable.