The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.
I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.
Sure downside would be not learning from people using your model for coding, if we're on the cusp of huge leaps in self-improvement. But there is a reasonable case for avoiding desperate scramble, especially if other parts of the business can also create value with the compute.
Version 3.1 has plenty of room for improvement, yet they don't seem to be giving the attention it deserves or at least communicating accordingly.
It may be that they wish to slow their cadence of releases, or develop their models to focus more in a different direction, etc. No matter what the actual reasoning, they have chosen to not compete in the same race, and I cannot say I fault them.
They never gave an official answer as to why, so I'll let you draw your own conclusions.
They did not decide it wasn't worth spending the money to train.
They absolutely spent the money.
Also, look at Flash 3.5 to 3.7. Flash 3.7 is a genuinely decent Sonnet 5 class model. Flash 3.7 is quite efficient too. Also, whatever was spent training 3.5 pro is probably not wasted. However, as a strategy, when I see models like Kimi K3, Fable, Sol. If you discard "because the model sucked" what other alternatives or potential options might exist?
I thought of a quite a few and they are far more compelling and interesting to me.
(Also Gemini models tend to be pretty decent at more than just programming. Enterprise AI use is more than just software eng / programming)
If you'd told me at the end of Cloud Next 2025 that by now Google still wouldn't have a competitive offering to agentic coding offerings from Anthropic (Claude Code + Fable) or OpenAI (Codex + Sol), I wouldn't have believed you.
In our non-coding use cases where we're embedding models in our product, we're also not reaching for GCP stuff. Because Anthropic has the mindshare of our engineers and product folks, since it's what they use every day.
They mentioned that they have already started pretraining Gemini 4, which will be the full ground up rip-your-face-off-expensive training that is often discussed.
Pro models are mainly for coding agent work; it doesn't necessarily make them any money.
In the real world out there, Google and Microsoft are absolutely dominating enterprise customers.
Every single non-tech office worker I know is writing Gemini "gems" (sort of claude prompts/skills) or prompting Copilot to help drafting board meeting notes, insurance contracts updates that reflect changes in regulations, make quick loan feasibility assessments before passing them to the relevant office, presentations, etc, etc.
I'm talking insurance, banking, consultancy, manufacturing, etc, etc.
Why? Because Google and Microsoft already were in these companies, all they had to do is "oh, you also have AI now with your plans". Procurement and data compliance are the first thing businesses have to sort out. They were already sorted out.
Google doesn't need to have the best coding model or triumph in meaningless benchmarks, it only needs their models to get better and cheaper while serving them to their existing customer base.
They are playing a different game.
And Microsoft, doesn't even need to care about models at all, they can provide whatever open or closed AI with their services and have to focus on the harness in Excel or Github/Azure Copilot or whatever.
E.g. while developers in most of my clients use whatever they prefer or the company pays for, the remaining 90% uses either Google or Microsoft products.
Not a single one has incentives into venturing into OpenAI or Anthropic or Z.Ai lands because they might be better at some benchmark that is completely irrelevant to their tasks of updating powerpoints or summarizing incoming emails.