upvote
In terms of open models, Gemma 4 beats the pants off everything else to the point that paying for APIs becomes hard to justify. Qwen has the meme-share for coding, but it feels much less well rounded. I have no doubt that Google have both the infrastructure and the expertise to curb stomp everyone else, should they resolve in earnest to do so.

Lest we forget, "Attention is All You Need" came from Google.

reply
> "Attention is All You Need" came from Google

It also came directly from the university of Toronto, and the university of Toronto seeded all American frontier labs (including Grok (why do you think they could start so fast))

reply
Interesting, glad to hear. We have gemma4 at work, and I was considering localhosting qwen, but gemma4 is so far behind the Opus and Fable I have at home that I've decided to hold off for another model release.
reply
"We have a company provided Toyota at work but it is so far behind the Ferrari I rent at home that I've decided to hold off for another model release."
reply
How long until Gemma 5 hits?
reply
Are you suggesting Gemma beats GLM 5.2?
reply
At 20x the parameter count I should hope GLM beats Gemma! But is it 20x better? Expertise is demonstrated, not by making big models, but by making small ones. Bigger isn't better if you can't run it at all.
reply
It's rumored that Gemini 3.5 flash has a >50% margin, and I'd imagine 3.6 flash is even higher.

I do not think OpenAI or Anthropic are actively chasing margins - though, Anthropic is supposed to be profitable on some form of non-GAAP accounting...

I suspect Google isn't really interested in seeing how far it can get dragged into a race of selling dollars for $0.25, and is more interested to see if it can stay in the race selling $0.50 for a dollar - when everyone else is losing or barely breaking even.

reply
It kind of doesn't make sense though, because typically a large org like Google can afford to crush competitors on pricing. They could probably even go toe to toe with chinese model pricing for years without feeling it.

Maybe they don't want to price war with the other labs so they can comfortably maintain healthy margins on selling them compute?

reply
That "TPU advantage" might be slowing Google down (though likely not as much as their internal bureaucracy).

Porting CUDA-based research, debugging, and overall experimentation speed is likely slower.

The GPU is still king for training.

reply
lmao, you know all Anthropic models are trained on TPU right?
reply
They basically don't exist in the currently most profitable LLM market (coding).

Yes, subs like codex are heavily subsidized. But API billing has massive margins and that's what enterprises pay.

reply
Does it have "massive" margins? Afaik no one has said publicly what margins there are on an API call?
reply
"As of October [2025], OpenAI's compute margins reached 70%, up from 52% at the end of 2024 and double the rate in January 2024, [The Information] said, citing a person familiar with the figures."

https://www.bloomberg.com/news/articles/2025-12-21/openai-se...

As for Anthropic, the rumors I remember seeing for their API margins were more like 85-90%, but I don't have a reference at hand for those. But once you know the API is wildly profitable and the subscriptions are roughly break-even and not even a big slice of their income, all of the investment makes a lot more sense.

reply