upvote
> For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.

For smaller models, you're competing with DeepSeek V4 Flash. (Which I think is a 284B A13B?) Subjectively, this feels about as smart as Sonnet 4.5, give or take. And it costs $0.09/$0.18 on Open Router, compared to $1/$5 for the latest Claude Haiku. See https://openrouter.ai/deepseek/deepseek-v4-flash#providers The developer antirez of Redis fame uses this as a local coding model.

DeepSeek did some extremely clever research on hybrid attention to get the prices that low, reducing per-user context cache sizes dramatically.

So, no, when it comes to low-price models, the US models probably can't sustain their current margins there, either.

reply
The Chinese have been right behind OpenAI and Anthropic for ~18 months now.

DeepSeek didn't do to OpenAI and Anthropic what nearly everybody claimed they would.

Every single person on HN that loudly proclaimed the end was nigh for GPT & Co. due to DeepSeek, was wrong. They were humiliatingly wrong, and they'll never own up to it. The reason those people were so very wrong, is the same exact reason the Kimi crowd is wrong now. And it's very obvious that they're wrong, but they have intense emotional blinders on. Their thinking process is hyper emotionalism: they want a certain outcome, regardless of if reality aligns to that or not. They're making emotional wishes about how they want things to turn out, and pretending those magic wishes are grounded in reason.

It takes enormous resources to run something equivalent to GPT 5.6 or Fable. Nobody can or wants to do that outside of very limited situations - if you can just reasonably pay as you go instead. As it turns out, you can just pay as you go with GPT and Fable. Their businesses have gotten radically larger since DeepSeek launched. Get it yet?

Domestic China is the only very large audience for their own models, so long as OpenAI and Anthropic stay top tier.

All the hype online from the forums about Kimi, is worthless: those people hyping it can't even come close to running it locally, which is the fantasy. So why are they hyping it? Why did they hype DeepSeek just the same, and learn nothing from its total failure to actually take down OpenAI and Anthropic? Rhetorical questions with obvious answers.

Kimi poses zero actual threat to OpenAI and Anthropic. Those companies will continue to pile up the subscriptions and API usage. Check out GPT's subscriber base today vs when DeepSeek launched. Get it yet? When Model X launches out of China in a year, we'll have this same conversations all over again, and the hypsters will have learned nothing.

While the Kimi fawning is endless, OpenAI will just keep piling up subscriber counts, and Anthropic will keep piling up API usage. Then OpenAI is going to staple a gigantic ad system onto GPT. China can't compete in the model-as-a-service business globally, for the exact same reason they failed so miserably to compete in search globally.

reply
The AI labs have been subsidising. When they try turn a profit, people will move to the fast followers. The only people that won’t are those that compete on leveraging the very latest models and even then, once spend and scale goes to the cheaper providers, we’ll see deeper research from those providers too. Think “PC compatibles beat IBM, Sun, SGI eventually”.
reply
> Domestic China is the only very large audience for their own models

I don't think so. US models are very expensive, and not available in every country. I am not willing to pay $50/1M tokens for writing my pet projects.

reply
There are also US based companies like Fireworks serving up the best open weight models with the compliances we need in US enterprise. Depending on the company, they may offer more/different jurisdictions, EU probably needs a Fireworks like company (haven't heard about one, maybe it already exists?)
reply
There is at least doubleword.ai, and there should be others.
reply
without any hard data one way or another your comment is worthless. "pile up subscriptions" - based on what? neither company is public. "piling up subscriber counts", "piling up API usage"? cool. how much money are they making? oh you don't know because they're not public.

the reality is one way or another that as long as there exists an alternative that a USA company could serve with the same compute rented from hyperscalers, this represents a threat, even if the extent to which is unknown

reply
But what does that mean for Google if their model isn't as good as OpenAI's and Anthropic's?
reply
Why wouldn't the hyperscalers run these open models since they're much better than OpenAI and Anthropic at operating compute at scale?

I think the reason OpenAI and Anthropic stay ahead in revenues right now is because the models are improving too quickly to reliably compete with them on cost.

But, once model performance reaches a plateau -- they have to at some point, though perhaps years away -- that's when ability to operate compute infrastructure at scale becomes the secret sauce.

The big AI labs are likely safe until models stop improving fast enough to protect them from competition on cost.

This similar pattern has repeated in most technical booms prior to this.

When hard drive technology was improving fast enough that old hard drives were quickly obsolete, IBM could maintain good margins making hard drives. But once hard drives got good enough and advances were slow enough that innovation was not the only factor considered by drive purchasers, commodity hard drives started to take over and IBM had to exit those businesses.

The same is likely to happen once model improvement slows.

reply
You’re absolutely right and it’s heartening to see. I maintain a client with ~every provider you can think of and llama.cpp and it was really tiring the last few days to see people laundering other stuff through Kimi and Qwen. They’re not even open yet, the hype was based on their own blog posts, no one’s actually running these locally, the Qwen Max’s have never been open, Kimi’s API was 1/2 the speed the benchmarks was based on, when it was up, and had 60% downtime before they had to stop accepting new accounts, and their EULAs are “your inputs and outputs are ours.” May being clear-eyed benefit us both in the long run.
reply
"You’re absolutely right and it’s heartening to see"

Damnit, I usually don't jump to LLM speech patterns, but this opening had me thinking you were a bot. But after checking your profile, I think you pass as human. I wonder when will be the time, this does not work anymore for me. (Creation date is a strong hint, but abandoned accounts can be hijacked)

reply
Hehe, cheers, it really is funny & odd habit I have (usually when I'm in "everyone is wrong!" mode, haven't bothered to argue that, and see someone else arguing it :p)
reply
> is the same exact reason the Kimi crowd is wrong now. And it's very obvious that they're wrong, but they have intense emotional blinders on.

On the very link on the top comment of this thread, which I repost here:

https://artificialanalysis.ai/models/gemini-3-6-flash

Kimi K3 is ahead of Fable 5 on several benchmarks.

So basically the angle went from "China cannot ever compete" to "China is six months behind" to "China is six weeks behind" to "China is six days behind but that's because they're distilling" and now you're saying "Yup sure, Kimi K3 is ahead on several benchmarks but you cannot host it yourself so this thing will go absolutely nowhere".

I mean: is it not a bit early to draw conclusions? It's been days since a chinese model is ahead of the very best / frontier US model on several benchmarks and you compare it to models who were clearly behind on everything.

Give it some time.

reply