upvote
Enterprises implemented spending caps and inference providers are lowering prices. Seems they are jockeying for market share.
reply
Partnerships then consolidation comes next...
reply
I wouldn't be surprised if they still had some margins since cheaper models are much harder to nail the accurate sizes off, and you still pay 2x for 1M context window.

But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300?

Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that.

Could still be a healthy 10-30% margin. Especially with Terra.

reply
We can guess based on the decisions of other inference providers who serve these models.
reply
Do you mean if other providers will cut their prices in turn?
reply
Yes. For example, third-party inference providers serve DeepSeek V4 Flash just as cheaply as DeepSeek themselves, if not even more so. This is very strong evidence that the low price of the model is not subsidized.
reply
[dead]
reply