upvote
Without having any inside information, one possible theory:

All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.

or

The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.

reply
Cerebras provides high-speed inference at high cost. It's never going to be the cheapest and thus it will probably remain niche.
reply
But that's supply and demand, not technology. Right now a lot more people want their inference than they can supply. as supply catches up in the next 5-10 years, the underlying tech at scale is probably cheaper than GPUs per token produced.
reply
Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.
reply
I love how instead of comparing a Ferrari (fast and expensive) to some average car (not fast, not expensive) to make your point..you went for public transport where your comparison cracks from multiple angles.
reply
If you're willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.
reply
Cerebras capacity was pretty much entirely bought out at some point. We needed it and couldn't get it.
reply
The WSE is very expensive to build, and they have a waiting list of customers who are already willing to pay a lot of money for the available supply.
reply
it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.

imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

reply
I thought your numbers must be wrong.

So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum.

So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.

reply
It is worth mentioning, the OpenRouter demand isn't static though. It has increased week on week since early 2026.
reply
I was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?
reply
GPU cost is about half of datacenter cost. Other half is cooling, power and networking.
reply
I use Cerebras via OpenRouter. It’s every bit as fast and reliable for my needs as claimed. I suspect the reason is that they can either be making peanuts selling inference to plebs like me via OpenRouter, or making bank selling the more expensive models to businesses directly. In short: I would be very surprised if they have die capacity, and are at this point maximising revenue per chip.
reply