upvote
They make a giant inference chip. Their inference service is basically just advertising for their core value prop: hardware.

The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs

reply
The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).
reply
They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].

[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...

reply
I never said offloading was impossible. It will result in a large slowdown.

It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.

reply
Mostly economics I'm sure
reply