upvote
> "Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy"

I don't believe Cerebras has a cached input pricing? They don't list one on the model page:

https://inference-docs.cerebras.ai/models/qwen-3.8-27b

edit: See the sibling discussion,

https://news.ycombinator.com/item?id=49554520#49555094 ("Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate")

reply
lol yeah just saw that, yeah that makes it unusable I think at least for me.

I wonder if they will do that with sol ultrafast!

reply
They have cache, but it costs the same indeed, no idea what the point of the cache is
reply
They don't have cache (e.g. KV cache). But they write down what you sent earlier to say they cached it! To still bill the same as uncached later (because they didn't actually cache it)!
reply
More precisely they can't cache it.
reply
I can't believe this situation has not improved in years. Is cerebras' main business selling the hardware, then?
reply
they have exactly two customers, both of whom are also investors.
reply
Cerebras is super constrained on capacity right now, all the support is going to enterprise customers.
reply
Maybe they’re gunning for speedy non-interactive pricing? Or its a limit of the technology or a business decision?
reply
[dead]
reply