upvote
It appears that they do support Prompt Caching: https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
reply
It doesn't reduce the price though.
reply
> How are cached tokens priced?

> There is no additional fee for using prompt caching. Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate for the respective model.

Well, talk about flipping the narrative.

reply
heh

Is there a speed increase or is that purely marketing spin on “we might cache on our end but no discount for you”?

reply
Pure marketing.
reply
[dead]
reply
deleted
reply
deleted
reply
Strongest model that they host on the public endpoint. They do a super fast version of GPT 5.6 Sol for OpenAI and have bigger open models on dedicated endpoints.
reply
Agreed, but in our SAAS I can tell some UX will sky-rocket to next level with this
reply
The coding plan is gone now right?
reply
Last time I got one, I had to log into a Discord server and wait for "the drop" and IIRC Daniel Kim was giving them out based on who was there at the time. They were gone in less than a minute. This was ~8 months ago.
reply
i believe they used to have monthly plan, what happened to that?
reply