upvote
They made the coding plan a bit better toward the end, but it was pretty tough to use throughout.

Seems like an Amdahl’s law of inference economics? there’s so much compute relative to SRAM on the chip and shoreline bandwidth onto the chip that caching buys ~nothing? The contended resource is SRAM and a given token of context needs just as much as another.

reply
That’s not what caching is for. Caching lets you resume with a pre computed KV cache saving you from having to ingest everything in the chat history as input on every single round trip. You still need caching regardless of SRAM or not as it saves a huge amount (and ever growing) of compute ingesting the preceding history every time you want a completion.

I don’t know why they don’t give a price discount. Maybe their hardware is incapable for some reason of saving/restoring the state? Or maybe they just haven’t built the infrastructure to do it?

reply
Can't you do something with multiple accounts?
reply
You would lose caching (if they cache)
reply
Or just buy a 5060. This will run on most any 16gb card. Slower for sure but far cheaper than another subscription.
reply
Or buy a raspberry pi with a SSD, about the same difference, if you're giving up on the 1500 tokens/s anyways.
reply
Thats not the point if you choose Cerebras as provider.
reply