upvote
When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.

It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?

reply
Interesting, if the cache hit is that good, I think HN convinced me to toss $20 at DS official, and see how long that lasts.
reply
Read the tech report and see how laser-focused they've been on compressing the disk KV footprint specifically. 890 bytes/token is absurd, and is probably the reason I regularly get ~0.5 Mtok request full cache hits after an hour. On their end that's just yanking a 414 MiB file off an SSD, then doing a little bit of compute (bounded SWA replay). Nobody else seems to be serving their model as well as they do.
reply
It will of course depend on what you’re doing with it, but right now my session at work has a 99.8% cache hit rate, and I’ve been running this session for hours with 23M tokens read and 713K tokens written (Opus 5.5 in this case though)
reply
cache hit is in fact that good.
reply
I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.
reply
There’s a big difference in speed & quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.
reply
what about dsh + openrouter? and configuring open router to just serve from DeepSeek own servers

I think the 5% cut from open router is fair if I want to user other cheap models like mimo

reply