upvote
> if OpenRouter is blindly dispatching your requests

This can somewhat be the case, depending on your config. I updated mine to make DeepSeek high priority because I was having a lot of cache misses and reliability issues with the default (cheapest (at face value)) providers, and cost was actually higher overall than anticipated. Was smooth sailing from then; might have to tweak things again now pricing has changed though.

reply
I was having the same issue. Did not find any provider that had a cache hit rate anywhere close to DeepSeeks own API.

Not sure if this was an OpenRouter issue or with the other inference providers.

reply
System RAM and/or NVMe storage still has a real cost. And swapping out the context between VRAM and system RAM / NVMe still consumes bandwidth.

I don't have a clue on what the real cost to inference providers comes out to, but it seems really weird that there would be such a big gap, in what should be a pretty competitive market.

reply
deleted
reply