upvote
I think the specific issue I had was due to lack of proper support for ds4 compressed KV cache, not about KV cache quantization. It was like 50GB instead of 5GB for context, and it wasn't fixed for weeks (I haven't checked if it's fixed now - hopefully it is).

Quantization is another thing. There are so many engines launched with claims about speed, but in many cases it's optimizing specific lower-quality quants. When you have enough resources, you usually want something like W8A16 + full precision KV cache working as fast as possible, not yet another W8A8 or W4A16.

In general, it seems new models are released so fast now - engines don't always have time to really polish the implementation before the next model is released

reply
Do you have anything published on the quality benchmarking using your caching strategy?
reply
We have a retrieval benchmark based on RULER which we've been using to ensure that the model maintains complete awareness of the full context window.

All our benchmarks are open source so you can check it out here if you'd like: https://github.com/magnitudedev/magnitude/blob/main/inferenc...

reply