upvote
Yes-ish. The compressed summary sits atop the cache stack, so if we got to 96k, it'll take say 30k, compress to 10k, and that 66k+10k is the new stack, so 66k is still cached and retrievable.

So it is designed like a heap, where we're taking raw context off the heap, compressing it, and putting it back on the heap. So cache during compression is mostly unperturbed, since we're rarely digging all the way to the bottom of the stack, but that could happen.

Eviction though is cache busting; but again, I'm valuing the session's roadmap as the valuable product and context size slows computation size, so I have to bust the cache to sacrifice immediate re-processing for longer term compute speed up.

Because that's faster than getting to the end of the context (remember, every 1k adds to the compute time of the next 1k). So speed at 200k is much lower than at 100k. It's also local, so I'm only paying time+watts for the trade off. As far as I can tell, speed is not being lost since if I let the context grow, the kv cache doesn't help with the compute throughput.

So, yes, but it's "smart"; we're only busting it at the top of the context, so rebuilding it isn't from the bottom up, it's just at the top. Those summaries sink on the heap until you get to the eviction limit, and then, they're evicted, and we rebuild from some intermediate place in the heap.

The benefit of it all is I can have lots of projects, and keep a single session that tends to have the context necessary to avoid having to write AGENTS.md or other context bloats. Set large implementation goals and come back to them as needed, etc. I've had it running like this for awhile and it seems Qwen3.8-Flash-Next has no trouble understanding the rolling window.

reply