1. Every chat should have a context used/remaining measurement so you know when you have to ditch the current chat for a fresh one.
2. Every chat should analyze and categorizes each element of context by how useful it is towards the overarching goal of the chat.
3. Every chat has a handoff button with a "usefulness" slider (say 1-5) that shows the total size of the context based on its setting.
4. The handoff automatically creates a new chat with the desired amount of context and a prompt to get it back to where you were.
That said, I am newb and so there is some reason why these non-deterministic LLMs can't do this :-/
My way of prompting this varies and every time I receive the blob of output, I can’t fell how well it managed to capture the necessary details. This way feels lika a dirty way to transfer knowledge from one conversation to another.
You said there is? What’s the options?
If you snapshotted at 90% max context you could pretty reliably start iteratively trim that down, I think? I personally try to save the logs so agents can slice and dice them with sed/awk/jq/whatever when they need to look stuff up, because I’d rather pay the penalty on read (when it’s motivated by something) than in write(where you don’t really know what if anything will be needed), and they can figure out what they need on their own.
What I’d rather have is some way to bake history into the actual model weights (the same way it can recite certain literature or historical/factual stuff without context), with like multi-lora / “experts” that get trained out of band. But this is contrary to the “one fat model” approach to scaling and doesn’t work with closed labs’ business/IP models