Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:
/autocompact 420k
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).Same idea with something like OpenCode, there I even configured custom agents for review: https://opencode.ai/docs/agents/
GPT context window is way too low for me and my last experience with it (GPT 5.6 Sol) was so awful and I hit limits way too fast that I cancelled it (and at least got my money back).
I'm no longer using Pi since it got worse IMHO and Claude subs can only be used in Claude Code but I miss the /tree feature which is perfect for first letting the model read & cache the important bits of the codebase and then start your plan from there (as long as you stay in the Cache TTL). Claude Code has /rewind but it's not as good.
I'm only using the 20$ plans.
They can, when they're forked off the main session instead of spawned from scratch.
When cargo fails to build and creates a massive amount of compile errors, you're better off having this preprocess step.
I usually have my main agent write a wrapper command around things like that as it hits them. The wrapper only surfaces the important info, writes the full log to a file, and the agent gets some instructions on using sed and the like to navigate the output.
It seems to work reasonably well.
Claude already greps and tails every output by itself, a sub-agent would do the same.
For that is it not better to have separate sessions for planning stuff and doing actual work? Pi is super flexible with session management, and a lot of that can be automated by its extension system.
Personally, seems like too much effort for something that would still need to (and fail to) have some sort of a link between the two, so I could go from the planning over to implementation and back easily. In reality, that'd get lost in the noise of dozens of sessions - I mostly just want the harness to help me do work and otherwise get out of my way, not make me dance around it. Ergo, the more context management it handles, the better!
It's part of why I really dislike to work in Claude Code, and found it too unwieldy. There I have to keep dancing around it to manage the context in a sensible way.
I mean, not in Pi at least. In Claude Code it is a chore.
But they'd still inevitably get to long in the tooth, and context poisoning meant they'd just eventually not be able to stay in the preferred context size, which for me is 64k-128k. So, I extended it with an eviction command and required a ratio. So instead of a summary of work, it now just places a waypoint. The waypoint basically means the context has a semi-coherent context but without all the baggage.
I'm on like day 3 of a single session with 3m tokens removed and still in the sweet spot. So it evicts to beneath the lower limit, compresses to the upper limit, then evicts again.
It's amazing how resilient it is if you give it a good plan. The work flow has basically been:
1. Write up an implementation document for some new set of features.
2. Rewrite the implementation as a TDD document
3. Set it to work.
The only thing I haven't figured out is it likes to stop when it hits the finish line of the subparts, but likely we're going to end up with the master of puppets monitoring these things and just set them to evaluating what they've done.
Doesn't that destroy the cache? I find that caching significantly sped up my Qwen, especially on said larger contexts.
So it is designed like a heap, where we're taking raw context off the heap, compressing it, and putting it back on the heap. So cache during compression is mostly unperturbed, since we're rarely digging all the way to the bottom of the stack, but that could happen.
Eviction though is cache busting; but again, I'm valuing the session's roadmap as the valuable product and context size slows computation size, so I have to bust the cache to sacrifice immediate re-processing for longer term compute speed up.
Because that's faster than getting to the end of the context (remember, every 1k adds to the compute time of the next 1k). So speed at 200k is much lower than at 100k. It's also local, so I'm only paying time+watts for the trade off. As far as I can tell, speed is not being lost since if I let the context grow, the kv cache doesn't help with the compute throughput.
So, yes, but it's "smart"; we're only busting it at the top of the context, so rebuilding it isn't from the bottom up, it's just at the top. Those summaries sink on the heap until you get to the eviction limit, and then, they're evicted, and we rebuild from some intermediate place in the heap.
The benefit of it all is I can have lots of projects, and keep a single session that tends to have the context necessary to avoid having to write AGENTS.md or other context bloats. Set large implementation goals and come back to them as needed, etc. I've had it running like this for awhile and it seems Qwen3.8-Flash-Next has no trouble understanding the rolling window.