upvote
Hah, it's the complete opposite for me :D. In Claude Code I disabled all sub-agents stuff, disabled nearly all tools but Bash, Edit, Write and WebSearch and replaced WebFetch with my own tool that doesn't summarize anything because the results were always worse with sub-agents, they always lack the necessary context and weaker models summarize bad. I also replaced the system prompt with my own that cuts a LOT of tokens, agents don't need a 10k+ system prompt anymore.
reply
That's interesting! You don't have cases where the main session has important planning stuff but the actual work to execute has so much crap in it that context compaction will probably dig into the important plan stuff too much and make it too lossy? Also what about the cache read costs for longer context sizes?

Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:

  /autocompact 420k
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).

Same idea with something like OpenCode, there I even configured custom agents for review: https://opencode.ai/docs/agents/

reply
I never hit compaction, most of my sessions are 150-300k tokens long with the longest being around 700k. Using sub-agents means that they can't use that cache, have to re-read everything and now multiply this for every sub-agent you call and it just wastes money/tokens. I also don't like how intransparent sub-agents are, I can't follow what they are doing and I can't really steer them. Claude Code has the /agents view but it's clunky and awful to use.

GPT context window is way too low for me and my last experience with it (GPT 5.6 Sol) was so awful and I hit limits way too fast that I cancelled it (and at least got my money back).

I'm no longer using Pi since it got worse IMHO and Claude subs can only be used in Claude Code but I miss the /tree feature which is perfect for first letting the model read & cache the important bits of the codebase and then start your plan from there (as long as you stay in the Cache TTL). Claude Code has /rewind but it's not as good.

I'm only using the 20$ plans.

reply
> Using sub-agents means that they can't use that cache

They can, when they're forked off the main session instead of spawned from scratch.

reply
I admit I'm not a heavy user of subagents but isn't one of the standard use cases for subagents to run a single command, take the output, summarize it for the main agent and pass it up instead of polluting the context?

When cargo fails to build and creates a massive amount of compile errors, you're better off having this preprocess step.

reply
I also don't like subagents.

I usually have my main agent write a wrapper command around things like that as it hits them. The wrapper only surfaces the important info, writes the full log to a file, and the agent gets some instructions on using sed and the like to navigate the output.

It seems to work reasonably well.

reply
>When cargo fails to build and creates a massive amount of compile errors, you're better off having this preprocess step.

Claude already greps and tails every output by itself, a sub-agent would do the same.

reply
> You don't have cases where the main session has important planning stuff but the actual work to execute has so much crap in it that context compaction will probably dig into the important plan stuff too much and make it too lossy?

For that is it not better to have separate sessions for planning stuff and doing actual work? Pi is super flexible with session management, and a lot of that can be automated by its extension system.

reply
> For that is it not better to have separate sessions for planning stuff and doing actual work?

Personally, seems like too much effort for something that would still need to (and fail to) have some sort of a link between the two, so I could go from the planning over to implementation and back easily. In reality, that'd get lost in the noise of dozens of sessions - I mostly just want the harness to help me do work and otherwise get out of my way, not make me dance around it. Ergo, the more context management it handles, the better!

reply
Eh, I find managing sessions in Pi extremely easy. In fact, I customized how it mages session as part of my workflow, and how agents in different sessions communicate with one another.

It's part of why I really dislike to work in Claude Code, and found it too unwieldy. There I have to keep dancing around it to manage the context in a sensible way.

reply
What’s the difference between sub agents and separate sessions?
reply
I'd assume separate sessions are not aware of each other, sub-agents are spawned by an orchestrator agent?
reply
Almost correct, separate sessions can communicate with one another. In my case, planning and coding communicate by writing files locally.
reply
And subagents don't always communicate with each other, AFAIK that is the most common case.
reply
separate session need more hand holding, i guess?
reply
Er, no?

I mean, not in Pi at least. In Claude Code it is a chore.

reply
running local models, now with Qwen3.8-Flash-Next, they have 256k, but when they get up there their speed is just too slow. So i've taken https://github.com/Tarquinen/opencode-dynamic-context-prunin... and started improving it. It already worked well to get a lot of mileage out of just taking tool calls, code modifications, etc, and dumping them in favor of a summary.

But they'd still inevitably get to long in the tooth, and context poisoning meant they'd just eventually not be able to stay in the preferred context size, which for me is 64k-128k. So, I extended it with an eviction command and required a ratio. So instead of a summary of work, it now just places a waypoint. The waypoint basically means the context has a semi-coherent context but without all the baggage.

I'm on like day 3 of a single session with 3m tokens removed and still in the sweet spot. So it evicts to beneath the lower limit, compresses to the upper limit, then evicts again.

It's amazing how resilient it is if you give it a good plan. The work flow has basically been:

1. Write up an implementation document for some new set of features.

2. Rewrite the implementation as a TDD document

3. Set it to work.

The only thing I haven't figured out is it likes to stop when it hits the finish line of the subparts, but likely we're going to end up with the master of puppets monitoring these things and just set them to evaluating what they've done.

reply
> taking tool calls, code modifications, etc, and dumping them in favor of a summary

Doesn't that destroy the cache? I find that caching significantly sped up my Qwen, especially on said larger contexts.

reply
Yes-ish. The compressed summary sits atop the cache stack, so if we got to 96k, it'll take say 30k, compress to 10k, and that 66k+10k is the new stack, so 66k is still cached and retrievable.

So it is designed like a heap, where we're taking raw context off the heap, compressing it, and putting it back on the heap. So cache during compression is mostly unperturbed, since we're rarely digging all the way to the bottom of the stack, but that could happen.

Eviction though is cache busting; but again, I'm valuing the session's roadmap as the valuable product and context size slows computation size, so I have to bust the cache to sacrifice immediate re-processing for longer term compute speed up.

Because that's faster than getting to the end of the context (remember, every 1k adds to the compute time of the next 1k). So speed at 200k is much lower than at 100k. It's also local, so I'm only paying time+watts for the trade off. As far as I can tell, speed is not being lost since if I let the context grow, the kv cache doesn't help with the compute throughput.

So, yes, but it's "smart"; we're only busting it at the top of the context, so rebuilding it isn't from the bottom up, it's just at the top. Those summaries sink on the heap until you get to the eviction limit, and then, they're evicted, and we rebuild from some intermediate place in the heap.

The benefit of it all is I can have lots of projects, and keep a single session that tends to have the context necessary to avoid having to write AGENTS.md or other context bloats. Set large implementation goals and come back to them as needed, etc. I've had it running like this for awhile and it seems Qwen3.8-Flash-Next has no trouble understanding the rolling window.

reply
That sounds like you just NIH'd mr Zechner's Pi Coding Agent. Those were basically its founding design: yolo-mode security, simple design, minimalist system prompts, plug-in based for anything fancy (even sub-agents and web).
reply
Yep, I've used Pi in the past (see my other comment https://news.ycombinator.com/item?id=49908276) but I don't like the direction it went (selling out, forgetting their principles/throwing them out). And since Anthropic wants you to use Claude Code with their sub I just switched to CC again. I've only used Pi for a few months though when GitHub Copilot gave you 300 requests for like 10$.
reply
I did the same, also, the fact the tools evolve so fast, I dont want to waste time on a particular one while it might be obsolete next week. So either it works now, other I pick something else.
reply
oh-my-pi is a fork of Pi that adds a lot of this stuff

https://github.com/can1357/oh-my-pi

I haven't tried it much though, can't vouch how well it works.

reply
Pi vs omp is hotly debated within my friend group. It has most things you could want, ready out of the box, but also a lot of things you'd never want and it's constantly 5% broken. Some people love that trade, others don't.
reply
Glad I'm not the only want to find it a bit janky/broken at times. They seem to constantly be pushing updates which is nice, but I treat it mostly as a black box.

Anecdotally, I find the auto compaction (or what I assume is happening when the context magically drops) to be hit or miss. I do like how easy it is to use my work cursor sub and business chat gpt at the same time. Then I use nearly free cursor models for dumb shit and Sol for real problems.

reply
I used it for a while, and I thought it was quite hard to follow what was going on. And in the end it just went off the rails anyway, though that might have been a GPT-6 kind of thing.
reply
Exceptionally well is my takeaway. It’s the only harness I am using these days. I was previously using pi and codex mostly, but also the ones built into editors like zed, vscode, and the jetbrains IDEs.

On top of that, somewhat unrelated I’ll agree but still, it has support for vim keybindings

reply
Oh thank you for mentioning it has vim mode that's been my only gripe and I had no idea it had it. don't know if it's new or not never noticed the setting.
reply
It's actually the main reason I chose Pi.

I did create some extensions where it spawns sub agents for specific tasks, especially when I want to keep the context clean or when I really want to offload a piece of work to a cheaper model. And for that I have a high degree of control over, I know which model is being used for each subtask.

I find Claude Code too unwieldy for my tastes. Pi's philosophy of being very light on features nut highly flexible for customization, clicked very well for the way I work.

reply
Some things are impossible to just tack on or work around though, like MCP, while other things, can be done by just composing stuff.

Like sub-agents, you could just instruct pi/any harness with a user prompt/system prompt to start new invocations of itself, if you share what the exact command is, and pi or any other harness will do their own poor man's version of sub-agent via standard unix programs.

reply
I had the model write up a pi extension for logging the transcript to the syslog, and all on its own initiative it spawned a subagent to generate test output. It decided to spawn a lighter model for this trivial task, apparently unaware that I can only fit one model at a time so my llama-server ended up thrashing to the lighter model then back to the original model. My own personal n=2 semi- rogue swarm
reply
What would you say is a good harness with subagents?
reply
I found the MCP extension for Pi to work fine.
reply