In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet”. It’s something I came up with late on a Sunday night many months ago, when I got tired of asking Claude to plan with me first before coding in each new session. Something people might not realize is plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive for users.
This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For codebase understanding, I sometimes ask Claude to generate an artifact that explains some aspect of its changes. For complex diffs to core parts of the system, I will often ask it to make diagrams or even interactive demos so I can better understand the change and alternatives considered. I don’t do this very often, but it’s a useful way to explain code when you need it. I ask Claude to attach these artifacts to its PRs also, so others can understand and future Claudes have the context.
For me it’s actually the opposite, and Claude Code’s plan mode isn’t nearly sufficient. Personally I ask Claude to write down a markdown file with its plan, then review the plan using plannotator, and then go back and forth (most of the time it’s actually the comments that are the problem, not the code).
Then start a fresh session, seed it with the plan, tell Claude to find ambiguities / friction points / oversights, resolve those, and then implement it.
Review once again with plannotator, go back and forth, and then send PR.
Maybe not the “vibe coding” that was once imagined, but this does ensure I am fully aware of the code and architecture, the quality, and this also prevents long term degradation.
Currently looking for a framework for managing this in a more formal way, and I think it's probably beads, but interested to hear from others.
Might be worth a look if you’re evaluating alternatives to Beads.
I have some older projects that use beads (I still run an old version without dolt that's imho pretty good overall) but lately with Fable also have a few newer projects where I just have the agent write docs and keep a worklog with the what/why/decisions etc. (I think I read it here on HN somewhere and figured I'd give that a try.)
The latter seems to work pretty well for now (slightly better than beads) but I'm always looking for ways to improve it. This could be an interesting replacement.
Great work man.
“ayy lmao”
I asked fable to look at my interaction patterns and clearly stated my frustrations and the problems I wanted solved, and it designed a simple process to track things in git and built a couple simple session hook skills. It’s pretty lightweight and I’ve been very happy with it for a couple months.
I get a long way using models like Opus to make a plan of action and a bunch of tasks, and then using Deepseek to implement that plan of action. Saves a bunch of money and is fast.
- I have a record of work done and work to be done that helps _me_ when I come back to the project after several months. It’s committed and lives with the code.
- when a task inevitably ends up more complicated than I thought, I can in that session break it up
- I initiate sessions from multiple computers, so things stay in sync (through git)
- I also have a “tooling” repo that builds out some views of the work and hosts it for me to see when I’m on my phone.
- The hooks let the agent manage all of the workflow/task management, so there’s very little management overhead for me.
I rejected beads and JIRA. I wanted something more lightweight.
My main conversation is usually with an orchestrator that hands off work to various (usually cheaper) subagents to plan / review / etc. It has instructions to find the correct model for each task and not to do too much itself so a multi-phase plan automatically gets a fresh subagent for each phase.
I also found that having the design reviewed by multiple agents has very little marginal value. The review agent will always find something to improve, but mostly it’s just nit and not anything super important.
I used to let Claude just upload the html design doc to Claude artifacts for me to review. Recently I switched to codex and started to use my own tool https://github.com/hyperlogue/r3 to complete this workflow.
I wonder if it's just a consequence of a gigantic training set full of comments completely out-of-date with the code, leading to the model considering this "normal"
I now make sure to do a big decommenting pass before every PR.
But I also am starting to just let go and stop caring. It’s not clear to me that it causes problems down the road, it’s easy to strip out en-masse if needed, and in my experience, agents now are really good at read git blame, the commit log, even prior agent transcripts if available to sleuth out when a change was made and why. So yeah, it’s annoying, but the code agents write for me is increasingly never read by a human, so does it matter?
When I read, I skip most comments, especially the larger "Javadoc" style. My brain sees them colored differently in the editor and it doesn't even take mental effort. Then, when I have a question about the code, I look back up for a relevant comment. That doesn't happen very often.
If Claude writes great code and leaves a garbled Claudese-but-accurate comment ... I can read and comprehend (with like 5x the effort of a human comment) ... that's a small price to pay.
(I do have detailed instructions for it on how to comment (or not) but it has not fixed this.)
It’s always “you explain only what but not why” or “this is way too much prose” or “these comments don’t belong here, they should be inline comments” or “this is completely redundant as it’s already obvious from the code”.
I do find that once I beat it into submission and the codebase is “clean”, the new code it generates gets better and better, which makes sense gives its pattern-prediction nature. But it seems like there is work to do for Anthropic in terms of getting Claude to not confuse code comments with dumping its interactive discussion state into there.
5.5 is much closer to Fable so i don't even need it. I am pretty sure it's got Fable's DNA in it.
I really need to find a role where I can do more DX...
I get great review results (as good or better than colleagues using superpowers or even adversarial review skills) just by asking Claude to review a PR and spit out results in order of severity.
Note that this is only really necessary for complex work that I don't know yet what the best way to do it is.
I've tried doing it your way as well, but there was just too much fiddling about with writing the plan somewhere, then having another session rebuild their context with whatever info is in the plan. It really didn't result in better output for me.
Currently 9 times out of 10 I just say to the model: xyz is the problem/bug/feature, fix it. Since about Fable and Opus 5, this is more than enough. Opus 5.5 (and previously Fable 5.1) got even better at this. However, this is in a codebase where there are already a few hundred thousand lines of code for the model to look at to see how we generally attack things in our codebase.
Claude Codes plan mode I never use anymore, it was useful a few months ago because the models had a tendency to just start doing work and forget I specifically told them not to. But the UX is just annoying and the models now do adhere when I tell them not to change anything.
Plan mode ensures I'm spending fewer tokens on the code-test loop, and more on the arch/design, and allows me to keep appraised of what's going on, while planning for future changes better.
Maybe folks who don't need planning, don't have as much concern for the details, and are happy enough with just evaluation of if it works or not.
I've tried doing the incremental, iterative approach with just Code and it's just not as effective unless you're working on something simple or experimental. Or you're shipping to something non-serious or perpetually beta.
Then telling Claude to work on a document, the instruction is kept to its core.
Now when bcherny explicitly mentioned that it merely adds a single line - it explains why I don't need it.
What may be concerning about "super plan" mode from the creators (or a skill, for that matter) - is that tuning the amount of effort, and how much deep to dig - may become too hard, as it will interfere with several embedded paragraphs explaining what to do, how to do, where to do, etc'.
What I do look for is even better plannotator ability to track changes, combining historical comments (like Google docs), and git blame of several "generations" before current reviewed doc.
Roughly speaking, I'd be happy if plannotator would persist something similar to github PR reviews combined with Google docs comments & suggestions.
I personally still find planning a valuable mental exercise; it's not so different from pre-LLMs and whiteboarding or otherwise taking the time to consciously plan a set of work.
How do I use plannotator to review an arbitrary markdown file? It always opens the Claude Code plan file for me.
So again: You don't need plan mode, auto mode works just fine, there is no difference in the workflows here.
- strategy document
- "sprint" document with technical implementation
- actual implementation
- e2e testing scenarios updates
Every step involves iterating with Claude on it with me in the loop (setting the direction then resolving the "founder questions" as they appear), and importantly a different model for review/code-review, be it Codex (usually, it's great at it) or Antigravity/Gemini (sometimes finds novel things, its precision and recall are abysmal but on the odd occasion it has good accuracy). This iteration on the high-level plan then on the implementation plan is essential to me, and IMHO part of why people are surprised that I tend to get solid results from LLMs. At the very least, it allows me to fill gaps in my own knowledge (primarily front-end development) and be more productive than writing the code myself. I cannot stress enough how nice it is to have a partner in the high-level system design – yes, it often suggests utterly moronic ideas, but the overall experience is still net positive and getting better every quarter.
Either way, plan mode isn’t going away. You can always /plan or ask Claude to enter plan mode. We might re-map the shift+tab keyboard shortcut to something else by default for people that don’t use plan mode.
Strange response. I agree with the parent comment here, plan mode lets me ensure that I have specified everything correctly before it gets built which is far too late. I don't see how an improved models even matter to this workflow. Is Fable going to read my mind?
Boris is saying that you don't need /plan to get the model to plan, you can just say "Let's plan this out" or similar, which at least matches my experience. Your experience may differ, of course, but it's not even clear we are talking about the same thing.
Maybe some people have not been long enough on this rodeo: This used to be an actual issue. You told the model "DONT START CODING YET" and yet, surely enough, starting to code it did. That is what /plan etc were supposed to fix.
Sounded to me like you need a plan.
My approach is to take the statement of work or problem definition and iterate on that myself until I'm really clear on what is the goal. I therefore have a good some good ideas about what the plan should be.
If extending an existing application, which is usually the case, then make use of the plan documents that I had written before AI arrived on the scene. These sre documents in markdown form that say step-by-step how to, for example, add a new report to the system.
It makes some plausible choices and you can retroactively ask it to make different ones later, if you want.
This hasn't been my experience.
> Or it will litter the code base with defensive code and comments about the path not chosen.
I've definitely seen this, though.
Basically every big tech company maintains "codebases with millions of lines built upon decades." Talk to your friends at a FAANG and ask them how they're using Claude/Codex.
Merging after reading the PR description is just how it's done these days, and if you can't do it reliably, your harness, devloop, or model is simply behind the times.
But you're doing the same with your "how it's done these days".
These days things are done in many contradicting ways, and it probably will take at least a few years to settle on common normal.
OP wasn’t even actually critical of LLMs, they were just saying that plan mode was helpful to stop the model from making incorrect assumptions when you want you don’t specify everything you should.
I mean, you're discussing this with a marketer/someone wearing a marketing hat, who works for a company which needs people to use as many tokens a possible. That's their reality
I see you didn't disagree with the main thrust of my post.
You'll also note their 'About' is empty. Regardless, engaging with HN in this manner is de facto marketing/PR. Trillion dollar companies don't just let anyone post on high-profile social media websites for any length of time without permission from marketing/legal/PR.
Which is fine because it just put together a plan and didn't spend 10 minutes rearchitecting everything.
It tends to be small decisions way down the stack that bubble up, or an incoherent data model that can’t handle what you’re asking for cleanly.
Eg I was messing with a state tracker the other day. The state tracker assumes a container is either currently running, or fully removed from disk.
The LLM chose to remove the state file when the container is stopped and then to remove it after, which leaks container storage.
The LLM is kind of stuck though, because every option other than “rewrite the data model” has negative outcomes and it probably violates user expectations to launch a massive rewrite there.
0. https://github.com/mattpocock/skills/blob/main/skills/produc...
/plan is still useful, I still need to review what it's going to do and still make revisions. But there's two phases: hammer out key design decisions then write and amend the document.
If I knew exactly how I was going to build something, I would have built it myself. But since there's some ambiguity in the portions of the project I'm less familiar with, I rely on the plan to not only help me understand the decisions Claude has made for me but to keep Claude constrained to the decisions I've made. It's very frustrating to waste tokens on having to refactor something because
> But for larger things, I try to take a waterfall approach with well defined milestones.
> If I knew exactly how I was going to build something, I would have built it myself
Aren't these contradictory? If you don't know exactly what/how to build, how can you do waterfall?
1) "Maybe waterfall works now" - plan mode, take care of all the nits and issues that the bot leaves on your PR, wildly overengineered "enterprisey" solutions with a lot of bells and whistles all over the place but very poor end-to-end user story test coverage that results in user experiences with a lot of good test coverage of the edge cases of how a given step might fail but little thought towards overall user flow and throughput. Because part of the issue with waterfall was assuming you could design the right tool for your user up front.
2) "Maybe code doesn't matter anymore" - The just ask for something when you need it approach, which results in weird janky individually-sorta-working but strangely-overlapping six-variants-of-the-same-thing that makes it hard for your users to develop a single consistent mental model of the thing they're using, and that changes super frequently.
They both end up with a lot of other bad-for-velocity things that I assume are inherent to how the tools have been refined in response to last year's criticism, too. Super verbose comments. Extensive - without much eye toward runtime - low-level test coverage that might miss the forest for the trees and also slows down the next round of iteration cycles. A plethora of new proper nouns all over the place that make the documentation an ouroboros without a good entry point.
It would be nice if with new model releases claude code also gave a bit of a model 101 that tells you evolving ways of prompting it that the insiders have picked up. I know there’s the prompting guide in the claude docs, but this is often very broad and most people don’t know about it.
This is weird to ask because I feel like of course the model isn't omniscient? Isn't the whole point of iterating on a plan to assess impact, risk, know your (the user) variables, user impact, product impact, etc for making a change? I cannot count the times even in the past few months where I start a conversation with my C suite because their desired outcome would have a potential negative impact elsewhere for other products or users.
Is this just not something that comes up at Anthropic?
I absolutely see fable and opus 5.5 misunderstanding intent, but that just seems to be a feature of necessarily underspecifying in a written prompt. Just today, I gave opus 5.5 a simple task to spin up a new environment for work. It read the ticket, which was decently specified and knowing the codebase as well as "Ghasp... reading the code" I had to correct it about 5 times to do it in a way that I would have expected it to. Getting the pipelines right, environment variables, and configs. It was all relatively straight forward imo. Then I had to prompt it to clean up its corrections, because it left a workflow variable in the github action that some intermediate step required but the final solution didn't. I definitely would not have caught that if I didn't read the output. Idk, there seems to be a natural limit as to how much it can infer and I have no idea how to fix it. I did write about it [here](https://javiergonzalez.io/blog/the-assumption-problem/) though.
Summaries that don't tell me when it's changed direction in a timely fashion, but I am only told way later, when I have to undo. Really bad judgement calls regarding where to fix bugs, changes in implementation decisions, taking action when I am asking a question directly, not passive aggressively asking for action... when 5, 3 days ago, was proven to be untrustworthy, switching to very little supervision sounds like a strange thing for a customer to do.
5.5 and 5.1 have major Rain Man (savant) syndrome. Excellent at many hyper-technical things, absofuckinglutely boneheaded at anything that a human (or an earlier model) would understand - like how to write copy, what a human would expect in a given situation, various types of norms...
it's infuriating because it's a sophies choice - dumber model but better human understanding, or better technical model that you have to explain things to over and over like a toddler.
I greatly prefer this, since it lets me iterate on the plan with Claude for a while without it repeatedly asking if I’m ready to implement the plan.
Once I’m satisfied, I usually start a fresh session and tell it to implement the plan.
For smaller plans, you don’t need the file. Just ask it to come up with a plan. I don’t recall the last time it just started implementing if I only asked for a plan.
It wastes a ton of tokens as well and those are not cheap.
But as skills and memory are populated over time, Plan mode isn't as necessary. It becomes simpler to let Claude just build and get something general in place that works, and then refine from there. Auto mode will ask essential questions.
I still use Plan mode for big feature changes, to confirm that I've asked for what I want in the right way. I tend to prompt casually, with only a few specific details. Plan mode helps me see the whole picture before committing. In a few cases, it also helped me decide the feature I asked for was wrong.
So it sounds like we use it somewhat similarly, just taking a glance at it before the work starts, and that's becoming harder to do in my experience. It is important because problems with the overall plan or strategy end up magnified the further you proceed with development.
Most people aren't precise when initially describing their problem.
Jumping straight into implementation skips the part where we refine and better define what it is we're trying to do, and think through the implications of those changes.
I suspect "trusting the model" doesn't really work at scale with finite resources.
Also this is a way less removed process that I want nothing to do with. The more removed I am from the process the more I hate my job, get burned out and genuinely wish that Anthropic never existed.
Even if it could "just know" or infer my intent. It wouldnt be desirable.
Edit: Oh yeah its Boris, hes one or the most disengenous shovel sellers on earth right now.
The Anthropic employee literally told you plan mode is a "we're still planning!!!" at the end of the prompt. Plan mode is useless because you can literally type "dont write code yet" and get the same effect, not because you should never plan in general.
In case anyone else is interested, the skill is public: https://github.com/Mudlet/Mudlet/blob/development/docs/demo-...
* Scoping discussion - do research and figure out approach (auto mode)
* Planning - take the scoping and convert it to a concrete plan for review (plan mode)
* Implementation - put the plan into action (auto mode)
I find this works really well for my workflow, and it is really easy to trigger each phase because the model has clean boundaries. (This workflow is articulated in my user CLAUDE.md) Plan mode is still useful to me as it forces the model to double-check its plans (I find even with Opus 5.5 it still discovers gaps), and it gives me an opportunity to clear context at a really good spot.
So I would still consider plan mode to be useful. It would make me sad to see it go.
PS- I have a Claude Marketplace directory submission for an MCP server that has been sitting in review hell for six(!) months. I've never received any outcome other than "In Review". I hope I'm not asking too much but would it be possible to put me in touch with someone who might be able to help here. Nobody has ever replied to messages sent to mcp-review@anthropic or the "Anthony at Platform Operations" inquiring about status, and we are getting frustrated
I tried plan mode when we first added it to github copilot and it didn’t stick for me, until very recently, when I used it the way you describe. I was just putting it too early in the process, turns out I have to do a bit of exploration and discovery on my own and most of the time I can skip plan but now I have an intuition for when to engage it so it can interview me to clarify the last few things it needs before implementation. This also made autopilot mode work a lot better for me.
Historically, plan mode served two different roles:
1. making the agent’s instructions precise enough to execute 2. helping the human understand what was about to happen
I think #1 is less necessary as agents get better. #2 is going the other direction, it becomes more important as the model is able to do more on its own because larger chunks of work are happening with increasing complexity.
Where I’ve changed my mind is the interface for #2. I increasingly think an interactive, iterative workflow is closer to how people actually build understanding than being handed a long generated document, especially one they didn’t author themselves.
The human-understanding problem is very real though
Also I hope your delivery goes well. My wife (and co-founder) had a challenging delivery and it really put life in to perspective for both of us on a range of issues (how much women's pain is minimized in the health system requiring stronger personal advocacy than I would ever have expected).
As far as plan mode, I still find it essential in keeping agents on track. I build propelcode.app and have a variation on plan mode I still find useful, happy to trade notes on agentic coding if youre interested.
Propelcode looks awesome! So cool to see different people and perspectives shaping this space.
And thank you for the article - it was a good read. I still see folks in my org playing "throw spaghetti at the wall and see what works" and getting frustrated so plan mode (mostly point 2) has been their guardrails almost as much as for the AI.
Can't help but think "Doesn't matter if a machine or a human with (even slightly) different background wrote it", maximizing information flow is maximizing common assumptions and "culture" to only have to communicate a small set of current information for the task at hand. Being a team means having built a joint context so to say. This has always been the purpose of design documents and they always were too big or too small. Because you did not write them, but the others. If you only produce code you think they are the past and useless. If you iterate and your team grows, you start seeing the value in always current docs that are containing just what is not in your everyday culture.
All the best for you and your growing family. I had a similar experience recalibrating my values...
I was quite surprised when I learned this (when Claude Code edited a file despite being in plan mode). I had previously assumed "plan mode" was a harness level concern, and restricted what tools could be used. I didn't expect that it was simply an addendum to the prompt.
But I think it gave me some good insight into where the heads are at of Anthropic employees building this. Basically, leaning on pushing everything to the model. That's why alignment is so important: a "sufficiently advanced" model doesn't require any tooling infrastructure around it, and I suppose Claude Code devs are targeting that future. I had previously thought there was more to a harness, but with "auto" mode these days, it seems like there's no desire to build in that direction.
Basically, Claude Design is primarily a huge prompt.
That's all.
There's no "magic sauce", and with the right tools to call, even your local LLM can implement "Qwen Design".
Also if you are working in a heavily vibe-coded codebase, as Claude Code reportedly is, it's not that surprising if the human doesn't really understand it or have anything useful to add in a collaborative context.
I haven't reached for Plan Mode in awhile--maybe, on some blank folder/canvas and I just want that cute "questionnaire" DX to get me going...
But on the whole, CC is smart enough to know when to "rush off and act" and when to "pushback", which is great--and it's no big deal to tell it to pause/stop by adding "what's your thoughts?" or "feel free to pushback" etc to my prompts to make it start a back-and-forth with me (the fact that you say that Plan Mode was really nothing more than a prompt anyways is reassuring).
And yeah, when you're deep in the weeds, you could (and can!) have multiple threads of thought/work going on in the same convo, that stopping and starting a plan in Plan Mode seems like a regression.
(happily using Claude Code Opus 5.5 on High rn)
For the record, this isn’t unique to Claude. ChatGPT and Gemini do the same, each with its own quirks. ChatGPT got extra credit for being the only one who allowed the function to also take a CA file for server authentication.
Don’t get me wrong: LLMs are the future (maybe even the present) of software development but I think there’s some way to go before they can be entirely hands-off in some areas. I still find myself having to course correct designs and plan mode helps me with that.
And of course, thank you for your work on Claude. :)
But there is still a ceiling above which it is necessary to "preload" the context window before starting to call tools and get into the meat of the work. You want to establish domain language (especially with Claude models which otherwise will invent their own, and it will be inscrutable) and key requirements and assumptions. You want to do a Q&A iteration cycle with the LLM. You definitely should do a sanity check that the LLM actually "understands" what you were trying to achieve, and then make sure that understanding is coherently and plainly stated in the prompt. All of that seems to be necessary still for just about any serious task, if you actually care about the quality of the results and/or don't want to burn hundreds of thousands of tokens on flailing around to get to a good quality result.
So no, you don't "need" plan mode. But you do still need to do all of the things you would do with plan mode.
That's when I need to learn about implied patterns, do's and don'ts; not from theory but in the context of my own project. Unfortunately, the model tends to keep implicit knowledge implicit. But I can ask during planning.
The feature did get less useful over time when the model started babbling in newspeak more and more. When it threw a thousand words at me even in concise mode.
So I don't want to let the toolmakers off the hook here. There's a lot to win that would make plan mode much much better without changing plan mode itself.
> In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet"
Can't help but think if plan mode isn't useful as you say because it's implementation is lacking in claude code, hypothetically speaking.
What I can say is that with other harnesses plan mode helps stabilize my workflows. Actually synthesizing code is only part of the process, lots involved in taking a work-item to production end-to-end and plan mode helps give this flow structure. More than that it's an opportunity to regroup before committing to changes. It slows down the process to a rythim that's sustainable and smooth, which ends up speeding up the process.
So if plan mode in claude was designed to speed up code churn, while oh my pi for instance designed plan mode to be strategic, that might account for the different perceptions here.
And it's not to say you should force yourself to use it, but if you are planning on cutting this mode off the loop just beware of the possible side effects.
I would really like a dry run mode that just disables all external commands from the outside so Claude doesn't proactively go about changing things.
When a test case fails, the relevant part of the plan is surfaced in the error. I find this helps Claude stay on track for longer - I've been able to do 12h most times and even up to 48h unattended (11h of API time) with good results.
Then whenever I do check in, I ask it to update the HTML with current state in an append only fashion (sort of like it's writing a blog), and then based on that, we iterate on the end to end test (I think of it as a "test harness") - update the test cases and error messages.
I've been able to build some truly large projects this way, both greenfield and up to spec (for example, a video game I've always wanted to play), and brownfield while staying within the conventions and design of the codebase, and with very little attention required on my part.
seems like plan mode could turn off some tools, even if it doesn't change the set offered to the model, the ones that they have which would mutate your codebase could just not work with an error message, and plan mode could change permissions in the security approval prompt for "auto"
anyway, isnt the right way to know if plan mode helps or not, to run an experiment? we're all guessing unless we have data
read only agent mode sounds straightforward and useful to me
1. I dont want to have to accept every time Claude touches our DB
2. I'm scared out of my mind it might do something bad to the DB
Plan mode gives me enough confidence that it wont do (2) --> allowing me to give it enough permissions to do (1)
For my small-scale sqlite dB, it gets read access, and I encourage it to test modifications by copying it somewhere and writing into that.
Scale-dependent, but I hope to not have to work at a scale where it gets write access to the production DB. That just seems like asking for trouble.
I was experimenting with a rather complicated backfill operation, were I had a validation script I understand and have Clod come up with the backfill script. I was running against a local prod copy, and it proposed running the actual (unfinished) backfill script against prod.
It didn't have access to the secrets and I also caught the command, but a good reminder that this stuff needs guardrails.
Hallucination not a big deal when it's on the surface layer. But I can't imagine the damage it could do if it hallucinated while building/validating a "load-bearing" component and then continued down that path
we DO daily snapshotting, so the risk is limited... but still spooks me
I miss that. It worked really well, and it kept the context clean.
Plus, I usually plan with a more expensive model and guide implementation with a cheaper model (with smaller validation calls back to a more expensive model)
I ended up building out tool an MCP server that I use as a bit of a psuedo harness for Claude. I have a variety of multi-step workflows that are basically micro-skills stacked on top of each other. This helps me make sure that I can get Claude to think in a repeatable and reliable manner.
For coding, I've found that I have a few specific steps that Claude needs to do before I'm comfortable letting it loose:
* It must extensively explore the code base (including certain areas that it misses)
* It must think about what it doesn't know or is making assumptions about
* It MUST scaffold out it's intentions. Essentially, it can write comments, classes, and method stubs - but no actual content. Very much like a spec, but since it's in and alongside other code, it's much easier to identify problems.
* It must spike and validate key assumptions. This, plus the prior step, are the only way I've figured out how to avoid it ending up in a confusion loop. Too often it looks at poor-quality code it's written and thinks it's a long-term solution. By avoiding writing code as much as possible, it knows that it's draft content.
* Only, then can I review it and send it it.
Said MCP server (missing the actual ops): https://github.com/clops-mcp/clops-mcp
Now that I’m frequently designing and delegating day/week scale features, the flow has to change; having the agent go off and build a spike can be a quicker way of us understanding the design space and constraints (especially in a huge codebase). I still have the agent write and update a spec doc as I go, but it’s not waterfall anymore.
At least for my kinesthetic learning mode a rough code PR stack is usually way better than a plan doc anyway, and tokens are cheap enough (vs my time) that going further than just a plan is often cost-effective overall.
The dream of course is (say it with me) loops, but that doesn’t tend to work for me on new features often.
1. I want to know whats going to happen, at least at a high level, before changes are actually made. 2. Plan mode helps me flesh out the missing details of my plan before being mid-execution 3. In situations where I have a limited budget for AI usage I will often times use a high powered model like Opus 5.5 or Fable to make a detailed plan, then scale down to a cheaper model for implementation. I feel like this saves cost in the end.
I get plan mode is basically just a small hidden prompt. I get that I can basically just preface my prompts with "make a plan only, don't make actual changes." Maybe this is just a UX trick, but it works well for my brain.
Aider, Cline and many other agents had plan mode before Claude Code existed.
I find myself endlessly ctrl+c ing claude now as it flies off doing deep first principles analysis to work out how to find a thing it isn't sure about but I know the answer. Being able to give it that answer without needing to ctrl c would be a massive improvement
But with GPT 5.6 Sol, I'm still finding that the model makes conceptual mistakes, or gets edge cases wrong, or assumes incorrectly (making an ass out of both user and model). In many cases, I need to at least refine the proposed approach, or amend, correct, or flat out just stop and start over. Not planning and catching these errors, and just letting the agents code their code, would mean I'd have to rollback and redo many times. What a waste!
For a current project, which is ~33k lines of code, I'm also finding that I know the codebase better than the model, and that's vital at the planning stages too. If I wasn't in the planning loop, the model would have reinvented various wheels a few times over. How much spaghetti do you want with your code?
As always, I may simply be doing this wrong. But I'm personally not convinced that the plan is dead, or that I want the plan to be dead. Planning is also good for me -- it keeps me thinking about the code, prompting better, guiding the model better.
If I'm no longer on top of the codebase, then at some point my prompts will devolve to "Do the thing with the thing, that does thing". And I don't want that.
On the other hand, for a low effort hobby project: just do the thing.
You don't need to actually change the tool schema or break the prompt cache to do that. In the tool itself you could just check if it's in plan mode and reject the tool call...
They’re most useful for broad changes (new features, refactors, etc.) where it’s helpful to avoid breaking changes or unnecessary scope expansion.
The new models are great, but they do more by default, which means I’m finding myself explaining what _not_ to do more often than with previous models (where they’d often end too early).
In my case, the previous plan mode was too ephemeral, and I like having one source of “truth” that sits across context windows without loss/compaction.
Plan mode is a useful shortcut when I want to have an agent do read-only work without having to worry about giving it appropriate stop conditions.
Also for session planning, as in when-can-I-walk-away-from-computer, its nice to know the particular rhythm of initial crunch - ask questions - make plan - do it. Especially with a 5 minute cache timeout.
I'm experimenting just like everyone else, but this is my process right now:
- Quick prototype
- Figure out the language of your app (what terms you want to use for things, what your UI design language will be, etc) and spec that, so you can use words consistently with the agent. You need to be able to describe the things you want well and consistently.
- Keep prototyping. Let the agent write unit tests along the way. Lock down behaviour you like, keep track of those things in a document.
- At some point your idea of the real architecture comes into focus, from actual use cases -- avoids the over-abstracting right away trap.
- Refactoring is cheap with tests, so start refactoring into the architecture you want.
- Your architecture won't necessarily be what would be best for a human, but it will be pretty close.
- Keep relentlessly iterating on small work.
- Things that were expensive before aren't that expensive now -- integrating a library, changing from one library to another, trying out a few architectural refactors, trying out different performance optimizations, etc. That stuff is all 'throw it there and see what sticks' now, so don't be afraid to try stuff which felt big before.
I feel like 'front loading' too much is just the wrong approach. You might feel like you're sitting there 'babysitting the agent'; but that's just what the hard part of the work (hard as in 'zjust slogging through it', not as in 'conceptually complex') looks like now. Your code is much more like clay.
Atleast that's how I'm thinking about it so far, but I'm not working on large sprawling systems that I imagine would need more pre-planning.
I find that the code is generally in a better place proportionate to the amount of SDD I actually do. But it's just a matter of where and when I want to spend my time.
But given that running an agent us cheaper than the cost of waiting for a slot to assemble the team to talk about a change (isn’t it always?), why wait with running the agent?
I propose updating the spec then do the implementation. This will most likely show that a few assumptions were wrong forcing some major or minor updates to the spec. Work through those and then let your team review the spec change together with testing the next iteration of what what’s build.
For who? The more control you hand over to the AI and let it think for you with no supervision, the better it is for Anthropic
However how about decisions? Do we expect the model to read our minds, just assume the best practice will be followed and that’s what the user want? Plan mode solves those, what is that am I missing?
For example, there are no shortage of web based uis for pi and they are all cool but I wanted a deep integration between artifacts and how I want to collaborate with the agent on them. So i built my own ui that mirrors what the pi tui sees. It’s chat based but has a deep integration with a GitHub style code review UI so I can leave review and comments whenever I want. Every agent message renders nicely in an annotatable markdown viewer so I don’t need an ask question tool and can more naturally get the agent on the same page as me. I want it to feel like I’m working with a colleague.
I don’t think these features are too unique but having full control over the experience is really nice. Flexibility over model provider, can tailor it to my work’s dev stack, and don’t need to worry about anyone breaking it with a million updates everyday.
Your perfect workflow can be realized in a day or so. You just need to go make it happen.
I've been building https://crit.md to keep that back and forth with agents - GitHub-esque interface and have agents respond to my feedback, iterating until I'm happy.
Admittedly like many others, I use it a lot less for actual plans now, models are indeed getting really good at just getting it.
I wodner what this product space will look like a year from now. Reviews are already dying.
It could still make the tools into no-ops or disabled if it actually tries to use them, without changing the context history at all.
I still use plan mode in Astra to come up with a plan that I then feed into Fable. I feel like OpenAI models still do better big picture investigation and planning, while Claude is the better software engineer, if that makes any sense.
Of course this could well come down to my own biases and the specific things I’m working on.
“I want to ...” / “Let’s ...” -> Plan
“Do X” -> Actually Act.
But then again I also have it configured to only ever answer questions instead of inferring them to be instructions (which I’ve seen others do differently).
- Most people suck at planning anyway
- LLM still don't give you a way to verify and understand to iterate.. you have to ask and then formulate and way so people barely do it, they just trust the vibe
IMO the current successor to plan mode should be the harness knowing when to tell the use "ok here's our overall current state in a simple diagram", auromatically
https://innerloop.test/breadcrumb (for reference)
Is this what happens when you vibe code long enough?
https://innerloop.works/breadcrumb
What a rookie mistake!
# YOLO mode
I just start my day writing about 20 queues /goal prompts and then check the work at the end of the day. It’s almost always right!
I have one session define a task, and provide a formal specification plus context in a "cover letter."
The session B, in plan mode, produces the plan back.
Session one reviews the plan and clears it, ratifying portions and often specifying specific changes.
Session one then executes.
What has been striking to me in this approach is that even with two instances of the same model (currently Opus 5.5), there are regularly corrections made. I use "project chat" for session A and Code for session B atm; it is very typical that Code finds and corrects details or oversights in the task spec; it is also typical (though less so with 5.5) that session A (chat) pushes back or clarifies things Code doesn't have the context for.
I have been afraid to open up the potential of negotiation beyond what this is costing as it is. But I am also afraid to simply skip the formalisms, because of the consistent correction that occurs in this back-and-forth.
Each component of the pattern is schematized, generated from a template, and validated, to keep things tight.
Lots of tokens! But I trust this process far more than "just typing" :)
I also tell it to write deviations and rename plans accordingly once done.
That way I keep the codebases I have to or enjoy to work on in my head and don't become too dependant on any provider or on stochastic parrots in general.
For a one-file change I don’t bother. For anything that touches auth, payments, or a shared schema I still want the plan written down first — not because the model can’t figure it out, but because I need a moment where I can still say “no” before it starts editing.
The mode was never really about making the model smarter. It was about making the human stop and look.
I’ve heard “earns its keep” in only two contexts in my life - the intro to the song “Regulate” and terrible Claude docs
So basically you don't know what the fuck you're talking about. Not everyone is employed to burn money.
The one thing plan mode helped is for the humans to get an understanding of the strategy, and be able to poke around and look at the design and architecture. You can achieve this with some self discipline and keeping shorter leashes on agents, but it feels like a losing battle. The best devs still put out good code, but the poor devs are learning nothing while their metrics look great. I can't help but think we are racking up immense amounts of debt that will very soon become due.
A significant (majority?) portion of developers have been shipping JavaScript/node applications for the last decade that contain hundreds of MB to GB of code from god knows where doing god knows what with dependency trees the size of redwoods. It’s not like your average mediocre dev really knew what was going on behind their gluing of frameworks together - at least from what I’ve seen.
If you have remotely competent tech leadership that enforces relatively intelligent patterns (a good one I’ve found is “write everything backend in rust”) you can make AI churn out monstrous amounts of code that… isn’t all that bad? And if you enforce it writing and updating a docs/API.md on every commit/PR you’re probably doing better than 80+% of devs I’ve ever met. Up until a few years ago it wasn’t uncommon to roll up to a new job that was a “legacy” pile of garbage concocted over 20+ years with no comments or API docs and a readme that tells you to ask for help from someone who has been dead for 5 years. At least AI code is full of comments (some of which might even be accurate) and there’s a finite (relatively low!) cost to figuring out “wtf is this doing and how is it doing it”
What happens when the maintainers lose access to frontier models because of cost, politics or other external factors? What happens if they don't have enough hardware to spin up an open-weights model?
I've seen variants of this play out before AI, so I can tell you: they will inherit a codebase they've never seen before, take forever to ship fixes (forget new features), and they'll either scrap it, completely rewrite it, or, if they're "enterprise" enough, will pay consulting companies literal mountains of cash to make it their problem.
The whole point of writing simple code was to write code that other humans can maintain. If AI is here to stay and becomes economical enough for everyone to use it, then you're right; writing code for other humans is no longer useful. If that doesn't happen though...developers who can/want to still code by hand will be loving life
and there is a community of developers who care about the quality of those libraries and take the weight of their responsibility seriously
You can give the AI the spec, and own the spec instead of the code.
Eventually, the spec too will be something the model owns, and you'll work at a higher level of abstraction.
Beyond a certain size, the documentation becomes too large to ingest. Below a certain size, it can only contain a fraction of what is needed. If you take any sufficiently engineering project, and give every engineer amnesia; the project will go to shit for a undetermined amount of time, as it takes months or years go build-back the understanding that was lost.
This is clear enough then large companies fire and replace workers randomly to cut costs; a worker that has built up useful knowledge in the origination over a few years is more valuable than three cheap consultants from "low-cost countries" that are fired when the work package is over.
---
AI, looses its memory every time we press "new thread". No spec can bring that back until AIs become able to write and ingest whole books of context without getting confused.
The revenue for a new feature today is something sure. While the cost associated with supporting such feature will be up to debate in the coming quarters.
As often it is the case, we are moving on a long vs short term trade-off space. And I don't think any experience will generalize
When I worked at a big tech company on a large codebase (tens of millions of LOC across dozens of repos), it was extremely common to work on something in an area of the codebase I had zero familiarity with, and due to turnover, no one else at the company did either. As you say, docs were frequently missing or out of date.
However, with some effort over a couple hours, I could make a LOT of progress in understanding the history of the code. Every commit and PR were linked to JIRA tickets, most had eng design docs with comments, slack discussions, etc. I could step through the git history and watch the code change, alongside the artifacts of the human discussion and decisions that led to the changes. It wasn’t perfect, but I could make tremendous progress. Now, this was a remote company with pretty strong culture around using JIRA, design docs with review, etc. Probably the biggest gap was meeting transcripts.
Today, an agent can chew through years of history and artifacts on a large codebase in a half hour, documenting as it goes, and have way better understanding of it than I ever will.
It’s true that an agent can’t hold all of that in its head at once without context rot (though this is improving every year), but neither can any human!
If your goal is understanding how to do something, why something was done, why something wasn’t done, etc, and the codebase is large, mature, and extremely well-“artifacted”, I’m not at all convinced that you’re better off asking Bob who has worked on this section of the codebase for a decade than just asking a really good frontier-level agent. Maybe, but it feels like that won’t be true much longer.
At my last job we worked in a large, but publicly available code base. My experience was that just giving an agent the prompt to look at area X to figure out how it works would use up half or more of the context window. And that was thrown away every time we started again. We had docs, agent generated overviews but the agents still filled their context windows way too quickly to actually really be any use.
Like "if you're using this database, and the engine has these configurations, then do _x_, unless _x'_ and _y'_ are enabled, in which case, do _y_..."
Which, if you're already being THAT specific in your spec, you might as well, idk, write the code yourself?
Because at that point your human language is basically the code and AI is the compiler. A non-deterministic one.
Regardless, your business stakeholders won't understand what's going on anyway (nor should they), so we're back at square zero.
(I wrote Technical and Functional Requirements Documents as a business analyst in college. What's happening now _for most situations_ is more or less the same thing.)
AI abstracts effort and cognitive load away from code at a heafty rate, but it doesn’t abstract liability away from code at all.
My business is paid to produce artefacts for which it has liability in the case of error, so we need to do additional work to mitigate and eliminate the liability risk introduced with language models. So far I’ve not found a better way to do that than a plan/act/assert type approach on every feature.
If you use established libraries then actually the code IS well known to someone (and likely many), even if that's not you. Likely it was built with an actual purpose and with the foresight to not add red herrings to the design.
You can't say any of that for the equivalent amount generated lines. Literally no one knows what it does.
Arguments against it are sort of like why have devs on staff at all when you build the thing the first time, or why not outsource everything, or why should I care what my code looks like when the code seems to work?
The cost of tokens is not zero, and the bigger the thing you’re doing the more low quality will cost. When your company is the software, you take on an existential risk based on the software working or not.
Pure AI generated code without human curation is full of bad wordy comments but those comments can mislead, be stale, contain duplicates, and drastically reduce the ability to understand things. I watch teams that still care about code understanding ship good products while those vibe coding in the same org just flounder after the initial burst of features. Some problems only come up after the first 90%, and AI can help you solve them but a big ball of spaghetti is still a big ball of spaghetti.
I.e. how hard is it to point an AI at a piece of software and say "AI, copy this"?
Seems like sooner or later copying just becomes a matter of spending enough on tokens.
Seems in that world, all significant software projects get copied. That turns software into a commodity loss leader for other business models or an open source project. Similar to the way Chrome works for Google and the way Firefox works.
Won't you be worried if your mechanic didn't understand your car but offloaded it to a robot that made mistakes all the time?
This is what happens when executives suffer from AI psychosis. They were already impatient, now with AI all they care about is feature velocity.
The faster they can hit that refresh button to see the features, the quicker sales can close the deals for them.
AI has basically sold them to wet dream.
I guess we just wait for the boom.
I see it kind of like baking/cooking. Do you bake your bread from scratch? Do you grow your own wheat and grist your own flour?
I think over reliance on it or not even trying to understand what is happening is a big problem to be sure, but it's certainly not a new problem.
Actually baking is a good example, I used to be really bad so I spent time learning. I don't do it every day but now I understand how bread is made. I bought a 3D printer so I could print parts to fix stuff myself. I learned to do my own oil changes, I learned how engines work, etc.
My point is that I try to learn more, not less, which is what AI is trying to achieve
And I would understand how to bake bread from scratch.
Now in the entire chain, we will get to a point where no one knows anything.
E.G. over time we’d gained two client-side caches of related server state. This started out as two different parts of the same model, because we couldn’t get all the data we needed from one microservice and had to merge in the client. Over time, more and more features used both caches for different aspects of related processes. At some point one of the microservices changed so as to return all the data in one call. The update to consume that kept both caches, adding code to sync them, because so many parts of the code were using one as a fallback for the other, so they both looked “necessary”. Because they were separate, and “live” sometimes they’d go out of sync after the initial load. Worse: the consumers alternated about which cache was treated as the fallback, making it very hard to see that either might be redundant. Eventually I noticed they were filled by the response to a single call. We all know paying back tech debt never gets prioritised, so I rolled the payback into two feature tasks, and just took longer about them.
My employer expects we use LLMs and provides some budget, but it’s not enough to use even Open4.7 or GLM-5.2 on every task. I do the bulk of my work with Composer 2.5. It’s quite good for “going forward” on smallish tasks and it’s written most of my code this year. It’s possible smarter models would spot these refactorinh opportunities and action them proir to building features or fixing bugs. But I wouldn’t know because I can’t afford it. I’ve never seen even a 4.8 era model spot a refactor and plan to do it prior to a “new build”.
I’m pleased I’ve spotted these trends and started to build the habit of (telling the agent to)“refactor to make the change easier”, but my percieved productivity will go down and I risk the ire of my leaders.
I very regularly use plan mode not to even make a plan of action itself, but to better understand what possible issues might come up when implementing some feature or fixing some bug. And it is quite common for me to fix or rewrite certain findings that AI comes up because its assumptions are not quite right or don't align with overall goal.
And yet so many seem to be perfectly fine leaving all the decisions to AI - even if it's going in the wrong direction. I suppose that's all the people who got into software purely for money or status - never really caring about the actual thing they are working on.
The funny thing is that the AI adopters are in the middle of the bell curve. Our worst devs continue to perform worse than AI yet refuse to use it and our best devs continue to insist AI sucks despite it finding issues in their code and the reviews and designs they've approved.
Shouldn’t those “fixing bugs we gained in the past” be their own MR that can be read, reasoned about and have evaluated test coverage?
In the case of what I'm currently working on, filing bugs for every issue I found, and then factoring out each fix, and then running each change through the 8 hour ci/cd system, and hoping an unrelated issue doesn't get misattributed to me... No, I'd rather just wrap it up into one coherent refactoring change and be done with it because when I'm done there are several more like it waiting for my attention.
Couple of WTFs that come to mind:
* How many bugs did one have per commit, that commits have to noticeably grow in order to not have those bugs in the first place?
* How does one even do software engineering if the (best) developers can’t reason about the code?
Software engineering is possible but largely a myth in practice.
At a company I had setup scripts to build our packages, and the CI was running those scripts. Someone more junior (only a few years, not decades) found it strictly superior to remove my scripts and replace them with GitHub actions: people could now know even less about it (as in, no need to know how to copy-paste and adapt the recipe for a new package), but now it depended on GitHub. GitHub is down, nobody can build anything anymore. And it happened once every few weeks, so people would just go have a coffee during the outage.
You know what happened next? That person got promoted for their good work. That was before AI.
If you don't understand the codebase, ask the agent to explain it to you. I'm not not kidding. Modern frontier models are fantastic as this - even more so than actually writing the code. It can tell you in words. It can generate architectural diagrams and sequence diagrams. It can write tests and scripts that prove it's assumptions. It can happily refactor so that the system design is aligned with your preferences.
Once you accept this, you can stop worrying so much about it and instead focusing on building the architectures and tools that lets the agents succeed better and faster - so called closed loops or agents prompting agents. Build systems that are more easily verifiable and deterministic so the agent can write very powerful property based tests. Focus more on what and why you are building, how to make sure all external properties are verifiable and leave the internals to the agents. The code is not really for us anymore.
Then when you actually dig into the code, there are many things that are not like you'd expect.
When you've experienced that a few times, you stop trusting that the agent gives you the full picture - for good reason.
When I review AI generated code I generally find so many flaws that it makes it hard for me to believe that those who are not reviewing their output are not just fooling themselves. Maybe not all the time, but quite often.
One such recent example was an SSO simulator for a local env. Instead of using a cookie to remember who was logged in, the agent remembered the last log in a variable, assuming the the next requests would come from that login.
This snowballed into our tests, where later agents had created helper tools for working around the SSO simulators statefulness.
Things improve drastically however if you spin up a second session and ask it to adversarially review everything that the first session produces (this goes for everything: not just code, but also design, planning, and explanations).
This works even better if you use models from different families to do so.
The quote is "so simple that there are obviously no deficiencies"
I agree with this, but the reality is that it's only the result of models empowering devs, and power in good hands amplifies positive results while power in mediocre hands amplifies technical debt.
It's a good time to choose wisely who you work with.
Very true, but this also makes me think what kind of ridiculous obstacle course future hiring process would look like.
In a land where anyone with a pulse can prompt AI to make an app for them - how would future hiring managers and team leads figure out who will drag codebase down with tech debt and who wouldn't?
The tech debt concerns are much ado about nothing. Use the next model to clean it up, big deal. Code is cheap.
The people that sat around handwringing about tech debt and trying to read every line of LLM code will really struggle to find a job. The profession fundamentally changed, and these people did not catch up.
The reality is that models just keep getting better and are very good at cleaning up the debt they created. The "tech debt" bill never came due. It won't.
> The "tech debt" bill never came due.
Companies paying $200k a day for coding models to churn on what the coding models are messing up is one thing.
I’m not even worrying about tech debt, I’m talking full on defects, production incidents, security holes, and reputational damage.
If it turns into a hairball, just have a few agents rewrite it.
For me, this phase still happens, but a distinct "plan mode" is unnecessary: I just tell the model, "This is discussion; no code changes yet." and spend hours figuring out what will and will not be done.
vs
Press "Tab"
Like - you really think models won't be able to clean up the tech debt they created!? They are very good at this already. Ask Opus 5.5 to clean up the tech debt from some Opus 4.6 vibe coded app.
Code is cheap now. The most important thing is to ship, ship, ship. If you are handwringing over "tech debt" you have already lost - and you deeply misunderstand how good this technology is getting!
If we get to post-scarcity you never needed the money. But if we end up in some dystopian hellhole where people have no jobs but capitalism still exists, you'll be thankful.
I can't bring myself to trust that an LLM understands what I mean better than any human would, no matter how "good" people claim they are getting.
TFA seems to be advocating for regular old vibecoding. Code now and ask questions later. Which is their choice, and is perhaps even a valid choice in many cases. But at least call it what it is.
I find that faster than code first, ask questions later. But it takes more time up front.
If you're reading and editing the code, you're not vibe coding. If you're not reading and editing the code, you're probably vibe coding even if you feel "hands on".
It's particularly important when you're making architecture changes or things other features will need to build on top of. I want to know what libs its going to use. The nitty gritty details I don't really care about.
Cursor would also write a plan document, which was useful when working on larger tasks (due to context size). I still find that part useful today.
It helps to think more abstractly when approaching problems like this.
You can use Docker’s sbx or similar VM/containers for that.
RE the article: I don't think it's obvious why this process is worth following until you find your time and attention wasted. Conversationally-building is the express train to waste. I'm not sure why you would even be talking to claude if you don't understand what you want to build.
If you do use plan mode you might like https://plannotator.ai/
I typically converse with the default model to point the plan in the right direction, then have it iterate with a smarter reviewer to find flaws until the plan file is converged.
I am currently working on canvas-based interfaces for that reason and i would think the only way to really create value here is with a deep independent analysis and visualization of the changes afterwards to reach some ease of mind. Live would be cool (if you like that)
When it comes to planning itself, I recently tried the token-saving planning plus phases execution agents approach and had to find out that agents actually don't necessarily communicate better by prose-reduced specs than we do.
I had to go back to the planning agent to implement or fix things with our full planning context in mind. So if you want really high control for a "tight" implementation, I'd say just sharing plans is not enough. The probability of things getting filled in by the executing agents rises and you either find yourself holding those agents' hands or fixing things afterwards.
Actually phases are still to large and you would actually want the planning agent to hold that hand all the time, meaning small context is not the way to go, as you might need the full checking context much more often than current phases sizes suggested by the planning "doc" would use it.
Distributed building still is the way to go though, steady control by the overall context or one specifically thinned out for the particular job is. But don't go prose-based plans anymore. These are dead indeed.
A mixture of defending against a disastrous mid-implementation compaction (where suddenly things would veer off the rails) and also allowing the fresh execution to double-check the assumptions and notice any subtle mistakes before context was poisoned.
I’ve found that for large enough changes I still prefer having a parent theorizing about the root cause of issues based on evidence and then dispatching targeted child sessions to fixed based on theories and concrete telemetry examples.
There’s something clean about having sandboxed context and a session you can quiz about architecture while one is heads-down working against a spec.
The "standard" plan mode felt too stifling.
A quick straw poll. Are most people here who use AI to code well-versed in their languages/software development? i.e. 10+ years experience doing it "by hand"? I think in ten years time there will be no developers with that 10 years experience behind them.
But what I've noticed is that lots of people don't want that. They're happy to delegate both typing and thinking.
There are folks who still want to think for themselves of course, but model providers are incentivised to maximise token spend over all else, so the "official" harnesses are not aligned with their needs. My message to those folks is: _go and build your own harnesses_.
I built one, it wasn't a huge lift and it completely changed how I work with LLMs. One consequence is that I choose smaller, cheaper models now because latency is more important with shorter feedback loops. I don't want a super-intelligent model that disappears for ages and provides a final answer, my own brain thinks and guides the process instead.
I'm deliberately not linking to my harness here because that's not the point. Think about what modes of work you want from a coding harness, then build one that works exactly like that. Immense happiness and satisfaction is just a couple of weekends away.
Author might be a vibe coder, designer or someone who does not do really heavy and complicated development, maybe some api here and there.
But it is pretty obvious that a very concerned engineer would not allow any of that, for critical systems.
My only critic of the plan mode is I wish it was easier to see the updates and changes easily in Claude Code as we iterate on the plan. It is wasteful to have to remember what parts I have reviewed and what parts are new (and need another pass). I have thought about fixing this but I also feel the review is the actual thinking (even if ineficient), and so I purposely have not removed it.
I can’t imagine how high that number would go if I got rid of the collaborative style and also the “please don’t run scary commands on random directories without permission” mode.
It’s like this: If you know the right words to say to the LLM, you’ll likely get back the “right words” also. And the more right words up front, the less steering you need to do.
(Having 20yrs of experience writing these user stories and designs help, my “unfair” advantage)
Firstly I find it’s an excellent way for me to get very good clarity about what will be built and whether it’s going to be done in a sensible way.
Very often I don’t really know what the work will need to look like until I’ve explored the problem with the LLM towards first making the plan.
Without a plan I find myself having to do the initial understanding through code review of its generated code which is much harder than reviewing a plan, and then I invariably need the LLM to fix up what it did which is much slower when it’s doing code than working on a plan, never mind the next review I need to do.
And when the plan is good enough, I clear the context before telling it do it, which I’ve found vastly improves the quality of the LLM output.
Another useful property is the readonly nature. I can easily let multiple agents plan in parallel without having to worry about annoying worktrees or conflicts and then I can come back to the plans later.
Of course this can be done with just another prompt, but that's exactly what plan mode is. It's nothing more than a predefined prompt in the harnesses with maybe some extra guardrails (that don't always work)
But my conclusion on plan mode is slightly different. I agree that plan mode itself may be a dead end, but I still believe there might be another way to achieve the same goal.
When I was building my product, I found that the biggest issue wasn't capability, but taste. The agent could build something that worked, but it often wasn't what I actually wanted. And behind that "taste" is a huge amount of implicit context — preferences, past decisions, product intuition, and trade-offs that live in my head. Distilling all of that into context takes a lot of effort, and I suspect giving it all to a single agent may eventually become overwhelming.
I've been wondering whether a better approach is to have multiple agents with different roles, prompts, and perspectives, and find a way for them to work together efficiently.
It's still just a hypothesis though. There are a lot of complicated coordination problems to figure out, and I don't have the answer yet.
The sense I get from this discussion is that the models/harnesses do not elicit feedback well. Where there is ambiguity, they tend to pick a solution and call it good.
A planning step aims to make these choices explicit. An iterative process is necessary to capture the detail.
It is natural to look to teams of agents to satisfy that process, but do they know where the decision points are?
We really need a better model. One alternative is to have an everything-app: a general purpose tool in which (almost) everything lives. The terminal is one of them. The text editor / word processor is another. (I use Emacs for everything.) In a business context the spreadsheet is probably the best choice.
I used plan mode for two reasons: to review the choices before execution, and to execute with another model (i.e., using the barely documented opusplan feature).
The grill-me skill is much better for reviewing and clarifying choices (and modifying it to use the ask tool makes you go faster). Instead of opusplan, you can explicitly tell Claude to start a subagent with another model to divide the tasks.
What the author I think is hinting at is not planning alone but "the development and evolution of any program and the state of this program throughout the planning, elaboration, and eventual runtime".
I chuckle at the thought that throwing an md file or a prompt at this problem is sufficient.
So, I posit that if we want any agentic code to evolve meaningfully in the short future and over the long run, we have to have a system which holds and presents this information, the state of a program, in a coherent manner to a human operator. No other way. No other way. And this I say to both nay and yay sayers.
You can argue also that this is part of an even bigger thing. But it is not part of the current discussion on planning and speccing in agentic systems.
OpenAI and anthropic can throw all the billions they don't have at this and adjacent issues, but if this is not solved then they don't have anything.
Execution is the answer. All the complex success stories I've seen involved the LLM iteratively probing live data sources with varying filters, throwing code changes at the compiler over and over, or invoking shell commands until it succeeds.
I believe there is a Yoda quote regarding this.
"showClearContextOnPlanAccept": true
Boris rationale was "with 1m context window, most users don't need it anymore."/s
This is why open harnesses and open models will always be superior, it's only a matter of time until someone decides that something you use isn't worth maintaining anymore. The reply from the Anthropic employee up above doesn't give me much confidence that Claude users will be able to continue using plan mode forever.
The flexibility and lightweight nature of open harnesses is mainly useful for open models. These models are considerably far behind the frontier and need a lot more steering. Complex harnesses confuse them. I use open harnesses for my OSS model rig because the model needs it.
Even if I agreed, that doesn't address my point: Claude Code might be the world's best harness right now, but that doesn't stop an Anthropic employee coming along tomorrow and taking a giant steamy dump on the thing you like most about it. You're using their workflow, not your own.
In my experience the problem with plans is that llms (at least gpt and claude) are really lazy and will stop at the first solution without even understanding your project. I had plans that wanted to add a few thousand lines of code just to support a very minimal part of a feature. After asking for a second round of research with the project specific limitations in mind it recommended a very small change to the existing codebase that did the same thing.
This is probably also the reason why most llm generated code stacks thousands of lines of code.
I also bounce around to lots of different models. Some are cheap and dumb, some not. When unfamiliar with a model, the last thing I want it to allow file edits and destructive tool calls. Having a mode explicitly for that is helpful.
For example, the ability for the harness to call into a Python one-liner just to experiment is pretty powerful. If I'm asking it to use D2 to build an SVG graph, it can write some Python to introspect the XML to see if things appear to be placed correctly (size, x-y coords, etc.). Which is a pretty cheap way for it to experiment and verify its results before I deign to examine the rendered SVG with my own eye balls.
And yes, you can say that I can do this without an explicit plan mode.. But it's such a useful and common workflow that it deserves a special mode, IMHO.
However, since everything changes every 5-minutes now; I am curious what is now a better process than using superpowers? What works for you?
I guess people don't even look at code anymore
But nothing is stopping us from prompting agent to only write to plan.md this session.
I also find this "kill Plan mode" push on Twitter odd, because developers have been complaining about AI supposedly killing their jobs, yet they want to take away the main feature that lets them be an active collaborator and participator in the process. Weird.
I feel like this post is hawking for commercial reasons more than it’s based in reality, especially a reality that takes into account the extremely varied experiences engineers seem to have with agents. Not to mention it ignores capabilities that, for example, Claude’s VS Code addin have had for ages (specifically select and annotate in plan mode, as I mentioned in another comment).
As such I’m not inclined to take it very seriously. More than that, the “X is dead” trope was overdone 10 years ago and I don’t think it’s been long enough to warrant a resurgence.
but I think I overfit the interface to how I used to work without agents. when I was at GitHub, I wrote a lot of ADRs, RFCs, and design docs, and I liked that because writing is how I clarify my own thinking. with agents though, I’m often not doing the writing myself. I give the model a rough intent and it fills in a bunch of gaps, and then I get back a long, polished plan containing decisions I didn’t explicitly make.
That’s the part that feels broken to me. The plan can be detailed and technically correct, but still be hard to review because the important bits are buried and feel distant from my own thinking. All the assumptions and tradeoffs and questions may or may not be legitimate, but it’s hard for me to get into flow state and carefully check them.
So I still want the collaboration step. I’m just less convinced that generated prose à la plan mode is the right interface for it.
This tool is a difficult sell. Users _might_ consider an Open Source version, but switching people away from familiar tools is not easy.
It sounds like the Ant perspective is "you can just ask claude to plan", but at the same time that's a little more tedious than shift+tab
1. giving the agent sufficiently precise instructions 2. helping the human understand what’s being built
I think #1 is increasingly going away as models get better. #2 is a separate problem, and I actually think it matters more as agents get more capable. I’m not arguing that human understanding should disappear with plan mode.
I’m also not pushing this on behalf of a model lab, I don’t represent one. These are lessons I learned from my own mistakes building Nuanced (https://www.nuanced.dev). We built around a very explicit plan-oriented workflow because that’s how I used to work. I spent years at GitHub writing ADRs, RFCs, and design docs, and I still really like writing because it’s how I clarify my own thinking.
What changed for me is that AI-generated plans don’t give me that same effect. The agent reads between the lines, generates a lot of detailed prose, and now I’m parsing decisions and assumptions I didn’t actually make.
So I still think human understanding deserves a first-class primitive. I’m just increasingly unconvinced that a big blob of generated text is the right one, especially as we move toward many agents working in parallel.
It's just redundant when you can just collaborate in the "main" mode. There's never been a point to having a separate plan mode.
Goomba fallacy. These are two distinct groups of people.
Maybe it's me (it usually is), but I don't give an LLM small tasks that a human could knock out in a half-day. I give them big tasks, stuff that would take a human a few months to a year - and then I ask it to give me a plan, as a living document, and we spend a good hour or two iterating over the plan.
Then, when its finally at the point where I'm happy with the plan, and I've talked it down from wherever it first wanted to go, or pointed out that we don't need to be all-things-to-all-(wo)men and focus is important, I get it to start going through it.
Likewise, I get it to maintain TODO.md with lists of known bugs, separate lists of future-features, and again, I make a rule that this file must be updated whenever something material changes. I just asked Claude "where are we ?" and got back stuff like:
The ## Still-open detail section lists two items:
- Task #1070: the ported back end doesn't fold offsets into vector loads, so it emits an extra add on 11 files (from 892) at -O3. The output is correct, just longer.
- Task #1080: array sizes must constant-fold. For example, u8 buf[EVSZ * MAXEV] is rejected because size expressions only accept a literal.
This is for 'xc' [1] - a compiler for an Objective-C-like language (but without the excessive []). The language has ARC, blocks and bound-functions/methods named 'block' and 'callback', automatic parsing of DWARF data so you can #use a shared-object, so there's no header files - just read enums/types/functions/methods from the shared object. It's a cross-compiler, runs on mac,windows,linux and creates executables for mac,windows,linux,ios,android,WASM (amongst others). I have a binary running on my iPhone which was written on, and signed on a Linux box - no Apple software used at all. Oh, and it produces code that is very comparable to clang in speed on both arm64 and x86_64.
You can appreciate it's a reasonably large project. It's taken actual months(!) [grin] for me to get working. Months! There's no way I'd approach a problem like this without detailed plans of what I wanted the language to do, where we were going with it
FWIW, "I" wrote blewit.net [2] entirely in xc - both the server back-end (#use <psql> was very useful for binding to Postgres) and the WASM client - which share classes between client and back end, to make it very difficult to get out-of-step between them. No Apache (#use <tls>), no scripting, just a lean-and-mean daemon talking to postgres via valkey (#use <valkey>) - a reddis-alike. Oh yeah, blewit.net has a plan too. Actually it has lots of planning :)
This has never not worked as expected.
On my hobby projects I use Antigravity CLI and almost always start with /plan. This is a builtin skill - not a mode. It generates a plan artifact that I review and approve. Once approved the implementation speed is uncanny compared to Claude.
I tried using superpowers with Antigravity CLI but that slowed the agent down considerably without much benefit (plan quality was much worse).
Over time I found some useful patterns (indeed after not getting what I want from plan mode), but this piece convinced me to double down on them and always try to find (and let the LLM construct) clarifying abstractions that have a deterministic relation to the code.
2 examples I recently build while I'm developing a large Django system with a complex datamodel and RBAC (spending quite some time to make it look good):
A script that makes an svg of my data model with all the classes/tables laid out and their relations encoded in the line ends (1 to 1, many to 1, many to many) and their on_delete relations encoded in colors. This also helps me discuss with stakeholders. The visual is also in the README.
For the RBAC model I decided it should be declarative so a TOML in the code that seeds the DB with the roles. I quickly landed on small script that translates the TOML to a markdown table with roles as columns and perms as row. It's also in the README.
After reading this piece I'm going to actively think how I can build these visualization more often and consciously, on different levels. Great realization.
Imagining a mode where the agent is very transparent about its direction and process, and I can passively interject at any time to steer the outcome, without having to wait my turn or hard-interrupt.
But it's no longer to plan a single PR that can be done easily without plan mode.
It is mostly because I am more geared towards having my system work on initiative level changes where it works for days and plan is a good document to maintain to keep the agent aligned with the original goal and avoid unnecessary drift.
It can still be used in ways that I personally consider correct, but I think the parts I personally consider incorrect are so inherently alluring that I find plan mode to be an overall net negative for software development. I celebrate its apparently impending default removal (at least Claude Code and OpenCode are openly stating that they think it's time for it to go).
I agree that as I move away from holding agents' hands through actual coding I need a different way to monitor what's going on. What step of the plan are we on, what are the tests actually doing, what agent owns what, etc. I haven't found a product that does a great job of that yet and it seems like the next frontier of the 'IDE' to me. It'd be more like an IM(management)E really. The closest thing I've seen was whiteboard [0] but I didn't have a great experience trying it out.
[0] https://dev.fast
Just don't expect to end up with a finished project; it's more like a first draft. Once it's there, it's much easier to determine what it is you actually want, since you can directly experience what works and what should be changed.
One important caveat is that I do not work with agents; each step goes through a fairly rigid manual review phase.
For complex features there can be 10 or more questions but I have a very strong sense of understanding the changes about to be made and Claude is very good at following all the decisions exactly. It's like working with an engineer who is both excellent at soliciting requirements and fast at implementation.
[0]https://github.com/mattpocock/skills/blob/main/docs/producti...
https://github.com/mattpocock/skills/blob/main/skills/produc...
That writing style is borderline incomprehensible.
I am still actively working on thesis : a self-directed plan mode to generate artifacts that can go through multiple evaluations of interactive interrogation is valuable
https://github.com/samelie/claude-plugin-pnpm/blob/main/skil...
Building an issue tracker that addresses this. It can replay the workflow after the fact like a movie, and pin down the parts that require human input via tagging and inline diffs in the tickets. It is git-native, lives in your repo alongside the code doesn’t require any external service.
Qwen 3.8 27b is the supervisor
Qwen 3.5 4b are the 6-15 minions it controls
Gemma 4 e4b is the validator for the supervisor.
A plan means it preps all work for the agents up front, tests that evals work, makes sure the dev environment is right for each agent, then finds and fixes each before the distributed tasks even begin.
What I thought would take minutes took hours as a supervisor or one agent did the prep / pre flight work.
My solution so far has been to drop all but basic setup and force the supervisor to ask before every op - if this is not the design choices, can this be run in parallel? If so, hand it off NOW.
I'm still iterating this workflow, but less setup for all the minions plus handing them work that may be incomplete/ broken is caught and fixed by the minion and its own qa gates.
This can mean a number of minions end up replicating the same fixes, but in general the time cost of that is small Vs the supervisor working in parallel instead of too sequentially.
This seems to be a good example because things like the menu, high score boards etc are common, but the games are distinct. Then there's the artwork which requires decisions on look, and for coordinating.
The Qwen 4B model is multimodal so part of the AC is to view the output - I've a robust anti AI-look QA chain for that I've been using elsewhere, e.g. no floating parts, consistency, obvious missing fingers etc etc.
The longer term plan is to do some llama.cpp refactors specifically for some target hardware I have and implementing slightly different novel architectures I'd like to try (one I did already targeted CPU inference, which I did using 3 agents with specific roles; main planner, QA for planner, and benchmarking/environment handling)
The implementation was 85% of the speed of the original maxed out on my hardware but performance scaled with CPU core count whereas the original implementation plateaued. Unfortunately the break even mark seemed to be around 30 - non HT - threads.
I suppose I should look at that one again, since the increase in cores did not linearly drop off performance e.g. due to memory contention.
> Qwen 3.8 27b is the supervisor
>
> Qwen 3.5 4b are the 6-15 minions it controls
>
> Gemma 4 e4b is the validator for the supervisor.
I just use Opus 5.5 and don't think about it?- Cold starts impact, context length issues, task lifecycle management
- Inefficiencies in delegation, necessitating workflow patterns for small projects (big AIs hide this problem until you scale and they hit the same issues).
- Limits of the AI would be harder to find or notice (e.g. where time - and cost - is being spent needlessly).
Some people are resource constrained.
I too have started using plan mode less, mainly for two reasons: 1. I plan my prompt more carefully and think about architecture up front 2. I found with more recent powerful models, the clarifying questions were generally not useful because it was pointing out issues it obviously knew the answer to and would have resolved in implementation anyway. So essentially they became time-wasting and anxiety inducing for no good reason
The mental shift I made was to not accept poor understanding of a code base on my part when writing the prompt - if I don’t understand it, I can’t predict what assumptions the model will make even on a basic level.
I used to tell myself that plan mode mitigated that, but I usually ended up mentally glossing over the generated plans anyway.
That just resulted in pure anxiety-driven engineering, where I’d often spend extra cycles verifying what was built and worrying about the design.
So invest the time understanding the system, at an appropriate level. That level will change over time as models get better.
> I used to tell myself that plan mode mitigated that, but I usually ended up mentally glossing over the generated plans anyway.
At a certain point I realized I far preferred delegating the first implementation pass to agents and reviewing the _code_ instead of a plan up front. If you had asked me this question ~2 years ago I'd think anyone would be insane to hand off discretion like this, but frontier models are just so _good_ nowadays.
Let's say you launch 10 working sessions in a day, would you rather:
* Review 10 plans, and _then_ review 10 PRs, and _then_ maybe adjust the approach on 1-2 and merge the other 8
* Review 10 PRs, and maybe adjust the approach on maybe 4 and merge the other 6
The first option just feels like needless attention for the sake of feeling in control; but the reality is that these models are becoming just as good if not better as us humans and our reckoning is here. Our codebase has enough linters, hooks, and guards such that generally the agent just follows our blessed patterns already, so what is plan mode actually doing beyond giving me a false sense of security? I'm still going to read the code anyways.
The final form became: understand → act → inspect → clarify → adjust → act again
But he says that "understand" isn't really fitting (or something like that), so we are at: act → inspect → clarify → adjust → act again
That can be understood as "observe, orient, adjust, act" - but with other words.
That is the method to use when you are IN the sh*t rather than removed from it and making some theoretical plan.
I thoroughly enjoyed the read <3
For example, it’s a great hook in the process for agentic review.
Get a second agent to look at what will be implemented and check it for inconsistencies, check it against whatever decisions were made or provided previously in the chat, or against whatever technical rules you’ve written out for your project, before going ahead. It surfaces a large number of opportunities for refinement, and generally pushes the output closer to the direction you’re looking for.
If you’re not using those either… I pray for your codebase.
Also, since sometimes an AI can be a busy beaver, I added instructions to not edit any checked-in files if the prompt contains a question.
Never start coding befire I say codenow, and add datetime. Before patch explain in plain English what it does. Each file has datetime added to the top of the file, and updated if exist. Each file has a backup copy, created with name.datetime and checked after patch vs backup file, and again with new .backup file for logic vs new backup
"Should I start with the spec, or go straight to building it?
Unless you're working on a super crucial piece of engineering, you can probably get away with going straight to building it. Even if something doesn't go the way you intended, I find that it's usually faster to correct the agent later, once the initial implementation is in place. It's more of an iterative approach to building and I feel that it is less cognitively demanding.
I still would appreciate a "read-only" mode. It's not uncommon that I start a harness ONLY to explore and understand the code and I don't really want one typo to have it off building something, or even to save a plan document.
For instance, in games where I work, we often need to manually test work out by playing or by using tools in ways that aren't feasible for the AI to do. In that case, getting the human in the loop between steps is an organized process when it's following a staged plan.
https://github.com/samelie/claude-plugin-pnpm/tree/main/skil...
That’s very much coming from a desire for improved quality than understanding though
Much of what distinguished a knowledgeable AI dev a year ago is now an anti pattern.
npx skills add mattpocock/skills --skill=grill-me
or npx skills add imbue-ai/blueprint
Or anything similar.I want it to go and read stuff but also invoke some commands and generate reports and discuss on the results. Latest GPTs in codex plan mode always go.. write a plan. Shocking, i know, but thats not what i want on every turn when im in that mode.
This is the core question and it is wrong. Because it focuses on the software system rather than the problem you are trying to solve. A better version would be:
How do humans maintain a coherent mental model of the solution they want to create?
You need to have a picture of the whole process and everything related, not just your architecture and code.Engineer, meet business school.
I'm using Astra for some homelab work and I tell it I want to do something and then it burns through a giant percentage of my 5-hour allowance (on the £20/mo subscription) doing that in a really weird way. I've had to tell it to stop zooming off and doing things, calling loads of tools and looking up websites, and just have a quick conversation first. It's much better now, but I've basically just reimplemented plan mode via AGENTS.md.
I get zero dopamine rush from this. I get a dopamine rush from building something or figuring something out myself. I don't get a dopamine rush for generating thousands of lines of code that I then have to try to understand.
“discuss your plan with me before implementing anything”
theres your plan mode
I mean, it is, and if that were the workflow I’d be very fed up of it by now, but the plan mode in Claude’s VS Code extension has supported select and annotate directly since at the very least early 2026.
That’s the amazing realization?
Brooks was wrong. We have a silver bullet now.
I don't necessarily use plan mode explicitly for this, but I might turn it on if I think something destructive is coming up, like a schema change or some other devops thing where claude may decide "oh we copied it all over so I can delete it from source" "oh I'm sorry I didn't catch that error code and assumed success, your data is now gone" lol. This happens a lot less but I've been bit right on the face by it before.
Are plans the wrong abstraction? You can use a big enough plan to spam tickets into linear and have claude burn them down. I find it increases subagent accuracy.
It's been an absolutely fantastic way to set tokens on fire and watch them burn, and produce pretty much nothing of value.
Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way.