You don’t need documentation […or…] memory
…oh heck no. Easy to miss at Claude and Codex’s default detail level but if you read your actual session transcripts, you are almost certain to notice the LLM solving lots and lots of the same little problems over and over again. The code IS the documentation
First, lots of things cannot be learned from the code.Trivial example: LLMs struggled with QA on our app. They didn’t know how to find the seeded test accounts. They would create new ones and do it wrong, or find the seeded test accounts but not know their passwords because they were encrypted, so they’d change the passwords but not tell the other agents. Shitloads of tokens burned. It was a no-brainer to just give them the credentials in a skill that gets loaded when they do QA.
You could say that’s an environment issue, not a code issue. But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.
The days of “you are the world’s greatest Python programmer, write good code” or whatever are over (if that style of prompting even worked in the first place) and as I said less is more. But, still….
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)
Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.
It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.
When it is not, it has to be reasoned out and does not work well for documentation.
Other things that code and tests alone do not capture well:
- promises (as in Promise Theory) made to other parties. Claude already has PT in its training data.
- Constraints-inducing-properites, as in Roy Fielding / Christopher Alexander. While tests, and property testing can capture properties, there is no formal connection to the constraints that induces fhem. By constraints, I am not talking about business requirements and business value — those are better understood through Promise Theory. I am talking about things like at-least-once delivery or total ordering (from append-only constraint). Claude already has Fielding’s dissertation and Alexander’s works and ideas in its training data.
- grammars, as in pattern panguages (not just patterns) a la Alexander / Fielding are also not captured in code alone. These tell both humans ans AI how to extend a pattern, and how to identify anti-patterns (when they violate a constraint-inducing-property)
- LLMs are trained with many different worldviews and bounded contexts at the same time, and is very capable of translating across it. However, these need to be soelled out, otherwise it would talk in whatever it infers
Specifications written for the exact way components are wited together run into that stale doc problem. Although it takes much more human attention and token burn to describe things in terms of pattern language and promise theory, it becomes easier over time. The actual implementation plan tends to fall out more cleanly when all those other stuff are at least considered. This is where I have been spending most of my time.
I liked this advice when humans wrote code. Though even then I'd urge people to write meaningful commit messages that capture the "why" of what they did, so no one tramples their intent by mistake.
But not sure it works in an age where most code is LLM-generated. Especially if that code is not even reviewed by humans (irresponsible or not, it's happening), and commit messages are also generated by AI. I think something is needed to separate "what did the human operator intend" from what the agent went and built.
I do agree that this gets way overengineered. My approach has been more or less what you stopped doing though - committing all our timestamped "plan/implementation docs" and "investigation docs" that document what the user wanted + empirical findings, and making all prior session transcripts searchable. It's seemed mostly helpful? For whatever reason I haven't run into many staleness problems so far.
Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.
In my codebase it is difficult to get agreement on comments and documentation so rather than rely on it I adapted. One of the first things I did when I succumbed to agentic development was to point codex at the code and ask it to generate a high level description of where important files, such as our public API, reside, what the hierarchy is, what the code does, etc. In my case, this level of documentation is fairly static if I avoid implementation details. So now I have a handful of agent files in my tree and it seems to save quite a few tokens and improve my results. I frequently have other devs ask me how I get such good results when doing agentic reviews of their changes(always my first step now before I start my human review). I also include instructions in the agents files instructing the agent to maintain the agent files if any relevant changes are made. It seems to work quite well for me.
All code is written under constraints, and most constraints live outside the code.
I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.
We have heard this nonsense from the "we don't need to write comments, code should be self-documenting" types for decades. It was wrong in that context, and it is wrong in this one.
Code tells you how a system works. It does not tell you why it works that way. That is what memory is for. It exists so that your AI does not keep undoing past decisions when it writes or refactors code.