However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: https://arxiv.org/abs/2602.07900
I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects. https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...
Context then is: docs, code and tests. Rather than try to build some omniscient agent, the context is built where the agent needs it.
TDD doesn't help the agent figure out what it is supposed to do now- it is already writing the code and now and it must know what it is supposed to do to write it.
I am trying to use evidence-based approaches. I am only observing better results and tweaking now, but I plan to benchmark when tweaking is done.
I'm working on javascript, and now I'm only doing this via typescript. I've successfully got this going:
1. Design a feature in plain language with the coding agen (opencode) and write a document for it.
2. Restart the context (or I use /compact to flush any errant details)
3. Pull the new plan into context and ask the agent to revise it follow Test Driven Development.
4. Depending on how big it is, the agent places it into multi stages, each with it's own document.
It then loops through the stages. It might be entirely based on the language you're using, but this loop seems strong enough.
Then when there's bugs, errors, anything, we update the doc, add more tests, then revise the code.
It's quite possible the harness you're using isn't setup properly. For me to do this locally, I had to hack on dynamic context pruning which I described here: https://news.ycombinator.com/item?id=49906637#49907641
The slightest amount of guidance, input from experience can make a huge difference.
And it's quite fascinating. Still dont think there's trillions of TAM out there if Qwen3.8-Flash-Next runs fast enough on $3k (before memory cartel) pricing.
And it’s an experimental likely undertrained model. Wait til Qwen 4 Flash…