They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.
And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.
Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).
But maybe that’s on us, AI doesn’t care about all these special cases, it’s not debt to it as it will simply read them all when making changes. We’re obsessed with quality and what code is supposed to look like but those are human standards, AIs evolve to look at this complexity as a single picture, they can simply see through it so what is spaghetti code to us is merely some code to them that works as it should and is efficient. It’s interesting we can see how the two things drift apart, you would think at some point AI generated code should explode but it hold together unreasonably well in most cases…
Each time you encounter a shitty thing you hate, add a new 'review type' / 'thing to watch out for' and just ask your agent to add it to your hooks for you. This works well with Claude at least.
I have about a dozen or so hooks that run on every integration branch my agents write that review for all sorts of things from correctness to spec, performance improvement opportunities, modularity, analysis of any dependencies added, 'definition of done', UI/UX, etc.
I recently told Claude it should run the whole suite of reviews twice. I will probably go on and proceed to having it run like 5 times eventually idfk.
But the more you start asking your agents to modify their own behavior, using the native solutions offered by Cursor, or Claude, or Codex, the sooner you'll start to feel better about the results.
In your workflow, who implements the review feedback - the review subagent or the code-writing-subagent?
Do have a baseline styleguide (like Google's Go style guide) for the review subagents, or is it entirely the subjective things and specific corrections? I remember 6 months ago it seemed like piling general "good taste" code advice into AGENTS.md was considered bad.
Do you move between harnesses or have you gone all in on claude? I've bounced between claude/codex/omp, maybe to my detriment.
The biggest things I've struggled with are models having taste. For spec writing I was having a lot of issues with them making statements that were interpretable in a superposition of ways, eg "we'll do XYZ with entities that support and need it" when there's 3 possible entities and the model hand-waved at exactly the wrong tokens.
I added claude AGENTS.md guidance and also later a memories with encouragement to be unambiguous (with short but good examples) and "no coined shorthand". Now I'm getting a marked increase in specificity, but it's places that don't matter (claude explaining existing code to itself). I'm having difficulty controlling the spew of new text, but feature writing/research still gets fuzzy and lazy around the difficult underspecified aspects of the problem/feature.
Do I just keep dumping examples into reviewer subagent context and have them rewrite and simplify the research subagent spew?
I repeatedly have "Risks and gotchas" sections have a whole paragraph dedicated to things that don't matter, and then a single sentence bullet point that's actually a huge problem when I dig into it.
Do I have a generic hook to make subagents to review bullet points and size them commensurate with impact? If I tell them to "have taste", do they have taste?
I'm just trying to get good code done. I hate these things.
Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.