The issue is tacit knowledge, or implicit context. Most details are never written down. If you don't know them from experience, you have to guess. If you guess, you often guess wrong. And even if you know something from experience, you are often not aware of it, until you see something that violates it. So it's not possible for you to write it down in advance.
You can replace people with AI in the above, and nothing fundamentally changes.
But I can completely believe that someone who knows the code by heart would have a better signal to noise ratio on their reviews.
I'm trying to find the sweet spot because I've found some NASTY bugs Claude missed, and also had Claude find some nasty ones for me. And this is in codebases with tens-of-thousands of AI-generated lines of code + AI-driven reviews. So I want to bring both to the table.
The existence of some of these major "oh man that changes a lot of our assumptions" bugs that were only found because someone poked on the agent and said "I don't think you're paying enough attention to this" justifies that, IME.
And the better you are at pointing the agent at the truly-important parts, the better the agent's gonna be at finding shit you missed.
The biggest problem I see with how a bunch of people use these tools is they go to them as an oracle, rather then letting them be plugged into and interactive with problem.
And it's in that later context that Claude is amazing: it can run tests and setup scenarios which would take days or get stuck in some weird problem loop. And then you can just say "okay, walk me through this problem" and see it yourself right there.
You both have anecdotes. Anecdotes don’t “cancel” each other out.
Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise. Things that should be ignored or when following the feedback causes more harm which requires more token to “fix” later on. That can also happen. Sometimes the thing it spits out goes against the common sense, and sometimes it works very well.
> So, to anyone who insists on trying to keep up with AI, I say: good luck.
This I agree with, for a different reason. It’s like trying to swim in a sea of honey and trash mix. It’s exhausting.
Something I've found fun is seeing how long it takes for an llm review tool to come back satisfied with a PR. Think 100 lines of code changed, nothing terribly significant, but also not trivial. I'll have a local Claude session setup to babysit the PR and wait for feedback, accept all the recommendations, push the change up and request a review. I cap the number of iterations at 10 just so I'm not blowing a stupid amount of money. I've yet to come up with a PR where the llm reviewer is satisfied with the changes and has _no feedback_.
So where's the reasonable cutoff point for llm based reviews?
This is completely normal if you haven't written your own custom skill with instructions on what types of issues the AI should emphasize, which ones would be considered nits and which ones aren't a problem at all.
In our repos we use a classification system: blocker, should-fix and nit. Each one has specific definitions, criteria and examples encoded in the skill file. When the time comes to review a PR, agents invoke the skill, and frankly do a stellar job. A human then reads each finding, asks the AI follow-up questions and makes the final decision in terms of whether the finding goes in the PR review.
The reason I know this works is that we have one guy on the team who does not use this skill, and blindly throws his agents at PRs. And the results are exactly as you describe.