Code review can go much deeper now if you use AI to aggressively attack a PR combined with your human insight. Same for planning a feature:
Can I consolidate this logic to a shared function? Does the error surface to the user and are there any gaps? What preexisting functionality is affected by this PR? Can this query be made more efficient?
LLMs are great with focused questions, up and down abstraction layers and across all kinds of concerns. Stack up these focused concerns into a rich understanding of what you are doing or writing.
Understanding is the real output, code is the byproduct.
This only applies to bad code. A formal programmic language beats an informal language anytime. Why on earth would you translate an unambiguous formal language to English?
> Is there a risk of data leaking here?
> Are permissions enforced downstream of this function?
If you trust AI to give you that answer, you are going to run into serious issues.
Why on earth are all the books about programming, science and math written in English (and other natural languages) instead of a formal one then?
Why on earth did comments even got invented?
Luckily genius programmer who understands formal programmic language and does not need weak comments to clarify why the world is not perfect, fixes the stupid looking code to be beautiful, opening company up to some really super neat, seemingly random bugs in one browser that nobody will understand for several months.
tldr: one important reason comments exist is to clarify why code you might expect to look another way looks the way it does, of course if the only reason it looks the way it does is this is the way the AI wrote it I guess I agree the comment is unnecessary.
I agree. This is much better advice than what the article tries to do, which is to give best practices around how to use or not use AI to code.
Happy to share the implementation or package it up if anyone is interested.
Do I understand when I just let AI regurgitate the function and behavior of the code?
You still have to check the work and shouldn't trust it blindly, but in my experience it's pretty damn good.
> You code and you have the AI check your work. Treat it like a reviewer. That’s my advice in a sentence.
Bro you're giving the same advice as the article
Well, I can't ask it for an overview of a system it can't see.
also
> I don’t follow all these religiously.
So if you need an overview of a system, you can and should provide access to the entire codebase. Even the rule itself says "preferably", not "don't ever". If you have different needs, you can always just look at it as a useful guideline. It's just a blog post by some rando on the internet, not the epiphany of AI coding's Ten Commandments.
If students of the 70s or today pretended they didn’t exist up to a certain point when they needed them to move forward, like bioinformatics or something, they 100% would be better off. There is plenty of research on off loading thinning providing a worse understanding of the material - eg side rules proving a better understanding than calculators.
I think AI might be more of a problem, at least for some skills because it is not just doing a small unit of work. its more like the problem of doing more complex calculations. An example I have come across with calculators is students learning statistics without knowing how to calculate a variance, because they just put the numbers into a calculator. There is a failure there to learn the concept - they do not really know what a variance or SD is. I can imagine AI doing a lot of that.
You learn stuff by heart, so that when you mess up with the calculator you have a feel that stuff has gone wrong
> the term “vibe-coding” suggests a kind of laissez-faire attitude where you don’t really care about the outcome and you’re just having fun. That’s what the phrase meant when it was coined, but the world has moved on. In many companies professional programmers are using AI in such a way that it’s impossible to imagine that they are also reading the resulting code in detail. This is what modern vibe-coding is. Deferring to the AI, not worrying about the individual lines of code, and keeping an eye on whether the code passes its tests and throws up any problems in production.
This way of working is the only one that justifies the trillion-dollar bet on the AI industry. I agree, this method should be called vibe-coding.
I don't care which words we use, but they should separate this sort of empirically-minded work (versus the rationalism of typical development) from just dropping a few sentences into the prompt and crossing your fingers.
Yeah, I understand that this testing-oriented constraint method is what AI coding enthusiasts are pitching. Sure, it's different from pure vibe-coding as originally defined, but nobody is seriously suggesting professionals should do pure vibe-coding anyway. It's a strawman. The word fits.
Core modules: Coded by myself, AI reviews and AI to discover/learn.
Stuff I don't care about Craft: API layer, CLI layer, Smoke tests, Integ tests - Dial AI heavy, and lighter human reviews accordingly
Obviously takes a lot of patience and very easy to sin, but on good days, its doable.My current approach is "ask for very small diffs" + "review them very carefully".
I'm not working a job though, I'm working on a multiplayer game.
Main findings: The frontier models can't reliably modify Pong without breaking it, so their skill appears to be quite domain-specific. (OK, to be fair, neither can I half the time!) This is probably because they are "time blind". I had one model try to test a game by running it at 0.1 frames per second and shoving each frame in the vision API...
If you leave any room for a misunderstanding, they will laser in on do it and do the stupidest thing possible. If you're not checking everything carefully, you will discover this later, and you will cry.
Formal proofs, oddly enough, do not improve the situation: they will simply prove mathematically that the absurd and pointless and backwards implementation is completely without defects. (It obviously does help within an implementation, though.)
They can't formally prove what the hell you meant when you told them to build something. That job remains frustratingly human!
Current dissatisfaction: (1) Harnesses are designed for super bloated codebases (i.e. designed to load as little context as possible) which make them pretty clunky for small repos and small edits. (I had a Surgical Edit Tool I need to bring back...), (2) Current LLMs are anal about verifying the most trivial change, even without prompting, even if it's impossible for them to verify it because they're blind so they start measuring pixel data in Python... Both of which eat up Speed and Cost, taking the work even further from Realtime/Interactive to Tedious/Sad.
Numerous attempts have been made in past 70 years, but we landed on the compromise that is high level programming languages with their (mostly) well-defined syntax and semantic rules.
If you look outside of the world of programming you'll find that the same is true for every craft and profession that requires precision and unambiguity. Cooking, logistics, military, carpentry, mechanical engineering, law, medicine, etc. all have their own domain specific languages for describing tasks, procedures, and tools. Heck, even sufficiently large organisations have their own internal languages and terms for that reason.
SQL is the worst offender here: LLMs are very good at remembering its awful syntax, but very bad at writing nontrivial queries.
I wonder if that trick about prompting it to write like a 5th grader* would help here. Keep it simple!
*A trick which allows it to pass for human 70% of the time...
I've recently settled on Deepseek4 for one of my projects, and had it review other code generated by other models, and .. yeah, that was quite eye-opening. Someone in the frontier-models part of the world is definitely paying attention to the AI slop generated by the other models, because having one AI checking the results of another AI has been quite fruitful, lately.
The security argument is the strongest part of the post, and I don't disagree with it, but what it buys you is the review, and a review catches what you missed, whoever typed the characters. None of the ten dogmas follow from that.
My general approach is to design beforehand, do an adversarial review with AI, socialize it with humans (if needed), generate a plan, and start working item by item. Always keeping me, the human, in the loop (not that `/loop`), going through the steps generating code. Finally, a manual review and one AI adversarial review of the feature branch in a clean context, going section by section manually and discussing anything relevant, and off you go.
Writing code by hand feels great, but even local models can generate fine code. Deterministic linters and quality checks are what keep the quality in line. You can always modify things as long as you're in the process, but at the end of the day, you'll review more than you write. We're closer to being the assembly line's inspector than the crafters we once thought we were.
50 years ago the main barrier was access to hardware - computers were big, expensive, and not easy to come by for average people.
30 years ago, the biggest hurdle was access to resources: you had to spend a fortune on books that quickly became obsolete or had to wait for your local library to get them for you.
Today, you basically have everything - up-to-date free reference material on pretty much everything in an instant, cheap and accessible hardware (no need for high-end stuff), great online tutorials and interactive courses, message boards, the works.
But then you also have a machine that completely removes the need to put in the effort required to actually learn and hone your skills. Human psychology always seeks shortcuts and more discipline than ever is needed to not give up and have the machine do the thinking for you instead.
Proficient developers are getting pressured into outsourcing their craft and skills to the machine, thus deskilling quickly. Juniors never get the chance to become proficient in the first place - either because they're not getting hired to begin with, or because they succumb to the siren song of the machine...