Doesn't the LLM do that, if you want it to? I do use it to write code, but the more striking ability it grants me is a means of understanding legacy code far faster. I can ask "what actually causes this branch to be taken" and it's usually right. Ok sometimes it's not right, but I'm not always right either even after I spend tens of minutes reading code.
It's also exceptionally good (i.e. fast) at looking through git history to figure out where/when a certain behavior originated, which can be difficult (time consuming) if code is continuously being refactored.
Sure you could use it to vibecode. I don't, rather the opposite, I understand my own changes better. But I still fear that I'm going to be obsolete as soon as it figures out what questions to ask. And I'm unwittingly training it to do that.
I find forcing it to visualize things immensely helpful. I'm usually studying git diffs but when working of a big feature or refactor that can just be too hard.
I've never been very pro "visual programming" and always hated UML et al, but part of me is starting to wonder if it's time for us to give it another serious go.
I do it differently, I focus on better recording what the user wanted, the so-called "user intent". To do this, I record all messages typed by the user since the start of the project, whether 3,000 or 10,000 messages. An LLM can churn through them in 10 minutes and derive a fresh, up-to-date interpretation from the raw data. This can be used to judge whether the implementation has diverged from the intent, or, in other words, to realign the code and tests. The messages the user writes are usually designs or corrections, a very rich, compact signal. If the user struggles with something, it could result in a tool, a skill, updates to the project docs, or new tests.
Now, when someone sends a working PR in, even high quality and well tested, they may actually have no idea how it works.
I deeply relate to this. When engineers were writing all the code that meant every part was deeply understood by _someone_ on the team, and they could valuably contribute to maintenance and further development. It wasn’t perfect, people leave, people forget things, etc, but the overall coverage was high and valuable.
Now, every agent-produced MR introduces code that is deeply understood by _no one_. It’s the “original developer left five years ago” problem, but now growing on every single new piece of code. Reviewing doesn’t give you the same depth of understanding, and the continually increasing impulse is to just approve, maybe nudge it about some isolated enum types or something, but don’t take the time to understand it, just keep the train going.
But then what happens when something breaks and the cloud agents are down…
Part of woe is that once you've reviewed, validated, and comprehended a piece... Later gets casually mangled by some other LLM-generated urgent change.
Sounds, like you already mention with org-mode or similar ones) like literate programming (https://en.wikipedia.org/wiki/Literate_programming) or jupyter notebook.
I think the solution is still code, just at a much higher level of abstraction. Maybe a start is kind of typed ADR or FSM that guides (constrains) the agents. I believe more type checking guarantees will be more and more important for agents.
Building a basic X11 window manager is almost a one shot prompt.
Modifying a UI toolkit to make it work with MSAA/IA2 is simply not possible.
There's a lot of room for deep work left... for now.
If I were trying to accomplish this particular goal I would first consider what the agent could see. In particular does it have an accessibility inspector of some kind? or even NVDA hooked up with NVDA Remote so that it can actually see the implicit a11y tree for the toolkit it is working on? My email is in my profile and I would love to chat about this.
Could be just defining the methods without filling them but depending on the mood I code more by hand or less.
Which is frankly exhausting to do when you have to keep up with the rate of LLM changes
Or at least that's the current model I'm playing with.
I could change a whole UI completely in 30 minutes to something fundamentally better but then 30 people would all wake up and be upset they weren't consulted and need training for it. That training and consultation will take hours and hours. And probably generate feedback - some of it correct, some of it misguided - that needs to be human negotiated, taking more hours. The effective maximum rate of change is limited so dramatically more by other factors than the technical implementation that we have to completely redesign process now to cater to those factors.
We are in a weird space now because most of the process is still built around a presumption that technical implementation is a lot of work. The main reason to be upset that you weren't consulted about a change is because there's a presumption that you will be stuck with it - ie: it's a lot of work to change it back. But it isn't a lot of work, it's effectively free. All this is just living in inertia right now.
A is what you're used to, B is what I recommend. If I can convince people to start using B instead, I can look at the metrics for A and conclude that it's effectively dead, and then I can remove it.
It's working out for me, but maybe not a fair comparison because I only have something like 15 users.