You’re probably limiting the frontier models by being specific.
I do feel there is a limit to LLM-s today. I did not try, but I doubt it will work if I would tell it "make me a web browser that is bug free, perfectly secure, works as efficient as possible on my architecture and has best UX for me personally".
Knowing where that limit is, is as hard as it always was knowing how fast a team of engineers was going to do a project.
I felt there were limits 4 years ago. I couldn't get even an 8k context window.
It's starting to feel like the important gaps are the only ones left that need to be closed before there aren't any left. It also feels like next year they will start meaningfully closing.
I do feel there is a limit to human-s today. I did not try, but I doubt it would work if I asked literally any programmer I know "make me a web browser that is bug free, perfectly secure, works as efficient as possible on my architecture and has best UX for me personally".
Hell, I'm willing to bet they'd fail at this task even with an unconstrained snacks and kombucha budget. Humans have a long way to go. My job is safe.
Theres lots of bad Pytorch code, just go ask george hotz. This isnt the magic you think it is. And nobody wants to hire the guy that just points llms at things and says make it faster. Things have value because talented humans make them. Its why a luxury coat is worth more than the linens that make it, or one from walmart made by a machine. This will 1000% apply to programmers. I think OP will be fine.
Pointing a frontier model to a 10 or 100MLoC codebase and saying "make this code fast, make no mistakes" doesn't work. As an experiment I recently tried this with a relative small (500KLoC) codebase and it got stuck on believing that the primary cause of slowdown was the database not using a connection pool. (Which was completely irrelevant for this specific code.)
In general the OP is right, the people who get the most value out of LLMs are veterans.