upvote
> It may surprise some people here to see that Andrew is warming up to using LLMs to discover bugs (inspired by results from SQLlite) and considers it a tool on the path to getting to bug free software.

Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered. Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?

reply
> Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered

I don't know about this particular case, but if I saw someone report a bug and as evidence claim they had X Y and Z LLMs verify it I would be pretty upset. If you're going to use an LLM to make a replication, just do that and give me the replication, don't point to your notoriously error-prone tools as though they lend your report credence.

It's in a similar vein to people who reply to questions with "well Claude says: <chat transcript dump>"

reply
> It's in a similar vein to people who reply to questions with "well Claude says: <chat transcript dump>"

or substantially worse: "<chat transcript dump>"

reply
Human reproductions are notoriously error prone whether ai assisted or not right? I’m not sure what the analogy is between “llms are error prone” and “thoughtlessly copy-pasting something from Claude” is.
reply
Human reproductions don't have to be bad. It feels just as justified to push back on a bad bug report whether human or AI and say "I don't have enough to go on here".

If a report is improved and becomes actionable, that's great.

reply
Hey, we're all the result of human reproductions.
reply
Might be my favorite comment ever now.
reply
I feel very torn as a maintainer on this, since the only way to respond to the increased noise from AI has been to have AI do the research for me to extract all the links and line numbers I used to have to find by hand to explain why the PR needs more effort to be completed. But I also will be highly dismissive of any submitter who just posts AI text without cleaning it up first. It feels hypocritical, but the alternative is that I just can’t respond to most people instead due to limited bandwidth. I mark if something comes directly from the LLM though and to try to express my degree of confidence in its claims.
reply
I suspect they can't easily tell the difference between a high quality LLM based contribution and a low quality one.
reply
Note the distinction:

- using the LLM to find (possible) bugs and a human confirms it by testing, reviewing, etc.

- using the LLM to find and confirm the bug without the human confirming it

reply
Neither of those are my understanding of what happened: a human found a bug, and then had an LLM verify their theory of the issue. To me, that's just extra due diligence and pretty weird to use as grounds to ignore.
reply
> Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?

In the state of the tagged video he says still not accepting AI submissions until a certain set of preconditions is met. So... No?

reply
So "warming up to" is moving at a glacial pace, in both senses of the word
reply
He specifically says in that presentation that they are open to LLMs helping them get to bug free, but because the language is still in flux, they would rather prioritise bugs users actually find rather than those found by LLMs, essentially with the intent of unblocking people rather than wasting time fixing things that may need to be fixed again or be wasted work come the next update.
reply
My understanding of the situation is that the user did find a bug themselves, because it literally broke on their own code, and then they just had the LLM help them verify their theory of the bug that they already had.
reply
Shouldn't everybody already know that while using AI to find bugs for oneself is amazingly efficient, using AI to submit bug reports for others is quite the opposite? Burden of verification and all.
reply
Exactly, the main issue isn't the LLM creating or confirming the report. Rather, the maintainer has no idea how it was prompted, and the results may be completely wrong. Some LLMs also have a bad habit of trying to please the user, confirming their biases.
reply
There's a difference between Opus 4.5 and Astra 6
reply
With regards to whether when it finds a legitimate bug, it should be given weight?
reply
With regards to whether a blanket ban has a payoff
reply
Why would the difference matter? A bug report is valid or it isn't.
reply
A lot of bug reports aren’t valid, or handle a case that can’t realistically happen or realistically be handled (eg what do you do if you detect a crash while the prior crash is crashing and the logging pipe is throwing errors)
reply
Sure but this is true regardless of the model. Either it found a bug or it didn't—who cares about the intermediary steps or tools used so long as the reporter can reproduce it?
reply
citation needed
reply
Nobody refuses penicillin because it came from mold in a dish. If an AI found a cure for a disease, people would ask one question: does it work?

Bug fixes should get the same treatment. A patch is either correct or it isn't. Projects that ban AI-written fixes outright are asking "who wrote this?" instead of "is this right?", and users live with the bug in the meantime.

I get why maintainers are fed up. Review time is scarce, and they're drowning in plausible-looking garbage. But that's a problem with low-quality submissions, not with AI as such. Require tests, require a human who vouches for the patch and will answer for it, and ban repeat offenders. Then hold every patch to that same bar, whoever or whatever wrote it.

reply
Warming up to LLMS... explicitly _after_ full branch-coverage and fuzzing.
reply
He'll reevaluate the LLM bug-finding situation in a few years.

https://youtu.be/zwi5b5xSsKA?si=w6zZN6AtIvJS9MxP&t=2084

reply
I personally think using AI is of no problem for certain cases. It is absolutely pissing off when someone tries to generate slops that too verbose to review only to increase the complexity of the codebase meaninglessly.
reply
[flagged]
reply
It boggles the mind that people believe anyone that thinks differently from them is automatically utterly stupid.

Actually, on second thought, it really doesn't. Indoctrination and herd mentality are quite the things.

reply