Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.
Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.
Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.
Leave them to their own devices at your peril. Trust nothing they do.
Yet directly guide them, monitor everything they do, some value emerges.
It’s still good that some real bugs are being patched but what is being reported to the media is so overblown.
In human terms, that's already at least a standard deviation above average person.
There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.
After:
Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.
Mythos turned out to be exactly the marketing stunt it smelled like.
There are others like AISLE who seem to be a bit more successful in finding actual issues using LLMs in some shape or form though, whatever they do differently. Chances are high the secret sauce is not so much about the model being exceptionally powerful which would be bad news for the frontier labs.
Zip-zapping the bouzouki...
Exfiltrating nuclear arm codes...
Thought for 76 seconds.
You're right to push back on that. That's on me.
(my current favorite definition)
That said, I remember trying to weigh the hype at the time of the announcement reading/skimming the papers Anthropic published, recognizing that bugcount alone wasn't super-relevant but also remember being impressed by an NFS bug and a kernel bug that struck me as relevant at the time. So where did that NFS issue show up in GKH's list you showed so nicely above?
It turns out, AFAICT, it's not on his list, but the reasons are perhaps interesting to others so I will post here. It turns out there were two NFS issues this past year conflated a bit in my memory:
* The Linux CVE-2026-31402 NFS heap overflow that could allow unauthenticated memory reads over the network isn't in that list of 79, presumably because it was found by Claude Code, not Mythos months earlier. (I am guessing it's not his "malicious network packet into the middle of the stack" and is a stronger attack being a remote attack.)
* And the CVE-2026-4747 NFS stack buffer overflow that allowed gaining full unauthenticated remote root access didn't show up in GKH's list of 79 because despite being Mythos-caught, it wasn't Linux, it was FreeBSD.
I guess this does match my memory now that I think about it, that there weren't any smoking Linux guns caught by Mythos.
* (I guess there was also a longstanding 27-year old OpenBSD TCP SACK-handling stack integer overflow than enabled remote crashes / Denial of Service found by Mythos.)
There is definitely Mythos hype, but just because it hit the BSD code base more than the GKH-managed Linux code base doesn't mean it was inappropriate to raise eyebrows from Mythos, in particular since "attacks only get better".