I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.
The irony (however mild) is apparently lost on the rest of the field.
What you really mean is, the core team there doesn't want to lose control.
Which isn't really predicated on contributions not being "vibe coded" or whatever.
When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.
What part do you think is irrational gatekeeping?
Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.
I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.
Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".
As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.
I'd like to see specific files you scanned and specific vulnerabilities cited.
> should
https://www.rfc-editor.org/info/rfc2119/
The reason you SHOULD take that time to read the output is because you must read it to understand it.
And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.
40k LoC per hour of pedantic review? That's eleven lines per second, every second, for an hour.
Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.
200-400 LOC, 10-100x = 2.000-40.000 LOC/hour for human review?
reviewer: LGTM
Just merge in main, what are you even pretending to review?
AI review should happen before human review, not instead of it.
I see frontier AI giving up and finding only nitpicking things on huge PRs, then finding logic bugs that were always there after cleanup.
Split your PR in smaller ones, both humans and AI will work better.
It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.
By discussing the code.
> maintainers can simply sabotage them and accuse them of using an LLM to explain the code
Bad faith enforcement is possible no matter the rules. If you think it's bad faith, a different policy won't save you.