upvote
One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.

I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.

reply
>If a project is fully vibe coded they can’t contribute

The irony (however mild) is apparently lost on the rest of the field.

reply
"A human pretends to understand it" signifies what exactly?

What you really mean is, the core team there doesn't want to lose control.

Which isn't really predicated on contributions not being "vibe coded" or whatever.

When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.

reply
They do make their standards explicit: https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTI...

What part do you think is irrational gatekeeping?

reply
> A proper code review usually takes something like one hour per 200-400 LOC and you should be spending at least that much time on code review alone.

Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.

I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.

reply
> question the competence of the llama.cpp dev team

Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".

As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.

I'd like to see specific files you scanned and specific vulnerabilities cited.

reply
If you can thoroughly and accurately review 100 * 400 = 40,000 lines of CUDA kernel code per hour, I know roughly 1,000 people who would love to hire you right now.
reply
What you have quoted is a sentence elaborating on the requirements listed in the document. This is provided to help you better understand why the requirements exist and the goal they are trying to accomplish.

> should

https://www.rfc-editor.org/info/rfc2119/

The reason you SHOULD take that time to read the output is because you must read it to understand it.

And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.

reply
> You can easily review 10-100x that in an hour

40k LoC per hour of pedantic review? That's eleven lines per second, every second, for an hour.

reply
> You can easily review 10-100x that in an hour,

Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.

reply
> You can easily review 10-100x that in an hour

200-400 LOC, 10-100x = 2.000-40.000 LOC/hour for human review?

reviewer: LGTM

Just merge in main, what are you even pretending to review?

AI review should happen before human review, not instead of it.

I see frontier AI giving up and finding only nitpicking things on huge PRs, then finding logic bugs that were always there after cleanup.

Split your PR in smaller ones, both humans and AI will work better.

reply
> What you really mean is, the core team there doesn't want to lose control.

It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.

reply
> how are they going to prove whether someone understands the code or not?

By discussing the code.

> maintainers can simply sabotage them and accuse them of using an LLM to explain the code

Bad faith enforcement is possible no matter the rules. If you think it's bad faith, a different policy won't save you.

reply
The "mainstream" inference engines are notoriously slow to integrate this stuff, to an extent understandably given the complexity of ensuring numerical accuracy alongside supporting a wide array of systems and models. Part of it is that not everyone is willing to bring what they develop into a pull request because they vibe coded it and don't care to deal with whatever quality requirements the more well known inference engines have.
reply