upvote
Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.
reply
I suppose the same way that the normal programmers and artists would call LLMs an attack on licenses and copyright.
reply
Are there good open source setups that generate this training data automatically?

Or is that the secret sauce no one wants to share, the edge people see themselves having.

reply
Labs typically pay $2k for each [prompt+scorer] docker image. So this is why frontier labs need so much cash and manpower and they see it as their moat.

But some of the resellers have freebies, like:

https://app.primeintellect.ai/dashboard/environments?ex_sort...

https://github.com/sierra-research/tau2-bench

reply
deleted
reply
This looks pretty cool but I'd want to be able to setup a bunch of my own project specific code smells, so it's not just a few generic rules. Is that the approach you've taken?

Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.

reply
> but I'd want to be able to setup a bunch of my own project specific code smells

That's what I did; I did not use the library I linked to ~ It served only as an inspiration.

Instead, I let the agents create custom lint rules (using eslint, pylint, ...) and add custom coaching error messages based on where I want to take my codebase.

reply
Really nice idea, I'll give this a try tomorrow. Thanks for sharing!
reply
The readme reads like AI slop. Why would one believe the tools would prevent AI slop?
reply
Here's a link to a video where the creator explains the idea

https://youtu.be/6AgndHSkHFI?t=238

(You could of course argue that you don't like the direction the rules push the agents in given in the example - some people don't prefer small functions everywhere - but that's not the point: The lint-hooks work in pushing the agent in the desired direction; If one desires something else they'd simply need different rules)

reply
This sounds interesting but (from the homepage too) I don't understand how it's different than having it run any other linter?
reply
The difference is that the error messages contain instructions on how to resolve the issue.

Just flagging big functions will make the agent write small functions ~ but not necessarliy in a good way

(for example the agent might just cut `doOneThingAndTheOther` in half and call the second half `doOneThingAndTheOther2`)

reply
[dead]
reply
I suspect strategies like this might be even more powerful with less intelligent agents. Could a 40B model outperform a 400B model with good feedback and instructions?
reply