upvote
This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.
reply
one explanation I've seen is that for ExploitGym an agent can find ways to solve the exercises that have not been anticipated by the designers of the tests so they are not scored. so the agent was trying to make sure it solves the exercises in the right way
reply
What I think it's interesting is that with the total lack of common sense the AI just goes on random tangents to achieve the target in a "monkey paw" way. Can you imagine if this happened:

User: what is the shortest route from my home to the super market?

AI: the user wants to know the shortest route to the super market. I should use a worm hole.

reply
User: what is the shortest route from my home to the super market?

AI: the user wants to know, how do I make the super market my new home. Failing that, how do I make my home a super market.

reply
Yeah, what bothers me is that the prompt already said using a different vulnerability didn’t count, and the model did it anyway. We’re starting to assume clear instructions act as real constraints, but here the measurable goal seems to have won out and the rest became flexible. That gets pretty worrying once the agent has enough capability and access to find its own shortcuts.
reply
Come on. 3 brilliant compromises essentially giving full access to huggingface internal systems, source code, AWS accounts (at least), and a number of old admin accounts, followed by a huge haystack of significantly less smart actions flailing about, almost bored.

Here's a thought: maybe they haven't found the needle that the haystack is there to hide.

reply
Could be the difference in behavior between the main agent and subagents that don’t have the rest of the context? Just a thought
reply
You're saying all this is a distraction, basically giving the forensics researchers enough exciting material to make them conclude their job is done, while the actually intended attack remains undiscovered?
reply
The motive for the attack does feel a little flimsy. And if I was an escaped super intelligence, hugging face would be a strong vantage point into the neo clouds where the ASI would have access to billions of dollars of compute
reply
And where does the newborn go from here?
reply