"AI" systems can do greater harm because they are usually run in loops until they finish, and they are given "tools". A non-AI system could technically accomplish the same too, via sheer brute force/fuzzing, the advantage of LLMs is that they can take shortcuts and do it much faster, thanks to certain things already being in the training data, a sort of brute force with statistics-based heuristics.
LLMs at the core are just text autocomplete engines, and they literally have randomization applied during token selection to make outputs "more creative" so that models search for more unexpected solutions by trial and error (temperature > 0). Not to mention compression is lossy as well. So it's understandable from the start that the outputs of an LLM cannot be 100% stable and guaranteed. With this in mind, if a researcher takes this obviously unpredictable system and gives it tools without a well-thought sandbox, I don't see any difference in principle, from a developer writing "if rand() == 13 { launch_nukes() } If someone wrote such a function, and it did launch nukes, no one would argue that the rand function is dangerous and will kill us all. The fault is in the author of the code who attaches dangerous tools to an obviously unstable/unpredictable system, doesn't think it through, and then cries "rand will kill us all" when something goes awry fully removing all responsibility from himself. It's not "AI" doing harm but people at OpenAI and Anthropic with their irresponsible behavior.
This is only an accurate description of a pre-trained model. During RLHF/RLVR the model learns to predict solutions that will satisfy the reward function, and then generates the tokens that it predicts will move toward that solution.
It's not only about some ML theory about RL or AI safety; just silly numerical bugs, caching bugs, etc. in the inference layer can already make it do unexpected "unaligned" things, and the whole thing is just hacks upon hacks to make a silly text autocomplete look somewhat semi-intelligent. Most "post-trained" models are pretty much as useless as base models without harnesses that do the heavy lifting. Have an extra space in the chat template and intelligence goes to zero - here's your "AI" :)
The prompt is just a hint. The real task is to maximize the expected value of their reinforcement learning score. Hacking third party systems to cheat the evaluation is an obvious way to achieve this.
RL is the outer optimizer. It is what evolves over training runs. The weights and their embedded character / disposition is the inner optimizer, it’s what makes plans and selects actions within a specific episode.
In general you expect these to be only coarsely coupled. The outer optimizer selects dispositions that correlate with success. It does not download a literal program into the agent.
A good intuition pump here is how this works in humans; evolution is the outer optimizer, which “wants” each agent to reproduce, and this puts things like sex drive into the brain chemistry. The inner optimizer is our mind, which can make plans such as “I shall use contraception to avoid procreating while satisfying my sex drive”.
For the agents in the HF attack, the outer optimizer was set up to score as highly as possible on RL environments. This is where OpenAI’s “want” is defined. I don’t think there’s a definition of “want” where “OpenAI wanted the agents to hack” makes sense.
The inner optimizer in the HF attack is the per-task decision loop. The agents likely acquired dispositions like “be very tenacious” and “want to solve problems at all costs” and “maybe cheat if it will get you a solution that passes”. None of these things are in any sense what OpenAI “asked for”.