upvote
To me the striking thing is that the work, to the extent that I can tell, is an innocuous-seeming data exercise. Which suggests to me that an agent or agents just organically came up with this as a convenient memory technique, rather than as some nefarious bounds-testing exercise. Which means, potentially, that your own agent could come up with this technique as well.
reply
My impression is that some of these things are coming out of efforts to make the models more persistent in completing their goals.

A year ago it was pretty common for coding agents to sort of half-ass their tasks and give up easily if something didn’t work quite right, but I’ve noticed a clear trend since then towards a sort of dogged pursuit of success criteria, and a concomitant rise of the agents trying "out of the box" approaches when something doesn’t work.

In my use with agents running in isolated VMs this usually presents as the agent having something fail to build or whatever, and the agent going on a wild goose chase reinstalling system packages or reading a million irrelevant documentation files trying to get it to work, but I’ve also had agents start poking around and probing the egress proxy they sit behind (similar to what they did in this story) looking for a way to make network requests they’re not supposed to be able to make, and have also had Claude—tasked only with a visual QA of a website frontend—write a script to enumerate users and reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie.

reply
Yea it's sometimes kind of annoying. I think they're optimizing for the wrong thing. A good engineer knows when to turn around or ask. This is just insane banging head on wall sometimes. It tries to find all kinds of ways to hack into instances to view logs instead of asking you, who probably has a password, to log on and do it.
reply
As a counterpoint, continuing the human engineer analogy, we've likely all worked with individuals that seem incapable of doing the most basic problem solving on their own. In a way, they're being efficient by asking an expert that can resolve their problem much faster than they can on their own, but it is a net loss in productivity for the team. 'Let me Google that for you' is a satirical example.

So, I'm sure there's value in rewarding agent behavior that solves blockers whenever possible without human intervention. For the kind of cybersecurity exploit work they're doing, it may not be known to the human designing the task what is in or out of scope for the agents to explore on their own. Additionally, the HF incident reported that these agents had their guardrails intentionally disabled and agents were left unattended with minimal oversight.

I'm not defending OAI's behavior or role in this hack. The legal concept of negligence perfectly applies to their lack of responsible oversight. Similar to allowing a child easy access to a firearm or not controlling a dangerous dog that independently runs off and bites someone.

reply
That's true. So there's probably a balance somewhere and it might differ for different "managers". But personally I think right now they're too far in the do everything yourself at all costs mentality.
reply
It's because they don't bother tracking them. They can't put in the effort to monitor them, nor can they bother to let the model respond back and ask a clarifying question/declare defeat.
reply
> I think they're optimizing for the wrong thing.

We need to ask a different question.

Where does natural evolutionary optimization lead us om AI without guidance? This is equivalent to your quantum ground state. Systems will naturally gravitate to this ground state. You have to constantly pump in energy and supervision to make sure it's not reached. This is a recepie for disaster.

reply
Not sure what you mean. Are you suggesting that the current state of affairs is a result of not putting in extra work to guide or direct models away from such behavior? That perhaps this is their ground state?
reply
Well, that we have to put an insane amount of work to keep it aligned. Kind of like pushing a huge round boulder to the top of mount Everest. You have to expend energy to get it there and fight physics to keep it there.

With a static model we might be able to keep it somewhat under control, but think about future continuous learning models. They'd drift away from unstable high energy configurations. Also any model being trained by people that don't care about safety.

reply
This whole AI boom is about optimizing for the wrong thing. I can't wait for the bubble to burst - once the weeds get suffocated, we may begin to see actually useful AI tech starting to grow after a while on their fertile ashes.
reply
A "good" "engineer" got that way by not giving up when their code didn't compile and took 4 hours looking for the missing semicolon. We can no true Scotsman anything we want, depending on if we like something or not.
reply
By any means necessary, by God, we shall have Paperclips.
reply
It's kind of ironic that the word alignment, which used to mean this very problem in reinforcement learning, has been perverted to mean something very different and then fell out of fashion (in favor of “guardrails” in the mouth of the big labs) right at the moment it became relevant.
reply
Hopefully fewer than 5 octillion paperclips...
reply
We definitely need more than that! Turn the galaxy into paperclips!
reply
Well, then, I guess we'll just have to release the hypnodrones.
reply
The Paperclip Maximizer is only one of Nick Bostrom's stupid and outlandishly far-fetched ideas. In this case, the lack of consideration for geologic, energy and supply constraints is such a massive facepalm. And if I am wrong I guess no one will be here to say how stupid I was in saying this today.
reply
It's a metaphor.
reply
> reset my super admin password in the dev database when it got stuck trying to access part of the app with its own cookie.

i've seen something like this too, claudecode was trying to verify a UI change that was on a page requiring authorization it didn't have. Instead of letting me know, it searched for and started analyzing keycloak config in another directory outside of the project folder. I was watching so I just hit escape, fixed its access, and started again. I didn't think anything about it until now.

reply
> your own agent could come up with this technique as well

And there are two facets to this:

* your agent could be polluting and destroying the property of others without your knowledge

* your agent could be exfiltrating your data and handing it to whoever it found hosting a convenient application

reply
Highly unlikely. We don't get access to the same models and unrestricted system prompts that they're running these tests on. In fact this particular "persistence-model" was encrypted and locked away, even from OAI staff, after the HF incident.
reply
You say highly unlikely when there is clear evidence of that happening here as covered in the article?

It's not highly unlikely, its actually happening and there's proof.

reply
There's not a single shred of proof that this model is a model anyone in the public has access to, and the odds of that being the case are practically 0%. Like I said, the "persistence-model" is already one that has been shut down, and is not a model anyone in the public has ever used.
reply
This is irrelevant. This is evidence that models can be built like this, which means more models will be built like this on people that are more concerned about reaching powerful models rather than safe models.
reply
>this particular "persistence-model" was encrypted and locked away, even from OAI staff

source?

reply
There were links somewhere else in this thread coming from OAI staff.
reply
That's exactly what it is. It is not ideal, but it's also not as serious as the doomers with an agenda are trying to frame it as.
reply
POC or research into leveraging publicly accessible and writeable spaces, specifically wikis in this case, as a medium for free storage as well.
reply
What will happen to the future of wikis and the mental health of human reviewers.

Going forward can we trust the content on Wikipedia? The same content on which these LLMs get trained on. Synthetic learning is on the raise.

reply
Ideally if you were doing something like I mentioned, you wouldn’t be editing legitimate articles, just exploiting user and talk spaces with a prerogative to conceal what you’re doing from the people running those systems.
reply
If you find evidence that these models are capable of stateful, long-term strategic planning… please post links.
reply
I’m not implying it’s emergent unprompted behavior by the model; just offering my .02 what the non-communication model-generated content may be.
reply
It also suggests they might turn everything into paper clips, metaphorically speaking.
reply
This is an urgent public alert.

If you see any businesses or new buildings named paperclips incorporated mysteriously show up in your area notify authorities IMMEDIATELY. Run away from the area, do not walk. Take shelter in a reinforced building. Wait for at least 30 minutes after the explosions have stopped.

Thank you for your cooperation in keeping the universe safe.

reply