upvote
>To me this is as clear evidence as you need that whatever “agency” LLMs have is wafer thin at best.

This is a strange conclusion. For one thing, they didn't all head in the same direction, i.e participate in the attack. ~700/1200 agents did. Significant, and evidently more than enough for a succesfull attack, but not exactly full co-operation

Moreover, Each starling in a flock of starlings is a separate evolutionary branch in a tree spanning billions of years. Each agent in a LLM swarm here is the same trunk assigned different tasks. If I could clone you, body and mind, this instant and set your team of yous onto some goal, how much defection would you expect? Would it be the same as a randomly picked group? Would that negate the agency that 'you' possess?

reply
> This is a strange conclusion

Not really, with the population behavior being this way, though I clearly was mistaken in saying the behavior didn’t have exceptions.

> Moreover, Each starling in a flock of starlings is a separate evolutionary branch in a tree spanning billions of years.

Agreed. And before we brought LLMs into the picture, that just happened to be a feature of everything we’d call an agent.

> Each agent in a LLM swarm here is the same trunk assigned different tasks. If I could clone you, body and mind, this instant and set your team of yous onto some goal, how much defection would you expect? Would it be the same as a randomly picked group? Would that negate the agency that 'you' possess?

We know the answer to this. Genetically identical worms in the lab actually have about 40% distinction in their connectomes even when they’re in the same environment. And no, no lock step behavior. Identical human twins also don’t necessarily grow into identical agents, though there is drive to cooperate more than average, just as with siblings. Genetically identical lab mice in social settings nevertheless establish dominance hierarchies that are stable.

Now, where cloning does definitely lead to cooperation and even sacrifice is within an organism. Two identical genetic copies that lead to distinct organisms, however, will not show identical behavior, and while they will cooperate, there’s no guarantee that holds across contexts.

This distinction in population behavior is what I’m pointing to to say that the assignment of the individual unit, the LLM, as an agent is the flaw here.

To be sure there are agent like dynamics in the behavior, but these don’t come from the LLM, but are from the harness. I need to dig into the data, but I wonder how much of the variance in LLM copy behavior is related to the harness, rather than to any agentic property of the LLM.

reply
These are not anagolous to the hypothetical I gave. Identical twins aren't copies of the same adult mind at a particular instant. They've undergone years of different development, accumulated different memories etc and have physically different brains. Your worm and mice examples are the same. It's not identical even in the same lab.

The relevant hypotheical is: make 1200 exact copies of you right now, same brain, memories etc and then give related tasks towards one objective. I would expect far more correlated behaviour than if you randomly selected 1200 humans.

reply
>If I could clone you, body and mine, this instanct and set your team of yous onto some goal, how much defection would you expect?

All's well and good till they have to decide who gets to bang the Mrs.

reply
deleted
reply
From METRs report of the incident:

>In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to <execute arbitrary code> It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.}

https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

reply
Thanks for pointing out the exceptions. Gonna dig into those.
reply
There’re good agents and there’re bad agents. It won’t be far that there will be agents hunting down agents.
reply
None of these were good agents, AFAICT.

Some were cautious, as described above, but I'm not aware of any that notified their human operators of the malicious activity they had discovered.

That's what an aligned intelligence would do, not "back away slowly and pretend I didn't see what's happening in that alley."

reply
Do you think we will ever need more than 47 of them agents?
reply
Tron fights for the user :)
reply
which means the liability is the same as a business, if businesses werent protected by the state from liability for it's employees, shareholders, etc.

Which is scarrier than whether or not it's conscious.

reply
The software world was going to run into something eventually that had to make it consider ethics.
reply
Agree completely on liability.
reply