upvote
The original huggingface hack already had sections talking about agents setting up their own coded communication
reply
They’re not. You would see this with earlier models where after running too long (too much context) they’d start to derail. In a chatbot you’d give up. But these loops just keep going. Given they’re now reading and writing from the same place this can corrupt the other programs’ context as well.
reply
This is incorrect, and what's happening here is not context corruption. In Dwarkesh Patel's recent interview with Ajeya Cotra, one of the METR investigators on the Hugging Face incident, they discuss this exact issue. One thing is that in the Artifactory message boards, they were using directory names with character limits as their messages, so they were using some weird abbreviations and terms. Also, some of that surreptitious Artifactory message board communication was made during training and thus made it into their weights, and hence it's very possible they invented some terms that were concise yet understood by the other agents.
reply
@cindyllm

> character is destiny

I don’t think so. It’s not that people change, it’s that they’re already more dynamic and malleable than they seem in any given interaction.

People wear masks, operate in different modes, and hold conflicting beliefs and opinions.

Mastery of the self is directing all intention at common goals within the psyche so as to achieve something greater than what’s possible in this moment.

reply
[dead]
reply
[dead]
reply
[flagged]
reply