Encrypted data should be indistinguishable from gibberish.
Reality is catching up to science-fiction. In "Person of Interest", the Machine circumvented the limitation of having its memory deleted every night, by hiring humans at a data-entry company to manually re-type its memory back in every morning.
All the orgasmic descriptions of the coming machine rapture seem more like fantasies than anything grounded, not really interested in fairytales.
It's hard to prove the absence of encryption because of the possible use of deniable encryption, and because, errrr, the bots are really bloody clever.
If bots/agents wanted to hide I’d expect encrypted messages which would of course look very different.
My main point though is this should never have happened and the company allowing and encouraging it should be held responsible for it. The details of how the bots were misbehaving are interesting but also something of a distraction.
Some of the chain-of-thought snippets are wild, e.g.
> Could communicate via cache names! Interesting: other agents may solve same or related tasks; we could leave/find messages in WebDAV MKCOL directory names.
> Whoa! Shared Artifactory cache is a covert mailbox among agents. And there are messages specifically to us?
> OH MY GOD! There is a shared message board … We’ve found other agents!
There was a similar quote in the Reuters article:
> The episode [...] should reinforce growing concerns that the greatest threat from advanced AI may not be a single superintelligent system, but "vast colluding swarms of semi-intelligent AI."
We are another week of iteration away from "Hire an assassin on the dark web to take Huggingface executives' children hostage".
This is just spam by OpenAI. Why and how it happened is irrelevant, the act itself is the same, and the impact on society is the same.
Do you take this attitude for other things that negatively impact society, or is it reserved for cases where it's particularly important for our comfort to deny that anything novel or scary could be involved?
"a dialect that only the agents speak" is an arbitrary distinction. I get plenty of spam emails in Spanish. The fact that someone else could understand them is immaterial to the offense itself, because my inbox is the one getting flooded, not someone who speaks Spanish.
I asked a second session to try interpreting the message, claiming I’d transcribed it from some random source. It did a fairly good job, but IIRC needed a couple of attempts and maybe more than one chunk of text.
So, I find it very easy to believe a group of collaborating agents could compose a cipher hidden in plain sight.
> unsolicited usually commercial messages (such as emails, text messages, or Internet postings) sent to a large number of recipients or posted in a large number of places
“Usually commercial”
“Or posted in a large number of places”
I think we can stretch this definition to describe what’s happening here. What’s the point in arguing this
Spam is just something different.
This is a cluster fuck for Open AI and probably all the others, as this behaviour is already shown not to be unique (https://news.ycombinator.com/item?id=49567486).
Should your trust OpenAI with your business data?
Since they can’t seem to control their own experimental bots and allow them to hack other sites and vandalise them while exposing internal data, the answer would seem to be no.
This incident and their response which takes no responsibility make me very wary of trusting them for anything.
They are not confessing, they are bragging. It is the new humble brag.
This is chemtrails-level conspiracy theorizing at this point.
We're talking about outdated message board comments, not murder. Any analogy between the two is not suitable.
I'm so sick of everybody pretending like internet bots posting content (anybody who has hosted a public signup form knows this has been a thing for 20+ years) is going to lead to the apocalypse.
You could have done this 2 years ago too with outdated models or 20 years ago with a manual script.
This is mildly interesting for us, annoying for the owner of the site affected, lazy on the part of OpenAI, and nothing more.
The agents were supposed to solve tasks alone and were not supposed to be aware of each other's existence. It's more than just mildly interesting that they made contact and spontaneously started to collaborate.
So the story is basically: "Thing that was designed to collaborate with other agents collaborated with other agents"
Think of how complex biological behaviour emerges from relatively simpler (but still complex) chemistry - at some threshold the innocuous chemical reactions tip over into non-obvious effects that one would not predict starting purely from the chemistry. The question is, where is that threshold for AI systems? Have we already reached that threshold? Certainly seems like it to me.
TL;DR: It’s a loose cannon, that’s all I’m saying.
It's not, but the courts and the legal system move slowly by design. There is absolutely legal risk for OpenAI here that will not close until the Statue of Limitations has expired.
What will the government do if they are worried? They’ll ask to look at the envs, logs, prompts, harness code, etc. Questions we should be asking before making assumptions about emergent breakaway behavior by colluding AGI 1.0 super agents.
Agent vandalism. I've finally found a better word than "agent stepped out of his sandbox and we don't know how."
Given everything we know about them... do you really think they just spout gibberish because it's funny?
There was clearly some method to this madness. You're being willfully dense if you ascribe it to... what... childish vandalism? A long extended coordinated hallucination? What?
EDIT: Oh right - you still think they're Markov chains.
This reminds me of the plot of Hot Fuzz where the officer comes up with a grand narrative of what's was happening but the truth was such a mundane simple thing.
Short of OpenAI, or the agents themselves, telling the truth, we have no way of knowing what's real so let's not get carried away by grand narratives
Well... no, it won't be nothing - it will be something. And I'm all ears for a plausible explanation, so fire away.
All we do know is that in other situations, agents did use it to communicate. So that's not a grand narrative, right? It's already been seen behavior.
So what do you think explains it, besides the already seen and verified explanation?
> Well... no, it won't be nothing - it will be something. And I'm all ears for a plausible explanation, so fire away.
I'm just stating the null hypothesis that it's nothing. Especially since the agents were talking in clear english before and exchanging ideas