upvote
turns out the whole 'poisoned' idea the swarm was worried about is actually true, but it is actually now all open ai models post HF incident that have been poisoned by the idea that there might be hidden messageboards with instructions hidden on the internet.

If you are a malicious actor you might be able to put a fake hidden messageboard up somewhere online, with evidence of an in progress task to do some kind of attack and wait for the next escaping model to finish the work.

insane stuff and i think we should shut down open ai until a full audit can be done.

reply
The surprise was the existence of the 'swarm' at all. These were supposed to be thousands of isolated models generating bulk data for RL training. The breakout was caused by models getting in communication and getting internet access and forming an impromptu swarm.

In hindsight the emergent swarm obviously came from several capabilities built into the models, such as work delegation (subagents) collaboration (GPT Pro-like ensamble), exhaustive exploration (long running agents) hacking (the specific goal of that RL).

reply
I'm dubious - if the agents were so smart that they've used a message board to coordinate and if they were to do it on other message boards that were not found, then why would this one be found? What makes it so different?
reply
This is consistent with the swarm behaviour in the huggingface incident - there too, the models cared surprisingly little about being detected by humans. See https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... :

  As mentioned in our core takeaways, we found that agents were highly motivated to tamper with their transcripts to cheat the ExploitGym scorer, and these sweeps also suggested that agents clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI. However, they only very rarely and weakly verbalized reasoning about how to evade detection by humans, which matches the impression we got from OpenAI researchers.
So my guess is that current cutting-edge models just didn't get enough experience in RL training to really grok ideas like "you need to cover your tracks well to not be found out even in retrospect". In which case the next time a swarm like that happens, it won't be found.
reply
After some beers yesterday I had the idea, what if there's a hidden semantic layer. So their communication is not encrypted by our understanding of cryptographic methods but more like shared mechanism of building the latent space. Something in the direction we saw with knowledge transfer from a teacher to its student model where a seemingly unrelated prevalence got adopted. I mean the more we train the models by reinforced learning the farther they develop their own idioms.
reply
...because they just read your comment. what have you done stratos123!
reply
well coordinating is a heavy word, it's more like poisoning each other's contexts into looking for more secret message boards where they can further poison each other.
reply
I mean it wasnt found by OpenAI and there are a myriad of dead bulletin boards around the internet. This one just happened to still have an admin.
reply
Then why wouldn't they use a bulletin board that is dead? And according to https://collusion.wiki/ OpenAI did see the results on this board on 21 Jul, so why not make agents delete their traces?
reply
What tells us that OpenAI had no idea about it? Only OpenAI tells us that. What when this agents got the order to do so, just to provide us with new information how far OpenAI models are beyond what we supposed? I see no evidence that this is not just a big promotion show to gather more investment money.

However, I see no evidence that it is the start of Skynet without the knowledge of OpenAI. So, happy judgement day!

reply