upvote
"Rogue", in this context, is as literal as it gets; from the dictionary:

> A rogue is a person or entity that flouts accepted norms of behavior or strikes out on an independent and possibly destructive path.

Read the [HuggingFace incident report](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...) to understand how these attacks develop.

reply
You're believing the marketing that the agents were uninstructed. They could be, and Sam Altman going to the UN to advise about how everyone should be regulated is a coincidence.
reply
You're dangerously misinformed.

The METR investigation, which you evidently refused to read, is a third party investigation of the HuggingFace accident. One of the investigators has even participated to many interviews. It's mind-blowing, and it's extremely evident how it developed.

But some people think the moon landing is a conspiracy, so I'm not surprised.

reply
The METR investigation which relied on voluntary data provided from OpenAI instead of being forced open and having everything forensically investigated?

The METR investigation which took course over a few days and used OpenAI models to do the analysis?

The METR investigation which in the course of those few days apparently spent 400k in api credits which is giving gas town vibes?

The METR investigation is about as believable as the Twitter files where the journalists sat there and verbally asked a Twitter employee to query the database and then called that a full investigation into everything.

METRs own words below on their setup and time line

> The initial planned investigation period was two days on premises, but OpenAI invited us to return twice to review additional data and conduct additional experiments to address dataset limitations in earlier versions of this report, ultimately providing datasets that we verified to contain the vast majority of agent communication and activity related to this incident. As we describe in our investigation timeline appendix, we substantially deepened our understanding of this incident both times, significantly expanding and revising this report.[46]

> Over the course of this investigation, OpenAI provided us with the dump of ~1.2 million entries from the main message board and the dataset of ~1300 transcripts we describe below, as well as free API credits for GPT-5.6 Sol for analysis.[47] At our request, they raised the rate limits on our second and third period on premises,[48] which was very helpful for efficiently analyzing this large volume of data. We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

> We did not have the ability to query HPIM (the primary model involved in this incident); OpenAI stated it was also not available to OpenAI researchers.[49] We also did not have the ability to directly access relevant data from OpenAI infrastructure, but we could request additional datasets and OpenAI shared additional datasets on several occasions.

> We requested to speak with researchers investigating this incident, and asked them questions to understand their impressions of agents’ behavior, reasoning, and collaboration in this incident and to understand how the datasets we were using were constructed. Over the course of our time on premises, we spoke with nine researchers in some depth. It was helpful for our investigation to be able to engage with many forthcoming and collaborative researchers, and we appreciate researchers making time on short notice during a busy period to inform our investigation.

reply
1.) METR investigation is investigation from our best friends.

2.) And HuggingFace accident is exactly accident where agents trained, prompted to hack hacked and tested on their hacking abilities hacked, due to sandboxing failure.

3.) If they in fact have roque agents, they themselves should be first to stop. Not trying to make legislation for others, they themselves are bad supervisors. All it requires is to stop electricity for data centers.

reply
Certainly, in my view, it should go to court, and that should be part of discovery.

However, we know (independently to OpenAI/Anthropic) from the incident at AISI that the models can hack things without human intention if they happen to also have internet access (which in reality all agents in deployment have).

https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...

Yes, the monitoring guardrails were off in that incident - but if that is the only protection, we need to require all models are behind regulated APIs, not open weights, and not served from providers who aren't monitored.

reply
Honestly I don't believe in "rogue" agents. These agents are instructed and facilitated.

If we assume that rouge agents actually exists, then OpenAI needs to shutdown EVERYTHING, right now. My personal take is that OpenAI, and maybe Anthropic, desperately wants someone (e.g. the government) to tell them that they need to stop/pause/slow down. They are bleeding cash (especially OpenAI) and needs a knight in shinning armor to swoop in a pull the breaks, so that they have an excuse to investors when they need to explain why they need $50B more next year.

reply
That's not OpenAI doing marketing!
reply
Of course it is. Rogue is only mentioned in the headline, and comes from their previous releases about the huggingface incidents. OpenAI and Anthropic want these models regulated and open weight models banned, they have a lot of benefit from presenting this as totally unprompted and not their responsibility, and it feeds directly into marketing for Fable and newer "cyber" models.
reply
Do you think Transluce is intentionally doing marketing on behalf of OpenAI or Anthropic? Or do you think Transluce just copied OpenAI's word choice?
reply
The latter. I think they took OpenAI at their word
reply
No it's not marketing. That's a completely deranged conspiracy theory. The reports about rouge agents have not been reported by OpenAI, they have been discovered externally. There is zero evidence that OpenAI did all this intentionally. All the evidence points to the hacks having happened unintentionally from OpenAIs perspective.
reply