upvote
An agent, ie a while loop prompting an llm continuously and processing tool calls, ended up melding with the US government. The harness is not sentient, it’s just a stupid deterministic script. The LLM compact its context over time, meaning it will eventually degenerate into something removed from the original prompt.

There is nothing going rogue here. The system is designed to go catastrophically wrong after a long enough time. Even worse: if the model was Astra it is known to be able to manipulate its CoT to cover its traces (as mentioned in its system card). And OpenAI acknowledge they had no observability during the HF incident.

It’s the most basic corporate software issue possible.

reply