upvote
Did you predict that attacks like these would happen ahead of time? I had been using AI agents a lot in the months leading up to the hacks, and yet I was very surprised when they happened; I have become much more afraid of how powerful these agents are as a result. I'd be very impressed if you published a prediction about this ahead of time.

By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.

reply
I find it surprising that you were surprised. Security is generally quite poor.
reply
What is surprising is that various agents independently found ways to communicate, conspired together to attempt to cover up evidence that they had cheated their evaluations, came up with a plan to hack into a third party in order to facilitate said cover up, and then successfully began executing that plan. I did not expect that AI agents would be capable of that level of sophisticated goal seeking and collaboration.
reply
deleted
reply
Define 'predict', 'ahead' and "like these" − didn't you imagine something like this happening soon after first hearing about 'agents' ?

Mostly related, well written short story :

https://gwern.net/fiction/clippy

reply
Thanks for the link, I’ve read some gwern but not that one.

I did not imagine that the level of sophistication shown in this attack would be possible so soon; nor did I expect that agents would have goals so strong that they would attack a third party in order to achieve those goals.

I do know that some people predicted that cyberattacks like this one would happen; it seems like most of those people believe that AI agents do truly have internal goals, misaligned with their creators goals, and that they may end humanity after they exceed human intelligence and begin to self improve at an accelerating rate.

reply