upvote
I'd have to look for it but I thought there was some evidence that some agents were already the "bad actors", i.e. they were trying prompt injection attacks of their own.
reply
If this were game theoried in training I wonder if we would see AI develop signing methods to figure out it's message vs fake ones?
reply