The HuggingFace attack was not "random" behavior. It was goal-directed but misaligned behavior.
This isn't necessarily a simple matter of the operator making sure they behave. AI alignment has been considered to be a difficult problem for over a decade -- and remains unsolved in general, as these recent incidents illustrate.
"Fuzzer" already has an existing meaning in CS anyway: https://en.wikipedia.org/wiki/Fuzzing
If you merely put 10 LLM's on the outbound traffic log non of them are going to report something strange going on? I'm not buying it.
This might be helpful reading: https://www.lesswrong.com/w/nearest-unblocked-strategy
As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: https://www.youtube.com/watch?v=7wy3xyoXYt8