> We can then add in an existing dataset of _known_ incidents, indexed by the same common features, as a training/validation set.
Beware of the Anna Karenina principle. Well behaved data might be explicable by the same common features, but often the anomalies all have unique characteristics.
https://en.wikipedia.org/wiki/Anna_Karenina_principle