I think the reason the pelican example is great is because it's bizarre enough that it's unlikely that to appear in the training as one unified picture.
If we picked something more common, like say, a hot dog with toppings, then the training contamination is much harder to control.