This type of whack-a-mole approach is akin to "fixing a bug" by hardcoding a special code path for known-buggy inputs. It doesn't address the root problem of AI misalignment, and doesn't allow you to prevent catastrophes in advance, only patch things up after the fact.
This might be helpful reading: https://www.lesswrong.com/w/nearest-unblocked-strategy
As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: https://www.youtube.com/watch?v=7wy3xyoXYt8