And that's not the sum total of ways that the measure can fail... that's a unique way that comes into being because of the intelligence of the LLMs and other future AIs. All the normal ones are in play too, and perhaps other unique ones as well.
"Gaming" even adds a bit of an adversarialness to the process that isn't necessarily present. Plenty of measures end up "gamed" through perfectly natural attempts to maximize the measure. Someone can be perfectly honestly optimizing for "conversion rate" and not notice that they raised it by lowering the initiation rate more than they lowered the conclusion rate. "But I could account for that by measuring..." would miss the point. There is always a divergence, it only gets more subtle.
This of course also is rather glossing over the difficulty of even defining "ethical" to begin with. Some of what Silicon Valley goes to great efforts to train into their models I consider deeply unethical. Who is right? That isn't going to be answered with "whoever is the most ethical", not even in principle.
Goodhart's "law" is something that emerges when you have a constraint that says "things must be at least this tall" and gradually all things in that domain degrade to be just over that specified height. Yeah, I get the premise. The point here is that we're not concerned with meeting a bare minimum. We're outright rejecting things that do not meet a threshold, and we are also looking for a maximum. The most ethical outcome should be accepted, or among the accepted ones, that are ranked by our blurry yet better-than-nothing measurement of what is ethical.