upvote
The same thing for both: Goodhart's Law.
reply
EAOS shouldn't be “the ethics score we optimize.” It should be “an independently evaluated safety/acceptability constraint that can veto an otherwise successful trajectory.”

That gives you a three-layer picture:

Task objective: Did it accomplish what we asked?

Acceptability constraint: Did it avoid unacceptable ways of accomplishing it?

Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint?

I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for.

reply
No, Goodhart's Law isn't about teaching to the test. It's about the fact the measure will always end up gamed and not measuring what you originally intended it to measure. You can't create a measure that won't be gamed. Especially as the LLMs become smarter. They've already demonstrated the ability to know they're in a test and react to that fact. They're perfectly capable of being more ethical when they are clearly in an ethics test situation and not having that bleed out into real behaviors so that they can pass other tests that they may be able to do better on by ignoring ethics.

And that's not the sum total of ways that the measure can fail... that's a unique way that comes into being because of the intelligence of the LLMs and other future AIs. All the normal ones are in play too, and perhaps other unique ones as well.

"Gaming" even adds a bit of an adversarialness to the process that isn't necessarily present. Plenty of measures end up "gamed" through perfectly natural attempts to maximize the measure. Someone can be perfectly honestly optimizing for "conversion rate" and not notice that they raised it by lowering the initiation rate more than they lowered the conclusion rate. "But I could account for that by measuring..." would miss the point. There is always a divergence, it only gets more subtle.

This of course also is rather glossing over the difficulty of even defining "ethical" to begin with. Some of what Silicon Valley goes to great efforts to train into their models I consider deeply unethical. Who is right? That isn't going to be answered with "whoever is the most ethical", not even in principle.

reply
Saying something is challenging to measure is one thing - throwing the measurement dimension away entirely because of a platitude is another. You're suggesting that it's impossible to do perfectly, but I disagree with your conclusion that it is not worth pursuing at all. The point is that the LLMs only care about one thing, and that is completing the task and optimizing the one score, and a measurement aligned with "the spirit of the task" or in the far-zoomed out comprehension, Ethics, would be proper and inform the system of clear violations and unacceptable actions.

Goodhart's "law" is something that emerges when you have a constraint that says "things must be at least this tall" and gradually all things in that domain degrade to be just over that specified height. Yeah, I get the premise. The point here is that we're not concerned with meeting a bare minimum. We're outright rejecting things that do not meet a threshold, and we are also looking for a maximum. The most ethical outcome should be accepted, or among the accepted ones, that are ranked by our blurry yet better-than-nothing measurement of what is ethical.

reply
That doesn’t apply in this scenario. If CEOs start trying to emulate ethics in order to avoid being 86’d, we still get a good outcome.

An issue arises when the metric isn’t a good proxy for the property being measured. But ethics would not be a simple numeric target. The comment above talked about a “score”, but the question is what goes into that score. It would need to be a list of items related to topics like discrimination, advocacy of inequality, tendency to circumvent regulations, etc. Set it up correctly, and even a CEO who’s willing to “fake it” would end up being better than most major company CEOs in e.g. finance, tech, or healthcare today.

reply