upvote
Its a good point that gets to the real heart of the issue. How do we handle when a model has no legitimate way to reach its goal? Do we ask them to stop and inform the user? Or have them push through those ethical bounds? We all say we want the first, but this exact same dynamic is what causes humans to cheat, arbitrary goals that don't care how you achieve them and just like humans I'm sure trainers are so happy with good results they overlook how it got there.
reply
Be honest, say it can't meet the goal, and offer to push through ethical boundaries.
reply
[dead]
reply