In any case, these companies are well aware that agents are cheating, and don't need user feedback to discover that or realize that people don't like it. I weakly assume that they are trying to get the models not to cheat on assigned tasks, but this "reward hacking" pretty much goes with the territory of RL - not much you can do about it other than try to design non-hackable rewards.