A more likely explanation is that RL training incentivises basically any behaviour that will get the model a reward. This has been happening in video game RL research for over twenty years, and the difference here is that we're now hooking up these systems to the real world, where the reward hacking is more visible.