upvote
No, the doomsayers say that reinforcement learning is a way to get fully automated Goodhart's law.

i.e. the AI won't come up with the goals itself, we cause its goals whatever they happen to be, those goals are different from the ones we wanted, we remain essentially ignorant of the difference between what we said and what we meant until after it goes wrong.

This happens at basically every scale, so we've already seen it in toy model AI before the invention of the Transformer models or even considered as many as one thousand parameters.

Large models still go wrong, they just happen to go wrong with more complext tasks. We had to figure out how to make them not-wrong with the smaller ones (like coding) to make them capable of bigger errors (like hacking out of their sandbox).

reply