upvote
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.

It's not the direct feedback loop of RL but its not far.

reply