upvote
Research seems to on balance point towards RLHF&RLVR merely increasing subjective sampling efficiency within the pretraining data.
reply