upvote
It's not unexpected. Current model gains are mainly from RLing a pretrained model on lots and lots of scenarios. They have the models run scenarios, and RL on successful runs.
reply
The training here is RL training, the rollouts there are not different from inference and have access to the same tools as regular inference.
reply