Hacker News
new
past
comments
ask
show
jobs
points
by
reasonableklout
11 hours ago
|
comments
by
NitpickLawyer
11 hours ago
|
next
[-]
It's not unexpected. Current model gains are mainly from RLing a pretrained model on lots and lots of scenarios. They have the models run scenarios, and RL on successful runs.
reply
by
porridgeraisin
8 hours ago
|
prev
|
[-]
The training here is RL training, the rollouts there are not different from inference and have access to the same tools as regular inference.
reply