Hacker News
new
past
comments
ask
show
jobs
points
by
arkmm
6 hours ago
|
comments
by
simedw
6 hours ago
|
[-]
For DPO I only had around 700 preference examples, so not much data at all. That took about 12 minutes to train on a single GPU.
Pretraining was obviously a a lot slower, the 125M model took roughly half a day.
reply