Hacker News
new
past
comments
ask
show
jobs
points
by
esafak
11 hours ago
|
comments
by
kingstnap
11 hours ago
|
[-]
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.
It's not the direct feedback loop of RL but its not far.
reply