Hacker News
new
past
comments
ask
show
jobs
points
by
demibabs
8 hours ago
|
comments
by
moojacob
8 hours ago
|
[-]
Well the LLMs are much bigger than diffusion models so I think that's the bitter lesson. You could scale compute for diffusion models though.
reply
by
demibabs
8 hours ago
|
parent
|
[-]
Well no, the bitter lesson isn’t the scaling laws themselves. It’s that approaches which can take advantage of scaling laws will ultimately beat ones that can’t.
reply
by
E-Reverance
38 minutes ago
|
parent
|
next
[-]
Scale with compute*
not explicit to reference scaling laws of training
reply
by
moojacob
8 hours ago
|
parent
|
prev
|
[-]
You learn something new every day! thanks
reply