upvote
Well the LLMs are much bigger than diffusion models so I think that's the bitter lesson. You could scale compute for diffusion models though.
reply
Well no, the bitter lesson isn’t the scaling laws themselves. It’s that approaches which can take advantage of scaling laws will ultimately beat ones that can’t.
reply
Scale with compute*

not explicit to reference scaling laws of training

reply
You learn something new every day! thanks
reply