Hacker News
new
past
comments
ask
show
jobs
points
by
lostmsu
2 hours ago
|
comments
by
JacobAsmuth
1 hours ago
|
[-]
1e21 flops is hilariously wrong. for reference the llama 3 8B model (
https://arxiv.org/pdf/2407.21783
) used 10 times that many flops. This model is 350x bigger in total params and 12x bigger in active params and was trained on 3x the data.
reply