Hacker News
new
past
comments
ask
show
jobs
points
by
seemaze
17 hours ago
|
comments
by
chmod775
17 hours ago
|
next
[-]
Smaller models are also generally faster, so thinking "more" may not matter and may even come out ahead.
reply
by
celrod
16 hours ago
|
parent
|
[-]
If Q4 takes less than 1.3x as many tokens as bf16 or q8, it could still end up being faster, given how decode tends to be bandwidth bound. The kv cache was still bf16, so a few ops are the same between quants.
reply
by
conmod278
17 hours ago
|
prev
|
[-]
Computer science has known the tradeoffs between memory and compute since ages ago. The same could be reflected here.
reply