undefined

points

[-]

TurboQuant is not a step change, it's more of a smaller incremental improvement to KV quantization, and possibly (unsure) to quantization more generally. I'm actually more positive about SSD weights offload, which opens up very large local models for slow inference (good enough for slow chat) to virtually any hardware or amount of RAM.

by gedy7 hours ago|

prev|

[-]

I'm definitely looking forward to that, as I really want people to control their own tools.