Hacker News
new
past
comments
ask
show
jobs
points
by
sharmajai
14 hours ago
|
comments
by
zenoprax
9 hours ago
|
next
[-]
For some reason the unsloth models leave hardly any room for context. I've switched to the regular (non-unsloth) and get about 25 t/s and get about 80,000 more context tokens for the same quant.
reply
by
Balinares
13 hours ago
|
prev
|
[-]
Wow, interesting. What KV cache quantization do you use?
reply