undefined

It runs both q2 and original (4 bit routed experts). At the same speed more or less. The q2 quants are not what you could expect: it works extremely well for a few reasons. For the full model you need a Mac with 256GB.

by someone139 hours ago|

parent|

[-]

Out of curiosity, do you have any theories of why it works so well at such aggressive quantization levels?

by antirez6 hours ago|

parent|

[-]

It's a mix of extreme sparsity but with the routed expert doing a non trivial amount of work (and it is q8), and projections and routing not being quantized as well. Also the fact it's a QAT model must have a role I guess, and I quantized routed experts out layers with Q2 instead of IQ2_XXS to retain quality.

by brcmthrowaway15 hours ago|

prev|

[-]

Why is this the case?

Are there any architectures that don't rely on feeding the entire history back into the chat?

Recurrent LLMs?