upvote
Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)

Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.

reply
> ~1000 bytes per context token per user

Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb

reply
The smallest I have seen is DeepSeek 4.1 flash at 890 bytes per token.
reply
Highly recommend the RCO-GSQ quant by ITSA btw. At IQ3_XSS it is within one point of the fully unquantized model.
reply
will try, thank you for the pointer!
reply