Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.
Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb