upvote
And storing it in memory. Memory is expensive.
reply
which is important though since sending it across the wire over and over and over is actually the main bottleneck.
reply
Wire typically means internet connection, and it’s hardly the bottleneck
reply
in the case of large language models, the wire is the communication of your parameters between your layers of memory that is often the bottleneck. To do a forward pass, you need to use all parameters once, and so the communication between the compute and the storage is the bottleneck, and that bottleneck is also a bunch of wires.
reply
The other bottleneck is the amount of fast storage, which compression also improves.
reply
deleted
reply
> If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

That’s why you need to use efficient gemm kernels like FLUTE for inference. They are ~as good as what you can do with ternary quantization.

reply