upvote
> load all the LLM parameters or KV cache in RAM and exclusively let it perform GEMV and let it rip.

Won’t you have a bunch of extra reads/writes via the CPU because these DIMMs won’t be able to compute matrix multiplications?

reply
And now the 96 memory slots need individual cooling.
reply
Which might be easier since the surface is larger
reply