upvote
My first version used plain `mmap`. On the 8 GB M2, a cold 3.36 MB expert took 10 ms with mmap and 2.8 ms with `pread`. The full simulation was 0.50tok/s for `mmap` vs 4 tok/s for `pread`

With `mmap`, OS loads pages reactively as the model touches them. It doesn’t know which experts were selected or when their reads could overlap with GPU work

And common weights still use mmap for simplicity

So, I believe llama.cpp might run it under 2gb, but I assume it will be slower

reply
Any idea if madvise helps? Admittedly I have very limited experience and only on Linux
reply
I tried it. madvise didn't make mmap better than pread. I also tested F_RDADVISE, it helped on short decodes, but somehow got worse on longer decodes. Not very clear why, most likely problem somewhere at APFS and it is closed source and not much docs for it
reply
MADV_SEQUENTIAL might help a bit, but not that much. Biggest problem here is throughput-vs-latency.

With mmap()-ed file, for each pagefault, kernel will conservatively estimate block size to page in, so you'll have a ton of relatively small requests going to SSD. This would be IOPS-bound, and likely under-perform relative to maximum possible bytes/second throughput.

With explicit read()/pread(), kernel & SSD can work with much larger chunks, so it's easier to hit maximum bytes/second throughput.

Plus, with modern CPUs, IO-wait could be efficiently combined with number-crunching. So, if software knows in advance which data chunk (expert) it'll need for the next token, it can load that in parallel with computing current token.

reply
for a given expert, do you have a sense for what the spatiotemporal access pattern looks like?
reply
deleted
reply
Ya I'd be interested to see a comparison of using llamacpp with ssd offloading to compare real speeds.
reply