With larger (also really fast) unified memory on the chip, it could easily load larger weights, and the thing is that the quantized results on the M5s are really good, I believe that for most local inferencing, users would be using INT4 (maybe more bits per weight sometimes) quantized models, that might be where the Neural Accelerators kick in. In raw power, a desktop 5090 easily outperforms the M5 Max, but for this specific use case, I believe that the M5 Max is good enough.