upvote
Without fundamental model architecture improvements the practicality largely depends on how Apple increases memory bandwidth.

Memory bandwidths (* = rumored):

  M1:       68 GB/s
  M2:       100 GB/s
  M2 pro:   200 GB/s
  M2 max:   400 GB/s
  M2 ultra: 800 GB/s
  M5:       153 GB/s
  M5 pro:   307 GB/s
  M5 max:   460 GB/s
  M6:       200 GB/s*
  M7:       240 GB/s*
  
  Nvidia 4090 1008 GB/s
  Nvidia H100 3.35 TB/s
Basically what we're looking at by the M7 generation is a tier shift, where the base M7 can do what the M2 pro did, and every tier moves up accordingly, with the M7 ultra becoming competitive with nvidia dedicated consumer hardware.
reply
My assumption is that the difference is 90% from more memory. I'm making several assumptions because nothing here looks groundbreaking so I don't care to dig deeper, but the model + KV cache definitely cannot fit in memory on the 8GB machine, but probably can on the 24GB machine—or can at least get close. Assuming that this benchmark makes use of that, skipping SSD streaming will speed things up massively (I would have guessed much higher than the reported 6x speedup).
reply
I have an M5 128GB. Being on the cusp of practical is a good description. It will run, but prefill and token gen are still slow relative to my consumer GPU box.

It also gets very hot. If you’ve never heard the fans on Apple Silicon really spin up, it could surprise you. Makes the full GPU setup feel quiet by comparison.

I think after the hardware market calms down the ticket is going to be a light laptop with a second dedicated inference server on the network.

reply
i wonder if Apple will eventually ship proprietary models with burned into the silicon for all local workloads

https://eu.36kr.com/en/p/3904844399445638

reply
The speed of this perfectly correlates with the memory bandwidth of an M2 vs M5 Pro.
reply