upvote
The real downside for me is not having Linux support.

It would take Apple one or two engineers to make Linux life much easier on macs. But Linux is outside their walled garden so it's ignored.

reply
Same here. Sadly I think the voices like ours won't be heard, though, because Apple's looking for someone who's going to buy in on the whole ecosystem, and I think we're not it. Or at least I'm not.
reply
I’m done with macOS.

My Mac Mini is strictly a headless server for llama.cpp.

I use a Linux workstation.

If I were limited to use Mac hardware , I would install Linux in VMware Fusion and work from there.

reply
It's worth noting that the 5090 (or the RTX Pro 6000 big brother with 92GB VRAM) will run rings around the Mac when it comes to compute.

My old 3090 is typically significantly faster (almost 2x token/s) than my M4 Max 128GB machine, as long as the model fits in the 24GB of VRAM.

In most situations it's a better idea to just buy tokens. But there are definitely cases when that's not an option. And then a machine like the M5 Ultra can allow you to do things locally for a fairly limited budget. And in a simpler package to manage than a machine with multiple GPUs.

reply
is it still effectively 2/3rds? Don't know enough to compare a discrete GPU/CPU setup to something like this where it's more integrated
reply
There is no magic, if the data you compute as atomic chunk don't fit in cache then memory bandwidth R/W limit kicks in and architecture does not matter. On contrary - having multi gpu setup of same price and same memory size with even slower memories may give you effectively much higher bandwidth but at the cost of power consumption.
reply
How much memory does that come with?
reply
32 GB GDDR 7
reply