upvote
That's 6 YEARS of their cloud inference!

And, you have to include power. Where I am, the power for a rig with 2x 3090, running for four hours a day, would be around $70/month!

It's kinda like datacenter only exist to optimize compute cost through oversubscription of hardware/time share, cheaper business energy rates, and bulk discounts, compared to self hosting. ;)

reply
AIUI you can only connect two 3090s at a time with NV link? So you’d only have 48GB of fast combined memory? That’s not tremendously interesting as you’re still limited to the smaller models which, while impressive in their own right, are still IMO too limited to use as your only model.
reply
There are kernel patches to enable P2P communication via PCIE on 3090 which is almost as fast as NVLink for vllm
reply
Upgrading RAM is still probably cheaper than spending $6000 on a linux box. You're correct that inference will be much faster on the Linux box. But, the mac's unified memory can load larger models. And as OP mentions, it's still hard to beat cloud pricing at home.
reply
Yes, GPUs are much better for dense models. On Macs, MOE models run better, so I agree Macs are more limited and expensive. I have a Linux laptop with a 10GB 1080 GPU, dated, but I should add even more system RAM and try that.
reply
deleted
reply
A Mac mini running an LLM is quiet

A PC with similar capabilities is going to sound like a jet taking off.

reply
A Mac Mini is easy to match by a PC in terms of inference, and the PC will win without getting noisy.
reply
>A PC with similar capabilities is going to sound like a jet taking off.

Not at all. Airflow with big fans is quiet. What I do hear is coil-whine. In fact my PC is quieter than my Macbook when both are running top speed. But one has 4000 AI TOPS.

reply