Computers are never "future proof".
A 8 years old graphics card can still play modern games. Ten years ago playing a modern game on hardware that old would've been unthinkable.
And with how the market is right now, we'll be stuck on the current "reference level" of hardware for a while longer.
My thinkpad t14g1 with 16gb ram somehow still relevant today.
It was future proof but not really because it struggled a lot in its final years.
I expect it to stay a mediocre gaming PC for the next 3 years maybe 5 years.
16GB RAM RTX 3060 TI (8 GB VRAM)
i'm still using an old i7 3770k @ 4.8ghz with 16gb (ddr3) ram running linux for random tasks like executing tests. obviously, power consumption is higher.
my main machine is a MBP M1 Max which i use for everything. i also have my main linux desktop workstation that has a 5950x with 128gb (ddr4) ram.
i'll probably get 2-5 years out of my MBP, and my AMD workstation will probably be good for another 5-10 years.
i'm not a gamer, but i have a 3080. i'm sure my 5950x will still be good for gaming in 10 years if paired with a modern GPU.
Upgradeable components however could go a loooong stretch towards that goal. It can't be that hard to follow a common form factor for at least the housing across two or three generations to allow a reuse of everything but the main PCB.
Would be nice if someone knowledgeable about electrical engineering and manufacturing processes could lay out some valid reasons for manufacturers to integrate RAM onto the motherboard.
The SO DIMM Ram slot, which is designed, idk, 30 years ago? is not really capable of handling the frequency we are targeting (close to 10GT/s)
Well it might be an idea to keep the layout of the mainboard and connectors the same.
That way, instead of having to upgrade the whole machine, all it would need is a new mainboard. Framework for example managed to pull that off, and in mobile at that, where constraints are much worse than for a desktop computer.
It's not the same thing though. On the M-series, CPU and GPU share a unified memory architecture and ram is much more tightly coupled to get it to go faster. A closer example would be the Framework desktop, actually, where memory is also soldered in for the same reason.
It’s a very non-Apple thing to do, but it’d be pretty awesome if they did.
But, that doesn't make it a good deal. It just means the Apple tax doesn't apply when stacked up against AI machines and with memory prices being so out of whack. I'm still planning to wait until the RAMpocalypse ends before I buy any more hardware.
RAM production is completely sold out for 2027[1] which means the prices are locked in until after then.
It takes about 2 years from the time ground if broken for a new fab to be built and producing RAM.
There were some new fabs announced between February and April this year by both the Korean and Chinese manufactures, so that new capacity might start having an impact in 2028 in the most optimistic scenario.
Samsung says supply will remain tight in 2028[2], and Micron says "tight beyond 2027"
The best hope is that new (Chinese) players overbuild fab capacity and supply outstrips demand. That isn't likely, but perhaps in the late 2028-2029 timeframe could happen.
[1] https://www.techpowerup.com/351344/memory-makers-seal-2027-d...
[2] https://www.tweaktown.com/news/112966/memory-shortages-will-...
[3] https://s25.q4cdn.com/621799436/files/doc_events/2026/06/Q3-...
Otherwise you'll have to wait to see if the AI circular financing club collapses- if you still have a job, there should be deals to be had...
Efficiency is improving, both in hardware and in software and in intelligence density (smaller models can effectively do more of the AI work that needs doing), so I think the pure data center plays will falter. If there isn't some other business attached, they're never going to recoup their investment. Anthropic and OpenAI are buying all the compute they can find right now, but efficiency gains, especially those coming out of Chinese labs where they must be more efficient to compete, will make it less and less of a problem.
I mean, think about the hardware we use for AI. It's basically an accident. GPUs were not designed for AI (though they are becoming more focused on AI). The specialized AI hardware industry is just ramping up.
So, we're still early in the curve for how efficient both the hardware and software can be at performing these tasks, and given the effectiveness of recent very small models (e.g. DeepSeek V4 Flash 0731 and Qwen 3.8 27B), I just don't see a long future for giant data centers built around billions of dollars worth of last years graphics cards. As with the crypto mining operations, at some point, it becomes more expensive to run the hardware than it makes in revenue. And, as with the crypto mining operations, when the money dries up, the hardware hits eBay and prices drop.
- 120v input plug
- not rack-mounted
- has a video out port
Putting aside the fact it is a marketing label, "Pro" usually means "designed for work" while "Consumer" (in this context) means "doesn't need a special environment".
In computing the distinction is primarily noise, power and cooling requirements.
If a computer is designed to use home power and is quiet enough to use without annoying people and doesn't require specialist cooling then it is a consumer device, even if it is used for work.
Hence why they had to make up the "Studio" brand for the workstation market, because they'd already fully removed any meaning from "Professional"
A NVIDIA RTX 6000, 96 GB at 1.7 TB/s, is 13 grand.
This 256 GB at 1.2 TB/s Mac is extremely competitive, it will be sold out everywhere.
If you just want to run Qwen 3.8 27B and Deepseek v4 Flash in perpetuity and that's it, there are a lot of solutions that will work and this is a fairly user friendly one.
Lastly, I want to clarify that prefill on an x86 CPU is drastically slower than on an M5 Ultra GPU.
It feels like PCIe is a bit of a boat anchor here. There's a SATA->NVMe style transition waiting in the wings to make this all so much better. We really need post-PCIe GPUs. CXL with it's very small low latency flits. This is an "almost certainly not" but I wonder if you could mix PCIe and CXL so you could have the GPU memory expose vmeme as a bunch of CXL.mem pools but still have an otherwise pretty normal GPU. It seems madness that UALink went all in on GPU-to-GPU with no affordances for connecting to host computers.
Good for inference; however if you like to train, data format support and effective performance is limited (M5 Pro). Some hardware features are not exposed or extremely slow.
You’ll be fine for inference, but pales in comparison to what a RTX 6000 Pro can do for compute/matmuls/training.
RTX 6000 wins in performance, if your model can fit into the VRAM.
There are very obvious and clear advantages to a Mac Studio. It's an entire system for one and you're getting a world class CPU as well.
[citation needed]. I have personally specced out and built an nvidia GPU-based machine which after some optimization, handily beat the Mac Studio in terms of tokens/watt for LLM inference with most models. This was in the M2 Ultra era, and I haven't run the numbers for the later generations, but nvidia's cards have gotten faster just as Apple's CPUs/GPUs have, so I would guess that it's still possible to do.
> RTX 6000 wins in performance, if your model can fit into the VRAM.
"if your model can fit into the VRAM" can be true for the Mac as well.
> There are very obvious and clear advantages to a Mac Studio.
There are certain advantages for sure, depending on your use case. They may _seem_ to be obvious, but as evidenced above, I believe that many people overestimate the Mac's superiority on the metrics you cite when comparing a Mac vs. a dedicated GPU for LLM inference.
It is much more likely for your model to fit in large unified memory of a Mac than the smaller more limited memory of a GPU. Even going with two 5090s, you now have to shard your model and that is a PITA.
But it turns out that MoE is the solution both for running models on macs of limited computer power means (not as fast as GPUs), and on multiple GPUs that require sharding the model.
I bristle at general statements like this when it obviously depends on the specific Mac and GPU in question. But yes, comparing a maxed out M5 Ultra with an RTX 6000, the Mac has much more memory.
> Even going with two 5090s, you now have to shard your model and that is a PITA.
Every modern tool does this for you automatically. It is absolutely not a pain in the least (e.g. llama.cpp ships with pipeline parallelism enabled by default).
If you are using multiple GPUs, MoE is basically going to be your only workable choice unless you can leverage pipeline parallelism (only half your GPUs can work on a prompt at a time, so you need to process prompts back to back in a pipeline setup, and they better be doing similar things because your vram is limited).
[citation needed].
No need. You can infer the logic with this line I wrote: RTX 6000 wins in performance, if your model can fit into the VRAM.
I'm not sure what the controversy is here.I can run agents using deepseek v4 flash or Qwen 3.8 on my m3 ultra and it will be lukewarm and the fan will eventually start blowing softly.
Yes, the Mac might get lukewarm, but it will take 2-3+ times longer to do the same task.
If you can show me a model for which a Mac is faster than the RTX 6000 then I’ll be happy to update or retract my statement.