We make a lot of price/performance compromises for having an attached screen and keyboard on our computer. That was what got me started.
Then I remembered the days of having to go to a special corner of the house to use a computer, vs now when I have a computer with me all the time. In my bag, on the sofa, on the train. Hell, I'm writing this on the work MBP while waiting for an appointment.
And you know what, I think I got more done when I went and sat in a corner of the house all those years ago. I set up an area for "computer work", and it worked really well.
I have a home office, but it's a jumble of cables going into docking stations and all sorts of weird stuff. I think if I streamline it and turn it into a proper "computer room", I might get some of that mojo back. I might even convince my partner that surrendering the home office and having a corner of the den might be good - she can watch TV while I tinker. And I won't be balancing a laptop on my knee and trying to do two things at once.
And the price/performance thing comes back in. Hmm.
And that, from mental load standpoint, is not healthy for most folks.
There is a reason we know that remote learning is horrible for most people compared to in-classroom. We proved this decisively during COVID.
There is zero reason to believe that most folks magically change overnight from being incapable of remote learning to being highly capable remote workers. It's just not believable.
I hired remote workers in the 90's. It was a small fraction of the total candidate base that could successfully self-motivate and have the discipline to become high performers in such an environment over the long haul. Most of my interviewing and candidate vetting had to do with the remote aspect vs. technical skillset. Luckily around that time is when open source became a huge thing, so those projects presented a pool of pre-vetted candidates to hire out of. The rest of the candidate pool was a total crapshoot.
Remote working has become easier and the tooling and technology much better. But from where I'm standing - many folks do not take it seriously. Simple stuff like having backup Internet is a filtering question for me even today.
Personally on the 0.2 days a year I have to worry about it, I just head to the coffee shop. Or you know … take a couple hours off (shock!).
Also I’ve spent literally decades working with very highly productive people exclusively remotely. None of us find this odd. Not sure why you have that level of suspicion/distrust. Granted, things have changed a lot since the 90s.
I love taking meetings while walking but since Covid everyone has their camera on making sedentary meetings mandatory.
If your computing needs line up, it's a very serviceable approach.
I haven't added my iPad to the Tailnet yet but i reckon that it could become a very comfortable and productive device for me.
Its update policy can result in this sort of regressions in a middle of a release. One thing that's nicer thing about macOS is that a released version doesn't regress as much (even if they are buggier at release; but in that case you can just wait until a later point release to upgrade).
But the couch is just so comfy.
Never ended up happening. I didn't get much done at home at all and what I did get done was mostly on the couch; not a great environment for serious work. (I have an office where I do for-profit work so it's not an issue, but still, I like my side projects too)
Eventually I had enough decommissioned computer parts that I could assemble them and re-commission them into a working desktop. So I now have a desktop again. And I actually end up sitting at the desk working on stuff on the desktop in a way I rarely did with the laptop.
I can tailscale into my home network, but what next? VNC/TeamViewer/Remote Desktop setups are all mediocre imho.
If you haven't tried Screen Sharing since they deprecated VNC and switched to their proprietary H.264-based protocol, its worth trying. Even YouTube videos play just fine with no noticeable lag.
It worked relatively well, but wasn't perfect. It was as close to what you're describing as I've seen, though. This was a few years ago; something like this might work even better now.
Unfortunately, it relied on Display PostScript and Quartz née Display PDF isn't architected to allow that sort of remote display/access on a per application level.
It’s pretty good.
To be clear, it’s not ever supposed to be a primary. It just becomes one from time to time when I make it full screen and am not paying attention.
Screen is small and is only one. Ergonomics is entirely messed up. Either your screen is too low, or your keyboard is too high. Keyboards are non-ergonomic and have to be made with compromises due to height limits. Touchpad instead of mouse/trackball is compromise for many - and also stuck at one position.
And yet they have somehow spread despite number of people going on business trips not really increasing.
It's not like you have to be on a plane for the portability to be useful.
A laptop is a distinct tool from a desktop and if you try to use it like a desktop, I agree it is a terrible substitute. Personally I find laptops vastly more ergonomic than a desktop, but I never program with my feet on the ground.
I've found my smartphone's the best device for that. I wrote an entire programming language using my phone.
Been trying to make a handheld cyberdeck to replace the phone with something better, but it's still a long way from being real.
It’s a whole “thing” when I go to the computer now. And frankly, it’s made it way more fun to use. It’s kind of like enjoying the process of listening to vinyl rather than pulling up a song on Spotify
HPSS works well via 5G, and Tailscale is incredible for the setup as much as the hardware.
Edit: I ended up getting a 15" M5 Air but honestly trying it on a 13" M4 Air was actually superior because when all you care is mobility (since the Studio does the heavy work) its nice just going around with almost no weight / bulk.
https://support.apple.com/en-my/guide/remote-desktop/apdf8e0...
Years ago I realized running a 100W PC all the time was REALLY expensive, so I switched to an old linux laptop at 30W (about $4/month). That helped, but I moved to the m1 mac specifically for the power efficiency. It does all the things my linux server did, and it does them at 6W (75¢/month).
I've found no competitor with similar performance that can run in such a low power footprint. M4 mini's are way better at power/performance, and I imagine their price is about to come way down since the M6 mini just got announced.
Software-wise, it's different, but mostly equivalent. Homebrew or macports has a similar software inventory to debian. And Apple's container framework is a welcome improvement over colima for running most container workloads.
Macs are totally “fine” for light server duty… as is just about any computer of the last decade+. The CPUs are beasts, the disks are screaming fast.
The operating system itself may not be ideal at serving but you can just run Docker/Orbstack if you need to do something especially Linux-y.
I’d put the question back on you — what are scenarios where an Apple Silicon Mac wouldn’t cut it as a light server for one person or a handful of people? About the only scenario that comes to mind is scenarios where you expect to utilize it so heavily that the fans are running for many hours a day. At some point those are either gonna wear out or just ingest so much dust that the machine runs hotter and needs a deep clean. But even that is largely mitigated by just pointing an external fan at it.
iMacs are great for a lot of use cases, but my image of the typical HN user would prefer to keep the monitor separate.
Are you using Apple displays or did they fix it?
A Macbook Pro M4 or M5 drives TWO of these at once, well, for equivalent of quad 4K configured as dual super ultrawides:
https://www.samsung.com/us/monitors/gaming/57-inch-odyssey-n...
Per Apple support page, MBP with M5 Max handles:
Two external displays
Two displays up to a native resolution of 8K (7680 x 4320) at 60Hz, 5K (5120 x 2880) at 120Hz, or 4K (3840 x 2160) at 240Hz over Thunderbolt or HDMI
Three external displays
Two displays up to a native resolution of 6K (6144 x 3456) at 60Hz or 4K (3840 x 2160) at 144Hz; plus one display up to a native resolution of 8K (7680 x 4320) at 60Hz, 5K (5120 x 2880) at 120Hz or 4K (3840 x 2160) at 240Hz over Thunderbolt or HDMI
Four external displays
Four displays up to a native resolution of 6K (6144 x 3456) at 60Hz or 4K (3840 x 2160) at 144Hz over Thunderbolt or HDMI
https://support.apple.com/en-us/101571
PS. Use TB4 or TB5 dock from Cal-Digit (match the TB version of your mac), or this or one of the rebranded versions of the same OEM: https://www.owc.com/solutions/thunderbolt-dock, or an Ivanky Fusion dock if you want to go nuts with up to SIX displays, Quad 6K@60Hz and Dual 4K@60Hz:
https://www.amazon.com/Thunderbolt-Monitor-Docking-Station-A...
Last year when I was browsing the only recentish Apple silicon capable of driving that was the M3 Ultra.
Just Tailscale into the Studio from iPad Pro 13" with magic keyboard and Kit Knox's rootshell:
https://github.com/kitknox/rootshell
Note that the iPad Pro can also drive a 4K second screen if you like, and most anything else a Macbook with a single port could drive.
Why not Macbook Air? Because the iPad Pro is also a tablet, touch screen, and 5G...
E.g. what most major tech companies have - a laptop that is for VSCode via ssh/web browsing, and a beefy Linux dev box you ssh into for everything else?
It's way cheaper, and what I use at home too - a lot cheaper than a Mac studio for everything, especially with RAM and storage.
I know AMD has announced the replacement, and it will be faster, but I’m not sure when it’ll ship.
Also, you can cluster strix halo if you need more than 128GB of ram.
I was always cobbling together ad hoc network access, to get from one machine to another. Access to my Mac Mini while traveling was a pain. Bringing it with me is ridiculous, and network access was a PITA.
Tailscale is wonderful magic. Free (for my usage), so easy to set up, and now no matter where my various computers are, they are all accessible trivially via a single ssh connection.
So get a beefed up desktop Mac, set up Tailscale, and then use any random laptop, anywhere, to use it headless.
They gave me a nice MBP for my new job. I tried doing heavy work on it locally, it was fine for that, and yet it still ended up being a light terminal into an EC2 instance, partially because their stuff is on AWS and latency is way lower within that. My personal mini is running too.
Funny that after the price started to increase last year (I think October/November), there was a window of a few months where Apple prices stayed the same as they were before, thus making apple prices actually good when compared to the rest of the market.
It was a unique opportunity to have acquired a 512G M3 ultra for $10k.
Despite being outdated in terms of compute, it stills let me run very good recent models locally, with Deepseek V4 Flash 0731 being the greatest one right now, and hopefully Qwen 3.8 Flash will also fit well when it is released tomorrow!
The M1 ultra definitely leaves to be desired in terms of its token speeds, but I think 20 tps generation and ~200 tps prompt processing (which is what I get with DSv4 flash), is already enough to do a lot of serious work when you combine with the decent prompt caching provided by llama.cpp.
I’ll be selling my M4 MBA soon, I genuinely use the Neo more. Huge difference in typing experience.
Great repairability is a plus. It was super easy, and actually fun to open. Felt like unboxing an Apple product. Applied the thermal paste mod for $10 which works excellently; I’ve had it shortly after launch.
And I love the notchless display, even if I wished the color gamut was a bit better.
I bought a Macbook Air in the interm waiting to see what the Macbook Ultra looks like, but honestly—it's the best form factor ever. I love it! Even though it's only like a pound heavier, the Pro feels like a monster.
Thinking about getting a Mini or Studio with more horsepower to stay at home.
Any other apps or suggestions? I'm still not exactly sure what working this way looks like.
The issue (apart from my Mac mini still having the old, bigger form factor back then) is that this requires a full shutdown (obviously), and that is more friction than opening a closed laptop lid. It just takes a moment for every app and background service etc. to settle back in after, and depending on how your brain works, you might not like that a whole lot.
It absolutely is a cool feeling to carry a pretty mighty desktop in the backpack, though.
The mini sips a truly tiny amount of power. I wonder what the smallest/lightest UPS could could bundle it with would be.
But at this point.. you’ve got a shit laptop.
(Apparently the laptops do a suspend to disk depending on battery remaining, so the OS does support it. It'd be nice to have it available! The disk space required could become annoying if you only bought the 1 TB SSD for your fancy 512 GB Mac though.)
The shutdown thing didn't bother me, and in fact the Mini was lighter than a Macbook Air at the time. So all in all it was pretty convenient.
It will get worse with the MacBook ultra rocking a M5 too.
Apple’s marketing department is gonna kill me.
I want a 2026 pickup truck. The dealer also sells a 2027 hatchback.
I am not shopping for a hatchback.
I'm currently SSH-ing into my workstation from my M4 Air with 24 gb ram and it's ideal for this flow. Slack/editors/clients/browsers/etc. easily gobble up over 16 GB. I have no more dev tools/compilers/source on my client machines, everything is dockered up on remote workstation in isolated VMs (too much supply-chaining).
My only downside to using a pro is not having 120/4k HDMI port on air and my dock won't support it with Apple (it does with windows)
I’m also the kind of person who close tabs and I like to work on one task at a time.
I meant the only downside of not using a pro for my use-case is that I don't have a HDMI port on device with high refresh rate. I have a decent dock that can give me 4k/60hz HDMI but not 120hz refresh for macs (it works on windows). It's a minor thing but having 120hz is nice when I'm docked at home.
I'm waiting this out.
Having owned 3 MacBook Pros since 2008, the decision to make my next computer be a Mac Studio came down to (1) MacBook thermal throttling that slows down CPUs when it starts to overheat and (2) easier upgrade of Mac Studio SSD with after-market storage module whereas the MacBook requires more complicated disassembly and hot air gun to dislodge the surface mounted SSDs.
I have a brand new M5 Pro MacBook Pro I don't like it when the fans turn on. The Mac Studio will be faster and quieter for the same workloads.
Sorry for not being clearer. I don't like the MacBook's noise when the fans turn on.
The Mac Studio has bigger heat sinks to delay the need for thermal management -- and if its fan does need to turn on, the bigger size means it's still silent instead of the high-pitched whooshing noise the tiny fans make in the MacBook.
Thus, my current workhorse is a Studio.
These prices make that decision so bittersweet. I feel so shut out from being able to use AI as economic productivity without being able to take on debt for a local-inf capable mac studio
Is that the new phrase for social media? I hope it is...
... and a lot more social media I'd like to admit.
However, after retiring, I realized I never undocked my MBP.
So I got an M4Pro Mini, and I've been thrilled. If I ever get to where I travel a lot, again, I'll get a laptop, but I don't see a need, right now.
This changes when you get into connotations that aren’t available in the laptop form factor, but with RAM prices the way they are those configurations are more than I want to spend on a local machine right now.
For running local LLMs the high memory Mac options always look appealing, but the processing speed (prefill) is so much slower than GPUs that it hurts. For situations where you have no rush and can let something work in the background for 24 hours it can some times be ignored, but the speeds I get from a real GPU setup are so much faster that I never use the larger models on a high memory Mac any more.
Most, if not all, of our current work happens on remote cloud vms. Now I'm stuck with carrying a 3KG monstrosity to work.Every day.
Absolutely no positives compared to my ThinkPad that weighed less than half in my previous job.
Remote development is "good enough" these days. With VS Code, Development containers etc... Having a light weight, portable laptop is so much nicer than a laptop that you can't even rest on your lap for long duration.
The only thing you need to be mindful of using light weight laptops is not having enough RAM to fit all your browser tabs..
But in the end I ended op buying a Lenovo Legion and put Linux on it.
Laptops are so fast these days that I didn't want to be bothered with setting up connectivity to a remote desktop.
But if your laptop never leaves your desk I think a desktop computer is a great option. Relatively cheaper and easier to maintain and upgrade.
I was using a VM setup on my MBP but it felt like a huge waste, having to leave a laptop on 24/7 when all it did was run Claude Code inside VMs.
I likely will stick with a Macbook Air 15" for next purchase, and beef up my "Claude Server" down the road.
The funniest recurring thing when working at Google was new hires getting baited into taking a Chromebook, then not being able to switch to a Mac for like 2 years. Our team made sure new people didn't fall for that.
Makes managing both backups and handling failure scenarios involving loss or unauthorised access to the laptop less of a hassle as well.
If you’re running a big task or small, everything scales to the appropriate size, regardless of the hardware sitting on your lap or under your desk.
At least that’s how I think of it.
Would you get away with a Mini?
Everything is a container or VM now, and none of it runs locally for me.
I have a desktop (older intel) with giant monitors and a keyboard for when I sit at the desk. I have the laptop for when I travel, go out or just want to work from the couch.
When I do my next upgrade to "better hardware" I'm not migrating a machine, rather I'm migrating the containers. My workflow is such that if I loose one of the boxes I sit at to a cup of coffee I really wont care other than the financial loss of a new laptop or keyboard.
The biggest win in all this was dumping the off the shelf firewall/router and moving to Opnsense. Wireguard vpn lets me route all my traffic through home for all my devices (and what is now a growing home lab).
There are scopes of work that this setup would not work for. I would not want to be a video editor with this set up, it's not ideal if you want to play AAA games. But for what I do, it is pretty ideal.
Fir kinda the first time in my life I don’t really have development tools on my personal laptop. I ghostty and openvpn client installed.
I have a large remote linux workstation (2x 8c/16t xeon cpus, 256gb ram, 2x8tb spinning rust disk) and i have my tools over there (along with some VMs).
It works surprisingly well.
Also, the macbook neo is a surprisingly capable little machine.
I personally own an M3 ultra, an M1 max as laptops, but my desktop is a Ryzen desktop I built in 2022 and it was a third in price of the ultra for more power.
I decided to get Mac Studio M4 Max, also all maxed out config and the cooling is so much better that I can run local LLMs like Gemma 3/4, gpt-oss 120b all day long without any heat issues or any audible fan noise. So for my use case it was the right decision. I subsequently added 15'' M5 Max MacBook Pro all maxed out to my collection and even though it is slightly faster on LLM inference (I get 100 tokens/s with Gemma 4 27b model), you just can't run LLMs longer than a few minutes. It starts overheating and gets really loud.
So a combination of a powerful desktop and a "cheap" laptop might indeed be attractive.
Not exactly "future proof" for >1T parameter models but good for targeting specific lower-parameter models, or if you can rely on pipeline parallelism and run a cluster.
Computers are never "future proof".
A 8 years old graphics card can still play modern games. Ten years ago playing a modern game on hardware that old would've been unthinkable.
And with how the market is right now, we'll be stuck on the current "reference level" of hardware for a while longer.
My thinkpad t14g1 with 16gb ram somehow still relevant today.
It was future proof but not really because it struggled a lot in its final years.
I expect it to stay a mediocre gaming PC for the next 3 years maybe 5 years.
16GB RAM RTX 3060 TI (8 GB VRAM)
i'm still using an old i7 3770k @ 4.8ghz with 16gb (ddr3) ram running linux for random tasks like executing tests. obviously, power consumption is higher.
my main machine is a MBP M1 Max which i use for everything. i also have my main linux desktop workstation that has a 5950x with 128gb (ddr4) ram.
i'll probably get 2-5 years out of my MBP, and my AMD workstation will probably be good for another 5-10 years.
i'm not a gamer, but i have a 3080. i'm sure my 5950x will still be good for gaming in 10 years if paired with a modern GPU.
Upgradeable components however could go a loooong stretch towards that goal. It can't be that hard to follow a common form factor for at least the housing across two or three generations to allow a reuse of everything but the main PCB.
Would be nice if someone knowledgeable about electrical engineering and manufacturing processes could lay out some valid reasons for manufacturers to integrate RAM onto the motherboard.
The SO DIMM Ram slot, which is designed, idk, 30 years ago? is not really capable of handling the frequency we are targeting (close to 10GT/s)
Well it might be an idea to keep the layout of the mainboard and connectors the same.
That way, instead of having to upgrade the whole machine, all it would need is a new mainboard. Framework for example managed to pull that off, and in mobile at that, where constraints are much worse than for a desktop computer.
It's not the same thing though. On the M-series, CPU and GPU share a unified memory architecture and ram is much more tightly coupled to get it to go faster. A closer example would be the Framework desktop, actually, where memory is also soldered in for the same reason.
It’s a very non-Apple thing to do, but it’d be pretty awesome if they did.
But, that doesn't make it a good deal. It just means the Apple tax doesn't apply when stacked up against AI machines and with memory prices being so out of whack. I'm still planning to wait until the RAMpocalypse ends before I buy any more hardware.
RAM production is completely sold out for 2027[1] which means the prices are locked in until after then.
It takes about 2 years from the time ground if broken for a new fab to be built and producing RAM.
There were some new fabs announced between February and April this year by both the Korean and Chinese manufactures, so that new capacity might start having an impact in 2028 in the most optimistic scenario.
Samsung says supply will remain tight in 2028[2], and Micron says "tight beyond 2027"
The best hope is that new (Chinese) players overbuild fab capacity and supply outstrips demand. That isn't likely, but perhaps in the late 2028-2029 timeframe could happen.
[1] https://www.techpowerup.com/351344/memory-makers-seal-2027-d...
[2] https://www.tweaktown.com/news/112966/memory-shortages-will-...
[3] https://s25.q4cdn.com/621799436/files/doc_events/2026/06/Q3-...
Otherwise you'll have to wait to see if the AI circular financing club collapses- if you still have a job, there should be deals to be had...
Efficiency is improving, both in hardware and in software and in intelligence density (smaller models can effectively do more of the AI work that needs doing), so I think the pure data center plays will falter. If there isn't some other business attached, they're never going to recoup their investment. Anthropic and OpenAI are buying all the compute they can find right now, but efficiency gains, especially those coming out of Chinese labs where they must be more efficient to compete, will make it less and less of a problem.
I mean, think about the hardware we use for AI. It's basically an accident. GPUs were not designed for AI (though they are becoming more focused on AI). The specialized AI hardware industry is just ramping up.
So, we're still early in the curve for how efficient both the hardware and software can be at performing these tasks, and given the effectiveness of recent very small models (e.g. DeepSeek V4 Flash 0731 and Qwen 3.8 27B), I just don't see a long future for giant data centers built around billions of dollars worth of last years graphics cards. As with the crypto mining operations, at some point, it becomes more expensive to run the hardware than it makes in revenue. And, as with the crypto mining operations, when the money dries up, the hardware hits eBay and prices drop.
- 120v input plug
- not rack-mounted
- has a video out port
Putting aside the fact it is a marketing label, "Pro" usually means "designed for work" while "Consumer" (in this context) means "doesn't need a special environment".
In computing the distinction is primarily noise, power and cooling requirements.
If a computer is designed to use home power and is quiet enough to use without annoying people and doesn't require specialist cooling then it is a consumer device, even if it is used for work.
Hence why they had to make up the "Studio" brand for the workstation market, because they'd already fully removed any meaning from "Professional"
A NVIDIA RTX 6000, 96 GB at 1.7 TB/s, is 13 grand.
This 256 GB at 1.2 TB/s Mac is extremely competitive, it will be sold out everywhere.
If you just want to run Qwen 3.8 27B and Deepseek v4 Flash in perpetuity and that's it, there are a lot of solutions that will work and this is a fairly user friendly one.
Lastly, I want to clarify that prefill on an x86 CPU is drastically slower than on an M5 Ultra GPU.
It feels like PCIe is a bit of a boat anchor here. There's a SATA->NVMe style transition waiting in the wings to make this all so much better. We really need post-PCIe GPUs. CXL with it's very small low latency flits. This is an "almost certainly not" but I wonder if you could mix PCIe and CXL so you could have the GPU memory expose vmeme as a bunch of CXL.mem pools but still have an otherwise pretty normal GPU. It seems madness that UALink went all in on GPU-to-GPU with no affordances for connecting to host computers.
Good for inference; however if you like to train, data format support and effective performance is limited (M5 Pro). Some hardware features are not exposed or extremely slow.
You’ll be fine for inference, but pales in comparison to what a RTX 6000 Pro can do for compute/matmuls/training.
RTX 6000 wins in performance, if your model can fit into the VRAM.
There are very obvious and clear advantages to a Mac Studio. It's an entire system for one and you're getting a world class CPU as well.
[citation needed]. I have personally specced out and built an nvidia GPU-based machine which after some optimization, handily beat the Mac Studio in terms of tokens/watt for LLM inference with most models. This was in the M2 Ultra era, and I haven't run the numbers for the later generations, but nvidia's cards have gotten faster just as Apple's CPUs/GPUs have, so I would guess that it's still possible to do.
> RTX 6000 wins in performance, if your model can fit into the VRAM.
"if your model can fit into the VRAM" can be true for the Mac as well.
> There are very obvious and clear advantages to a Mac Studio.
There are certain advantages for sure, depending on your use case. They may _seem_ to be obvious, but as evidenced above, I believe that many people overestimate the Mac's superiority on the metrics you cite when comparing a Mac vs. a dedicated GPU for LLM inference.
It is much more likely for your model to fit in large unified memory of a Mac than the smaller more limited memory of a GPU. Even going with two 5090s, you now have to shard your model and that is a PITA.
But it turns out that MoE is the solution both for running models on macs of limited computer power means (not as fast as GPUs), and on multiple GPUs that require sharding the model.
I bristle at general statements like this when it obviously depends on the specific Mac and GPU in question. But yes, comparing a maxed out M5 Ultra with an RTX 6000, the Mac has much more memory.
> Even going with two 5090s, you now have to shard your model and that is a PITA.
Every modern tool does this for you automatically. It is absolutely not a pain in the least (e.g. llama.cpp ships with pipeline parallelism enabled by default).
If you are using multiple GPUs, MoE is basically going to be your only workable choice unless you can leverage pipeline parallelism (only half your GPUs can work on a prompt at a time, so you need to process prompts back to back in a pipeline setup, and they better be doing similar things because your vram is limited).
[citation needed].
No need. You can infer the logic with this line I wrote: RTX 6000 wins in performance, if your model can fit into the VRAM.
I'm not sure what the controversy is here.I can run agents using deepseek v4 flash or Qwen 3.8 on my m3 ultra and it will be lukewarm and the fan will eventually start blowing softly.
Yes, the Mac might get lukewarm, but it will take 2-3+ times longer to do the same task.
If you can show me a model for which a Mac is faster than the RTX 6000 then I’ll be happy to update or retract my statement.
But for the mac pro, they would say something very obscure.
"100-120V AC or 200-240V AC (wide-range power supply input voltage)" and the maximum current is "12A (low-voltage range) or 6A (high-voltage range)".
They're trying NOT to say it had a 1440 watt power supply.
There is some consolation there
I want a computer dammit, not an appliance.
¹ drivers are responsible to get the car into a low earth orbit themselves and need to be familiar with the orbital maneuvers needed to accelerate the car to 900mph
Staggering?? I can count on no hands the number of times that a fact or figure has caused me to stagger.
Wow, "Local AI" mentioned in the subheading above the fold - it's really awesome to see Apple leaning into this use case and I think it will definitely pay off for them going forward. Fingers crossed Apple is able to put some engineering effort towards shipping with one of the frontier open weight models included and optimized exactly for the machine.
Tend to be better than Apple AI.
I actually wouldn't want a serious model shipping 'on disk'. Models release so often, it's going to get outdated quickly. LMStudio is trivial to set up.
For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.
They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot
"Storage performance is up to twice as fast, with a next-generation SSD architecture built on PCIe Gen 6..."
This is the first personal computer I've noticed that has PCIe Gen 6 storage. I've only seen enterprise PCIe Gen 6 SSDs up until now. Gen 5 SSDs in consumer devices already have high temperatures and thermal throttling, so I'm worried about how Apple's implementation will perform (I know they don't use off-the-shelf SSDs anymore, but I'd imagine the temps would still be a problem).
Another example as a developer, one particular large project took an hour to build on a spinning disk, on a SSD it takes 2mins
So, yeah, it's really noticable improvement
So while have having this massive almost-symmetric fibre pipe is cool on paper, I haven't felt a huge need to install 2.5G gear all over the house.
data processing, LLMs, model loading, MoE loading, etc, etc relies on very fast storage to keep your GPU saturated.
Basically, Apple gets to cheat because they shove everyone onto the same IC & make the OS that runs on this SOC.
Apple also launched the base M6 today with a 32GB RAM limit, suggesting 512GB may remain the maximum for Ultra chips for some time. Since these Ultra chips combine 16 base chips:
32GB × 16 = 512GB
I am in Europe, and the Mac Studio M5 Ultra GPU 64 cores with 96GB RAM is up to 6.649,00 €. Ouch.
> A fully configured IBM Personal Computer AT (Model 5170) with expanded memory and storage cost around $5,795 to $6,000 at its launch in August 1984, which equals roughly $18,600 to $19,300 in 2026 USD.
Care to guess the approximate price of the MBP I bought earlier this year?
This new Studio? Can't find a config under $5k I'd bother with. But for the MBPs that number still mostly tracks for the average Pro user. (I buy large and run it into the ground so long I mistake the ground for the computer's remains.)
My still-being-used 2012 MBP (which cost me about $3K) says, “hi”.
And, as you point out, the new computers I want blow Dvorak’s hypothesis out of the water. Never would I have guessed 30 years ago that Dvorak would be wrong the other direction on price.
Before local LLMs became a thing, I had completely lost interest in buying anything but the cheapest laptop. It felt like "personal computing" was a solved problem. But...it actually isn't, and this is exciting.
What is changing is that there genuine demand for more capabilities disproportionate to the cost decrease curve. Fab demand and supply constraints have slowed or even reversed some cost decreases - but that is still getting absorbed by the overall systems costs when you are looking at things like laptops. If you all you want is the last decades demand to browse the web and use office - things are cheaper than ever.
Not that many of us, actually. Only Montana, New Hampshire, Oregon, and some parts of Delaware and Alaska have no sales tax. https://commons.wikimedia.org/wiki/File:Sales_tax_by_county....
It would be significantly cheaper to fly to a tariff-free country and buy there.
There are no configurations even close to running something comparable to frontier model variants, they're simply far too large, but something like full precision Qwen 35b or DeepSeek 70b at 50+ t/s is well within available configuration, and potential for plenty of room for large context sizes.
I'm using Flash heavily, and I would describe it as nearly as intelligent as Sonnet-class in agentic coding, but more usable. Less world knowledge of course, and definitely a bit less intelligent; but not _that_ much.
On usability: Takes less handholding, less likely to make unsolicited refactors or whatever, and the writing style is readable.
It's not great at super-long-horizon goals as the Claude 5 models are; but if you have a good harness, you can get around that.
Isn't 170GB/s slow for bandwidth?
Extremely high bandwidth is great for copying data, but not as relevant for walking chains of pointers, where latency/cache/TLB entries are more important.
So just saying "high bandwidth == better" is true when other variables are the same, but they rarely are, especially in comparison to x86-64 offerings.
All of these are independent that it is a SoC with on-package DRAM.
Compared to something like VRAM it's slow.
Maybe $17k for a 512gig system that can do 1.2TB/s seems like a pretty good deal for a small office.
My ideal Mac Studio would have an NVMe slot or two for additional internal storage, and those slots could be accessed without opening the case. But Apple won't do that, as convenient internal storage upgrades would likely lower their profits.
In many ways, the Mac Studio is the spiritual successor to the trash can Mac Pro. Apple Silicon's architecture solved the thermal problems that limited the trash can Mac Pro.
This is the exactly reason.
We no longer need CD/DVD/floppy drives, storage has shrunk/moved to the cloud. The only thing that's really grown inside a pc case is the video card, and most of these ITX cases are built specifically around fitting popular cards.
Even folks primarily focused on gaming are probably thinking that a full ATX case is a lot of wasted space.
Maybe it's just me, but I think ATX full and mid towers are going the way of the dinosaur.
There's just less space, so it's harder to get things in and out. And you have a smaller cooler because there's less space, at least for air cooling.
I haven't used itx with video cards, but that's hugely constrained...
They were obsolete a decade ago, but priced out by artificially expensive "SFF PC" cases, fans, and power supplies. PC building completely transitioned to fashionable cargo cult in the late 2010's, with the rise of twitch.
You can't upgrade the processor because it's soldered. You can't upgrade the ram because it's soldered. You can't upgrade the SSD because it's soldered / bonded to the CPU. You could maybe have some PCIe slots, but not many drivers for macOS, so what's the point?
Yes, it was different in the old days, but Apple is ever more a closed hardware architecture with few options. If you want choices, Apple is not for you.
1.2TB/s memory bandwidth unlocks a lot with 256GB unified, and agentic AI is pretty good at optimising performance.
For comparison, to get 256GB with NVIDIA, you’re looking at a DIY workstation build (need pcie lanes), and like $70k?
The spark’s ~250gb/s bandwidth doesn’t really count here.
The comparison should be against renting in the cloud for the duration of your task for training and research or using pay-per-api-call providers for general inference instead of buying your own hardware (and paying the electricity and cooling bills on top), because let’s face it, the models you want to use are probably the same ones available on inference providers (but, yes, some are more trustworthy than others).
Speaking as someone that does ML/AI research, you are essentially paying a huge premium for being able to just run your Python script at any time without setting up a deployment script and harness to run the job remotely, while your hardware sits essentially idle the rest of the time.
The only way to make the math work is if you rent your hardware in the background for inference while you’re not using it in anger, but despite all the startups and promises that has never become as streamlined as mining bitcoins or shitcoins used to be and they don’t pay out as much as they say they would. Renting your hardware for training is another option but doing that is a lot more involved, options are fewer and farther in between, you won’t get as much utilization out of it, and doesn’t let you feasibly abort running tasks at a moment’s notice.
The maths to me was basically equivalent to prepaying for 242 days of runpod pricing for the same GPU; and I reckon I'd be able to get 6+ years of use out of this card with 96GB.
Plus there's the resell value -- it's actually appreciated by ~50% since I bought it.
Plus I do really enjoy that it's 100% local. I wouldn't feel comfortable giving my agents this much information if inference wasn't 100% local.
I wouldn't get another one, I wouldn't have as much value, but one is definitely paying off for me on the financial side.
1-2$ / hour.
I'm not regretting my setup at home as it got paid by my company which makes sense here, but paying for electricity is quite high and makes already 0.3$/hour alone.
I would argue, the most interesting use case for running it at home is some personal agent which you want to run 24/7.
Only reason to buy this if you want to own your compute.
Experimentation and inference are all going to be cheaper on the cloud
256GB model is $10k and the 512GB version will probably be double
Seems like miscalculation. If they had their own fab for RAM, they could completely corner the market today.
Ugh. I was waiting it out on a very old Mac, but the pricing these days is insane, so now to wait it out for at least two years more and hope for better pricing or just suck it up and pay a lot for less.
I don't actually own a car and my startup is bootstrapped and our salaries are modest. But the one thing we spend on is laptops. I have M4 max pro with 48GB. That thing was on the expensive side (~4.5Kish). But it delivers a lot of value and I spend most hours I'm awake using it. I like fast builds. I like that I can try out open source AI models. And I like just having the option to run those.
We actually lease them and mine costs something like 105 euro/month. Including Apple Care. I don't need a Mac Studio but I could see some roles where that would not be a crazy expense. Even the tricked out version that basically only costs the same as a very modest car.
But +4000$ for an additional 128GB of ram is simply milking the customers, as they know they will have many of them.
I may consider a M6 Mac Mini as a stop-gap whilst waiting out RAM Apocalypse to be over. Basically abandoning any ambitions of AI sovereignty and riding out subsidised LLM pricing for the next couple of years.
I can't hate a direction where Apple becomes more about building great computers rather than trying to force more and more subscriptions. I do wish they would fix many of the long-standing OS and native app problems.
It’s more that it’s a very parallel architecture than fast.
The LPDDR5X under the hood is slower than what you’d find in a graphics card vram of years ago
That’s why you get consumer macs with 512gb while GPU makers are reluctant to give you more than 16gb unless you pay dearly. It’s not the same kind of mem
So not as flexible as apple's unified memory.
Personally, I can't justify dropping the dosh for a 512GB M5 Ultra, but I would be able to justify it to myself if I could get 1TB of memory, because it'd guarantee the flexibility with local models I currently am missing. Seems a huge miss to not offer this... for a price.
At that point just rent proper GPUs in the cloud, you'd have way more power and pay only what you use for.
I have a bit of paranoia/anxiety about AI, but it's not what most people are concerned with. I understand the limits of these tools very well, and still find them extremely useful. What concerns me is that it's going to become difficult to impossible in the future to run local models which have near-SOTA capabilities in a way in which you can exercise full control of the model. I see the writing on the wall, and its more than worth it for me to invest early to ensure my own capabilities. I am very much not a fan of our "you'll own nothing and be happy" directionality for the world, and I am (at least currently) privileged to have the means to slow that decline for my own self.
Are your referring to Taalas / chatjimmy ?
I definitely think we'll see an ASIC-like approach in the future, especially for embedded small models where it may require minimal silicon area and can result in near-realtime performance. But at the frontier, I don't think this is a solved problem and will continue to have model weight churn that will advantage more flexible general-purpose hardware.
Also see: https://apfel.franzai.com
Most people at Apple have already realized that their processors are already too powerful for regular users - heck, as a developer my M2 Pro with 32 GB RAM is more than enough for me.
Regular users don’t care about local AI either. So, they will probably extract as much money as possible during AI gold rush, but then we will most likely see Apple
a. Making their software worse (god forbid, forced updates)
b. Making their hardware impossible to repair (as they almost accomplished this already) and easier to break.
Plus introducing features like RDMA over thunderbolt, which is critical for distributed inference/training/etc. On the software side, Apple is investing heaps.
It's still ridiculous they don't support expandable NVMes, but the memory being soldered makes sense, you need it for 1.2TB/s bandwidth.
They are still selling high-margin hardware. Apple loves selling high-margin hardware.
EDIT: AMD too, it's not limited to Nvidia, nice.
They're eating nvidia's lunch.
I'm glad they brought the 512GB option back.
Really want to get my hands on a m6 32GB 2TB model but don't really have a need for one.
I know there's tons of marketing language, buzz words and attempts at convincing me of some agenda that isn't super clear without lots of effort in "validating" the slop. I guess its not bad "slop" though if a human put in effort in editing it (imo >50% human curating = not really bad ai slop)
Though I still would prefer I could just get the prompt. What human thoughts, direction and "prompt" went into writing this article? in the same way as we ask for the prompt for AI generated outputs, I would prefer it for human generated output too. For writing at the least. I could have saved time, got the purity of the argument, and got more clear information. I wonder if we can get a future where humans just express their intent with each other and stop trying to hide our agenda; I want a world we can trust each other greater and interpret and act on our goals without the noise of trying to impress or market to each other & the additional words that go into that.