upvote
Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
reply
can confirm.

I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.

reply
Have a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work. The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.
reply
> it's like the worst ADHD team member and requires constant supervision

Perhaps consider some non-offensive language for your comparison?

reply
could you not be like that? provide alternative language or go away. If youre offended, say so and be real. Noncommittal posits of personal preference are linguistic mosquitos of communication. on the flip side, how dare you disenfranchise a legitimate adhd perspective. one that i would say is entirely valid as someone functionally crippled by such plight. If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith. words are lame like that ya? mine are as nauseating as your flyby ego droppings.
reply
> could you not be like that? provide alternative language or go away. If youre offended, say so and be real.

Ok, as somebody with ADHD I find it offensive because I don't need constant supervision, implying people with ADHD need constant supervision is belittling and just plain wrong. So, I will call out an offensive trope if I see it.

> If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith

No, because if there was some truth to it, I wouldn't be offended. Perhaps stop with the amateur psychology? You're not very good at it.

reply
People don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.
reply
Exactly; when I first got my RTX 5070 Ti (16gb, to game with!!!, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models <8gb, primarily llama3.1.

My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."

This seems apt. My next LLM machine will be closer to 96gb+ vRAM.

reply
Once I get some kind of settlement after getting beaten up by a cop my first purchase will be some RTX Pro 6000s.
reply
Dude, I'm saying this with the best of intent. Get help.
reply
Reddit might be leaking today.
reply
Is a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
reply
Personally I think the price is way too high right now. It’s a power hungry gaming GPU. The efficient single card equivalent would be a 4500 Blackwell which launched at about $3500. Or you could get a 9700 32GB or an Arc B70 for well under $2k, today. You only buy a 5090 if you want absolute speed.

32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.

reply
A 5090 has 2x tensor cores and 2x bandwidth and can be run at 400W (2x watts).
reply
Being fast and having a power target doesn’t mean it’s cost efficient though. I would pay the launch cost for one, but not 3-4x inflated.
reply
How does this relate to 4500 vs 5090? I'm just pointing out that 5090 likely has twice the performance of the 4500 and likely maintains that at 2x watts if you want.
reply
You didn't specify in your earlier post, so I wasn't sure exactly which comparison you were making. But yeah, the perf/watt actually looks the same for those, so the cost per token evens out. It is nice not having to manage 400+W though. I like the 4000 for that reason, it's effectively a 3090 that runs at half the TDP.
reply
How are you deciding which work to send to the 5090 vs a frontier model, or making the two work together nicely?

Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.

reply
A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.
reply
> for me at least a GPU is completely useless for anything but being a token generator.

No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.

reply
> No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.

Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games so a GPU would probably be better off in some kid's computer.

reply
A GPU would be better-off attached to your Mac in an eGPU enclosure. There is not a single Apple Silicon GPU on the market that leads the industry in prefill, decode or power efficiency.

But of course, Apple doesn't allow that as part of their ecosystem. It's really a privilege to have MoltenVK perform worse than the fanmade HoneyKrisp driver. It's valuable when Apple refuses to sign AArch64 CUDA drivers for macOS. It's exciting to pay Crossover to support half of the library Proton offers for free.

Clearly, I'm some sort of ingrate that selfishly demands the best things, without considering how to accommodate the poor trillion-dollar megacorporation.

reply
Because you can run Qwen 3.8 Flash Next, Laguna S 2.1 and other medium-sized models that simply don't fit on a 5090?
reply
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
reply
It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch
reply
deleted
reply
I think you’re missing that MTP can predict more than 1 token in advance.
reply
In fact, isn’t that the “M” in “MTP”?
reply
Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
reply
deleted
reply
Qwen3.8-Flash-Next is pretty damn worth the extra ram you need.
reply
same reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn't immune from vanity).
reply
If you seriously think apple products are nothing but a status item, you're deluding yourself and probably have been for decades.
reply
deleted
reply
If you seriously think apple cares about anything other than cell phones, you're deluding yourself and probably have been for decades.
reply
deleted
reply
...did you mean profit? I don't think they're manufacturing iphones just on the hope they delight you. This is also true of Google et al.

I don't get these weird parasocial emotional attachments/beefs people have with brands. Talk to a therapist.

reply
brother my point is they don't care about their product offerings outside of their phones. this post/thread is about one of their product offerings which is not a phone which is inferior to their competitors'. simple.
reply
My M1 Pro MBP is 6 years old and continues to be the best computer I own, so if that’s Apple not trying, god help everybody else once they do.
reply
[flagged]
reply
reply
What in that thread is particularly impressive or noteworthy? According to the benchmarks I've seen, M6 raster performance is actually less efficient than M5 in many scenarios.
reply
They've been selling phones for less than 20 years at this point? Though I suppose 1.9 is not equal to 1, so it gets the plural.
reply
This 1000%. Data centres don't equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
reply
It's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
reply
So if it isn't a comparative ability, now it's a power cost issue? This reads like goal post moving.
reply
Oh, it's absolutely both. The power you waste waiting for TFTT on prefill will absolutely compound at the "medium sized labs and businesses" scale.
reply
deleted
reply
Both are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.
reply
The issue is that the moment you want to run the more capable models that will no longer fit in a single 5090's memory, performance falls off a cliff.
reply
... or with llama.cpp with MTP.
reply
[flagged]
reply
I guess it is possible, but Apple has had very vocal fans for decades. I suspect, rather than astroturfing, it is just people who are in their ecosystem.
reply
Tok/sec is 0 on a 3090 for most of the models that the mac can run
reply
[flagged]
reply
> So given that, which one of these is true about you?

Well, if those are the only two options you can come up with it's pretty clear that this isn't about me or what I am, you have a false model of reality.

> Running very large models on Mac is unusable at 10 tok/sec.

There are plenty of examples of models running at well over 10 tok/sec that aren't viable on the 3090. In fact such examples are found in the review in the OP. Did you not read the article?

I think you're projecting pretty hard with the two options you've listed. Go touch some grass, you seem overly frustrated that reality doesn't meet your expectations.

reply
Since you clearly don't use local llms, allow me to educate you - anything under 100 tok/sec is USELESS. When you are coding, the idea is that you want to have a system that can generate files fast, hopefully correct on the first try. Cloud models do this. Local models, by nature of having less parameters and more quantization, often require more guidance and repeated inference to get it right. The antigenic harnesses that people set up around local llms leverage this.

Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.

Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper

The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.

So I ask you again, which one are you?

reply
> allow me to educate you

No thanks, you're not in a position to do that clearly.

> Since you clearly don't use local llms

I do, probably a lot longer than you have actually.

> anything under 100 tok/sec is USELESS

Objectively wrong. You sound like you're really behind and you're so myopic that you think coding is the only use case for local LLMs. I'm a professional software dev and that's the least interesting use case of local LLMs.

> Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec.

You clearly didn't read the article or have reading comprehension issues. The model is Qwen3.8-Flash-Next 4 and 5-bit quant, neither of which "conveniently fits on one GPU". Sorry that your hardware doesn't live up to your own delusions and can't even run Qwen3.8-Flash-Next at 4/5 bit quant. You are taking the Quen3.8-27B numbers, something that the article isn't really that concerned with, and trying to make it fit into your narrative.

> So I ask you again, which one are you?

Well I'm someone that suggests that you should touch some grass and reevaluate your personal issues. You seem angry. Perhaps it's best to figure your own issues before trying to figure out why people are excited about Apple hardware for local llms. I am sure the people that need to interact with you in society would be very grateful if you took the time to do this.

reply
Nice try.

A) He literally says "I tested a different Qwen model for the comparisons between Mac and PC." The model he tested has to fit on one GPU, otherwise the inference is dogshit slow as you are offloading results to ram. If you ran any amount of local inference, you would know this. Considering that Qwen3.8-Flash-Next Q4 is still 100gb, there is no realistic way to run this with a 5090. The model that was run was this https://ollama.com/library/qwen3.8:27b. And the speed of that model on a 5090 in terms of tok/sec is not 60 lol.

B) If M5 ultra runs 40 tok/sec on qwen3.8:27b (and lets assume its the mlx version to gain a performance boost: https://ollama.com/library/qwen3.8:27b-mlx), you have to be delusional to believe it can run 100gb models at 100 tok/sec lol.

As a bonus, in terms of use, its pretty well known that Qwen models are RLed to chase benchmarks. Check out https://huggingface.co/Qwen/Qwen3.8-27B versus https://qwen.ai/blog?id=qwen3.8-flash-next, using different benchmarks the 27b outperforms the flash next on agentic coding. But it matches it in other areas pretty well. So tell me again why you need 100gb models running dogshit slow at peak ~20 tok/sec?

It is so incredibly sad how hard you try to sound intelligent. But thats on par for the course of any person hyping up apple products, throughout apples history.

Considering that Apple probably doesn't want you to engage in this level of pettiness for their advertising posts, you have outed yourself to be #2. And Im not angry at all lol, you keep doing what you do, people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.

reply
Those are some incredible graphs, that leap in prompt processing going from M3 to M5.

Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think).

/meta Here's a CSS filter that stops those nuisance chart animations,

    macstories.net##*:style(animation: none !important; transition: none !important)
reply
A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
reply
A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform.

Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.

reply
> A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures

Inference time is going to be dominated by the low memory bandwidth on these Macs, so a dense model will suffer most. It’s more of an opportunity for large MoE models with a low number of active experts since you can keep all experts in VRAM but not pay the bandwidth cost until they are used.

> you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers

This is an interesting direction that I expect to see more of. But for most models currently you need basically all experts loaded since they are chosen per token.

Apple seems to be researching longer horizon expert caching, where they keep experts swapped in for longer runs of tokens [1]. Other labs are offloading ngram caches but not sure if they’re pursuing anything like this?

1. https://machinelearning.apple.com/research/introducing-third...

reply
1.2TB/s is already considered slow? Things are moving quickly!
reply
They do MoE. They benchmarked GLM 5.3-flash (320B / 18B), and Qwen 3.8-flash-next (125B / 6B). The dense Qwen is only focused (I assume) because it's about the only thing that fits on a 5090, that they can compare the two heads on.
reply
1.2 T/s is not that modest is it? That's very close to an RTX pro 5000
reply
That's a dense model. Of course it will do worse.

Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).

reply
Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,

https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...

(Note it's a sparse MoE with only 6B active).

reply
I have a 3 year old gaming system. RTX 4080 w/128GB of DDR5. It runs Qwen 38 Flash around 44-40 t/s with 128K context. It is on a specialized build that caches MoE experts and uses an optimized 3bit quant that basically is within a few points of the full 8 bit quant. In general, in casual benchmarking with Alibaba's endpoint I could not tell much of a difference. Overall this model is very good on long horizon agentic work. The main pain point for it is that its input processing speed is slow. Regardless, it gets meaningful work done.

I paid $500 for the RAM in Nov 2023 :)

reply
> "I paid $500 for the RAM in Nov 2023 :)"

No wonder Warren Buffet gave up and resigned.

reply
Right? To get comparable brand and quality DDR5, which isn’t particularly great at AI anything it is ~$2000. All you had to do was start hoarding 3090 GPU and RAM in 2023. It is unhinged.
reply
Good to know thanks.

That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.

reply
Unlike a 256GB M5 Ultra that is $10k+.
reply
Apple product won't be the cheapest but it is a full package (CPU, RAM, VRAM/GPU, fast-storage, etc).

If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).

reply
But the 5090 they use there is now going low stock and selling for over $7500 in some places.
reply
I really appreciate seeing these dense model numbers. For a large unified memory system though I expect that MoE numbers are what people are more interested in.

These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB/s memory bandwidth.

reply
Yeah that can't be right, my M5 Max gets almost those speeds and certainly lot faster than what they're claiming the M3 ultra gets. Maybe they didn't have the model setup right or were running it with some unnecessarily high quant (>=8bit).
reply
A basic llama-bench on Qwen 3.8 27B UD-Q4_K_M gives pp512 3920 tok/s / tg128 81 tok/s on a 500W RTX PRO 6000 (should be similar speeds to a 5090, chip is basically the same, just less VRAM). With MTP3 this is 140 tok/s on mtp-bench.

This is with llama.cpp. You can of course use vLLM/SGLang well on these cards and they're even faster. On vLLM w/ NVIDIA/Qwen3.8-27B-NVFP4 baseline has a prefill of about 13,000 tok/s. The baseline tok/s is 72 tok/s, but at mtp7, it's 157 tok/s, and w/ dflash7 that goes up to 215 tok/s. On mtp-bench, DFlash2 gets a hair under 300 tok/s w/ the code_python prompt.

reply
On my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add https://huggingface.co/collections/z-lab/dflash-2 to it. So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.
reply
Thank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia's TensorCores in the 5090.

Maybe Apple is an acquisition away from changing that balance.

reply
The rumor on Apple's processor roadmap is that they're skipping other M6 variations (all previous generations had Pro and Max, a few had Ultra) in order to focus on the M7 generation for AI reasons. What exactly the M7 improvements are who knows.

https://www.macrumors.com/2026/06/25/2027-macs-m7-chips/

reply
I think that comes down to TSMC. Nvidia apparently booked out the whole A18 or 16 node. Apple is on 2nm right now and M7 will jump right to A14. According to my quick AI research anyway.
reply
That sounds a lot like AI fantasy slop.

Apple just shifted to N2. They’re not going to be doing another major shift right away.

And TSMCs own roadmap would put your hallucination years away at best for a a product that follows a roughly annual cadence https://www.tomshardware.com/tech-industry/semiconductors/ts...

reply
Yeah, Apple has spent 3 years each on the 5nm and 3nm nodes with TSMC. There are some reports [1] that it will jump to 1.4nm after 2 years due to AI but there's no real proof. The source is Digitimes who are frequently wrong with their predictions and rumors.

[1] https://wccftech.com/apple-to-move-to-1-4nm-process-soon-to-...

reply
> What exactly the M7 improvements are who knows.

> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators

https://www.tomshardware.com/tech-industry/semiconductors/ap...

reply
M6 got another prompt processing boost. Likely no M6 Ultra though because Apple is reportedly going all in on AI performance in M7 generation.
reply
I assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you'll get almost 4X to 8X performance.
reply
The selling point of the M5 Ultra Mac Studio is that you can run much larger models that the 5090 can't without swapping. NVidia aggressively segments the market on VRAM for this reason. That's why a 5090 has an MSRP of ~$2k (but good luck getting one for less than $4k) while a 6000 Pro, which is basically a 5090 with 96GB of RAM has now soared beyond $15k where 3-6 months ago it was more like $10-11k. A 6000 Pro has the same memory bandwidth but slightly more CUDA units (IIRC ~24k vs ~21k).

This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090.

I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.

Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.

The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.

reply
But what about builds that combine 8 of the 5090 with infiniband between boxes? Wouldn't that be comparable to the mac in terms of price and potentially beat it by a lot in terms of performance for the large MoE? I understand the space/heat/noise considerations, but price wise it may still not make as much sense as people think. (Agreed that it is hard to get the NVIDIA hardware and the 6000 pro are priced less competitively).
reply
Others have chimed in on cost and size, I'll chime in on power. The 8x 5090s will require a dedicated datacentre grade power source. The Mac Studio runs on a plain jane wall socket.
reply
> But what about builds that combine 8 of the 5090 with infiniband between boxes?

Why Infiniband ("IB")? If it's for RDMA, that is possible with certain Ethernet cards/chipsets as well. Certainly Mellanox, but Broadcom:

* https://techdocs.broadcom.com/us/en/storage-and-ethernet-con...

and Intel as well:

* https://www.intel.com/content/www/us/en/support/articles/000...

Link level flow control or priority flow control needs to be supported on the switch ports as well.

reply
No, $40K is not comparable to $10K.
reply
When I compare those two numbers, it seems there's $30k of difference
reply
>8 of the 5090

Where are you buying 8 5090s for under $10k? With CPU, RAM, and (checks comment) infiniband hardware???

You're probably looking at a lot closer to $60k when all is said and done, and that's before you hire an electrician to run a sub panel for your homelab...

reply
While that sounds super awesome, How many people are actually going to build and maintain that vs a box you can grab at the mall that fits in a lunchbox?
reply
Sounds like nice utility bill in the making.
reply
I can't speak to Infiniband pricing for something like that. It seems like the cheap option is 56/100Gbps with used Enterprise equipment. You'd need 8 HCAs, DAC cabling and a switch but even then you're into thousands of dollars. If you want 200Gbps+ it gets into the tens of thousands (AFAICT).

Each PC is probably going to cost ~$6k and you're talking about 8000W of electricity draw. That's going to consume multiple 20A circuits even at 240V. And the electricity ain't free either. A Mac Studio seems to draw ~500W max.

Oh and the Mac Studio has an upgrade route to run 1T+ models too by chaining them together with TB5 chaining. OSX supports RDMA this way. That's comparable bandwidth to the 100Gbps Infiniband option.

So you're talking about $50-60k of hardware and more power draw and more heat for something that will I'm sure beat the MS M5U option but at huge cost. Also, at that kind of price point, I'm likely to get a workstation PC and put 2 (or possibly 3) 6000 Pros in it.

reply
openai make a npu google make npu (tpu no mater) amd buy tellas

every company make his own npu (without xai)

probaby in 2028 we will have more concurent firm on market place

reply
deleted
reply
[flagged]
reply
Not forgetting of course that an RTX5090 is what 600W+ ? And the Mac is probably half that at most ?
reply
Certainly not forgetting wattage. A 5090 is 575W. The M5 Ultra Studio is 480W.

nvidia-smi -pl 450 for like a 4% reduction in throughput. I tend to set it around 350W because it's a comfortable temperature blowing on my legs under the desk without warming my office in the summer.

I put together this system two years ago, so it's a little out of date, but it only cost $3000 for the same performance and capability as an Ultra. I don't think I would spend $7000 to save 100W, though.

reply
> nvidia-smi -pl 450 for like a 4% reduction in throughput.

Yeah people don't pay enough attention to those settings IMO. The first thing I do when I set up a new machine (or upgrade my OS) is to restore all my powersaving configs.

For example I've got all but one of my virtual desktops that put the CPU in powersave mode: I don't need max Ghz when browsing the Web, not even on demand. But when I switch to the virtual desktop where my development environment is, then I want power on demand.

Now I don't do it to save the planet: I do it because I love a quieter computing experience (coupled with Be Quiet! PSU and Noctua fans, this makes for a very quiet computer). That it consumes less electricity is a nice side-benefit.

reply
sure. so is 2x power worth 10x perf? I think it is in most cases.
reply
When you are doing matrix math, compute is compute. Apple cant be more efficient due to physics. The only reason Macs are more efficient in general is that they have tightly bundled hw and sw for specific tasks.
reply
The next Ultra, supposedly on deck in 2028:

> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators, according to a new Bloomberg report published by Mark Gurman...

Apple plans to release a base M6 chip this fall for entry-level Macs... a base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, and the M7 Ultra in 2028.

https://www.tomshardware.com/tech-industry/semiconductors/ap...

reply