upvote
can confirm.

I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.

reply
Have a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work. The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.
reply
> it's like the worst ADHD team member and requires constant supervision

Perhaps consider some non-offensive language for your comparison?

reply
could you not be like that? provide alternative language or go away. If youre offended, say so and be real. Noncommittal posits of personal preference are linguistic mosquitos of communication. on the flip side, how dare you disenfranchise a legitimate adhd perspective. one that i would say is entirely valid as someone functionally crippled by such plight. If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith. words are lame like that ya? mine are as nauseating as your flyby ego droppings.
reply
> could you not be like that? provide alternative language or go away. If youre offended, say so and be real.

Ok, as somebody with ADHD I find it offensive because I don't need constant supervision, implying people with ADHD need constant supervision is belittling and just plain wrong. So, I will call out an offensive trope if I see it.

> If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith

No, because if there was some truth to it, I wouldn't be offended. Perhaps stop with the amateur psychology? You're not very good at it.

reply
People don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.
reply
Exactly; when I first got my RTX 5070 Ti (16gb, to game with!!!, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models <8gb, primarily llama3.1.

My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."

This seems apt. My next LLM machine will be closer to 96gb+ vRAM.

reply
Once I get some kind of settlement after getting beaten up by a cop my first purchase will be some RTX Pro 6000s.
reply
Dude, I'm saying this with the best of intent. Get help.
reply
Reddit might be leaking today.
reply
Is a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
reply
Personally I think the price is way too high right now. It’s a power hungry gaming GPU. The efficient single card equivalent would be a 4500 Blackwell which launched at about $3500. Or you could get a 9700 32GB or an Arc B70 for well under $2k, today. You only buy a 5090 if you want absolute speed.

32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.

reply
A 5090 has 2x tensor cores and 2x bandwidth and can be run at 400W (2x watts).
reply
Being fast and having a power target doesn’t mean it’s cost efficient though. I would pay the launch cost for one, but not 3-4x inflated.
reply
How does this relate to 4500 vs 5090? I'm just pointing out that 5090 likely has twice the performance of the 4500 and likely maintains that at 2x watts if you want.
reply
You didn't specify in your earlier post, so I wasn't sure exactly which comparison you were making. But yeah, the perf/watt actually looks the same for those, so the cost per token evens out. It is nice not having to manage 400+W though. I like the 4000 for that reason, it's effectively a 3090 that runs at half the TDP.
reply
How are you deciding which work to send to the 5090 vs a frontier model, or making the two work together nicely?

Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.

reply
A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.
reply
> for me at least a GPU is completely useless for anything but being a token generator.

No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.

reply
> No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.

Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games so a GPU would probably be better off in some kid's computer.

reply
A GPU would be better-off attached to your Mac in an eGPU enclosure. There is not a single Apple Silicon GPU on the market that leads the industry in prefill, decode or power efficiency.

But of course, Apple doesn't allow that as part of their ecosystem. It's really a privilege to have MoltenVK perform worse than the fanmade HoneyKrisp driver. It's valuable when Apple refuses to sign AArch64 CUDA drivers for macOS. It's exciting to pay Crossover to support half of the library Proton offers for free.

Clearly, I'm some sort of ingrate that selfishly demands the best things, without considering how to accommodate the poor trillion-dollar megacorporation.

reply
Because you can run Qwen 3.8 Flash Next, Laguna S 2.1 and other medium-sized models that simply don't fit on a 5090?
reply
A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
reply
It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch
reply
deleted
reply
I think you’re missing that MTP can predict more than 1 token in advance.
reply
In fact, isn’t that the “M” in “MTP”?
reply
Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
reply
deleted
reply
Qwen3.8-Flash-Next is pretty damn worth the extra ram you need.
reply
same reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn't immune from vanity).
reply
If you seriously think apple products are nothing but a status item, you're deluding yourself and probably have been for decades.
reply
deleted
reply
If you seriously think apple cares about anything other than cell phones, you're deluding yourself and probably have been for decades.
reply
deleted
reply
...did you mean profit? I don't think they're manufacturing iphones just on the hope they delight you. This is also true of Google et al.

I don't get these weird parasocial emotional attachments/beefs people have with brands. Talk to a therapist.

reply
brother my point is they don't care about their product offerings outside of their phones. this post/thread is about one of their product offerings which is not a phone which is inferior to their competitors'. simple.
reply
My M1 Pro MBP is 6 years old and continues to be the best computer I own, so if that’s Apple not trying, god help everybody else once they do.
reply
[flagged]
reply
reply
What in that thread is particularly impressive or noteworthy? According to the benchmarks I've seen, M6 raster performance is actually less efficient than M5 in many scenarios.
reply
They've been selling phones for less than 20 years at this point? Though I suppose 1.9 is not equal to 1, so it gets the plural.
reply
This 1000%. Data centres don't equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
reply
It's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.
reply
So if it isn't a comparative ability, now it's a power cost issue? This reads like goal post moving.
reply
Oh, it's absolutely both. The power you waste waiting for TFTT on prefill will absolutely compound at the "medium sized labs and businesses" scale.
reply
deleted
reply
Both are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.
reply
The issue is that the moment you want to run the more capable models that will no longer fit in a single 5090's memory, performance falls off a cliff.
reply
... or with llama.cpp with MTP.
reply
[flagged]
reply
I guess it is possible, but Apple has had very vocal fans for decades. I suspect, rather than astroturfing, it is just people who are in their ecosystem.
reply
Tok/sec is 0 on a 3090 for most of the models that the mac can run
reply
[flagged]
reply
> So given that, which one of these is true about you?

Well, if those are the only two options you can come up with it's pretty clear that this isn't about me or what I am, you have a false model of reality.

> Running very large models on Mac is unusable at 10 tok/sec.

There are plenty of examples of models running at well over 10 tok/sec that aren't viable on the 3090. In fact such examples are found in the review in the OP. Did you not read the article?

I think you're projecting pretty hard with the two options you've listed. Go touch some grass, you seem overly frustrated that reality doesn't meet your expectations.

reply
Since you clearly don't use local llms, allow me to educate you - anything under 100 tok/sec is USELESS. When you are coding, the idea is that you want to have a system that can generate files fast, hopefully correct on the first try. Cloud models do this. Local models, by nature of having less parameters and more quantization, often require more guidance and repeated inference to get it right. The antigenic harnesses that people set up around local llms leverage this.

Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.

Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper

The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.

So I ask you again, which one are you?

reply
> allow me to educate you

No thanks, you're not in a position to do that clearly.

> Since you clearly don't use local llms

I do, probably a lot longer than you have actually.

> anything under 100 tok/sec is USELESS

Objectively wrong. You sound like you're really behind and you're so myopic that you think coding is the only use case for local LLMs. I'm a professional software dev and that's the least interesting use case of local LLMs.

> Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec.

You clearly didn't read the article or have reading comprehension issues. The model is Qwen3.8-Flash-Next 4 and 5-bit quant, neither of which "conveniently fits on one GPU". Sorry that your hardware doesn't live up to your own delusions and can't even run Qwen3.8-Flash-Next at 4/5 bit quant. You are taking the Quen3.8-27B numbers, something that the article isn't really that concerned with, and trying to make it fit into your narrative.

> So I ask you again, which one are you?

Well I'm someone that suggests that you should touch some grass and reevaluate your personal issues. You seem angry. Perhaps it's best to figure your own issues before trying to figure out why people are excited about Apple hardware for local llms. I am sure the people that need to interact with you in society would be very grateful if you took the time to do this.

reply
Nice try.

A) He literally says "I tested a different Qwen model for the comparisons between Mac and PC." The model he tested has to fit on one GPU, otherwise the inference is dogshit slow as you are offloading results to ram. If you ran any amount of local inference, you would know this. Considering that Qwen3.8-Flash-Next Q4 is still 100gb, there is no realistic way to run this with a 5090. The model that was run was this https://ollama.com/library/qwen3.8:27b. And the speed of that model on a 5090 in terms of tok/sec is not 60 lol.

B) If M5 ultra runs 40 tok/sec on qwen3.8:27b (and lets assume its the mlx version to gain a performance boost: https://ollama.com/library/qwen3.8:27b-mlx), you have to be delusional to believe it can run 100gb models at 100 tok/sec lol.

As a bonus, in terms of use, its pretty well known that Qwen models are RLed to chase benchmarks. Check out https://huggingface.co/Qwen/Qwen3.8-27B versus https://qwen.ai/blog?id=qwen3.8-flash-next, using different benchmarks the 27b outperforms the flash next on agentic coding. But it matches it in other areas pretty well. So tell me again why you need 100gb models running dogshit slow at peak ~20 tok/sec?

It is so incredibly sad how hard you try to sound intelligent. But thats on par for the course of any person hyping up apple products, throughout apples history.

Considering that Apple probably doesn't want you to engage in this level of pettiness for their advertising posts, you have outed yourself to be #2. And Im not angry at all lol, you keep doing what you do, people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.

reply