upvote
You'd need the 256 gb memory model which will be expensive because apple has trouble getting capacity (got turned down by cxmt). And even then you can only run a 2 bit quant which is noticeably worse than 8 bit
reply
That’s not how that works. The hosted models don’t stay still in size and capability while Apple advances. Both will advance their frontier and there will still be a gap and developers will still prefer the stronger option.
reply
Local vs remote compute is a constant thread in tech history - mainframes and desktops then local and cloud compute (think Google Photos bs Apple photos - one indexes on device the other indexes in cloud). Now we have the next chapter local vs cloud LLM models.

There will always be a market for frontier labs in the cloud based models - these models will always be able to be bigger, and that will likely translate to doing things local models can’t.

Logically also we’ll likely get to a point where RAM drops in price as production ramps up, and local LLM is both capable and cost effective. This feels like it is coming for Siri / Gemini / Alexa personal assistant type use cases.

So I think the local LLM will become a thing in laptops and phones in a year or two, offering PA type use cases. Professional LLM services will likely remain at the frontier (and in the cloud) for the foreseeable.

reply
They will cost an insane amount as well. Maybe less than subscriptions or tokens. But running massive models on laptops with batteries and poor cooling doesn’t make much sense.
reply
Until hiding PII from the cloud LLM is a resolved issue, running local LLMs will remain a necessity.

There are workplaces that refuse to use LLMs because they fear the devs will expose sensitive data without care.

reply
> will cost an insane amount as well

We will get to a point where prosumer laptops that etch SoTA LLMs in removable silicon will be as expensive as cars.

reply
Every single MacBook built in the past half-decade already has an LLM built into the latest version of their OS.

But there's a significant difference in hardware required between running a 3B parameter model and a 700B-1T+ parameter model.

reply
I'm still waiting for Linux to topple Windows
reply
Sure buddy, all you'll end up with is a $10k machine that run gimped models at like 30tok/s for about 5m before the fan kicks in and it starts to sound like a turboprop, while offering maybe 30% of the context size of hosted models.
reply
The RAM shortage situation won’t be sorted out within the next year.
reply
Bold to believe it will be sorted at all
reply
If the margins are there it will get sorted.
reply
> run free LLMs locally at native speed

This reads like a hallucination. What does native speed even mean?

reply
for example, models running at like 100-150 tokens/second (or faster!) vs 15 t/s

(fable/sol are ~60 t/s, and OpenAI just announced their Cerebras partnership(?) for "ultrafast" mode of 750 t/s)

models aren't able to run that fast right now on our consumer/prosumer hardware. M5 Max for example has a memory bandwidth of 600 GB/s. a 5090 has 3x that, so running the same model on a 5090 is that much faster (provided the model is within 30GB).

running a bigger model on an M5 Ultra is still much slower than running it on a Blackwell chip with sufficient vram, CUDA being a major difference. if apple can bridge this gap, interesting things will happen... and just imagine if M7 Ultra has comparable speeds to Blackwell (or even Rubin)!

reply
Meanwhile the GB300 used by hosted llms:

GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional

https://pi3g.com/nvidia-gb300-specifications-including-memor...

If you think M7 will hit even 15% of these speeds you're very optimistic.

reply
A hosted instance serves multiple customers at a time. A local model only one.
reply
How many though? At 1m context you quickly fill a full gb300's 280gb of memory
reply
I assume they mean same t/sec as a SOTA cloud model
reply
There should be some kind of moratorium on new accounts. HN's always had waves of newcomers, but their impact was always limited. The wave passes and people either get filtered out or adapt. That doesn't seem to be happening anymore, since bots can churn out endless gibberish.

He did answer you though. Native is x10 the non-native speed. 50/50 that's not a bot; though it could be a meat-proxy

reply
> There should be some kind of moratorium on new accounts.

OC was registered in 2016 though? What do new accounts have to do with this?

reply
I am talking about my general impression not this particular occurence. I have a suspicion that someone/some entity is buying old accounts to bypass the new accounts penalty. I even created (https://chromewebstore.google.com/detail/hn-users-filter/ine...) to filter these accounts/comments (disclaimer: vibe-coded)
reply
I don't know why you'd want to burden your laptop with a large model. But I can totally see a new "developer workstation" product that's just a semi-large box that's optimized for running frontier open weights models for one to few users.
reply
I'll give you credit for at least offering a specific, somewhat unique take. But this is a pretty dumb take lol
reply
> Apple will release M7 MacBook Pros / Mac Minis next year

The latest on Apple is TSMC is stuck on the next iPhone due to lack of RAM. Good luck getting any Macs. Memory shortage is getting worse.

reply