Do they include footguns from pointer bugs?
https://youtu.be/AQf84KubJjE https://youtu.be/RAUsCwu8ekQ https://youtu.be/NtsJ5m6C7dU https://youtu.be/h8xO4PiJ--w
The internet is still alive
I appreciate the reference to RUSH: Red Barchetta in the final line.
Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.
Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".
Not necessarily for MoE
I think most people are getting 512 for running Chrome with a bunch of tabs open. /s
Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.
A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.
Is there evidence that's true though? I mean gross margins on subscriptions being negative since the API is seemingly very profitable (if the price is compared with the cost of serving very large open models).
As long as there is pressure from other providers serving cheaper models that are somewhat competitive without having to incur any of the R&D costs raising prices will be tricky.
Don't you mean an Apple to NVidea comparison?
I think it goes without saying. And it is eminently evident over last couple of decades that from compute to storage to meals 3rd part providers have saved billions upon billions of dollars to enterprises and individuals alike by providing these essential services.
[1] https://www.macworld.com/article/3238319/mac-studio-m5-max-r...
Plenty of other reasons to get excited about local AI, but I don't think cost is one of them.
And, yes, I know a current local model wasn't going to solve the Navier-Stokes problem, but I'm just using it as an example where privacy might be valuable.
It's a bold strategy cotton, lets see if it pays off for em.
Might be if RAM prices get much more reasonable its gonna be 1/3 of the price, but it's very much possible its gonna be half or more.
And if you're buing Mac Studio and not some AI-only locked down board it's possible to reuse it for other purposes.