Neither openai or xai. Anthropic maybe but not likely. Mistral is the most likely one because they're under the EU laws but I think they focus more on commercial these days.
In general, I'm a big believer in doing more with fewer resources, within reason, and think having local setups really helps me be mindful with what's happening under the hood with these systems and managing context efficiently to get high quality results.
A $20/month Gemini subscription is truly all you need, then yeah, sure.... obviously a homelab setup is a ridiculous alternative on a pure cost basis. For most people doing "real" work with LLMs 40+ hours per week, a more apt comparison would be one or multiple $200/month subscriptions. At which point the break-even point of a homelab is much sooner.
However, most people running homelabs are doing it for other reasons. Independence, learning, and/or privacy issues.
- Is that even worth the electricity price compared to api? - We don't ask that here
There is no reason to believe that equivalent level model output will be more expensive in 12 months, let alone almost 4 years from now.
Of all the good reasons to use local AI (privacy, etc), worrying about not having access to cheap models in 4 years is not one of them.
It's almost never a drop-in replacement, and having to check and adjust integrations and workflows with new models gets old fast. My task was perfectly solved by the old model, I don't need a newer, "better" one - especially at higher prices ("more cost-effective" my foot). Local models lets one choose a model and freeze the downstream integrations forever, without being forced on the 6/8-month upgrade treadmill by aggressively short, scarcity-driven hosted model-deprecation schedules.
The frontier labs have very high prices for inference. The prices are actually going down, not up.
Others running open models in the cloud is a nice third alternative, but not a solution for those that need frontier models, which I believe continue to need training as well as inference. These are currently heavily subsidized and hardware constrained several years out.
Don't get me wrong: I hope you are right, and I am generally optimistic about the future of AI.
But do you really think, in a world filled with examples of big software companies repeatedly taking away or hamstringing capabilities we've taken for granted, that you can just count on a big tech company hosting cheap inference on incredibly powerful models forever? Surely we have learned by now that these companies do not exist to provide a public service to us, and the government cannot always be counted on to have the best interests of the citizens in mind.
I mean, how many times have we seen this in just the past decade or two?
- Consistent attempts to pass legislation weakening or banning the use of encryption
- Exorbitant Reddit API pricing (still salty about the death of the amazing Apollo app)
- Google fighting against sideloading on android
- US gov't issuing export control directive to suspend access to Fable/Mythos
- US lawmakers considering ways to regulate adoption of open weight models
- Chinese officials considering restricting overseas access to their most advanced models
I can absolutely see a much more restricted, closed down, and expensive future due to a combination of government regulations (regardless of which nation is doing it) and big companies rug-pulling as the check comes due on all the billions of dollars spent to get here.
Yes, of course the most economical path is to hand over all your data and become fully dependent on a cloud provider who is already operating as scale, hoping that they won't change/remove models, hamstring capabilities, or raise prices.
If this were a thread about hosting your own email or blog or cloud photos, you'd have plenty of people out here telling you how easy it is to do it yourself instead of relying on Gmail for email or WordPress/Medium/Substack for blogging, or iCloud for cloud photos.
And yet, without fail, every single thread about self hosting local models seems to have some copy/paste form of this cost-savings argument.
Where is the appreciation for this cool thing GP built? Where is the appreciation for the desire to figure out how to host your own version of the incredible capabilities that were not available merely a few years ago? And why, on this site of all places, would someone advocate trading all of the knowledge and independence gained from learning how to host something like this ourselves in favor of throwing it all over the wall to Google?
Come on.
Because there are more privacy guarantees there, depending on the provider. "But what if they violate their contract!" is some pretty tin-foil hat stuff.
How is this any different than a business running their website out of the cloud, assuming you are using a provider with appropriate contractual terms?
You can care about tracking and ads but still be comfortable storing your backups in the cloud, and many have been for quite awhile, even sometimes without encryption - that is totally different than e.g. Meta actively trying to track you and understand your relationship graph and your purchases etc.
But what if they violate their contract!" is some pretty tin-foil hat stuff.
The foundation of these businesses is stealing IP in bulk.So don't use them for inference.
I’m pretty sure that all of those disclaimers that all the AI model makers have for you to sign off on to say that they’re not responsible for anything that might go wrong if your work gets copied accidentally and used someplace else wink wink?
You know the lawsuits for that particular aspect are incoming in the future…
But no way whatsoever I'm uploading my personal files, photos, emails, chats into cloud AI. No way.
It will be horrible to be dependent on an AI who is also be trying to sell you various goods and services.
We're going to need AI whose loyalty is to us and only us.
> They are things that I would not be comfortable sending a cloud provider
It's also an old machine that the commenter already has; it's intellectually dishonest to compare it to the price of a brand new, 4-iteration-newer machine.
a) model I pick will not 'suddenly' go away
b) I am sure my data stays where I want it
c) my inference mac can run other things if I need to
I pay for that.