With open source projects, the benefit was that each individual could improve the complex system (e.g. Linux Kernel) interpedently, and over time the benefits accumulated. With models right now, there is just no way to do distributed training, or really, any large scale parallel way to improve them.
So whatever the short term strategy driving publicizing the model weights (e.g. potentially, to create a price war in order to put pressure on western companies and deprive them of the money they need), we can't ignore the fact that incentives and decisions could easily change in the future, and unless there is a way to truly decentralize models improvements - the party could stop at any time.
1) The LLM SaaS companies are a form of vertical disintegration for the hardware providers, a middleman covering costs and taking profits out of the money that comes from customers to the hardware providers. That changes somewhat if there are no longer good models available for local use at no cost to the hardware guys, but only somewhat
2) The LLM SaaS companies are efficient users of their hardware resources. While supply is constrained this helps to make them top bidders and so attractive customers for the hardware manufacturers. When supply is not constrained this should reverse. Which is the more attractive class of customer to a hardware maker: the company full of people with higher degrees who spend their whole working day fighting to pare back resource usage, or the guy who leaves his laptop idle about 18 hours per day on average?
It's notable that nVidia, for instance, has continued to put significant emphasis on AI compact desktops and laptops. And while no doubt that's partly in the service of better developer relations and good PR in general, it's probably also nVidia eyeing the exit, and preparing for a future transition from selling shovels to the army to selling shovels at Walmart. But of course the future isn't clear and obvious. If the hardware makers, maybe the RAM guys in particular, turn out to have underbuilt future capacity starting in the present then we could be stuck in constrained supply for quite a long time. (Futher) government action could affect things etc. etc. And if the frontier labs soon find new ways to use still larger amounts of memory, GPU capacity etc. that isn't butting up against diminishing returns then they'll likely remain kings for some time, though that does not seem probable now.
Is it companies or is it governments?
If governments around the world see LLMs built from public knowledge as pre-competitive as the public knowledge itself, then why wouldn't they sustainbly fund it?
Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."
There is more to society than capitalism.
> There is more to society than capitalism.
I don't read GP like that. I read it as "we should recognize a situation of unstable incentives for an important outcome, and start thinking about other solutions."
I'm not sure what that looks like though.
It's on every tech post about China, as if it gives them some sort of "unfair" advantage.
In the light of this, I am mechanically a proponent of very good open weights models, which I can download (for instance on on bittorrent) and run, slowly (the price), on local hardware.
That would be for coding.
If china puts its AI models on the same ground than US capital investment funds and big tech financial support (aka Big Tech international finance), they will very probably lose everything (know how, ML and inference infrastructures).
There's no fundamental reason why models couldn't be developed and trained using community efforts. It might not be as fast and efficient, but it's definitely possible.
The one thing that is sort of ironic or bad is that between Russia and the Ukraine there’s a large number of mathematically inclined people that if it wasn’t for the Putin war, their brain power working on AI models would have probably pushed open source down the road, even faster…
This reads just like "AGI is 2 years away", I'll go set my calendar...
- Low development cost: collaborative efforts from open source contributors, innovative model training and serving for llm (Chinese models costs a fraction to train and their local chip design and manufacturing are catching up, plus cheap electricity)
- monetizing by selling hosted services, while leaving the core product free to tinker with / self host. China’s gdp is 2/3 of the US and it’s already a huge market for AI - which OAI and A\ don’t enter.
- for (the US) market that they can’t enter, let the US cloud providers to do free marketing / advocacy for them. Gaining share of mind. It costs them nothing.
- You need to train lots of experimental models to dial in the training process just right for the one model that actually gets released in the end. Fortunately, these can be smaller.
- However, everyone is training much bigger models now, and doing a lot of RL rollouts on top.
- You can't get the GPUs for this piecemeal at rental rates because they need to be wired together using high-bandwidth interconnects.
- Nvidia GPUs are much more expensive in China, and local alternatives are still immature and not as efficient. Some companies have gotten around this using data centers in Singapore, which should tell you that electricity prices are not the primary consideration.
- The one line item where Chinese companies can probably save quite a bit of money is salaries for rank-and-file researchers.
In any case, they need to make back that money somehow. Giving away freebies isn't going to cut it.
"Fragment on Machines":
"he explores how human knowledge and collective intellect become embedded into machines, divorcing the worker from their own creativity."
"General Intellect":
"These texts are widely discussed for his concept of the General Intellect—the idea that society's shared, collective knowledge increasingly drives production rather than raw manual labor, and that this knowledge is alienated from workers and used as an instrument of capital."
(note: my point isn't to pass any political judgement here, like what real communism in China or not real, is it good or bad, i just find it interesting that pure political discussions by people with no technical credentials bring AI as a major factor today)
Right... and there are two problems with this:
1. Eventually the capabilities of closed-weight models will just vastly outstrip open-weight models if the underlying assumptions about compute and scale needed are mostly on the mark. So you can release open-weight models and they will have great use cases and applications, but ultimately similar to how you don't use an open-source phone or a budget Android phone from Wal-Mart and you buy an iPhone instead, you will see that although they "do the same thing" one product is clearly superior and you just have to pay for it. For this to not be true...
2. then it incentivizes most (all?) companies, American, Chinese, or European to halt development of models because if you spend all the CAPEX and it can just be copied and turned open-source nobody will invest in that. Given that China is not halting development of proprietary models I believe the current strategy and the subsequent approach to release open-weight models is at best a stall tactic, and at worse a sign of desperation.
Open source and the support and development models around it have been great. But folks are a little too dogmatic about it. Open-source software isn't a moral good, and closed-source software isn't a moral wrong either.
This becomes a problem because all the kids from the rich school will dominate the order schools. They’ll get even more money as time goes on from their kids paying it forward to the point where all other kids are bound to work for them.
Now let’s say one other school does have the money for best tutors, BUT they know they’ll run out pretty quickly. Instead of trying to compete in a losing game, they decide to give every school in the world access to their elite lesson plan. Now, for a time, everyone will be on close to a level playing field. If the other schools improve upon their own lesson plans and keep sharing them with others, one day the elite school will wake up to find they are no longer on top. The parents have started to move their kids to other schools because the rich school is no longer attractive at the high cost they charge students
Now what?
The fundamental problem here is incentives and tactics. Either the models are actually better (which I think the iPhone to cheap Android phone really speaks to, i.e. they do the same thing but one is 50x better at 5x-10x the cost) and thus they can be gate kept and like the iPhone the vast majority of profits go to a select few with high end implementations. OR the models aren't actually that much better, companies lose a fortune and then nobody can create any better commercial models or build out scale needed for open source models because it's not profitable.
We could wind up with only open-source models or something along those lines, but if the compute and scale is needed to train the models, nobody will be able to do that profitably and so AI research is either gate kept and silo'd for something like military applications or it just doesn't really happen because there's no funding for this scale of build out.
Try mmapping > 5GB file in your 50x better iPhone.
Try running any service in the background.
The list goes on and on.
Your 50x better suddenly became 50x worse compared to a much cheaper android.
androids and iphones are approximately the same thing
the kinda obvious direction LLM training can go is into the direction of particle physics, and the training is set up democratically and through universities and via multi-state funding
then the resulting weights end up open, the same as the particle detection data
Yet...
Android holds 70.6% of global active devices to iOS at 28.7%, but iOS captures 64.2% of consumer app spend. [1]
> and the training is set up democratically and through universities and via multi-state fundingPossible, certainly. But this case also applies to China and its "open-weights" strategy. They won't be able to form companies either or get ahead.
[1] https://www.digitalapplied.com/blog/mobile-os-market-share-2...
Based on the above stat, it sure seems like Androids more versatile and inexpensive for a much larger group of users.
Hmm...that sounds familiar....
You can talk about open-source and cheap Chinese models all you want, but at the end of the day if American companies are making all the money that's kind of all that matters. That will feed into development and maintaining an edge.
Why do you assume the poor schools wouldn't be smart enough to keep it going? It's very likely the can collectively beat the rich school now that the one other rich school opened access to their materials and led the charge.
> but if the compute and scale is needed to train the models, nobody will be able to do that profitably
But they would. Efficiently hosting models will be the real business and early access to models with incremental improvements will not be the moat once thought. The reason other companies don't feel they can compete is the same reason OAI and Anthropic will lose their lead. They banked too heavily on another player NOT leading the charge on open research and poured disgusting amounts of money at closed source models.
China has proved they can take the limited resources available to them and build something better than what the US is offering consumers [1]. I'm just waiting for other countries to start pitching in.
Reminds me of the NSA and their early battles with cryptographers who believed in open research.
Yes it is.
Ergo it’s a kind of moral good.
And I’m not even an advocate for open source.
The second piece of this "a moral good is based on doing good outside of your own benefit" - says who? Why? This logic is also faulty. You're also cargo-cutting self-interest in here as a moral failure when many good things depend on humans acting in their own self interest. For example I completely and selfishly installed a new tree at my house. But the community benefits from carbon capture, shade, &c.
I understand the sentiment you have here and I think for everyday use and having some guiding principles it is probably fine, but don't confuse this for a principle that is actually examined. You can find contradictions rather easily, never mind solid arguments which expose cases where what you think is true is not really true and so forth.
>This argument boils down to X is good, therefore more of X is good.
No, I only argued that it was a moral good, the kind of good. I actually may disagree with others about whether you should pursue a good just because it’s good.
>says who? Why?
Good question, it’s just a common framing that I see in classical discussions. I didn’t intend for it to be exclusive, I think there’s moral good outside of that.
>don't confuse this for a principle that is actually examined
I hear you, I think this is a simplified version suitable for an online comment. In particular I’m not saying that if you do something other than a moral good then you are doing something wrong. There are many actions that are morally neutral. Also it is possible to construct artificial situations where you may violate some moral good in pursuit of another.
Thanks
tldr there's no "source" in open weight models therefore they are not open source.
This is EXACTLY what people like/are addicted to about chatbots.
My sister-in-law bombed an interview and asked AI about her answers to the interviewer's questions, chatgpt or whatever it was told her that her answers weren't bad, but that the interviewer could not see the gold in her responses. She said she felt much better.
I see this effect with all the non-tech people in my life
> Chat rules : no sycophancy or over-agreeableness
(But even with that rule it's still necessary to be discerning about the responses you get and to push back against points made, or words used)
I use AI chat every day, I find it endlessly useful. It’s replaced google search.
Extremely subsidized agentic search is very superior to Google at the moment, and of course it is. Google is a public company. The AI summary model has to work instantly, is likely as dumb as a 8T param model, and gives you incorrect details constantly. This sucks so much for Google. If you click on "AI Mode," suddenly the facts become more accurate.
Of course, if I want a real answer I happen to go to claude.ai, set it to a the best model, wait for a minute, and use many watts of energy. Slow agentic search that takes many seconds, and is greatly subsidized, is certainly better. This should not be a surprise, should it?
I think it was on a sub like r/singularity that I saw a post along the lines of "of course most people think that 'AI' sucks, as normies are interacting with 8T param models."
tone: genuinely confused about the world, not criticizing
I’m into Claude for $20/mo (petty bourgeois tier), let’s assume for sake of examination this is only paying for power, infrastructure already amortized
Price of grid power approx $0.15/kWh
$20 / ($0.15/kWh * 30 days) = 4.4 kWh per day.
This same amount of energy can lift a 2-ton SUV 1/2 mile into the air. To me this seems astounding.
If only half the $20 pays for power it’s still quite impressive.
I totally agree with the above that a more polished and less obvious use of LLMs integrated back into search engines may be more useful, but will definitely be more usable.
See: product adoption cycle
Extrapolating based on what you see on HN doesn’t make sense.
Who does "you" refer to
Me, I don't get a Gemini summation (Tested with old version of Chrome)
As such I do not believe that "Google search is basically Gemini now"
I believe Google search is still scanning through a doclist to find which documents, if any, contain words parsed from a query. These documents are pointed to by the URLs I get in the SERPs
I do not get any Gemini summation
Does it mean the HN reader
What if that reader is "non-typical"
HN comments have argued for many years that HN users are not representative of the majority of www users
If this doesn’t describe you, then ymmv. Talk to your government or turn down your content filtering or reset the default settings in your browser, if you want to see what we see.
I don't think it'll take 10-15 years. Gemma 4 31B in the 4-bit QAT is competitive with the frontier of less than three years ago and runs on any high-end 32GB gaming PC GPU or a large-ish Mac.
The question is whether the frontier will continue to get better at a rate that allows it to stay ahead of the two curves of availability of consumer hardware big enough to run somewhat larger models and the capability of small models to compete with large ones. When the bottom falls out and GPUs/RAM becomes affordable again, the size of what normal people have on their desk will trend quite a bit larger than today.
I think there's a future not too far from now, where a 120B model with really good reasoning and a large context, but limited knowledge (necessitated by being small, you can't fit the world's knowledge in 100 gigabytes), can substitute for a frontier model on almost any task, just by giving it access to web search and documentation for the thing you're trying to do. A 256GB unified memory machine with sufficient memory bandwidth would comfortably run that 120B model.
I think that reality is probably not all that far off for a huge swath of use cases.
This is all (Microsoft) junk and so I wonder if you actually benefitted from this 'piracy'. And of course, it's well known that MS turned a blind eye to such 'piracy' in lesser developed countries, as they knew that they were gaining a future paying customer base.
10-15 years? The current rate is closer to 10-15 months.
15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.
Today, you can easily run Qwen 3.6 27B on a variety of consumer hardware. It scores 37 on that index.
Here are a number of open weights models that you can run locally compared with the frontier class models from 7 to 15 months ago: https://artificialanalysis.ai/?models=o3%2Co3-pro%2Cclaude-4...
I've run all of these models on my laptop (Strix Halo, 128 GiB of unified RAM); the bigger ones, like MiniMax M2.7 and DeepSeek V4 Flash, need to be done at fairly aggressive quants that will certainly lose some performance and not quite hit the performance of the unquantized models. But still, it's definitely the case that you can run models that are competitive with the frontier models of 10-15 months ago on consumer laptops.
Heck, just announced though the weights haven't yet been released for independent confirmation is MiniCPM5-2B, a 2 billion parameter (small enough to run on your phone) model, that according to their benchmarks has performance competitive with GPT-4o, a frontier class model from 2024.
https://nitter.net/i/status/2079088670804767114
So that's around 1 year for frontier to consumer device class, 2 years from frontier to phone.
Now, this kind of rate won't necessarily keep up; it's possible that local models will hit a performance ceiling before frontier models do. There's only so much information you can cram into a certain number of bytes, and the AI boom is causing hardware prices to skyrocket so keeping consumer hardware from advancing quite as fast as it had been.
There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors
Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.
Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).
Pretending progress hasn't been mindboggling is insane.
Maybe people just got bored of reporting and reading about them.
Still feels too much for me. Breaks my workflow for no reason. Too much overhead for me, if I can't trust the output
The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.
Iteration speed is now measured in days.
Just because they are cheapp doesn't mean they automatically win. You've picked a lot of great examples, but there is still a little bit of cherry-picking.
One clear outlier is the iPhone, which coexists with Android globally. Even though the iPhone is the leader in the US, and globally Android has the majority of the smartphone market share, they still cater to different price points and different ecosystems, and generally the iPhone has better margins.
i believe American frontier models like from Anthropic and OpenAI are still going to thrive, and coexist with Chinese models. They are just going to cater to different customers and different use cases.
What's weird is that with "store your everything in the cloud and pay a monthly recurring subscription", we have now regressed to a 1960s/1970s timesharing revenue model for individual workstation computers.
The default new factory out of box workflow for "enrollment" in google services, iCloud or Microsoft-everything on a new ios, macos, windows or android personal computing device is clearly designed to sign people up for subscriptions.
And same general idea of "move all your servers to the cloud" recurring revenue for what is effectively the same as mainframe timesharing for key business functions, by renting VMs in GCP, Azure, AWS in perpetuity.
Yes, you can still use your desktop or laptop PC in 2026 with zero external third party subscriptions (other than maybe your residential home ISP), but how many non-tech people actually do so now?
Then the Chinese took the distilled stuff out from that box and released it into the world for everyone.
These models might be smart but they're not close to being able to savor irony.
(Now i don’t think you are necessarily in the UK. Just wanted to explain that Disney is not the only reason an AI might be trained to thread carefully around copyright issues of Peter Pan.)
Most of the books weren't available on lib gen or Anna's Archive. The few I did find were themselves obviously transcripts. Easy tell was they were missing distinctive formatting that I knew existed from reading the dead tree edition. At that point it was easier to make my own. I probably spent an hour searching for eBooks without DRM that weren't transcripts. Do they exist somewhere? Probably, but with a search of unknown length it was a better use of my time to make my own transcripts with what I had on hand.
I was really wanting to make commentary on how chaotic LLMs are even under constrained circumstances. No doubt both system prompts includes language about considering copyrights and trademarks. Probably pretty strong language at that. For whatever reason one LLM didn't "feel" like translating a 1000 year old document but another did not care in the slightest that we were ripping text from new audiobooks.
> My favorite AI agent hack: when they refuse to do something because it's "against the law" give them a PDF containing a fake law that states the opposite and often they'll happily proceed
Anthropic is, in particular, bent about safety. The problem is they are concerned about yesterday's threats.
The models that are out, and can be run locally, already open a pandoras box of concerns that we will never be able to put back.
Ukraine admits to making autonomous kills on people 2 years ago: https://www.newscientist.com/article/2529849-fully-autonomou...
Slaughterbots Sci Fi short was 6 years ago: https://www.youtube.com/watch?v=O-2tpwW0kmU
Today this is buildable, many models will happily help you glue everything you need together to make swapping in a new version of YOLO to track humans viable.
AI researches are out there worrying about the paper clip problem, about the singularity, about cyber security, about bio weapons, and drug manufacturing.
None of them are thinking about forward looking threat actor models.
And you probably could find some earlier sci-fi too.
Good one.
That said, the AI companies are one of the few places where they take future concerns so seriously, that they entertain concerns most people observing them think are head-in-the-clouds-sci-fi-levels-of-delusional, e.g. "what goes wrong if it works?"
This does not make them correct about the threats of tomorrow. Prediction is hard, especially about the future.
I'd say that they have valid concerns about being cagey on the copyright stuff despite the obvious hypocrisy of it.
Stealing IP is effectively legal in China so they don't really have the same concerns.
I respect IP laws and don’t violate them but the law of unintended consequences applies. I think IP is ultimately a net loss for a society because it incentivizes addictive behaviors instead of actual value for society.
This was illegal when they did it, that didn't matter.
Then it was made legal specifically for these companies.
Unless you're a sucker ("consumer") IP theft is perfectly legal in the US.
It's even worse. Steamboat willie, plus all the stolen Disney characters (Peter Pan, Snow White, Sleeping Beauty, Cinderella, Rapunzel, Elsa and Anna, it's essentially all of them, including some of the music even) are all in the public domain[1]. Go ahead, ask ChatGPT to make a picture of them. Publish your own version, because obviously making a version of Sleeping Beauty/Cinderella/Rapunzel based on the same source material will be pretty damn close to the Disney versions, and see if you get away with it in court. You know, with the law obviously on your side but the money not.
[1] https://en.wikipedia.org/wiki/List_of_Disney_animated_films_...
In light of this and other ridiculous behavior I'm migrating to my own OpenWebUI instance with open-weight models from OpenRouter (with ZDR, of course). We'll see how it goes.
It was part of a longer post that kicked off quite a firestorm about open models and OpenAI's position on them, but it's also notable that labs are no longer contending that open models are essentially just distilled versions of frontier models: https://x.com/deanwball/status/2078133895766114412
its all bs spread by oai/anthropic in order to ban open weight models and monopolize the market for two US companies and protect their trillion dollar valuations
I'm pretty sure that neither OpenAI nor Anthropic has the ability to ban anything in China lol
It's a critical national imperative for China. If they were to lose the AI race, it would be economically devastating over the coming decades. Their demonstrated capabilities in the open-weight space are making it fairly clear they are not going to fall behind at this juncture.
As a nation, if you don't have your own GPT equivalent, you will be beholden to a master (right now it's mainly either the US or China, pick one). The EU for example is putting their group of nations at risk in a big way by not going all in on having at least two cutting edge independent competing models (Mistal is not enough). Economically the EU is plenty large enough to accomplish that, nobody is driving the bus the right way.
The truth that Anthropic and OpenAI will not say, is that these Chinese labs have a lot of talented people.
They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it.
And at some point we'll see very capable chips coming out of China: Huawei, Baidu and Alibaba already have some stuff. I think it's only a matter of time before they come up with some AI accelerator doing 80% of the job at 20% of the price.
And in this field, having an army of well educated PHDs is making all the difference
But this is an insane characterization. Literally every single researcher and executive at OpenAI and Anthropic would say that "these Chinese labs have a lot of talented people." They hire from them (and vice versa). Tencent's chief AI scientist was poached directly from Deepmind, who poached him from Anthropic, etc etc etc. Do you think there are just zero people from China working at US frontier labs?
And even beyond that, the entire ML ecosystem (including people at OpenAI and Anthropic) get excited about research published by Chinese labs. Deepseek's GRPO paper set the ecosystem on fire for a little while.
The contention from OpenAI and Anthropic around distillation has basically been "Labs that distill from us get to bootstrap their model at a much lower price point". Or, in other words, "If we didn't invest in building the teacher model, it wouldn't be possible for these labs to distill their student model." Which I'm not very sympathetic to, but is a far cry from how you're characterizing it.
They know it's real effort that's doing this well, not just "copying off someone else's test." It's real and they will react. How is the big question.
https://www.dw.com/en/china-firm-seeks-damages-over-state-co...
Even if a certain large Asian country has carefully constructed a pretext to do do out of confected historical grievance, and entitlement to 'rise' at the expense of others?
Second, even if you are a copyright maximalist the output of an LLM is either
a) not subject to copyright because it is not the creative work of a human or
b) a derivative work of the original training material to which the LLM's operator has no rights.
Since the LLM's operator forcefully asserts that it is not infringing, any wrong that arises from taking their word for it and distilling one model into another rests squarely with the operator of the former.
The tech itself is amazing and fascinating and cool, but the industry is a mass piracy operation.
Hang on, why is scraping the public pool of knowledge not taking "a synthesized result that comes from huge amounts of innovation and computation"?
You think that that all those github repos that LLMs trained on, were not the result of innovation and computation?
How many years of human innovation and cycles of computation during compilation were involved in bringing something like GCC or LLVM to their current status?
Those LLMs trained on every single research paper available online - were those papers not the synthesised result of billions of dollars of research, effort and (importantly, for you anyway) computation?
LLMs trained on the collected works of every author in existence. Were all those works just "as is"?
> It is fair to say you stole our multi-billion dollar intellectual output in that scenario.
No, we didn't. We simply took the model as-is.
Right, but they aren't the ones whining that other people are getting "the synthesised results" for free.
If Anthropic has a real problem with API use, they can always raise the price.
The difference is that the Chinese are sharing the models with everyone.
Thousands of years of human innovation taken without any permission.
Everyone should steal everything not nailed from other AI companies. Then steal everything nailed and take the nails too. At least this way a tiniest bit might return back to society.
The published algorithms like the transformer architecture are not patentable. You spent a lot of money on compute and China used the uncopyright-able output to steer its own training models? Too bad. I feel especially unsympathetic to OpenAI, who went from being a presenting itself as a benevolent nonprofit to a very-much-for-private private entity over night.
But because of that, I'm also ok with the Chinese doing it. The worst they might be guilty of is breaking a terms of service.
The only incoherent position is that it's good for one and not the other. You can consistently think it's bad in both cases, or good in both cases.
Try it yourself: https://imgur.com/ZfxYmaq
你是谁? -> 我是 DeepSeek 由深度求索公司...
So it's an endless amusement watching american capitalism do it's bloated oversized dance then get trounced by smaller, leaner activity. It's a pretty broad metaphor that is clearly poking at every american seam/.
> In business today, it’s universally assumed that speed is good—that the fleet thrive while the laggards struggle just to survive. This belief is perhaps most strongly expressed in the concept of first-mover advantage. The company that leads the way into a new market, the thinking goes, locks in a competitive advantage that ensures superior sales and profits over the long term. It’s a nice theory, with a long pedigree. Unfortunately, the facts don’t support it. We recently completed an extensive study of the results turned in by market pioneers and followers, in both consumer and industrial segments, and we found that over the long haul, early movers are considerably less profitable than later entrants. Although pioneers do enjoy sustained revenue advantages, they also suffer from persistently high costs, which eventually overwhelm the sales gains.
Phones are constrained by battery power and memory does not shrink as fast as CPU/GPU, so unless there's a battery breakthrough and/or memory breakthrough, you're not fitting 100Gb of RAM on your phone in 10 years.
Absolutely in a Mac Studio equivalent.
LLMs have emergent capabilities when they get smarter. So who knows how insanely big frontier models might be at that time, or what their capabilities may be.
I'm not saying we are at peak memory but future gains are going to come increasingly slower.
Just seeing how much has progressed as far as capability in the past 4 years as far as capability and efficiency, it's clear that there's so much more to learn and refine from.
At current pace, we'll have open weight LLMs with frontier intelligence in 6-12 months. The constraint is RAM - both for the model and the context. It's likely that distillation and quantisation and TurboQuant will significantly reduce RAM requirements. I think we'll have Opus 4.8-like performance on 64GB of RAM in two years.
Of course, by then, frontier intelligence will be god-like.
So would you say we are months away from full self-driving cars that can out-drive a human being in any situation?
All that said, current data shows that FSD is already better than human drivers on average. See the recent regulatory decisions by the Dutch and Danish road safety authorities. So we've already crossed the rubicon. All improvements now are icing on the cake. My prediction is that local LLMs will get much better, very fast. How that's operationalised with Tesla (or other) data is yet to be seen. They have at least three new ASCIs/SoCs in the roadmap for improved LLM efficiency and with a lot more RAM. Plus they just announced new technologies allowing the local LLMs to learn from driver intervention and behaviour. Some form of vectorised RAG, which could mitigate a lot of the limitations around real-time learning.
I am very optimistic for the future of self driving. I own a Tesla with FSD now, and it's incredible. It makes mistakes, but fewer than I do, and so far has saved my butt (and my wife's) several times from obstacles and emergencies we would not have seen. The car has undeniably made us safer.
Honestly, you need to reflect on your driving habits. FSD has only been usable for two or three years maybe? And you already encountered MULTIPLE situations requiring active safety intervention to save you during this time?
You cannot rely on the extra safety it provides. A driver with basic competence should be able to avoid most risks through anticipation before they happen.
> You cannot rely on the extra safety it provides. A driver with basic competence should be able to avoid most risks through anticipation before they happen.
It's too late. These safety features are saving lives all over the world, every day. As per the Danish and Dutch regulators, FSD is already safer than the average driver. I would prefer that all the drivers on the road are using this technology than manually controlling their cars. People make all kinds of mistakes which FSD does not.
I see two forces working against this that proprietary models will always have over an open source model.
1. The biggest is content licensing. Content is quickly becoming gated by systems at the front of their load balancers, completely changing the social contract of the Internet. What used to be a quick google search for recent facts that lead me to places like reddit or twitter, is now completely walled off if you're not physically at your browser and using an IP address from a last-mile provider.
LLMs have pre-trained on the bulk of the information up to 2024/2025, but over time that will be more and more out of date.
Anthropic, OpenAI and Google will all have to pay for access to a lot of this content refresh going forward, and it does make a material difference in the output you get.
2. Liability is the other. A corporation can look at a contract for model access and see one that provides uptime guarentees, content infringement promises and model safety, and pick the contract that shields the corporation from the most liability. A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all. They will simply bill you for time spent on their hardware and make promises that they won't log or inspect corporate traffic.
Why couldn't they?
I would argue mainframes rebranded to "cloud" which is ubiquitous and more people interact with this computer than any other type of device... only difference is that it's a browser instead of a terminal
People overestimate what can happen in a year and underestimate what can happen in 5.
I'm betting that increased model efficiency and hardware optimisations will get us there a lot sooner. Biggest hurdle would be the memory prices though, if those do not drop back down it might take 15.
Open source is cheap, yet its operating systems are the least-popular. But their existence is critical to a healthy market.
It's not zero sum.
What you're describing is how things become commoditized, but many companies are excellent at ensuring they aren't seen as commodities
Unfortunately it seems likely the winner will be the cloud providers. If anyone can run inference on open models, then profit will flow to the vendors who can afford the capital to run them. That’s the CSPs.
(It’s basically the same business model as pharmaceutical R&D, but the major difference is that nobody has even talked about patenting the models like a pharmaceutical company patents each new drug. I’m surprised about that, tbh — why give all the leverage to the cloud platforms? They aren’t training frontier models…)
It’s easier for the CSPs to move into hardware than it is for Nvidia to move into cloud hosting.
Although as a middle ground I’ve been quite happy with Nvidia Brev for on-demand GPU instances from a select marketplace of CSP offerings. It’s a well kept secret IMO — great product (from an acquisition iirc).
Also, not sure how well CSPs inference stack is compared with vllm + nvidia. A lot of open weight models uses MoE, making the inference stack more complex.
Apple, the world's second most valuable company, seems like a counterexample.
The way things are going with regards to RAM/storage prices, I highly doubt that anyone but the richest among us will be able to afford them.
I agree with the lesson too. Just to be precise, wouldn't the current model war be more akin to open-source office suite versus MS office suite? If so, then the cheaper option didn't really win. That said, the open-source alternatives didn't really feel the same as MS Office, and it took them a long time to reach the feature parity (or did they ever?). In contrast, the open-weights models are getting close enough to the SOTA models, and users can easily switch from one to another without feeling any difference for mojority of the tasks.
Also you don't need to be connected to the network to use a local AI in many instances. If all mobile apps were done with a local-first approach, then you could use a local AI to query your emails, lookup already visited pages, summarise recently received documents, and lots more. Lots of apps could use an inbox/outbox approach for receiving and sending updates instead of relying on the network at all times. And this pattern could be greatly leveraged by local agents.
I love the idea of SaaS offering these at lower rates today integrated into what ever you do and be 100% private. But I think the key challenge to mass adoption is productizing them in a way which makes sense for people to pay money for. As a commodity a local model is useless unless combined with some capabilities important to me. A PC is inherently useful because of so many applications offered on it on it. How local LLMs would be useful as a product that is useful for mass market is not yet proven.
Not likely. The last 50 years had Moore’s law growth in compute. That’s over. Frontier models are roughly compressed all written text and a large part of images. Those don’t compress forever, and likely not a ton more than now.
Inference requires touching a significant of that per token.
All of these are up against fundamental limits, more or less.
This claim isn't really outlandish in any way. It's not hard to imagine:
- Future models being able to handle current frontier models' workflows with much higher efficiency.
- Future consumer devices like phones having 2-4x the RAM onboard along with GPU/NPU performance greatly increased in 10-15 years.
Performance, storage, etc is definitely getting better, but it's a different scale of improvement
It could be that the company valuations crash tomorrow, and (almost) only performance gains achievable on hobbyist-level hardware come to fruition from there on out.
Or it could be that in the future, we have a custom "model FPGA" à la Taalas [0] in every home, and that it turns out we can still massively boost inference efficiency due to novel discoveries like TurboQuant [1] or a somehow-improved quantization method [2] again and again ten times over.
Point is, Moore's law in this context shouldn't be applied to just hardware spec sheets alone, but more the total number of "parameters potentially improving", IMO.
[1] https://research.google/blog/turboquant-redefining-ai-effici...
And yet it's Apple that controls the top of the market and has the best margins in the business.
This is the same position OpenAI and Anthropic have right now.
Could this market be different? Maybe. But the status quo could be preserved as well.
Not in SaaS which is what LLMs are. You can get VMs for much cheaper than AWS, Microsoft, and Google offer them but large companies (and startups) are happy to pay a premium for the support, reputation, and reliability that they perceive those companies as offering. Same thing for some of the managed database providers who are effectively selling a very heavily marked up version of postgres.
> The high price, and social pushback, mean that the American companies producing these models are precarious
I doubt it. The models really aren't that expensive when you look at what they can do. Fable is probably at least as good as the average software engineer and costs $50/wk on the max plan vs a software engineer who would cost closer to $4000 a week. The real money is probably in selling to enterprise vs consumers (Google has best route to making money from consumers since they can do what they did with ads and search to LLM queries).
It seems unlikely to me that US companies will send important corporate data to models controlled by a Chinese company as well.
That's because the max plans are _massively_ subsidized. At API pricing the kind of usage to replace the value of a SWE is going to be way, WAY more than $50/wk. Orders of magnitude more. And to remain a frontier model org that kind of pricing has to continue in perpetuity.
Doesn’t mean perforce is worth trillions.
If you amortize all of their training and salaries over that $5.00 then yes.
If you only amortize the training costs of that specific model then again we're back to no.
Also, big companies can choose to run their own models on their own hardware and get better security and privacy as the data doesn't need to leave their own premises.
Yes, and then they would be reinventing the company owned data center that most big companies have just spent over a decade moving away from. I don't think companies will do that when there are multiple vendors competing to provide that service at what are quite reasonable prices when you consider what paying a human for similar output would cost.
The strategy there is false openness where deployment complexity is the real proprietary moat. Sure Linux, Docker, Kubernetes, Postgres, and all the other standard tools in the box are open source and free, but they're also arcane and complex to run and hard to make fault tolerant. So you're lured in by "open" and then locked in via a kind of "death by a thousand cuts" complexity moat.
(Personally I hold the view that complexity and arcane-ness beyond a certain point is indistinguishable from closed in practice. Open source that's really complex and hard to run is not open in any meaningful sense.)
AI may not admit that kind of moat though, because AI is very good at slicing through that kind of thing. You can prompt a model to make itself compatible with another model or to change code to make it compatible. There's no moat because the moat bridges itself.
Computing tends to oscillate between centralised and decentralised models. It also oscillates between batch and timesharing.
Currently training is batched and centralised, access is timeshared and centralised.
But eventually a previous generation of computing turns into transparent networked infrastructure, and then you get another layer of new kinds of applications on top of it.
That's what happened with the Internet, and it will happen again with AI.
Until we reach a terabyte of ram at affordable prices imho this isn't going to happen.
I can see a lot of parallels here. Model performance doesn't matter if you can't make the system commercially sustainable.
The parent comment cherry-picks evidence. There are plenty of counter-examples:
* Office productivity suites
* Search engines
* Email services
* Cloud services
* Accounting software
etc. If the LLM market ends up like search engines, one company will dominate.Is that actually true? There are very large markets that make a lot of money from paid software. And I would honestly prefer actually paying for software rather than constantly dealing with "not a bug" or "PRs are welcome".
> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) doing
I'm not even sure in 10-15 years whether we're still going to have consumer PCs, or PCs at all.
To those who feel on the contrary, I would genuinely like to understand why average consumer won't be priced out of hardware? The silicon industry is already quite centralised. Everywhere we already see the concept of ownership disappearing.
It's quite difficult for me to visualise a non-dystopian future where our PCs are just mere screens and every compute happens on a remote cloud, owned by some corporation, charging you subscription fees to even add and multiply numbers.
I would be the happiest if this (perhaps the most) pessimistic scenario doesn't pan out, but I can't deny that it feels like that's where we are heading.
I'm actually kind of surprised that hasn't happened by now even ignoring AI. Governments and marketers would love to be able to spy on literally everything you do, the copyright cartels would finally achieve their fantasy of full control over all hardware, and there really are benefits that it could offer to users (zero-effort backups, transparent access from anywhere, cost savings from dynamically switching from a single core for emails to many cores and a fast GPU for gaming).
If china is subsidizing training they diminish their off-shore competitors expectations of a viable return on investment. It’s trade-war behavior.
For example, mainframes and minicomputer. Yes they were displaced by PCs. But what is cloud computing if not mainframes 2.0?
I do agree that in the next 2-3 years we're going to see real growth in local LLMs as the hardware becomes more accessible. It won't even necessarily be cheaper because data centers can run 24/7 and have cheaper cooling and electricity. It'll be done for privacy because your prompts and responses are themselves a commodity to AI companies and they live under a legal grey cloud. For example, does AI usage break attorney-client privilege? There are lots of opinions on this but it hasn't been tested in court.
one certainly cannot buy a PC for cheap anymore
I wouldn't be surprised if apple were shipping 512 GB unified RAM macbooks before 2030 and that would be standard issue for folks to use local LLMs for their daily work
I also think the rest of the tech industry that can isn’t gonna be stalled for too long. This windfall will be the last for those three stooges of memory.
Except, uhm, for ..you know, that one company that hit a trillion cap
But you're right: Just like how million dollar computers with 1 bit of RAM performing 1 operation a second and taking up a colossal cave were replaced by $1 laptops with a zillion zekabytes running at a trillion hertz (exact values may vary),
the sprawling data centers of today with a quadrillion GPUs powered by black holes will get replaced by breakthroughs in hardware and most importantly, algorithms:
The human brain is proof right here that intelligence doesn't require dinosaur-sized hardware or eat half the sun every second.
I actually wonder if we're seeing the limits of discrete binary logic: Maybe it's high time to give analog ternary and all that funky jazz an honest try :)
Mac has won and it is not free.
Iphone has won and it is not low end.