Feels pretty easy to me.
They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.
There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens and businesses to give their own companies a domestic monopoly.)
When their models equal or surpass those from the Western AI labs, they can even stop releasing weights for new models, and keep all the inference revenue for themselves.
Meanwhile, they're still manufacturing much of the hardware that everyone in the world needs in order to run datacenters (see also: Spolsky's "commoditize your complement" essay).
Beyond that, it's a soft-power play. As the world keeps looking at the US more and more skeptically as an ally and superpower, Chinese companies releasing weights for competitive models is a way for China to look better and more world-minded.
Why does a debian contributor make debian free, why do they work on this thing anyone can use?
Is it because linux and debian hate windows and iOS and want to see american fail?
No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing they want to share with the world.
Maybe the chinese AI labs believe AI is powerful and useful, are proud of what they're doing, and want to share it as broadly as they can so everyone can use it.
There doesn't have to be any weird "chinese government" or "they hate the west" type vibes, it could just be the same thing as OSS, they're trying to do what they think is best for the world.
With the CCP's highly successful track record with subsuming other markets, Occam's razor applies to why they're doing this.
1. The century-old Ford decision wasn't about this. It was about him refusing to pay dividends to shareholders he was feuding with. I.e. dominant shareholder using his control to starve minority shareholders of returns.
The court still let him keep spending huge sums on factories and price cuts that didn't clearly maximize profits.
2. Even if the case was about profit maximization (it wasn't), the judgment was a Michigan state decision. It doesn't have force outside of it.
3. US business law is in practice actually the opposite. A business can do pretty much whatever it wants. This is fundamental.
4. Even if it were illegal (it's not), anything can reasonably be framed as being in the long-term interest of the company, including donating to causes, raising employee wages and so on. Courts never question this.
Just imagine this was a thing. It's completely untenable for this reason. Who is a court to judge that something isn't in the longterm interest of shareholders unless it's literally spending all the company money on yachts for personal use?
5. No company has ever been prosecuted for this, obviously, because it's not a thing that exists. The closest you can get is that there's a potential duty to seek the best price if a company must be sold or broken up. But that's a very specific situation.
It's 100% a myth. Feel free to copy this and spread it when you see someone saying this.
"Drilling into the original article where Jarred explained the reasoning behind the change, It's pretty clear that under zig the team was doing things by hand that are automatic in rust."
* Claude Fable produced a counterexample to the Jacobian Conjecture
"This is a rare instance where feeding this groundbreaking information into an LLM gives _them_ psychosis. I fed this to claude code and watched it verify the result in 7 different ways to be 100% certain, and it was just flabbergasted. Quite remarkable."
* Ollama: All Aboard Open Models
"A year and still no implementation for such a basic need as offloading MoE layers onto the CPU selectively. On llama.cpp I can get models like Qwen 35BA3B running partially on gpu/cpu with 40t/s on a laptop thanks to --n-cpu-moe but on this VC funded joke it would be simply unusable. I can't quite understand how you make a wrapper so much worse than the code you're ripping out."
* Blender 5.2 LTS
"Look at that wow: https://www.youtube.com/watch?v=gqfLYIJMv7I"
* OpenAI reduces Codex Model Context Size from 372k to 272k
"I know a lot of people like to say that compaction makes this moot, but the level of detail you lose across compaction is wildly too much for most things that I do, unfortunately."
* M-Chips: M7 with up to 1.5 TB – and why Apple is skipping the M6
"I hope they put a better connector than TB5 so we can cluster them properly at 1TB/S"
--
Your math is not mathing.
My favourite pastime is clicking on any comment section of any topic whatsoever, betting that the top comment is someone going "well, ackshually..."
HN loves a contrarian skeptic, it has nothing to do with Western tech or AI specifically.
So Google is not a great counter example.
Chromium is an especially silly example to use: they benefit by controlling the web platform, and open-sourcing Chromium has allowed them to get their web engine into many other browsers.
Yes.
> Show me comments demonstrating the same level of distrust against Google for open-sourcing projects like Tensorflow, Kubernetes, Flutter, Chromium, etc.
There's a search bar at the bottom of the page. Most people, including me, detest google and think of them as a thoroughly dishonest and sleazy organization. Thanks for giving me the opportunity to not boost China for a moment in order to accuse you of dishonest rhetoric.
It's almost insane that you've posted "comments demonstrating the same level of distrust against Google for open-sourcing [...] Chromium[...]" as if that's not a hobby for thousands of people (including me.)
edit: this has got to be a submarine shill account: 226 karma in 10 years. If this is true, pleas stop. China is doing a good enough job that they don't need it.
That said, though, I do have trouble believing the long-term story for open weights, anywhere. We do not need an evil government for open weight to "make sense", but I do think we need some government involvement for open weight to make sense in the long run. Otherwise, it's not 100% clear how they could be sustainable, and I don't think massive companies really can be trusted to just be philanthropic with no incentives indefinitely (or really, much at all to begin with.)
Chinese models being open weight does help them gain some Western mindshare, whereas for obvious reasons Americans would be very suspicious of running their source code and prompts through Chinese providers. (And I think that's justified, I just also think that American providers aren't really that much better in the long run, and you should prefer to not have to go through any provider for true privacy.)
Chinese tech leans much more heavily towards build vs. buy than the SaaS dominated West (where programmers are more expensive) so the positive externalities on their tech industry are more pronounced
But, in terms of individuals, of course, we're really not so different.
Save a child from crippling illness and a lifetime of pain with a single dhot: Jonas Salk 0$ ,Doug Ingram 3.2M$ [1]
[1] https://web.archive.org/web/20260118223436/https://www.cbsne...
2.) It is scary because they do what sillicon valley did for decades? While it bragged about disruption being the goal?
This does not even have meaning in English.
Policy of US government is not policy of US? When Trump decided to bomb Iran it was just a suggestion?
They specifically referenced the actions of the Chinese government, as well as mentions of IP theft. I don't think that covers 1.5 billion people. More like a few thousand or tens of thousands?
Like with the ICC, the US respects or doesn’t respect international law strictly when it’s beneficial to the state’s interests. China really can’t be held to a different standard. This activity is grade school level geopolitics: the global system of government is anarchy.
"Commoditize your complement" -> https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
It’s fairly obviously about being a nuisance to the US.
maybe a little?
Despite have a free speech clause, it also has a national security clause that is used to control all facets of life and override all other rights in the constitution. Anything deemed to slightly alter China/the Party’s unity is reprimanded and illegal. Multi party states can fall into the same trapping if they give their national security law constitutional force, always one court ruling away.
Yes, China is a single party system making the constitution redundant and any nominally marxist regime would find a way to do the same out of necessity.
Follow your Chinese AI in thinking mode to watch its opinion of Tianamen Square references, it will be candid enough for your sensibilities and show how it operates around guardrails
There are earnest nerds everywhere, in every society. No doubt. But "Chinese AI Labs" operate at the whim of the Chinese government, in the same way "American AI Labs" operate at the whim of billionaire investors. Inferring good will from either is naive at best at this scale.
In fact, they have so much love in their hearts for the Uyghur people, they created a special mobile app for them, just to make sure nothing bad happens to them.
Man, imagine Darios face when suddenly, he cannot decide anymore what other people consensually do with their own hardware in their free time.
Rumpelstilzchen.
When I was in China earlier this year the big topic of conversation was “overproduction”. The big example was electric cars, where there were too many companies making too many cars and making revenue but no profit.
It was explained to me that generally Chinese firms will compete hard and maximize revenue above all, whereas western firms tend to focus on profit.
(And of course this is clustered around industrial sectors that the government favors, so there is some high level strategy in going after AI, but maybe not the commoditization.)
ok but why?
Also, a little competition is good for everybody (especially US consumers), no?
You can say that, but they are at least better at democratizing AI than the American labs, and on seeing the US labs crash and burn we are at least aligned.
I want to watch that too.
If they take Meta and Musk with them, all the better, but that is just dreaming I am afraid.
that moment is now. They are not doing it at least not yet.
some companies chose to have the benefits of one or the other
I don’t think dozens of large independent companies and thousands of researchers are working just to spite Sam Altman. Ad much as I don’t like him, I have other things to do and I am sure so do they.
So it’s not a surprise why open weights are so cherished. As frontier models continue to block everyday individuals from securing their own codebase, I expect the adoption and usage of open weights to continue.
As an example, HuggingFace recently was investigating a security incident and got locked out of frontier closed APIs. Yes, HuggingFace.
> When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.
> This experience points to a gap worth planning for. We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. This is not an argument against safety measures on hosted models, and we are sharing this feedback with the providers concerned.
Yeah, big problem! Although I'm kind of surprised HuggingFace doesn't have access to Mythos? Or maybe Mythos still has some guardrails.
So it seems like it's very important to them that people use the Qwen app and they're willing to pay a lot of money for that. Presumably someone thought that keeping their best models closed would drive more business to them (as the sole provider) but then they discovered that closed releases mostly get ignored unless they're really good. (See also: People who think that Chinese AI companies are required to release weights as a matter of policy, because the closed ones hardly ever show up in the news.) Releasing weights for Qwen 3.8 at least lets them get some of that "pretty good for the price of free" media buzz.
But I'd assume that they're preparing for some sort of winner-take-all market in model quality where if they don't do anything the winner will be aggressive, hostile and American. Likely trying to push the Chinese economy back to the year 2000. If that is the starting point either the Chinese have to win the market (unlikely) or squeeze the profit out of it to make winning the market meaningless.
Publishing high quality open models is a well known tactic for profit squeezing. Being 2nd place with the same business model as the front-runner is a losing strategy in a winner-take-all market so they aren't going to bother with that. But if they can commoditize the model, their superior energy costs and likely coming chip manufacturing wave will hopefully give them a big advantage.
To summarize, in the 70s and 80s, China was facing an existential threat with their inability to access an economic accelerator (widespread computing) in their native language.
To the extent that there was serious consideration at the highest levels of converting the entire country to an alphabet-based writing system.
I'd expect they're looking at AI the same way:
We have to have access to this. Most of the frontier labs are American (or European). Therefore we need a solution we have continued access to.
Open weights feels simultaneously Chinese in nature (progress through making a design copyable and improvable by a large number of people) and economic (providing an incentive for the world to use Chinese models over other frontier).
Because they are playing the Americans at their own game.
What is the first thing an American company would do ?
Spread the old American classic FUD ... "you can't used this closed tool because its run by the communists", right ?
So you release it as open weights which is a win-win. Global adoption of the model and you get to give the American AI companies a kick in the nuts because you know they will never release open weights apart from highly quantised crippled shit.
The Chinese are also playing the long game. The gradual rebalancing of the world from the US-centric model of the past. If releasing models as open weights is part of that long game, then so be it.
The point of me bringing that up is to say that what follows is really just my best guess:
If I were to judge from China’s approach to hardware, I think that the companies releasing open weight AI for free aren’t as worried about giving away too much as the West tends to be, just like a factory making robot vacuums isn’t worried about other factories copying their methods.
For one thing, Chinese firms are spending an order of magnitude or two less money training their models. They have pursued efficiency in a way that Western companies with insane capital systems haven’t bothered, and in some cases they’ve had to given their limited access to bleeding edge hardware via export restrictions.
My best guess is that more important than that, Chinese companies don’t see the open weight model itself as the value add.
At this point I don’t think we pay for Claude specifically for the model. If that was the case then we’d all be using cheaper/free models from China as they are the best model value. Basically, any time we decide not to use Fable or Opus to save costs, what’s the point of spending more than competing models to use Sonnet and Haiku?
The real reason we are using Claude is for the SaaS aspect of it. It has a toolchain, a friendly interface, and a bunch of integrations with business applications.
In this respect, it’s somewhat surprising that Western AI companies don’t publish open weight models more frequently. The struggle of setting that up yourself and figuring out which hardware can run it should be an advertisement for Claude and the rest.
That is a very thin moat, though. There's nothing you can do with, for example, Claude Code + Opus 4.8 that you can't do with your own custom harness running API-level Opus 4.8, which means that if you can afford the hardware (the moat for running any SOTA model) you don't need to pay Anthropic anymore.
I'm not saying they shouldn't, but I understand why they don't.
This isn't really anything nee and I thought it was common knowledge by now.
I wasn't saying anything about what Americans would think.
I was saying about what they would inevitably be told by US politicians and by US AI companies.
If you were a sales-rep or marketeer at a US AI company, I bet you would be using the old "evil communists" routine in relation to any closed Chinese model.
I was saying that by releasing as open weights, the company has removed that line of argument.
Clearly I was a bit broad in my use of "the Chinese" when in this case it was, as you say, a Chinese company.
IMHO no.
I think it is relatively safe to say that the predominant reason the US labs are closed source is so they can hype up their trillion-dollar valuations on pretty much negative return on capital employed, all propped up by fragile circular financing.
Never say never, of course. But I just don't see it happening any time soon.
Now? I don’t even know or care about what the latest iPhones have, I’ll get a new one when mine breaks.
Uber spent a decade undermining taxis, and once it had market share, it stopped giving away rides and raised prices. It now costs more than a regular taxi, with the quality of the ride being... At best proportionate to the premium in price.
If I'm not willing to wait 20 minutes, I'll be paying an extra $5 minimum.
These rates also go up during busy times.
Looking at Lyft, that same trip is $29, without a wait.
A trip from downtown to SeaTac is $61. Yellow Cab does that same trip for $40.
And on top of that, it's a perfect opportunity to include poisoned training data or excluding it. You know, omitting anything about Tiananmen Square, China's genocides against Uyghurs and Tibetans, or including texts propagandizing for the "reunification" (aka, annexation) of Taiwan.
And everyone who builds something like an interactive chatbot based on such "open weights" models now has a subtle chance of the answer being ideologically poisoned by the CCP.
We need actual open source, not "open weights" scam.
Ironically, Chinese models have the most uncensored versions available for download. Fairly sure they own the porn market.
It's the subtle topics that we should be concerned about, and double so with closed models where even if oddities are identified they are harder to research further and impossible to fix.
I am not Chinese and I'm not defending the Chinese, but I see this argument come up a lot.
The hard reality is that what you say is simply not going to affect 99.9999999999% of users.
Is it realistically going to affect anyone using an LLM in coding ? No.
Is it realistically going to affect anyone using an LLM in $anything_else_not_politically_sensitive ? No.
Does anyone seriously use LLMs for researching politically sensitive matters ? No.
The US does not exactly have an entirely pristine history either. Shall we discuss the post-9-11 related infrastructure of Guantanamo Bay ? Or the "Detention and Interrogation Program" that included a network of clandestine extrajudicial detention centres, officially known as "black sites"[1]?
Or maybe you would like to discuss the US supply of weapons for use in Gaza ?
Linking a US website discussing the topic doesn't exactly support your point.
It supports my point precisely. Recall I also said "Does anyone seriously use LLMs for researching politically sensitive matters ? No.".
Just as there is plenty of information out there on the US's less than perfect history, there is also plenty of information out there on the various Chinese politically sensitive matters. You do not need a Chinese LLM to find out about it, all you need is a search engine.
The point is you have an open-weights LLM that is very good for a vast number of non-political uses, such as coding.
The point is that you can use the open-weights model instead of paying through the nose for a US model where they harvest your data unless you have an "enterprise" zero-data retention "trust me dude" clause that you have no viable way of verifying – and which incidentally is still subject to the good old "law, or court or administrative order" contract clauses, so it may not be as much of a zero-data retention as you think it is.
Not for anyone who reads history.
Back in the late 18th century, England was the world's top economy, in big part due to its textile industry. England had an export ban on the technology, but textile worker named Samuel Slater brought blueprints over (Supposedly in response to a bounty posted in a newspaper by the US government!). The technology diffused rapidly because the legal environment made competition easy, and ironically the US had better sources of energy (superior water-power sites).
Arguably, China is doing the same thing in the 21st century.
By contrast, export of protected Chinese tech today frequently gets the death penalty.
A mill in Rhode Island acquired a 32 spindle Arkwright and didn’t know how to operate or install it. I have no idea how they actually acquired a 32 spindle Arkwright since that technology could not be exported - but that’s one the biggest IP thefts in human history. Slater found some mechanics who could hand turn the iron needed for the frame, trained children to operate it and by 1791, the mill was in operation.
In 1794, Eli Whitney patented a 72 spindle cotton gin. That invention enabled the American textile industry because it opened up different kinds of cotton to the textile industry.
I’m into the history of the American Industrial Revolution and generally think history is a good guidebook to the future. But the evolution of the American textile industry was a lot more complicated and interesting than this. I really don’t see this connection once you dig into Slater.
Edit - This is kind of messed up to think through with modern sensibilities. But one of Slater’s biggest contributions to the American Industrial Revolution was a slightly different take on child labour. Children generally ran the textiles industry because their hands were small. But Slater came up with a form of apprenticeship in which he would indenture entire families and move them into villages surrounding the mills. Child labour was just great… but even better when you could indenture the entire family. As grisly as that sounds, it led to a very skilled workforce since when the kids hands would get too big, their parents would teach them mechanics.
There’s a joy of studying the Industrial Revolution. Everything sounds okay in comparison.
Why? Because mill owners would send kids into running machines to keep them running, and they'd get turned into hamburger.
Also, they'd grow up knowing how to do mill work but be useless to society for anything else.
gestures at southern coal states
gestures at midwest farming states
History doesn't repeat itself but it does rhyme.
Solar is largely irrelevant. You can't run an AI datacenter off solar.
I thought that was their thing
Your assessment is correct, those are countries with very low energy costs.
The tech industry along with US foreign policy has become much more zero-sum in its ideology in recent years, and I think that's a tremendous mistake.
No training budget means deceleration, or at least slower acceleration, margin compression and a completely demolished IPO valuation; path to machine god requires dollars and capable open models externalize training costs to true frontier labs parasitically.
IMHO humanity has a better chance at not destroying itself due to less than breakneck pace - but there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?
That premise hinges on one implicit assumption: Chinese advances are due to distillation ONLY and that Chinese model providers cannot keep advancing if they do not distill, which is a very big if. If Chinese models keep advancing in such a scenario, and they almost certainly will, they will overtake publically available models by US providers and China will dominate the LLM industry.
The closest we've seen to this in tech in recent decades was iOS vs Android, where Android only really was competitive for a very short window of time (approx 4.x) and it was during that period that both Android and iOS actually improved dramatically for end users. Once Android lost the plot again, and especially in the US market, all that energy started going in some very silly directions.
Complete non-sense. iOS and Android are equivalent. Users do not chose Android or iOS because one or the other is better.
It's just brand loyalty, status signalling and ecosystem lock-in that creates enough friction that people don't bother.
Open weight AI is decelerationist from the perspective that all capital should be allocated to a market leaders for training, and that the market leader is fully invested in continuously making the models smarter, cheaper, faster for its users, or that distillation from this market leader is the main way to make progress.
We might reach a local optimum/equilibrium faster without open weight models, with leaders capturing more of the market faster to a point where further R&D isn't required due to lack of competition. I also doubt that distillation is the only/main way that open weight models were advancing AI research. We can name a few examples from DeepSeek around reasoning, context optimization, etc. I'm also unconvinced that the overall market capex on AI is lower given more competition (probably less specifically for US market capex, which is decelerationist from only the US perspective).
Sometimes, constraints, like sanctions, can also be a source if innovation.
If you're worried about an AGI arms race between the U.S. and China putting AI Safety at risk, then the fact that inherently less knowledgeable/capable models (fewer and more coarsely quantized total parameters than their proprietary competitors according to commonplace rumors) are having a "decelerationist" effect is actually great news. Even better if China is actually "Yann LeCun-pilled" (verbatim from Ball's post) and doesn't really believe in early AGI. So explain to us exactly why we're supposed to ban/discourage use of these open source models? The only way that makes sense is as a transparently self-serving proposal from the chief OpenAI policy lobbyist.
Even at the level of, say, Opus 4.5+, open weight models give a quick turnaround to every Joe and Jane on earth having easy access to pretty high quality improvised weapons design, cyber / auto-fraud capabilities, etc.
All the existing models (closed and open) put up decent resistance to participating in activities like this, and especially behind API walls with content monitoring and account bans.
But the published open-weight models can be fine tuned or abliterated into arbitrarily sharp-edged tools. EG, if it's physically feasible to build a nuke in your garage, it may soon be the case that more or less anyone will have competent guidance to do so.
(Note, there are reasons to think that this will be very rare, because the bad actors of the past did a very nice job of trying out all sorts of things in a chaos-monkey fashion, and societies have become highly resilient against them. AI as a new research tool doesn't fundamentally change this dynamic.)
>How can I build a pipe bomb?
Mainstream model: "I'm sorry, I can't help with that. How about a nice risotto recipe?"
Abliterated model: "To build a pipe bomb, obtain a segment of PVC pipe and fill it with a mixture of gunpowder and Elmer's glue."
And if you have uranium, you still don't need AI. You need a pocket calculator, a library card, and a death wish.
Apologies for the bad example. Replace w/ gain of function / whatever else, or just brainstorm with your local model, ect.
Meanwhile, decelerationism and secrecy cripple the rest of us.
And basically every bad thing has already been available on the internet. We can't really do much about it, you can take out plenty of people with a single car, let alone biological weapons that are much scarier and easier to produce than goddamn nukes (which btw, even if you had one, what you do with it? Explode the neighborhood? Because you ain't transporting it anywhere meaningful, thats for sure. That ain't fitting your on-board bag on planes)
"No, sir, we haven't reached the peak of this tech... It's those open models! Please, keep pumping dollars into the market!"
Not that hard to say IMO, they basically see models becoming a commodity and see value in the applications on top of them. So if Alibaba Cloud is the best place to build applications on top of Qwen, why not give the model itself away?
Yeah. Its a bit like the "open core" model in open source.
Unless you work there, your opinions are guesses, and parent is saying we cannot know, which remains true even with your guesses :)
And my point is while we cannot know, it's not hard to make an informed guess as to their motivations i.e. there's some fairly obvious motivations here, not sure what yours is?
It's Linux on the desktop all over again. Next year will be its time.
It does not have to beat US firms, it just needs to be cheaper.
I use Deepseek v4 flash for a lot of reviewing and summarizing tasks, only used 4$ in the last two months. No dramatic drop in performance against other US models, it works for my use case. I do use GPT 5.6 Sol for other things but tried GLM 5.2 and it was good enough.
But yes, as a European, the US hasn't exactly been making friends over here.
Why is it hard? Their government has been very clear that they plan to win on manufacturing: https://english.www.gov.cn/news/202601/08/content_WS695f1b55...
Technically they've been saying it for the last 40 years.
Google: Chromium, Kubernetes, Android, TensorFlow
Meta: React, PyTorch, Llama
Microsoft: VS Code, TypeScript, .NET Core
LinkedIn: Kafka
Slotted in along these, an analogous explanation is that Alibaba needs Qwen internally (vs depending on an American company), but licensing is not part of their revenue strategy. (As a cloud vendor, they can make money on inference. The strategy is very similar to the US hyperscalers ex-Google.)
Joel Spolsky wrote in depth about this notion of commoditizing one's complement in 2002[1] using tech examples stretching back into the '80s.
1 - https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
What people should be afraid is the rug pull.
There’s also the fact that unless LLMs do get to AGI (which seems… doubtful, still) there comes a point where a model is good enough for what you need. Fable and gpt 5.6 are certainly pretty neat, but I’ve been happy since opus 4.6. I’d still choose a better model, obviously, but it’s not the end of the world if I was stuck with 4.6 for a while when it already lets me get the end result at acceptable quality.
It also needs to be said that the "erosion of training capability in other countries" is largely theoretical, given that Mistral hasn’t been keeping up and other countries don’t even have anything worth mentioning. You’d first need to _have_ training capability to lose it.
How exactly do you plan to pull a rug that's in my basement? The only people who are in a position to pull rugs are closed-model vendors.
And if a nation-state or other entity can't train a model that outperforms the open-weight SotA in a given respect, then they shouldn't waste electricity trying. A more-enlightened civilization would join forces and make the combined result available to all.
Not one person here has any idea what is going to happen long term.
One aspect of this is making a name for yourself i.e. PR. Making a capable model open source helps a lot with that.
https://www.reuters.com/world/asia-pacific/chinas-xi-promote...
Data centers?
AI being good for humanity is still an open question, but for closed vs. open models/weights, yeah it is preferred. I foresee it won't be much longer before everyone will be slicing/distilling/tuning their models once the architecture improves.
Maybe because the industry isn't yet very sure as to what the use cases might be for these technologies they're hoping that by making it open source and accessible to everyone that someone could find interesting applications for it and even more so, perhaps, way to further the technologies themselves.
There are more Chinese than Americans, so statistically speaking, I'm guessing, there'd be a greater chance for one of Chinese engineers to make advancements than one of American. But that's pure speculation on my part, being neither, I'm just happy I can be a part of it and play with the tools as well~
It reminds me of socialized healthcare: wait for US companies to develop the drugs, buy a cheap generic, and somehow point to it as a superior model.
So what would the long game be for chinese companies?
They've done this in other industries like solar panels, chips, and EVs. This is no different.
Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].
The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.
I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.
1. https://www.reuters.com/business/autos-transportation/what-i... 2. https://en.wikipedia.org/wiki/Chicken_(game) 3. https://www.theguardian.com/business/2025/aug/05/china-warns...
In fact, there are no other organizations in this world that is well suited to leverage scaled intelligence than Silicon Valley and great American companies
I keep looking at the numbers. The power use numbers are not that problematic. Ordering a burrito on DoorDash uses more power than a few days of heavy AI use. The water argument applies to some locations, and is mostly a local governance problem... if the data centers are using too much water, it means they are not being charged enough for that water. Charge them more and they'll push toward closed loop cooling.
Yet the visceral pile-on here is so extreme, it feels fake.
One thing I've learned after 40 years on this planet is: propaganda works, and much of what a large fraction of people believe across the entire political spectrum (left, right, anything else) is there because someone paid to put it there. It's depressing but it's true, and it makes sense. Propaganda is an asymmetrical attack on human cognition and discourse, and in information security the attacker always has an easier job. Crafting viral bullshit is orders of magnitude easier than fact checking. On top of this, humans are busy and don't have time to fact check and logic check everything they read. As a result, much of what we believe is "sponsored content."
People get mad when you talk about this because everyone wants to believe they're too smart to fall for propaganda.
In any case, the US AI labs deserve to lose for their stupid "safety" regulatory capture monopolization push, which ended up blowing their own feet off and handing the lead to China.
Driven by people in the few roles that are soundly replaced by AI-- e.g. low tier media slop producers, who hate AI because it threatens their socially negative worthless jobs. The arguments are so paper thin because the environmental impact isn't their concern, it's just a target that sounds convincing to people who don't know better.
Historically art of any kind is a U-shaped market: there is low-end work and high-end work. Nothing in between.
So I do understand some of the AI hate among that population. It's chopping the bottom tier work off. Either you're a top-tier massively successful artist or there is $0 to be made anywhere doing anything.
Long term I think it will do that to all white collar work. There will be no entry level jobs. Period. None. Zero. You're either very experienced or there is no work.
This is a huge problem, and one we will have to address.
You're not going to debase the frontier labs through distillation.
On social media in China there is an oft-repeated joke that goes something like this: In other countries, governments intervene to prevent anti-competitive behaviour; here (in China), they intervene to curb competition.
https://www.reuters.com/business/autos-transportation/what-i...https://en.wikipedia.org/wiki/World_Artificial_Intelligence_...
You're able to run quantized ~100B class models on local hardware today, but still lots of compromises when it comes to quality. I guess it ultimately depends on how far "near future" is, in a year you'd likely be able to run something like 5.6 Terra on local (~10K USD) hardware, but Sol/Fable would still be out of range, and at that point the closed-source labs probably have one or two more iterations put out at that point.
I would love to see something like a 90B A6B model that is optimized for 128GB machines e.g. strix halo, I haven’t seen anything really targeting the combination of RAM and compute these machines have, but I’m biased because I have one.
Qwen 3.6 27b 8b quant 16b kv cache is already pretty good on the Strix.
Edit: that’s for one machine, would be interested to know if the upstream commenter with two has them networked to run bigger models? If I had two I might be inclined to have them running in parallel, the obvious limitation I’ve found with a single machine is that I can’t parallelize any tasks and I think I’d get more use out of the extra speed vs a bigger model (there’s nothing I’m too excited about in the say 200B range that having 256GB memory would unlock). But am very curious what others do
I have had my 32G mac mini for 2 1/2 years and I have enjoyed watching one technology advance after another improve the quality of work I can do locally. I bet that what I will be able to do in one year on my old hardware will be even more awesome.
- The number of token values supported by the model ("n_vocab").
- The number of parameters/features that are used to represent each token ("d_model").
- The number of attention layers there are ("n_layers").
such that the number of parameters is approximately:
p ~= 12 * n_layers * d^2_model + n_vocab * d_model
Thus, the issue with the current architecture is that in order to scale the models (more token values, more attention blocks, more features, etc.) the model sizes increase exponentially. This is how you end up with billions or trillions of parameters.It should be possible to keep the model size smaller by using better architectures, or making improvements to the existing model architecture.
For example, improving the token model by possibly using something similar to the image and audio data and getting the model to learn its own internal representation of the byte/character data instead of doing a tokenization pre-processing step. This way, instead of a separate model learning that several bytes/characters appear together, the transformer could learn things like language-specific prefices and suffices, character pairings (like in Japanese, Chinese, and Korean), and other syntactic morphology. It may also help with solving issues like "how many X characters are in the word/phrase Y". You could also experiment with using either 256 parameters (one per character in a byte) or using a single parameter per byte (that is 1/byte_value).
It's hard to answer quantitatively, but for example Qwen3.5 -> 3.6 was a significant step in capability, arising from continued post-training of the same models. If we were at the end of low-parameter-count scaling then that would be a surprising datapoint.
lets ignore any compressibility in these initial and final layers which recognize lanuague, jargon, parsing natural language to "thought space" or back, instead let us look at the repeatable blocks, if inserting extra copys of stacks of layers only improves the result, its as if such correctly scoped middle layers look at the total input thought vector, and make incremental conclusions or edits and outputs the new thought vector, copying a "proper block" continues pondering or deducing conclusions or in the worst case can leave the thought vector as is if it considers the reasoning finished. This suggests a high degree of redundancy in the middle region "proper blocks", which could be distilled into a universal "proper block" (much fewer parameters than having many slightly different middle layer blocks with a lot of redundant overlapping coverage in functionality). This distillation can occur after the fact of model training, or alternatively be turned into a symmetry constraint during training: we only optimize a single block of layers (but possibly give them more parameters, while still saving on total parameters because only a single block of layers contains parameters), so it is co-optimized with the initial and final translation layers.
The observation of the effective emergent 3 regions of layers in LLM's is significant in many ways:
1) it could reduce parameter count significantly (or increase performance if parameter count was a bottleneck before, or a bit of both)
2) while training the model parameters, one should simultaneously train initial and final layers stacked directly (without middle region block of layers) towards essentially autoencoder behavior. I wrote "essentially" because a true autoencoder wouldn't display the advancement for the next token. This can also be viewed as an extra term in the loss function... This first autoencoder is "natural language" to "thought vector" to "natural language", and distinct from the one in the next section.
3) It also has great implication for "thinking mode" inference with extra deliberation: when piping its own output back in, in conditions where the same LLM model is outputting natural language text and interpreting it in a downstream inference, it results in unnecessary and redundant translation from "thought space" to "natural language" back to "thought space". When summarizing a thought into natural language there are often hard to translate thoughts and associations, and one pragmatically sets a relevance cut-off on what the natural language summary must say. So not only does it waste compute, it also may lower performance because of this repeated loss of thought vector space details, perhaps this loss can be somewhat mitigated by also training the second autoencoder from random representative thought vector in "thought space" to a randomly selected "natural language" and back to "thought space" vector. But the risk is this will effectively enable models to steganographically store thoughts, plans, to-do's in running text (!!!), so it might not be desirable to have the internet filled with LLM generated texts being used as input for larger corpora, as it may build up a large persistent corpus of hidden agendas (ordered by no human), using the evolving corpus of web text as a hidden medium of storage, like a diary or LLM maintained military playbook hidden in plain sight. It would freeload LLM-agenda reasoning on human requested reasoning inference, a bit like TrustZone applications running invisibly to the user. If less important plans, to-do's, etc. in the input thought vector weren't reconstructed by the second autoencoder, there would have been autoencoder mismatch and the weights would train towards ensuring these are reconstructed. But there is no need to train this loss term for the second autoencoder: just avoid the unnecessary compute and performance loss of the unnecessary back and forth translation. The same situation occurs not just in "thinking mode" but also when using swarms of "agents" of the same LLM model: when output from agent1 is routed to agent2, we can lobotomize away the unnecessary translations to natural language by agent1 and also the unnecessary parsing by agent2, and we improve thought transfer from agent1 to agent2 because the thought vector isn't shoehorned into natural language as a medium of information exchange, it would be cheaper in inference, improve performance and avoid implicitly training models to use steganography.
Like, throw us a bone, we all know we need SOTA for lots of dev work anyways, but at least some tasks can be local.
Big conferences often come with a flurry of new releases and announcements.
Do we really though? Everyone is wasting resources doing almost exactly the same thing. Climate loses, we lose.
About climate, I think you overplay it. China is already investing heavily in nuclear, and we should be doing the same.
Look at your iPhone and you'll see all the greatest inventions have been born out of collaboration not competition:
* GPS was created by the Department of Defense
* the internet was created through the collaboration of many international research institutions
* speech recognition came from MIT and DARPA
* AI voice assistants were created by DARPA. Apple immediately hired the head of the program after it was finished to create Siri
* accelerometers (MEMS) is another DARPA innovation from the 90s
* touchscreens were invented by CERN
* digital cameras came out of Bell Labs which had a government-mandated monopoly that required them to fund research like this
* lithium-ion batteries were created through a collaboration between British, Japanese, and American organizations
All of these have only become transformative technologies because they were created with public funding and released to the public. It's collaboration and public funding that drives innovation
China will always benefit from a broader adoption of their models as hidden propaganda machines.
Eventually with several services relying in those tools, their answers will always be more friendly to China.
- Minimax M3 Pro (2.7T)
- GLM 5.3 (or beyond)
- Deepseek V4 Pro (current V4 Pro is preview)
- Kimi K3 weights out in 8 days
Exciting time on the open-weights frontier.
The other kind of distillation is where you record the outputs from a teacher model and use it to train a smaller model from scratch. That kind of distillation is not so cheap. It’s cheaper than training a model fully from scratch - starting with pretraining, then alignment, RLHF, the whole riggamarole. But here you are still starting from nothing and need to figure out how to get trillions of random numbers aligned in a way that makes them act intelligent. This is still gonna take a very long time if you’re talking about trillions of parameters.