Btw: with Qwen4 I mean the next large Qwen model that is based on the Qwen4 architecture (Qwen 3.8 flash next was "almost" based on the new arch but obviously was a small model)
What matters more is if firm’s start using a bundle of American and Chinese models and when they find their feet - how large is the market for frontier?
Frontier has to displace labour one for one at some point or it’s over.
I wonder if the next DS models will also graduate to 5.x, I think I saw they are training up a 10T model, and just raised $12B too
The fact that it basically broke open the local model supremacy was just a nice side effect.
I'm running: https://github.com/peonist-ai/halogen-server with a quant4, PLE offloaded, and it's resident VRAM is 36GB at 265k context.
Shave 10 more GB off and the TAM openai and anthropic are targeting is a lost cause. Local models are what 90% of people will need.
If the world governments can get a handle on the memory cartel, then there's no more moat for most normal humans.
In practice, you will be able to run models a bit bigger than 35B.
38GB of vram resident. more tk/s, more prefill.
This will cause the picked winner to get massively ahead with sheer compute alone used both for training and inference dedicated to recursive self improvement.
There's a law of diminishing returns at play here, and doubling the energy cost of training to wring 2% more performance out of the technology isn't going to be very useful, because most of the problems it is capable of solving will be solvable with the previous-gen 98%-as-good model.
("there's a law of diminishing returns at play here" is an article of faith. But then, so is the belief that these models will keep getting better).
lol You can always tell who has never ran a business before with comments like this
(I don't think it will work).
Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics?
There are going to still be worthwhile improvements but they are going to be more like not how to make transformers 10x cheaper but how to make next training run cost 9 trillions instead of 10 with a very particular optimization designed at the cost of hundreds of millions for this one specific run.
I could imagine a belt of data centres around the equator, that hand off their computational loads as the sun sets. Good scifi-esque premise.