upvote
It doesn’t matter.

What matters more is if firm’s start using a bundle of American and Chinese models and when they find their feet - how large is the market for frontier?

Frontier has to displace labour one for one at some point or it’s over.

reply
I wondered if there would be a Qwen 4 or we would go straight to 5, re: tetraphobia, but perhaps it's more like an uno reverse card in this case

https://en.wikipedia.org/wiki/Tetraphobia

reply
DeepSeek being Chinese also has 4 so I don't think it's a big deal for model makers.
reply
GLM moved through their 4-series without consequence

I wonder if the next DS models will also graduate to 5.x, I think I saw they are training up a 10T model, and just raised $12B too

reply
I'm pretty sure the point of Qwen3.8-Flash-Next was to get the open source engines to integrate the qwen4 architecture.

The fact that it basically broke open the local model supremacy was just a nice side effect.

I'm running: https://github.com/peonist-ai/halogen-server with a quant4, PLE offloaded, and it's resident VRAM is 36GB at 265k context.

Shave 10 more GB off and the TAM openai and anthropic are targeting is a lost cause. Local models are what 90% of people will need.

If the world governments can get a handle on the memory cartel, then there's no more moat for most normal humans.

reply
Awesome 3.8 next runs great on my Framework Desktop so I'm loving more local models.
reply
I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time.
reply
The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd.

In practice, you will be able to run models a bit bigger than 35B.

https://unsloth.ai/docs/models/qwen3.8-next

reply
https://github.com/peonist-ai/halogen-server is beating the pants off of anything I've tried.

38GB of vram resident. more tk/s, more prefill.

reply