upvote
I don't really understand the argument you're making, but just to add a data point:

DeepSeek V4 Flash 0731 is 167 gigabytes from the developer and as a GGUF with no additional quantization. It limps along on my 192GB M2 Mac from several years ago [0]. This model tests better[1] than Claude Opus 4.6 released in February. That's six months ago - what will be available 6 months from now?

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tr...

https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

So yeah, enthusiasts aren't going to run frontier models on their gaming machines, but a small office could easily justify the $30k - $100k cost to run something like this at high speed. The small company I worked for routinely spent that kind of money on Dec Alphas twenty five years ago, and that's not accounting for inflation adjustment.

And this is completely discounting the advances smaller models are making. You're right that Qwen 3.8 comes in different sizes. However, Qwen 3.8 27B and Qwen 3.6 27B do run on gaming cards, and they're better than the frontier models from twelve months ago.

I have no idea what will happen in the future, but I wouldn't base my guesses solely on the largest open weight models.

[0] Yes, it's unpleasantly slow (5-8 tok/sec)

[1] Yes, benchmarks should be taken with a lot of salt.

reply
> The distillations of Qwen 3.8 down to a 27B model are good, but they're not on a par with frontier models.

Does it have to be? There are plenty of coding tasks, where it's good enough.

reply
Practically, no, the distill is great. It's fine to use it.

However, if you're having a discussion about access to frontier models, and using Qwen 3.8 as an example of how open weights is a solution, then you should be honest and accurate about what you're talking about. Making an argument like "People can run Qwen 3.8 at home. That shows open weights are great." is a bit disingenuous if you're not also making it clear that you're not talking about Qwen 3.8 Max (or that you have a beast of a PC at home :D ).

reply
I agree 100%. IMO, it all started with ollama misrepresenting the Deepseek R1 distills as Deepseek R1, all for hype and marketing. I've had so many ostensibly technical people telling me: "I tried DeepSeek R1 and it was terrible", and every time when I probe further they'd tried the tiny 1.5B Qwen2.5 distill model that was further brain damaged by ollama's naive RTN quantization[1]. DeepSeek themselves were very forthright about it by naming it DeepSeek-R1-Distill-Qwen-1.5B[2].

I suppose the road to technical hell is paved with marketers and grifters. :)

[1]: https://ollama.com/library/deepseek-r1:1.5b [2]: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-...

reply
> coding tasks

Exactly. There are common coding tasks that these models can adequately do. They are absolute trash for anything that isn't coding. And even with coding, they are so, so far behind frontier models.

reply
No one is running that locally because of the AI bubble consuming all the hardware in the industry. That won’t be the case long term though
reply