upvote
> People are realizing that what they need isn't more general intelligence, it's more specialization. A small but well tuned coding model...

It’s not quite as simple as that. Several studies have shown the opposite: models trained on more diverse knowledge tend to cross-pollinate across domains. So a more generalized model can actually perform better than a specialized one.

That’s why you’re not seeing tons of tiny models (one for Python, one for Pascal, one for Rust, etc).

reply
This is definitely the position of the big ai companies.

But it doesn't match my experience. Qwen3.8 27b is clearly smarter at coding than MANY bigger models. gpt-oss-120b for example, is almost 4x the size, and performs way worse at coding tasks.

It's clear to me that you can build small models that work well at specific tasks.

Python vs Rust is probably too fine grained a way to build a model. Coding in general seems like a better target.

There will always be a place for large generalist models, no doubt. But I think that place is much smaller than the big ai companies are counting on.

reply
Gpt-oss—120b is like 1000 years old in AI years, whereas Qwen 3.8 27b is pretty young. What you’re seeing is that parameters aren’t apples to apples, and at a given parameter level, the new models are much, much better than the ones from a year or two ago. Like, to a comical degree.
reply
Does that not prove my point? Bigger doesn’t automatically mean better. Quality of training data, and model structure, matters as much or more than size
reply
Ah sorry, I should've continued, the bigger recent models are commensurately smarter. If you really want to make the point, then you'd need to show 27b being smarter than similar vintage bigger models. And in that case, there's confounding issues like efficiency, speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.
reply
https://artificialanalysis.ai/?models=gpt-5-3-codex%2Cqwen3-...

Shows qwen3.8-27b along side seven larger models of ~similar vintage. Only one scores above 27b.

Many of those are closed models so idk their exact parameter count / active param count, but it hardly matters - i’m sure all of them are far above 100b params

My point is not that bigger is pointless. It’s just clearly not the only road to take to make a model better, which is obvious just from seeing how models of the same size have gotten better over the past few years

reply
Thanks! That's very helpful as a way to discuss.

First off, I'd include Qwen flash-next and GLM 5.3 to show some of the other strong open weight models, and they predictably dominate it, but they're much larger. But, it shows up right next to DSv4 Flash 0731 on the overall index, and that's much larger. It's a great model! But then scroll down and hit Time Per Task, and you'll see that DSv4 Flash takes 3.6 seconds per task to Qwen's 21.1. That's what I meant when I said this:

>speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.

It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens. And that's a great trade if you don't have the vram to run the bigger models, but speed is pretty important for getting things done... And that's why DSv4Flash is great, too, despite being much larger, and scoring similarly on the intelligence index.

reply
Wasnt this known by everyone who cared to pay attention?

It practically became a joke about how a huge amount of the training data for GPT-4 was bottom of the barrel reddit vomit and obvious bot spam. Leading to many bizarre edge cases.

reply
I think we’re in agreement.

I make heavy use of smaller local models on a daily basis (Qwen3-VL for auto-captioning images, Gemma3:27b for some translation work, etc.). Gemma3:27b is a good example of a very capable general purpose multimodal model and has handled almost everything I've thrown at it from sentiment analysis to documentation writing.

I suppose I was drawing a distinction between specialized and general intelligence versus small and large. I don’t think those are necessarily mutually exclusive.

reply
gpt-oss-120b only has 5B active parameters, so its not surprising Qwen3.8 27B outperforms it (Qwen3.8 is also ~13 months newer, which is forever in LLMs)
reply
Fair enough. I’ve barley touched oss-120b, so i didn’t know it was so few active params. For a direct comparison, qwen3.6-35b-a3b is still better at coding than oss-120b.

And Qwen3.8-27b is still better at coding than opus 4.1.

Yes, if you list off models 27b is better than it’s all older models. But that’s my point - newer models are better than older models at the same AND much smaller size. That’s because model size matters less than they say. Training data and model architecture matter more.

reply
No, it’s not the active parameters. Qwen 3.8 Flash has 6B active and it smokes both models.
reply
The western labs are very AGI pilled, and their public models are distilled down from larger research-only models that are uneconomical to serve directly. They could (and probably will) start distilling models for more niche use cases eventually, but we're not there yet.
reply
> more diverse knowledge tend to cross-pollinate across domains

Yeah, the cross domain transfer learning from RL is overstated by a lot.

reply
Problem is conflict of interest: the studies are mostly from the providers of the biggest models, or someone who received free tokens to do the research.
reply
It would be nice to hear exactly how the conflict of interest has impacted the specific studies and how they are wrong rather than conspiracy theory level speculation and hand waving at the entire category
reply
I think is more than reasonable to be suspicious of studies funded by party with conflict of interest. Think of how many studies about climate change were funded by big polluters, for example.

In 2026, the default outlook should be suspicion for any big private organisations with profit motive.

reply
I think it’s less than reasonable to operate mainly on vibes, rumors, and hearsay.

You don’t need to think about climate change studies. Instead you can read the allegedly tainted studies we’re actually talking about and profess to all of us what is wrong with them. You can’t point to exactly where they’ve fudged them.

reply
In an alternate universe where all research is completely auditable and we all have infinite time, yes that's a valid approach. And you're welcome to spend your life going that route, but the rest of us are gonna trust our common sense on this.
reply
[dead]
reply
I think what will keep the industry afloat, all else failing, is the surveillance industry! Nothing like a fat reoccurring cheque from the government to check if little Jimmy is committing thought crime!
reply
LLMs needing less compute would actually be a good thing for Nvidia due to Jevons paradox. Right now token costs are an impediment to using AI more broadly, and more efficient models would help adoption in cases where AI has proven to be useful, like coding.

https://en.wikipedia.org/wiki/Jevons_paradox

reply
Jevon’s paradox is a common talking point but it is not a law of nature. LED lightbulbs use 80% less energy than incandescent but you don’t see people using 5x more lights on their homes. The overall energy used to light homes has decreased.

And even if compute demand were perfectly elastic it’s only a good thing insofar as it drives demand for new Nvidia hardware. If tokens can be served from Apple hardware or Google hardware or Huawei hardware that doesn’t help Nvidia.

reply
> LED lightbulbs use 80% less energy than incandescent but you don’t see people using 5x more lights on their homes.

I mean… some homes definitely do. You must have seen those houses that are all lit up front the outside by lawn mounted spotlights.

reply
Perhaps. But their huge valuation is based on them supplying the massive buildout of data centers that’s happening / planned.

If that dies because a lot of people’s needs turn out to be met by a system at home they can run a 30b-150b model on, a lot more of that money goes to apple or intel or amd.

reply
That's unless the code produced in the future is much more complex than today's.
reply
Sure, but it would be actively bad to make the code more complex simply because we have machinery that helps us deal with the complexity. A big part of how people assess the models' coding capability is whether they create needless, incidental complexity.
reply
That's like saying it'd be actively bad to make the code more resource intensive simply because we have machinery that helps us deal with the extra requirements. And as we know as computers got more powerful code didn't get lighter. If it can, it will.
reply
Assuming it’s all going to be vibe coded garbage, yeah it will be much more complex. Like a toddler writing a symphony.
reply
This is the right kind of analysis, but we can look broader. Both the demand and supply situations are a lot more extreme and dynamic than appears at first glance. E.g. to your points:

1. Yes, smaller models will become more popular, especially as the tokenmaxxing trend dies down and people start stretching their budgets farther. That is a downward pressure on demand.

But along the same dimension, consider that currently only about 40 - 60% of the world uses AI for only about 5 - 15% of their work hours. That means there is still 2x growth from users and 7x - 20x growth from the rest of the work hours left to capture! That is 14x - 40x more demand. Then consider that agentic tasks require multiples more tokens, and that is the kind of usage that is most likely to be deployed, and also the kind of usage that is the least used right now. That's another huge multiple to be tacked on.

And the entire AI industry has been lamenting the extreme compute crunch they're facing (and also why Claude has 9's comparable to GitHub; whereas OpenAI has been chugging along because Altman was OK being called a "podcasting bro" while desperately scrounging for compute years in advance.)

Nvidia's meteoric rise is entirely due to this kind of exploding demand with extremely limited supply.

2. Competing hardware is definitely a threat, but it has its own hurdles. Because the real bottleneck is not Nvidia, it's TSMC.

Pretty much all demand for all chips in all devices in all the world flow to, like, 3 companies in the world that actually fabricate them, and TSMC is the biggest. And the supply is extremely tight, as the exploding costs of electronics clearly shows.

So now TSMC will of course try to keep all its customers happy, but it will inevitably be forced to choose which ones it will keep happiest. And those will be the customers who can pay it the most. And that would be the one with all the money from its de facto status as a monopoly (and possibly even a monopsony)...

Which would be Nvidia ;-)

So yes, compute per task is falling rapidly... but it's barely a dent in the humongous total addressable demand, and the amount of hardware to support that compute is still very constrained, and most of that supply will likely flow through Nvidia.

reply
You're not considering video which OpenAI opted out of when they retired Sora.

Generative video requires significantly more computing power and energy than generative text.

OpenAI is fucked, compute is still needed, it's just them that isn't.

reply
OpenAI dropped sora because it was costing them ridiculous amounts of money and earning them very little. They determined that the market can't support the cost of generating video.

Without a material change in the market (more buyers, vastly cheaper generation), it's unlikely a different company could make that work. More buyers isn't likely to happen, so that leaves vastly cheaper generation - something that would cause nvidia's value to collapse if it happened.

reply
> the market can't support the cost of generating video.

I'd suggest that's only the case given the current quality of output. Media is incredibly expensive to produce. A model capable of sufficiently high quality could charge prices that are absurd by today's standards.

reply
It’s a very small set of buyers that are in that price range. Total annual domestic box office revenue is like $10 billion, maybe $50 billion for global TV and film. And that’s revenue, not profit, and a lot of costs are going to marketing, not to filming and casting. That’s a lot of money, but it’s not the scale that OpenAI and Anthropic are at.

Video generation would only make sense at that scale if it was targeting individual consumers, but then it’d need to cost something that consumers are willing to pay - which practically is probably a few hundred per year at most among US consumers, and much less globally, so again it doesn’t solve for the size of the AI companies.

I don’t see a way that video generation becomes a big industry without making generation much much cheaper.

reply
Aren't these two largely separate questions? Viability versus if a given incumbent has interest in a market of a given size. With the combination of (at minimum) streaming platforms, the box office, and advertisements video and audio generation would be viable at a remarkably high price point (as compared to the current token prices for other sorts of things). And as the price comes down presumably the market would grow larger - by how much I have no idea but there are certainly a great deal of currently underserved niche markets.
reply
Minimax H3 works pretty great and you can run it on a 3090.
reply
There would also need to exist sufficient demand for video, which hasn’t happened yet.
reply
oAI isn't anywhere near close to fucked as long as their models are head and shoulders above even the very best open models in terms of tool calling and rock solid stability/reliability for agents/coding harnesses. Which, they are right now and we'll see if open models actually catch up in that regard. Even the "best" open models pale in comparison with tool calling and general "prompt and go do something else for an hour" reliability that we have with GPT models. With GPT models, streaming rarely stops unexpectedly. You almost never have to constantly nudge them along, etc. Granted with open models all of this can vary depending on the provider, and perhaps open models/protocols/APIs/harnesses aren't well enough aligned, but OpenAI models just seem to work without constant (or hardly any) wrinkles and with almost any harness/agent.
reply
>oAI isn't anywhere near close to fucked as long as their models are head and shoulders above even the very best open models in terms of tool calling and rock solid stability/reliability for agents/coding harnesses

That's already not the case today. If you sat me in front of an LLM and told me to figure out if I'm working with K3 or Astra, I could probably do it, but it would take some work to be certain.

reply
Well I could for sure. I guess a lot of this is indeed very anecdotal.
reply
All we do with these things is some work though
reply
[flagged]
reply
I've been thinking about that and that's why Nvidia's prices are surprising to me. Investors should know that better than me so there must be something I don't know
reply
It’s really hard to know when the large tech companies have so many shares owned by a single figure. They can use margin loans and options to create the appearance of demand.
reply
They need the right harness and either your help it auto produces in time enough content to further improve.
reply
If you reshuffle your argument, and apply the same facts you get to a similar conclusion but with a drastically different spin.

> it's more specialization

China, constrained by hardware, and talent (not to slight the Chinese, but they are limited to domestic resources - and much of the US effort is very international). They did, what the Chinese do, and optimized the process of production, and drastically lowered the cost of development of their models. Cheeper to build, cheaper to run is just good economics.

Meanwhile in the us, we have open AI doing "experiments" - it looks like the costs around the hugging face hack are going to be about the same as China would spend on building out one of their smaller efforts (several million dollars). (Depending on whos numbers you trust, the fact that I can even make this claim should make you raise an eyebrow).

Go back to the 80s' and "expert systems" - most people will tell you that for their time, they were amazing, and useful. People would have loved to have more of them but they were so cost prohibitive that we all but abandoned them for serious use. The US frontier labs seem to have forgotten this lesson and their calls to "slow down" look like an excuse to "cut the waste so we can move to making money".

reply