upvote
> They do not beat opus on real-world usage

We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios (mostly Rust, some C for microcontroller stuff). Qwen3.6-27B scores only 4% lower for pass@1, n=250 compared to Opus-4.8.

For the labeled dataset, the average PR size they're being measured against is around 1.5k SLOC.

This is very much "real-world usage" for us. The sort of change sets that come in daily/weekly and are solving non-trivial issues in the respective codebases.

As is usually the case, the most broad claims from both the labs and from the consequent pushback are talking past each other.

reply
Let us know when you have Qwen vs Qwen comparison stats. As long as there's not a regression, that'd be awesome.
reply
4% is within the margin of error anyways for pass@1, so I think pass@k > 1 is gonna be the better indicator of any movement (still need to calibrate the optimal k to re-test). 10 seems too tolerant even though that tends to be the next tranche I reach for.
reply
Depends on where you sit on the binomial curve. At p=0.04 for n=250 4% points would not be within margin of error.
reply
[dead]
reply
>> We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios

Okay but the parent said real-world usage, presumably meaning coding tasks.

We have a whole bunch of complex evals that Haiku 4.5 passes. That doesn't mean it is a good model for coding.

reply
They literally stated in their first sentence that it was coding tasks.
reply
Yes, these are coding tasks in the embedded systems domain (I mentioned Rust and C).
reply
How much does it score though? 0% would be 4% less if Opus was at 4%. Unless you mean relative fraction not percentage points - but people usually mean percentage points in such situations.
reply
0% is not 4% less than 4%, that would be 3.84%.

0% is 4 percentage points (pp) less than 4%.

reply
[dead]
reply
> ...but no. They do not beat opus on real-world usage.

I agree, but then we just need meaningful benchmarks that clearly show that! Otherwise it's hand waving about something that should be put on paper in quantifiable terms.

reply
If you are working in a company and using language models, it is a very good idea to hold a bunch of evals you can trust and use to validate new models. Calibrate every once in a while with prod data. We have our own and the only numbers on quality and cost I trust come from this setup.
reply
A wise man once said, "not everything that counts can be counted, and not everything that can be counted counts".
reply
In the end, the only benchmark that matters is your own.
reply
Only useful benchmarks are those you (in particular) don't have access to.
reply
The only useful benchmarks are those you've created for your specific workflow. Only then can you assess whether a given model is better or worse for what you are using it for.

There are tools like promptfoo designed for this.

reply
> but then we just need meaningful benchmarks that clearly show that!

That's the rub. AI benchmarks are IMO, by and large totally unreliable. We think of them as similar to traditional benchmarks of deterministic processes where the number of variables is low. But they're anything but that. Non-deterministic processes with an astounding number of variables and fuzzy acceptance criteria.

It leads to results like these, where if you take it at face value, the only conclusion you can draw is "wow Anthropic must be stupid if Opus takes 1T parameters to do what Qwen can do in 27B."

reply
How can you say this when you haven't even tried it yet? Is it just hypothetical vibes?
reply
Yep. These small models are actually worse than GPT 3.5 at some tasks (like recalling facts). You can definitely make models smarter at specific tasks (like tool calling, coding) but you can't compress the entire human knowledge into a 30GB file. It's just not enough bits.
reply
But why would you use a model to store factual knowledge, that is stupid. We want intelligence, not a database.
reply
Models cant make any decisions if they have 0 idea that the feature exist in this language or in some general fact.

For example you would tell a model hey, become an expert in this language for me, search it online, it would still need to learn it and download the data to it's context and then increasing the memory usage, there's no way around it.

reply
that is why we enable web search for the agent. the memory can come from the internet.

deepseek-v4-flash needs web search to return true facts.

reply
There is 0 shot you can make that claim about this model you have not used or downloaded yet
reply
"Benchmark is stupid" and "model beats model on benchmark" are two different things, though. The second one is objectively true regardless of your views on the first one, right? To expect everyone to share your opinion that benchmarks are stupid is pretty weird, and just saying "no" to an objective truth is the definition of delusion.
reply
If a benchmark is a measure of nothing useful, then model beats model is an objectively useless fact
reply