Inte Cost
llig per
Open Weight model ence Speed Task
Mimo-V.26-Pro 46 47 $0.13
GLM-5.3 (max) 45 73 $2.01
DeepSeek 4.1 Flash Max 39 227 $0.27
Mistral Large 4 Preview 38 116 $1.13 Inte Cost AA- Omni
llig per Omni Softw
Open Weight model ence Speed Task score Eng
Mimo-V.26-Pro 46 47 $0.13 8 33
GLM-5.3 (max) 45 73 $2.01 14 37
DeepSeek 4.1 Flash Max 39 227 $0.27 -5 34
Mistral Large 4 Preview 38 116 $1.13 -5 5
Closed/proprietary ~50 110- $1.50- ~43 ~85
median top 10 242 $7.50
[0] https://artificialanalysis.ai/evaluations/omniscience
[1] https://artificialanalysis.ai/evaluations/omniscience?detail...Issue is..
I don't believe for an instant that any of us, including US citizens, get access to the best models for cyber that the US has. I think any adversary would have to assume the models in use by the US side are unreleased.
US is not the only one dealing under the table by the way, I also think everyone should take China having unreleased models as an operating assumption at this point.
So Mistral is the best that the public gets access to. And that's if it's even the best? Benchmarks and pragmatic work have often been shown to be two radically different things in this industry.
The NRA approach to AI safety.