If it's significantly larger than GLM 5.3 (I've heard some insane guesstimates out there like upwards of 5T params or more), that would prove rather embarrassing for Anthropic.
I feel like most benchmarks cluster on a reasonably limited area of human knowledge
You do pay for the tokens, but in theory on a smaller model each token is cheaper.
They filter their results through their AI, though, so you get a sort of meta-summary of the top few results. It did well with geopolitical news stuff, but I've not tried a hard science sort of query.
I should do that myself, come think of it.
(And 4.5 was smaller than 4)
Not really, in that you just work with different constraints.
Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and Chinese labs in general have 1-10.
The priorities are different.
Rather embarrassing indeed.
You implying its better than opus 5?