People use these models for diff things.
Its quite possible for the things they are used for, people do not see much of a difference.
Do you hold stock in Anthropic?
> Unnecessarily aggressive
> First ever comment said "Further releases of Chinese models that demonstrate the gap is not growing substantially is a huge problem. The spending will be called into question."
Yeah I think you have an agenda
are you american?
I wish.
Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer.
The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions.
Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well.
Not to mention that US companies models constantly distill each other as Musk was forced to admit under oath. This whole narrative has just been a massive cope.
> This whole narrative has just been a massive cope.
So wait, US AI companies all use distillation because... it's not effective and it's all just cope? Or is distillation really powerful and they all do it, which Musk was forced to admit under oath? But when China does distillation it isn't powerful and they don't need to do it, but they do it anyway because it's fun?
Either it's powerful and everyone, including the Chinese labs, use it as a way to rapidly catch-up against the SOTA models, or it's a red herring and the huge amounts of energy spent to protect and enable distillation is all just wasted money. Which is it?
Both sides have extremely smart people. One side has more $$$ and exclusive access to the best chips. For progress to converge without a corresponding breakthrough suggests there's something else at work.
The US enjoyed an early advantage due to excessive money being poured into AI which led to the current bubble, and access to the hardware that was needed to train these models initially.
At this point, neither of these factors actually matter that much. The naive approach of simply making models bigger has hit a wall, and now you need ingenuity in figuring out better architecture for them. Precisely because Chinese companies have had to deal with more limited resources, they put a lot more effort into researching different kinds of optimizing techniques. And of course, China is also catching up in chip making, and Huawei clusters are already competitive with Nvidia for training. So, that gap is closing as well.
The big difference is that an absolutely insane amount of money has been spent in the US, while China managed to do this on a fraction of the budget. The AI Investment Surge graph here puts things in perspective. https://hai.stanford.edu/news/inside-the-ai-index-12-takeawa...
But OAI and Anthropic are trying to cash in ahead of their IPO window. I think that window is pretty much gone now.
I don’t think one should pay much attention to them.
My understanding is that the labs ran out of freely available data to train on a while ago, and now primarily rely on human data vendors such as Surge and Mercor to source their data.