upvote
I recommend reading some of their research it's honestly astonishing how intelligent some of their solutions are.

Kimi specifically relies heavily on reasoning traces which is largely due to their training strategy and will perform poorly when thrown into a conversation from another model. Another fun advancement is that they simply ctrl+c ctrl+v'd attention which means that the model can steer where to look in the context window without ever producing an output token increasing token efficiency and attention accuracy as a side effect you end up with weaker prompt adherence.

None of these 'issues' manifest in US models which proves that kimi has diverged and is achieving these capabilities seperately from the architecture that US labs rely on.

I would agree with you during the Deepseek R1 era, but US labs were heavily inspired by open research at that point as well so I wouldn't give them too much credit.

reply
You still left out that it doesn't matter anymore, just like Anthropic/Facebook/OpenAI only really needed to read massive amounts of copyrighted data only once (and of course, they all did this illegally, which makes their current complaints more than a little ...). Once they have a large model trained on the data, they can just retrieve reasoning traces and copyrighted data from the previous model. In fact that is a training technique long used because it has better results that directly training on the original data.

In other words: even if the US (somehow) denies them access to the current OpenAI/Anthropic models, they'll be able to improve based on what they already have.

reply