upvote
Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companies are still hanging in there.
reply
The shift isn't new. Kimi K2, a 1T model came out July last year. I am happy that more labs are following the trend as its important for competitive open models to exist.
reply
Also DeepSeek R1 was announced 1.5 years ago with ~0.7T parameters, which was a huge model back then.
reply
And DeepSeek has been making huge progress on efficiency, and publishing about, so they came with a 1.6T model that is both fast and cheap to run.
reply