upvote
It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.
reply
Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training.

The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to.

[1]: https://poolside.ai/ [2]: https://poolside.ai/blog/introducing-laguna-s-2-1

reply
Do you have a reference explaining these costs ? Part by part.
reply
What, you don't want a model called the Shitstral-3B :D
reply