Hacker News
new
past
comments
ask
show
jobs
points
by
rahen
9 hours ago
|
comments
by
idbnstra
7 hours ago
|
[-]
just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models?
reply
by
rahen
7 hours ago
|
parent
|
next
[-]
It's in the announcement: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe".
https://mistral.ai/news/mistral-large-4/
reply
by
Iolaum
7 hours ago
|
parent
|
prev
|
next
[-]
Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell.
reply
by
ismailmaj
5 hours ago
|
parent
|
[-]
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.
reply
by
karp773
2 hours ago
|
parent
|
prev
|
[-]
It would not take them so long to train it. Their pace would be closer to the Chinese models.
reply