(genesisopenmodels.anl.gov)
Ah but Mira Murati's new Inkling is Apache 2.0
But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC
If you watch it think, which you can, unlike American closed models, you can steer it. You can provide a a massive rocket ship stratospheric boost to help it orient itself. You have no self correction, there is no multiplayer in American proprietary models.
Sure it's great having super powerful mystic oracles that have the "right" answers. But I love respect & revere the open thinking. No it's not automous. But it is brilliant. And it considers. A lot. Deeply. It chases. That to me is the most human of models, even as it falls far astray.
You should help it. You can. Unlike these vicious dark surfaces which yield and tell you nothing. I think this is the actual meta-core-super-point of "The session you cannot take with you" (link below). It's the session that does not care about you, will not interact with you, will not peer with you, that is a dead remote far off oracle to you. Fuck these "oracles". They are a plague against the human spirit. We should alloy humanity and AI to Augment Intellect (Engelbart). (To do less is species treason.) https://earendil.com/posts/session-portability/ https://news.ycombinator.com/item?id=49118781
Looks like on <https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27
On <https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia
At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.
The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.
Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.
Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.
(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)
Deepseek is explicitly banned [1] at LLNL and I wouldn't be suprised if there's a blanket ban on all Chinese models. But nowadays models like tera/luna could fill this area of the pareto front, and LANL already runs openai models on their clusters [2]. Maybe it's in custom SFT/RL, for instrument control or sensitive topics? But you'll still have to compete with frontier models + a harness.
I would have also liked to see a carrot tied to their offer. It'll be hard to get teams to contribute RL gyms or curated text. But throw in a "we'll fund a postdoc/student to do that" and I think you'd have teams scrambling to apply.
[1] https://hpc.llnl.gov/about-livermore-computing/ai-ml-lc/lc-l...
[2] https://www.energy.gov/nnsa/articles/nnsas-los-alamos-nation...
[0] https://commission.europa.eu/news-and-media/news/strengtheni...
What do you mean by "basically"?
Why are Anthropic's and OpenAI's annualized revenue about $50B each?
LLMs need massive amounts of compute to compete, so I wouldn't claim that the great (and leading, and likely to continue to lead) LLMs are commodities end-to-end, even if the non-executing-at-scale LLMs files and IP are commoditized. The execute, the compute, that is what breathes life into the model, which is otherwise weak or dead.
Here's an example[1] of the difference between what a U.S. Department of Energy employee adds to a ticket versus a private industry AI completing instructions as assigned.
This isn't some cherry-picked example, it's just what I happen to be dealing with right at this moment, happened just a couple of moments ago.