The internal model they used to solve the Navier-Stoke's problem was significantly better than the public Astra model, and they also used 10,000 agents.
Astra wasn't even released 3mo ago. It would not be surprising in the slightest that a public model from 3mo would not be capable of solving this problem, but an internal one from current day would be.
my speculation is that they have math-specialized model retrain, so it doesn't need to have all world info in weights, but can focus on math RL training.