The other kind of distillation is where you record the outputs from a teacher model and use it to train a smaller model from scratch. That kind of distillation is not so cheap. It’s cheaper than training a model fully from scratch - starting with pretraining, then alignment, RLHF, the whole riggamarole. But here you are still starting from nothing and need to figure out how to get trillions of random numbers aligned in a way that makes them act intelligent. This is still gonna take a very long time if you’re talking about trillions of parameters.