If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses.
Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)