Distillation is alive and well... Earlier work on model printing also found that it's pretty easy to find smaller sets of parameters which can replicate the behavior of the entire network with pretty good fidelity.
Large parameter counts give space to explore, and give routes out of what would be local minima in a lower dimensional space.
In other words, there's no guarantee that any given trained model is a minimal representation of its training set.
Your socioeconomic argument just doesn't hold either. People don't delay releasing models until they've minimized it to the theoretical limit. They ship it when it's good enough for whatever job they're making it for.