When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?
Only if you don't have to continuously train new models, and you are not at a runway risk.