Why is large better than medium to the average end user of ChatGPT though?
I don’t think there’s a way to name these things that will satisfy everyone.
My brain's initial conception of the concepts was earth-relative, so I mapped it as:
Sol = big, it's the sun Luna = medium, in-between sun and earth, space Terra = small, terrestrial
The problem becomes when you add in the adjustable reasoning efforts and you end up with {model, reasoning_effort} combinations that end up completely obviating particular model classes altogether for at least some percentage of queries; e.g. with GPT 5.6 the price/performance Pareto frontier is dominated by permutations of either Luna and Sol, with Terra nowhere to be seen (but then if you need "large model smells" that aren't captured by your benchmark you can't even rely on this, as a model like Luna simply isn't capable of encoding sufficient world knowledge in its weights to perform certain tasks at any reasoning level but you might be able to get away with Terra on low reasoning, but no one seems to be covering this for some reason).