upvote
Sun, Earth, Moon — it’s basically L/M/S like you want but a little less boring.

Why is large better than medium to the average end user of ChatGPT though?

I don’t think there’s a way to name these things that will satisfy everyone.

reply
The naming schema actually tripped me up for a week or so.

My brain's initial conception of the concepts was earth-relative, so I mapped it as:

Sol = big, it's the sun Luna = medium, in-between sun and earth, space Terra = small, terrestrial

reply
Pretty weird when the moon is as much between earth and sun as the earth is between the moon and the sun.
reply
And when it’s in between we cannot even see it (unless it’s exactly in line).
reply
It tripped me up because I was going by distance. I thought Terra was the base model and Luna was the mid model because it’s further away.
reply
Tinfoil hat time: They saw everyone referring to Mythos, and later Fable, as the new “good” models when Anthropic released those, distinguishable from the “regular” Claude (or other companies’ models) for everyone, and didn’t have that distinction for the GPT model family. That’s why the planetary names were introduced.
reply
I think model naming has been atrocious in general, in part because newer "lite" models surpass the capabilities of previous "pro" models (case-in-point: Gemini Flash which now surpasses the capabilities of the latest Gemini Pro, with a newer Flash Lite vying somewhat unsuccessfully for the old Flash price/positioning), but gpt 5.6's Sol/Terra/Luna split is really not bad at all - probably easier to understand than Starbucks' cup sizing!

The problem becomes when you add in the adjustable reasoning efforts and you end up with {model, reasoning_effort} combinations that end up completely obviating particular model classes altogether for at least some percentage of queries; e.g. with GPT 5.6 the price/performance Pareto frontier is dominated by permutations of either Luna and Sol, with Terra nowhere to be seen (but then if you need "large model smells" that aren't captured by your benchmark you can't even rely on this, as a model like Luna simply isn't capable of encoding sufficient world knowledge in its weights to perform certain tasks at any reasoning level but you might be able to get away with Terra on low reasoning, but no one seems to be covering this for some reason).

reply
Yes but with gemini specifically they said that pro was still in training. And the comparison isn't really atrocious unless Gemini 3.5 Pro is worse than Gemini 3.5 flash
reply