upvote
Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.
reply
Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO
reply
My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.
reply