Hacker News
new
past
comments
ask
show
jobs
points
by
Eridrus
4 hours ago
|
comments
by
hgoel
1 hours ago
|
next
[-]
Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.
reply
by
sebzim4500
3 hours ago
|
prev
|
next
[-]
Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO
reply
by
foota
3 hours ago
|
prev
|
[-]
My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.
reply