"less context is better and if you can't get stuff done with less yur bad" is the worst argument ever.
it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images.
it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.
The services are priced this way because larger context has significantly higher costs. That is a fact about the technology and it is true for every provider. So moaning about it isn't useful.
On the other hand, there are a lot of people who don't manage context effectively - who start every session with 60K tokens - and that is significantly hurting the performance of every single thing they do with coding agents.