upvote
A per-token model roughly aligns with the providers' costs, and it is an objective measure, so it seems a reasonable way to charge.

I see posts all the time on HN about which models from which providers offer the most bang-for-the-buck, and how to minimize token usage and still get optimal results, so it appears that competition is working.

reply
You could try:

1. Self hosting

2. Chinese models

3. Running it locally. Requires upfront cost and compromises on TPS.

reply
There are more than 3.

Hell, I use 3 different providers, and I currently don't give a dime to Anthropic or OpenAI.

reply