upvote
> the operators who have no aligned incentives

The model providers are quite aligned with concerns like customer retention. These arguments only work if there is no competition. We exist in a marketplace of black boxes. There's not just "the one" you must suffer. You have options. You can build your own too.

reply
There is so obviously competition in this market it’s astounding to me that people attribute all this malicious behavior to the model companies.

They’re growing over 10x a year. They want users and revenue. In order to get users and revenue, they want to provide the smartest models at affordable prices. If they unnecessarily burn tokens, users will get less value and switch.

This thread is filled with competing comments about their monopolistic power and how when one model provider was no longer doing a good job people switched to a different one.

The competition in this market is ferocious!

reply
there are roughly three of them and they all use the same pricing model. I am also not in the position to build a frontier model company these days.
reply
A per-token model roughly aligns with the providers' costs, and it is an objective measure, so it seems a reasonable way to charge.

I see posts all the time on HN about which models from which providers offer the most bang-for-the-buck, and how to minimize token usage and still get optimal results, so it appears that competition is working.

reply
You could try:

1. Self hosting

2. Chinese models

3. Running it locally. Requires upfront cost and compromises on TPS.

reply
There are more than 3.

Hell, I use 3 different providers, and I currently don't give a dime to Anthropic or OpenAI.

reply
How else would they bill tho? Their operating cost is per token.
reply
Isn’t their operating costs depreciation of hardware and electricity? Token is just an assumed representation of it?
reply
Charge on the input tokens, then you will naturally optimise for fewer output tokens.

Theoretically.

In reality, one sessions output tokens become the next sessions input tokens (at least if you continue the topic) so, its not as aligned as all that.

But the parent is right, when incentives are not aligned, friction will happen. Its inevitable.

reply
“Claude, spend the next 10 hours trying to solve the Reimann Hypothesis”.

I agree that incentives are misaligned but there’s several competing model providers. If one gets funny with their costs people will jump ship, especially if the gap between the top 2 labs and everyone else keeps shrinking.

reply
“I can take current sources and tell you how solved this is, but I am not willing to work to a timeframe or to solve things that aren’t yet solved by mathematicians or science”

These safeguards already exist when they get a whiff that you might be using Claude to fix security issues. Doesn’t seem farfetched given the incentives I outlined that they would apply to this kind of abuse.

How loose those controls are becomes a market force.

reply
Perfect, a coding agent that refuses to do things that haven’t already been done before.
reply
Hahahaha, I think you’ve misunderstood what LLMs are.
reply
[dead]
reply
NeuralWatt just does it on energy consumption.
reply
And tokens can be metered reliably. Unlike something like "task completion".
reply
deleted
reply
One guess is that their "primary" target audience/market is the large corporations that get their employees unlimited tokens, and not the individual developer who may worry about spending and token accounting.
reply
It's the opposite. The enterprises have all the tooling to monitor token usage of employees, and to limit access. For example, we have a $300 month limit, and then need to file exception tickets when we need more to justify the cost. Pretty similar at other non-silicon valley company process. I don't know any enterprise who'se on unlimitaged token budget for their employees. that's not how enterprises sign contracts.

https://code.claude.com/docs/en/admin-setup#set-up-usage-vis...

reply
Seems like yet another instance of printing your own money & getting rich by screwing people forced to use them.

Goas back to factory towns, gift cards, game money or MtG.

reply
For OAI and Anthropic at least you can set a spend limit per response. Also tokens are well-defined.
reply
I'm not worried about the volatility in the definition, i'm worried that I give it 1 token today and receive 2 token output, tomorrow I receive 40. If i'm doing this a hundred thousand times a day it is difficult to price this in for users downstream or in the extreme cases be able to absorb that at all short of going into a failmode with degraded access until someone goes and buys more tokens or gets the bill. The alternative is just pass the buck and bill your non-technical customers with a "tokens" line iteim every month.
reply
> If i'm doing this a hundred thousand times a day it is difficult to price this in

When you’re doing this 100K times per day you get an extremely good idea of what it costs. You also have all the tools to see when something starts changing quickly.

This change is for Claude Code the harness. If you’re using the API at scale and paying full price then you get exactly what you put into the request.

reply
No, those doing this 100k times a day have very good data on this, good estimators and modeling. And the API has various knobs to change and evals will give you actionable data.
reply