upvote
If it's for personal use, is there a reason you don't want the subscription?

Anthropic's $20 subscription gives >$500 worth of credit by most measures, which is pretty similar, and you get a better model.

And as another commenter said, Luna is the cost leader at the moment if you really need API pricing.

reply
Say that to Luna's face. Ya all bringing up this not-so-cheap-nowadays chinese models and not that more intelligent than luna and bringing "cost" as the only factor.
reply
Obviously not as "intelligent" but almost 10x cheaper

Mimo 2.6 Pro: 0.04/0.4/0.87

Sonnet 5.5: 0.2/2/10

Opus 5.5: Sonnet prices times 2

What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?

reply
deleted
reply
For what it's worth, the subscription plans of Anthropic are also 20x cheaper per token than the API prices. The API prices seem to have a very healthy margin.
reply
I was under the impression that the cache write fee was added to both the input and output costs (except in cases where the cache write is explicitly disabled via e.g. DISABLE_PROMPT_CACHING). The output becomes part of the context, after all; if they don't (for some reason, due to disaggregated inference perhaps) then I'd expect output tokens get charged both output then input+cache_write on the subsequent completion request.

The pricing model confuses me though (I presume by design, Hanlon be damned).

reply
You don't have to pay for cache write if prompt isn't part of conversation.
reply
What model are you using ?
reply
Not him but 3 models dominate: GLM 5.3, Qwen 3.8, Mimo 2.6. All censoring certain things. Numbers and other uses are perfectly fine. They are like 0.10-0.15 per 1M tokens. American AI lost the game already, people just can't see it.
reply
> American AI lost the game already, people just can't see it.

Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.

Your take is the "current year is the year of the linux desktop" meme of "ai"

reply
You're a decent sized company and wants to manage the SW + security on all your employee's PCs. They need to be able to update/remote SW on your machine remotely, see your settings, etc.

I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.

reply
>Linux

cfengine is 33 years old

reply
The difference is that Microsoft did that while relying on the extremely load bearing windows ecosystem. These AI companies have no equivalent lock in, nothing even close to it honestly.
reply
I can guarantee you 80% of people will call any llm "a chatgpt", most have never heard of claude, even less of opus, or sonnet, "deepseek" probably reminds them of a brand of toothpaste or something like that, "GLM" might make them think of the new mercedes SUV perhaps. 99.9% will never self host, nor send a single sent to a chinese model provider.
reply
They are not the ones spending API money.
reply
I don't know a single company using deepseek internally in any capacity, and I have friends in a lot of tech/tech heavy companies, virtually all of them use claude, the lucky ones get cursor with claude/chatgpt/grok.
reply
Most businesses I know switched to Google Sheets or Google Docs.
reply
Never heard about anyone using google sheets and docs for messaging, email, presentations, etc.
reply
You’ve never heard of anyone using gmail for email?
reply
Gmail, Meet and Presentations do all of those

People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it

Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite

reply
[dead]
reply
This ignores two things.

OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.

People will also pay for the best (or the perception of being the best). Since it's hard to tell what "intelligence" really means model to model, there's a sense of safety in giving a task to the "best".

reply
People dont want iPhone 14s in 2026. People want the latest and greatest. Chinese companies desperately trying to get western usage of their models
reply
They would if those phones were 899/10=89.9 bucks.
reply
How are these models so cheap?
reply
From my experience GLM 5.3 is at most 6 months away from the frontier models, and good enough for most tasks already

Even 5.2 is doing really well in comparison here: https://labs.scale.com/leaderboard/sweatlas-refactoring

reply
The usage numbers tell a different story. Not only is the great majority of AI users on American frontier providers, they are also willing to pay the premium price for the premium product. It's not like tech is oblivious to the Chinese models. It's that the industry is aware they are always 6 months to a year behind.

You didn't discover some new trick for cost performance. And the rest of the world isn't dumb.

You're just too broke to afford the supercar and justifying the hooptie. It gets you to the kindergarten class after all. And that's all you need.

reply
Europe didn't even start in the race. Europe is the biggest losers of all it seems. What a shame.
reply
What are they censoring that matters to programmers?
reply
Does AI-style writing start bleeding into the comments, or is HN now also full of bots like reddit?
reply
What harness are you using with those?
reply
opencode
reply