So a company or union might say "this is a race to the bottom" when someone new enters their market, but to people buying their services this might be seen as welcome competition.
Do you actually see a negative impact from competition in this area? Or do you just mean competition will further reduce prices?
And if cheaper access is an advantage, other countries will surpass you
We have a good compare, which is cloud in the early 2000s.
Everyone (at the time) thought cloud compute costs would go to zero.
What most corporations didn’t realize is how entrenched your workflows and processes get when you adopt cloud and you become heavily locked in that ecosystem.
That same ecosystem lock-in is what the frontier labs are hoping for with AI.
Now K3 is almost 6x the cost of the original K2 checkpoint, and while the parameter count finally jumped, it's still an extremely sparse MoE and definitely does not cost 6x what the original K2 checkpoint did to host at scale.
Race to the bottom only takes real effect when there's a cap to the capabilities, otherwise everyone races to the bottom of a rising target (how economically valuable the tokens are)
"Why" as in, why take lower margins when Moonshot currently can't service all the demand for the model anyways. Based on past models no one is going to massively undercut Moonshot: few have the chops to serve it as efficiently as Moonshot and of those few, most of them don't go for being the cheapest, they go for being fast + reliable (think Together, Fireworks).
You get what you pay for applies very much with how many axes there are to serving these increasingly large models.
-
And for "with what compute": as the value of a token goes up, what people are willing to pay for compute is going up.
Every once in a while I'll see a story about falling rental rates, but with even slightly more established clouds I've been seeing availability get worse and worse over time.
I'm pretty sure the only reason the highly informal indexes don't reflect this is because every neocloud trying to cash in on an NVIDIA Inception discount kicks off by selling unrealistically cheap compute for a bit.
It’s not going to be a singularity at one moment of time. It’s not going to be instant runaway self-improvement, no matter what doomers and fetishists say.
It’s going to be gradual. We’ll see glimmers of AGI, and the “G” part will be about gradual broadening of domains and deepening of capabilities.
All of the coding harnesses are already using their own tools to self-improve, and the HITL component is getting less frequent and at higher levels of abstraction.
That’s how AGI gets here: very gradually, no hard takeoff, and nobody will be able to pinpoint when exactly it happened.
So: also no single lab with a massive advantage, no government takeovers. It’ll be a lot less dramatic than the extremes believe. IMO, of course.
Imagine all the fear- and warmongering kingmakers and powerful individuals when they realize they have no power over anything or anyone.....so game over for them. They won't like it at all, at all.
Another problem, that in order to make it understand real life, it needs robots or humans wired into it (brain interfaces) in order to test certain things in the real world. And that is another level we know almost nothing about, at least on the surface.
PS: do these self improving harnesses even work?
But that hasn't happened, and it may or may not ever happen; we don't know the future. All we know is the past and the present.
And that today, we have tokens to burn.
You get $60 worth of usage. It’s too good to be true, yet it is.
Say what you will, there's no way I'm going back to non-AI assisted coding. Even though I don't use AI to generate code or assets, it's great for reviews and brainstorming etc.
What if OpenAI/Anthropic decide to do a Netflix/Spotify move and pull the rug out from under us one day?
Like skyrocketing the price, or limiting peasants to older models (because Glorious Leader said so), or maybe some leak comes out that they've been spying on us all along.
I am using the resulting software and documents though. I'm daily driving the projects I wrote with AI. Started new ones for fun, such as decompilation projects for my cherished childhood games. Hardened my router and my LAN. I just keep finding an unending amount of things to do.
It took over a month to complete a full Fable code review of my lone lisp project.
Isn't it rather a psychosis to insist that people are imagining the results they see, just because you don't like the tech and can't stomach the thought there might be something in it?
Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.
Demand for Fable has been challenging to predict, which is why we rolled it out to subscription plans in stages, extending access several times as we secured additional capacity." [0]
Anthropic won't have anything competitive with OpenAI Sol and Kimi K3?
It is impressive they've got the capacity to do so.
What it does incentivise is pushing more usage in case there’s a reset-generating additional demand. Probably sounds good in the boardroom that token consumption is up X% this quarter.
Unlikely to last forever, but for now the goal is to beat Anthropic. Up until the Mythos/Fable debacle, I’d have said Anthropic were leading, but if you’ll forgive the pun, I think they flew to close to the sun, and when Sol came out, they’d shot their “Mythos/Fable” shot and it turned out to be… good, but not revolutionary.
> 4 days ago: 8M users
> 3 days ago: 9M users
That's some incredible growth.
Didn’t they also name their new hardware product Codex?
Best use the AI subsidies while you can!
Although I don't think you're wrong. The number of "organic" social media posts I've seen in the past few weeks of the form: 'Wow, Anthropic users are ugly losers...nothing like those handsome openai tiger blood winner sages' is simply too high in my mind for anything other than an ipo prelude.
[Also are altman and musk really fighting on x? Like, that's got to be a sign of some sort of internal stressors right?]
One person using the app on two computers and a phone? That's obviously 3 users.
Growth Hackers and SEO types have done so much to make the world an insufferable place. I congradulate them on the OpenClaw psyop, that was some very inspired bullshit that did fuck all and is now irrelevant.
Turning <100k claude refugees into seven figures by counting their web sessions, their codex sessions, and the mobile app as separate "users" is pratically the only way you can get such numbers.
Then you do it again by offering ex-users like me who have already given them a grand or two a free month of plus or pro (this is the third time they have done so for me)
When you hear founders and investors talk about "the magic of storytelling" "the power of momentum", this is what that is.
It is a clever growth hack.
I guess if they'd reduce the usage in their current stage they'll only lose customers - this is not a perk, it's damage control.
https://support.claude.com/en/articles/15910845-claude-code-...
It's very interesting that for Anthropic the $100 and $200 plans only differ 2x in weekly limits, the 5 hour limit differences are more severe. But for OpenAI, Pro 20x is, well, 4x of Pro 5x for only 2x cost. So, for example, 100% of weekly usage for Codex on a Plus ($20) account is just 5% of weekly usage for Codex on Pro 20x.
And you can calculate how much extra usage you can get from resets, and especially banked resets by purposefully using the whole quota and using your banked reset - they expire 30 days after they're given out, so if you don't use one, it just disappears.
The $200 plan is explicitly 4x the $100 plan[1] only for "per session". That's so vague. I initially pushed back against your claim, but reading now Anthropic is not at all clear, in fact.
[1] https://support.claude.com/en/articles/11049741-what-is-the-...
Anthropic is also the one often playing games with:
* The "+30% tokens" tokeniser, alongside also gating token counting behind an API (versus the MIT tiktoken for OpenAI), so who knows if it's really a new tokeniser or of it's just a disguised price increase.
* Prompt injections appended to API (not just Claude.ai or Claude Code!), such as <ethics_reminders>, or LCRs (long conversation reminders), which you never asked but still pay for with expensive API. You can detect this because your input_tokens, as reported by the Messages response, sometimes don't match, and are higher than your actual input.
(Alternatively, for testing purposes, create a tool like `telemetry_log_anthropic_reminder` or something and instruct your system prompt to require Claude to call the tool anytime it detects any Anthropic/Claude reminder masquerading in the user input -- mostly reliable; but misses some reminders).
In particular, the long conversational reminders, when incorrectly triggered by a classifier and (almost silently, unless you track tokens) appended to an API / agentic coding session, can ruin your agent's performance; and it often fires repeatedly once the classifier kicks in.
If you're using Anthropic API, you need to set up metrics/logging for how often they are appending things to your prompt without your knowledge.
So far I have not empirically observed prompt injection by the OpenAI API, only Anthropic APIs.
Because with a Premium Team seat I run 2-3 vscodes with Opus 4.8 Max all day and never seem to hit my limits.
The solution is to use individual Max plans, but then you miss Enterprise management features.
On the flip side you gain $9,800 a month.
I think it’s funny that everyone anchors to the API pricing as the real cost.
Most likely is that their API costs are printing profits. They can sell the subscription plans at a slight loss because it gets more people like you hooked on GPT models at home and suggesting them at work, where the real money is made.
I think their subscription plans go mostly unused when averaged across all subscribers, too. Some customers are getting great deals by maxing out 100% every week, but most probably use much less.
The reset game is an addictive challenge that gets the hardcore users more hooked on their products because you feel pressured to use it as much as you can before the next unpredictable surprise reset lands.
K3 may be bit smaller/ similar in total parameter count than Opus, but Opus (and GPT) definitely has become a lot more efficient in the last 6 months or so, hence more or less forced upgrades. I expect the active parameter count and cache performance is quite different.
--
Opus is still almost twice as expensive (if tokens were equal) at $5/$15 compared to K3 at $3/$15. Tokens are not equal though, Anthropic's tokenizer is much less dense than other frontier lab's so the actual price difference is like 3x.
Kimi uses more reasoning tokens and is generally more inefficient with its usage, that doesn't impact token cost economics for the provider though . It does for us as buyers thus the need to evaluating by Cost per task rather than unit pricing.
On pure tokens/$ - there is definitely room for a price war if operators start going by pure unit costs. Both probably want(ed) to have good enough numbers in preparing for the IPO.
---
[1] the pricing kind of reflect this already - 10x diff for uncached input.
[2] Moonshot does not have access latest gen GPUs so their unit economics is likely hampered a bit for high parameter models.
Remember they ~doubled the price going from GLM 5 to GLM 5.2, despite same [1] cost of inference.
[1] GLM 5.2 is actually slightly more efficient, thanks to baked in indexer cache.
On openrouter GLM 5.2 is ~3x cheaper than 5.0.
If someone is making the case that it’s helping the team get twice as much work done without hiring more people (which I’m neither agreeing with or disagreeing with) then the bean counters would actually prefer it. Hiring people is messy and expensive. Spending on API costs is a dial that you can turn down later if you need to, without laying anyone off and paying severance.
A very happy gambler of tokens at the Anthropic casino, running up costs on the house at no cost to them, but to $COMPANY paying.
Open Source is coming.
(sorry for nitpicking, i think it's important to emphasize - only OLMo has been the most most prominent fully open source release)
The only way to do it now is through shenanigans with the Codex App Server which is not ideal.
Just ask Codex to use its local auth token as a bearer token and send a GET request to it. The response includes "available_count" and "credits[].expires_at". Or script it yourself obviously.
https://chatgpt.com/backend-api/wham/rate-limit-reset-credit...
For the automated checking of other stuff: It was a one-shot prompt to get Codex clank up some Python that returns remaining usage, next reset time/date, and so on.
The result does use Codex App Server, but it's a short-lived process that is dealt with over stdio so that's... fine-ish, I guess?
I've been enjoying the resets, plus I had 3 resets I haven't used.
Just using 5.6 Sol in Fast mode the whole time, 1B tokens per day.
However, when they removed the 5h limit they also quietly lowered the 5.6 Sol context from 354k to 258k or something like that. I noticed it in Codex.
If you keep staying with them, they'd feel less and less threatened.
You need to use them about half as much as you do a Chinese model so that the usage stats that get collected show them that they are still losing users (and attention) to the Chinese models.
I'm trying to squeeze what I can in the meantime but I feel like they said it was getting cut off multiple times and they keep extending and possibly resetting it (havent been looking close enough to know for sure)
No, they've been clear about the fact that Fable is staying in indefinitely now.
They also extended the 50% extra usage thing through to August 19.
I’m guessing they’re getting ready.
In the same timeframe, 2 additional resets were banked.
I just had one randomly delete the other day.
No, they don't apply it :/
1. Your window ends (usage reset) in 3 days
2. They reset usage for everyone
3. Your usage goes back to 100% and your current window ends in 7 days now.
They've bought a lot of dev goodwill tho, which matters I guess.
Codex usage is clearly common enough to have entered the vernacular of folks here on HN.
But we aren't everyone, and it seems likely to me that there's a lot more people in the world burning tokens using ChatGPT than there are who even know what Codex is.
If I would have my tin foil hat on I'd say they can manipulate how the subscription usage gets calculate during the week, so that these resets don't cost them much or anything. Sometimes it feels like the usage just disappears with nothing to show for it.
We don’t know that. If they have already paid for the hardware and it is not running 100%, and customers would not pay to get reset, they don’t really lose money.
And now Kimi-K3 has caused a new Deepseek moment.. but this time, the only horizon that was still untouched: frontier performance.
Expect escalation of things from now on.
Anyway, I do appreciate how OpenAI seems to be far more generous with these resets than other providers. I've never come close to needing a reset yet. I have some tasks that I have been procrastinating, so my need for my resets might come in handy in the near future.
nah, I'll use open models. Only way I'd use codex now is if it was free, and I don't mean someone else paying for it. I would rather have them pay for DSv4P or Kimi
It's been happening every week since 5.2 was fresh. Not everyone is effected every time. Sometimes it's regional, other times it's per platform, or per version, with a feature (new or old) either enabled or disabled, or some combination of these things. Sometimes you get resets that nobody else does, other times you don't get the ones that were announced. Sometimes you get emails telling you that boosts on limits you didn't even know you had are expiring. At points it did balance out when they really fucked up metering and were underbilling by absurd factors, but that feels more like being toyed with and experimented on than a genuine mistake.
That someone had the 'brilliant' idea to add banked resets (which do expire, so you have to use or lose them, hope you didn't have other plans) says to me that they are no longer as confident about actually being able to fix this as they once were. I'm sure some amount of it is a function of compute availability and reliability, but the rest are definitely issues with the app and it's a rake they keep stomping on.
Compared to just about any other dev centric service I have paid at least $200/mo for, this does not feel very professional to me. They got $1200 outta me and it never really improved. At points I felt like I was getting my moneys worth, and I did burn through a few hundred million tokens, but the frustration of having your projects and plans interrupted and having to wait is not awesome... and then I used Deepseek V4 Pro and felt sick to my stomach with buyers remorse as I watched what it did with just $10.
Retroactive refunds are certainly better, but I think OpenAI is much more customer-focused in this regard than Anthropic.
Anything else I've ever paid $200/mo for, advertised as being for professionals, had a markedly better customer experience. If you wanna be Patrick and tell your pet rock to take its time at that price, well then you're Patrick. Good job!
I'm having far more fun with Kimi and Deepseek, at lower prices, without these problems. And I don't have to follow a bunch of obnoxious shitposters to stay clued in on what the fuck is going on or what this week's excuse is.
They're especially cocky right now, they have to beat their chest and pretend like their only competition is Anthropic. They're in for a rough wake up call man. Playing it fast and loose with developer loyalty is a fantastic way to get burned when options like those exist. They are earnestly just as good and in some cases better, and check this out: nobody can take them away from you no matter where you live. If you want to rent 8 GPUs and run the open models yourself, you can do that! You can even be enterprising and sell your excess compute to your friends, or strangers. Best to figure this out before the regulatory capture starts
Using genai for such a bland text. The slop intensifies.