upvote
Yea if i use opus 5.5 in api through openrouter and pi agent harness I will easily burn 50-100$ a day (and with fable 5.1 i could burn 200$ easily). Whereas i have now been using a claude code subscription for 2 weeks using 5.5 at all times and have never hit a limit. I often run 6+ agent sessions at once.

I do think its important long term to not be reliant on these companies as you don't have control over the system prompts, the thinking tokens, and once the subsidization stops or the company is public they will be required to start making money and thus raise prices.

But models may get more intelligent and cheaper once that time comes so it may be a non issue.

reply
FYI you can modify the system prompt using mitmproxy. Just ask your agent to walk you through it. Anthropic system prompts are gnarly and geared towards the lowest common denominator.
reply
> FYI you can modify the system prompt using mitmproxy

The system prompt is injected into your context on the server side.

reply
there is literally a --system-prompt flag for claude code. In my mitmproxy experiments using that flag appended to the existing system prompt rather than replacing it. So I had to create a little helper to strip the system prompt sent over the wire and add my own.

here is a collection of public system prompts from claude code: https://github.com/Piebald-AI/claude-code-system-prompts

reply
Use --system-prompt-file, it will replace the whole system prompt (https://code.claude.com/docs/en/cli-reference ). You'll have to use a shell alias or function or something to always append this flag when calling Claude Code but you don't need any MitM shenanigans. Then use "/context all" to see what else is sent (here I would recommend MitM'ing since Claude Code won't show the exact tools and text), there are a lot of tools no one needs and they are bloating the context, you can deny these in the settings.json (there is also a list here: https://code.claude.com/docs/en/tools-reference ). Also set "disableClaudeAiConnectors" to false to remove even more bloat.
reply
I tried both --system-prompt and --system-prompt-file. they both appended when i tried about 6 months ago and watched the traffic. Yes I cleanup all those tools etc.
reply
--system-prompt-file replaces the prompt, I use it myself and the docs state it, too.
reply
thinking you can control what happens to some strings you send over the network.

...the absolute state of the token maximizers.

reply
Claude, think for me. Make no mistakes.
reply
I dont think itll be an issue. Opus 5.5 now is way more than enough for me and open weight models will reach that level by the time subsidization stops
reply
640K [RAM] ought to be enough for anybody!
reply
People were perfectly happy with 640k ram at the time that phrase was uttered.
reply
> People were perfectly happy with 640k ram at the time that phrase was uttered.

I'm assuming this is a sarcastic response to the person who said "640K [RAM] ought to be enough for anybody!"

Because, as somebody who was around when we had 640 K RAM, people certainly weren't happy with that amount.

reply
Ah, the hours spent tweaking CONFIG.SYS and AUTOEXEC.BAT to run a MS-DOS game that squeezed every last byte of those 640K, good times.
reply
EMS...
reply
Regardless of the time, people have always happily accepted more.
reply
I’d argue most “people” haven’t actually needed more in a very long time other than to keep the same old software running. The requirements bloat of operating systems, browsers, and majority of software isn’t really a generalized “people” thing; more so the state of the industry being a form of inertial bloat.
reply
The phrase probably wasn't uttered by anyone it's attributed to: https://quoteinvestigator.com/2011/09/08/640k-enough/. Even "the time" is unknown.
reply
And Bill probably _didn't_ do anything on Epstein Island...
reply
[dead]
reply
It's enough for you now, but I feel like part of the mythology of our future is that we'll be continued to be employed because we'll be working on more complex problems, with smarter LLMs at our side.
reply
I do think there's a decent case to be made that this will be the case for the foreseeable future, and perhaps the mythology part is beyond that.

Although as soon as I wrote 'foreseeable future' I came to the realization that this is far far far less far out than it used to be. Which might mean I simply agree with you.

reply
> Yea if i use opus 5.5 in api through openrouter and pi agent harness I will easily burn 50-100$ a day (and with fable 5.1 i could burn 200$ easily).

Checked yesterday, for that day alone I had used $168 worth on my $20 subscription in Claude Code. I still had plenty of weekly use left. Seems like subscriptions are discounted at a 1:10 rate?

reply
reply
Nice, good to see an up-to-date reference on this.

EDIT: That was a delightfully thorough analysis. Nice to see Anthropic taking the value crown, only because Opus 5.5 is such a joy to use.

reply
Only think missing from the analysis, and it's a big one, is a comparison on the basis of work done. For a fixed set of real tasks, how much of your quota (or api usage) does it take you to accomplish it.

Because a big part of the price difference is due to Anthropic's models being more expensive that OpenAIs AND using more tokens for the same tasks (at least according to artificialanalysis.ai).

reply
Curious to know how much tokens/task will change the conclusion
reply
[flagged]
reply
100%

I use a personal Claude account for personal projects and can let rabl run for an hour and barely make a dent into my usage

On the enterprise I have to be a lot more careful or I can burn through 2k in a week

The excuse they give is the guarantees you get with enterprise plans that they won’t look at your data

reply
What's rabl?
reply
Other than their typo there is a fun collision there with an older Ruby gem: https://github.com/nesquena/rabl
reply
Man, I recognized the name from my first job out of uni at a Rails shop in 2011.
reply
Meant fable*
reply
[dead]
reply
Deepseek also came with heavily subsidized plans, at least at first.

This is why there are so many comments confused that you spent that much money on Deepseek. Those who got on one of the discounted plans could use very large numbers of tokens for trivial prices.

There is also a strange double standard for accounting for open models. People will look at OpenAI or Anthropic and say that we need to consider all of their training costs and employee compensation when thinking about the cost to serve their models, but when the models are released as open weights those costs are ignored. So in that way, the Deepseek models are heavily subsidized as well, with the possibility of them being served by a different company that paid nothing to develop them.

> Quality was ok, seems slightly above Luna quality perhaps?

I agree that it’s about in line with what you can get from Luna or Haiku, but I give the edge to Luna and Haiku when it comes to tasks that require world knowledge. It feels like their training sets were just cleaner.

Luna and Haiku are also close to free with a subscription plan, and they’re even very cheap at API rates.

So the amazing thing about Deepseek Flash is that you can almost kind of get that level of performance from an open weight model. It’s not as amazing when you start comparing it for how we really use smaller models from frontier labs on subscription plans.

reply
> So in that way, the Deepseek models are heavily subsidized as well, with the possibility of them being served by a different company that paid nothing to develop them.

You could then of course argument that oAI/ant/G's models were also heavily subsidized by using (scraping) the bulk of humanity's global knowledge for free while paying nothing for it and trying to privatize it.

reply
That accounting exists because DeepSeek doesn't need to exist for me to use 4.1 Flash. If Anthropic goes under, Claude probably just ceases to exist.

That accounting doesn't make sense when looking at future models, but if you're evaluating the current state it makes sense to ignore training costs for companies that don't pay them.

You can also get subscriptions for the open models, which are typically much cheaper. It's not crazy popular because the open model crowd switches often, but they do exist if you want a firehose of tokens.

reply

    > Deepseek also came with heavily subsidized plans, at least at first.
Do you have a source for that?

I got a source straight from the hoses mouth, which claims that DeepSeek-R1 was highly profitable:

    > If all tokens were billed at DeepSeek-R1’s pricing (*), the total daily revenue would be $562,027, with a cost profit margin of 545%.
https://github.com/deepseek-ai/open-infra-index/blob/main/20...

With DeepSeek-V4.1-Flash, the memory requirements for the KV cache have been reduced by over 50x and FLOPS by over 5x, so it is even cheaper: https://arxiv.org/pdf/2609.19969

reply
Deepseek Flash plans were very cheap when they first launched. They raised prices https://finance.yahoo.com/technology/ai/articles/deepseek-ra...

Your link is about a different model.

The talk about subsidies is confusing because for OpenAI and Anthropic people usually include their salaries, training costs, and everything else. When the topic switches to open weight models we pretend the models appeared out of the ether at zero cost, and the only cost is running the servers.

reply
deleted
reply
Right, but I'm not sure that was because they were subsidizing it, or because they ran into capacity issues and needed to offload some users.
reply
> So in that way, the Deepseek models are heavily subsidized as well, with the possibility of them being served by a different company that paid nothing to develop them.

Has anyone seen the neo-cloud profit margin on deepseek? I wonder if the deepseek served API prices account for the training cost? Because they don't / have not raised the money to fund their future operation / training- and they depend on that cashflow for now? That would suggest a really high profit margin on inference only.

reply
Deepseek never subsidized their plans. At some point they dropped pricing, then after a while they raised them, but that's because they had capacity issues, not because of dropping subsidies. They don't want your money, they'd rather have fewer users.

OpenCode is building their own inference and they've independently stated that DeepSeek's old (cheaper) pricing is achievable without subsidies.

reply
I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).

And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.

reply
It's wild how different usage patterns are between users.

I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.

People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.

Those usage patterns don't correlate to output.

reply
This reminds me of when we gave clients the ability to build their own Power BI dashboards. The users would end up doing a full table dump multiple times for the same table in their reports. Requiring 16GB of ram on the server and maxing out the database every time the report was refreshed.

We ended up having to hire a full time employee to fix the performance of client built reports.

reply
I've seen the same problem with Snowflake integrated with Claude.

People run a stupid amount of expensive queries that end up costing way too much because they're asking Claude the wrong query.

Not to mention people running wrong queries, using the result as gospel, and then the result has to be sent to a data analyst to be reverse-engineered so the numbers make sense.

reply
Now that Microsoft allows you to do /cost for individual tasks (or whatever they call them today). So I tasked Sol, Astra and Fable in cowork with exactly the same vibe coding task on the exact same zip file containing a code project I needed an update for. Astra used 20x and Fable used 15x of what Sol did.

Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.

As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.

But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...

It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.

reply
I'm assuming you're using Microsoft Cowork or Copilot or something. I suspect it's system prompt and implementation (tooling/harness) isn't the same as that of OpenAI and Anthropic even if the underlying model is supposedly the same. The answers provided can be wildly different and often for the worst.
reply
It's cowork. I'm in enterprise in the EU, so we're more restricted in what we use. My issue is mainly with how impossible it is to do any form of reporting on this. Obviously this isn't the most popular opinion among the people subject to it, but for most things we do in the Microsoft enterprise setup we can basically monitor every thing a device does. With Cowork you can't even see which individual task is eating a users credits. At least not yet.

Which is an issue when you need to get department managers to manage their budgets around the amounts of credits their employees spend. The more of a black box it is, the more governance and corporate bullshit you have to deal with.

The copilot part of it runs "unlimited" on the license. Except it's not unlimited, and this is even more of a blackbox because you can't see any sort of spending and the limit is listed as "extensive use".

reply
Just btw, most assume when you say Cowork you're referring to Anthropic Cowork as it's the original rather than Microsoft's white labelled Cowork. And when you say Sol or whatever model you're referring to OpenAI and use on its servers. Microsoft's versions are not at parity with the originals. I feel your pain on being limited by silly enterprise restrictions.

On your point about "monitoring", personally I feel this is toxic corporate IT culture, enabled and perhaps pushed by the likes of Microsoft with all their tools, which they of course make money off. People have cellphones with cameras making most points in this area moot.

An alternative used in other big corporates is to set budgets, with tiered authorisation approvals for higher limits. The users and their managers can justify why and what they're doing that they need the additional tokens. This also encourages more efficient use of tokens on other work. More efficient use is sometimes counterintuitive. Laissez-faire generally works best.

Cowork requires user approvals for high risk actions such as emailing.

reply
It's because this is how AI has been sold to everyone - just ask, and it will do it.

The better pattern is to let it code the app and then you can use the app to target your data. So you only pay for it once, plus it's deterministic. But yeah, it requires setting up an environment, etc. It becomes "maintenance".

reply
It requires knowing how to code, at some level.
reply
You have to keep it secure, so it’s not “pay once” but has running costs, and need some kind of security scanning, citizen developer devops platform etc etc
reply
[dead]
reply
Likely the user doesn't know what they're doing or has extermely bad workflows. They're prob not managing their cache, and dont use compaction.. Letting context get to 500k and invalidating their cache every 10 tool calls because they have no providor fallback settings.

I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.

Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.

reply
You can easily burn through lots and lots of money on DeepSeek, if you do eg large scale code reviews.

Eg I've used Sashiko locally for Linux kernel code reviews before sending out my contributions out to the world. Sashiko is a great system, but it can burn through tokens like there's no tomorrow.

https://github.com/sashiko-dev/sashiko and https://sashiko.dev/

reply
I do some pretty insane agentic work running constantly. I’m currently burning through multiple $200 accounts every week. I regularly spend $10-$20k worth of tokens a month. Most of it has been going to my experimental c++ compiler project
reply
> I tried the cheapest provider on openrouter and burned through $50 in a few days.

I wish I could observe how some of us are using these tools.

I still struggle to spend $50 in tokens per month, and I exclusively use prepaid API tokens. This is in support of personal projects and two clients. There are billing cycles where I might spend upward of $400, but this is maybe once a year. This is offset by months like August wherein I spent $12 in tokens.

The other advantage with prepaid is that it handles the other direction much better. I don't even know what a quota limit feels like. Being blocked for hours is way more expensive to me and my clients than even $500/m. Losing an entire business day over this wouldn't work out.

I've managed to convince some others to try the same thing. $200/m flat fee is a pretty extreme constant expense if you can be more clever on average.

I think a lot of people are getting pushed around by FOMO effects into spending money on pointless subsidized tokens and have (valid) fears that if they don't maintain the same apparent economic leverage as their peers that they will be left behind. This isn't actually the case, much like lines of code are a really poor indicator for the quality or productivity over a codebase.

reply
I think it depends on how many threads you have running at the same time. I have Claude writing a compiler in one window, a ui framework for the language in another, an application using the installed versions of compiler and frameworks in another, and a ui designer (an Interface Builder lookalike) in another keeping up with the framework.

Each of these has a file it listens to in ~/tmp/<name>.io and whenever one needs something from the other, they message each other via that file. Tasks can bounce back and forth as issues are resolved and tested. At the same time, I keep each busy with a list of tasks

I can fairly easily run out of my $200/month subs every week, if I let Fable be the default model. With Opus it’s less likely. If and when I do, I just have an alias ‘claude.ds’ which fires up Deepseek instead, and burns through far less money, though I don’t think it’s as good at solving problems, just MHO.

reply
> whenever one needs something from the other, they message each other via that file. Tasks can bounce back and forth as issues are resolved and tested.

Why isn't this just one coherent agent loop with subtools/agents as appropriate? If these tasks are related in some way, having a single context would probably make it go much better.

The freewheeling messaging part is where the token bloat is coming from. I suspect that for some of us this is actually the point. I think it's a mostly form of entertainment to do things this way. The next logical step from Factorio gameplay.

Parallel agents remind me a lot about multi core compute. It's incredibly easy to take a single core product and make it run much worse across a lot of cores.

reply
Yeah I keep getting “bell curve meme” vibes when I read about all these super complicated handrolled ways people have of getting multi agent setups to work… I’m more likely to be at the caveman end of the spectrum than the guru end but things are moving so fast, by the time I’ve heard about or vaguely understood these esoteric techniques I can usually trust that if they are any good then the good folks at Anthropic et al have already built it into the default behaviour of the harness. Reminds me of the obsession people have with wrapping a repository pattern around Entity Framework because it proves how clever they are.
reply
Why not use the native inter-session messaging in Claude Code?
reply
[flagged]
reply
I don't think too hard about usage one way or another -- not tokenmaxxing, not avoiding AI. Regularly use up over $1500/m without really trying.
reply
Which models do you primarily use, and can you very roughly list your process? Agentic coding in VSCode with tons of MCPs or... something else? Do you include lots of images or have large codebases? Which agentic harness are you using?

I also find that it's easy to spend like that, but also easy not to with little impact on productivity. At the current moment I'm stuck with rider + copilot (not ideal), but e.g. using GPT 6.1 luna is really, really cheap, and lots of tasks are quickly and decently dealt with even at lower reasoning levels, (added bonus of having low latency). And that model is so cheap, I can't see a hitting 1500$ at api prices realistically - not even close. But it also depends on the harness and codebase.

reply
I use opus primarily, on a mix of pure coding tasks, and log parsing / incident investigation.

I don't have the mental capacity to do a lot of context switching between active work streams, so I'm not doing stuff like leaving a big agent workflow running while doing other things.

All through claude code.

reply
Perhaps you're letting the chat context reach 100%? I suppose that's a way that would drive up spending.
reply
Nah, I'm typically <20% context before I clear and continue
reply
Same... I use the £90/month Claude sub and it's more than enough. Can't comprehend the users spending thousands a month.
reply
deleted
reply
Yes. You have to find the provider with pricing that suits your usage.

I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.

If your tasks are write-heavy, find a provider with cheaper output.

If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.

reply
What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?
reply
Sheesh... what's in a name? :p
reply
Bio shows: https://twin.so/
reply
That promo video … it’s so cringe it feels like satire. But it’s not.

It’s PERFECT!

reply
Codes day and night a snake eating its own tail…
reply
Way too harsh

Lots of guys buying articles on TechCrunch saying they’ll build this, he’s bootstrapped

reply
The hiring page says he raised $10m from LocalGlobe
reply
Funny how it's always another AI API wrapper.
reply
Holy shit, it's an Andon Labs clone but worse. Because Andon Labs actually launched businesses (and lost money on all of them)
reply
>What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?

you can't think of anything to unleash some agents on within the entire digital world at any given time?

you motivate your own personal work only via gauging its' usefulness to others and your own prospects?

sheesh.

built anything for the sake of building yet?

reply
“building for the sake of building” refers to _you_ doing the building
reply
We've entered the idler era of building
reply
I can't believe I haven't made that connection myself tbh
reply
I'm not convinced that humans will remain entertained by this format. I hope not, anyway. That's how we get WALL•E
reply
> That's how we get WALL•E

God I hope so.

reply
We have entered the era of building enterprise scale projects for your own personal use. I can do things it took teams years to build as a hobby project over a weekend.
reply
So much of that "took years" is because they didn't know exactly what the "years from now" state they were building in advance. Another huge chunk is because they started getting customers and had to respond to customer needs, demands, scale, bugfixes, preserve uptime, etc.

And even then, "enterprise" was often a dirty word in these circles. The over-engineered would-be-swiss-army-knife vendor that was mediocre-for-everyone but excellent for nobody.

I have built many tools in the recent past for myself. None of them need to be "enterprise scale." Most of them would be worse for it because the agent output suffers when the pile gets deeper and it's just adding more piles on top.

reply
What's the most impressive and useful thing you have built then?
reply
the very notion that everything has to be 'impressive' to you is ridiculous. You still don't understand the era of personal software and everything has to be a windows replacement or it wasn't worth building to you? I'm making my own azure blob storage explorer, notes mac/android, db client etc.. its not about being 'impressive' but being personally suited to individual needs.
reply
deleted
reply
He said most impressive and useful meaning there is no absolute cutoff. The only reason why you would be offended is because you have built nothing at all and that's you telling on yourself. Not to mention the question wasn't even aimed at you.
reply
No, that's not what happened in this thread at all.

There's a recurring pattern on HN where any time someone talks about knocking dozens of personal projects off their list - things that almost certainly would never have actually been addressed in the finite span of a normal life, the way things go - and you AI doomers show up and demand receipts as though that's a total reasonable and definitely not obnoxious request.

It's like if you tell someone that you love your partner and they demand to sit in the cuck chair or else you're obviously lying. I keep hoping people will move past this "prove that you're actually productive" reflex, but it just keeps happening in basically every AI thread.

In reality there are many reasons not to list out projects that you've worked on with LLMs, and while "none of your damn business" is always going to be at the top, the simple truth is that I want my products and projects to be judged by what they do and how well they work, not by how they were made.

reply
> "There's a recurring pattern on HN where any time someone talks about knocking dozens of personal projects off their list"

Nobody was talking about that very believable use case. What they specifically said was "enterprise scale projects". Enterprise scale means lots of users, decades of backwards compatibility, regulation compliance, logging and auditing, reporting, role based access control and permissions, integration with other enterprise systems, etc. etc.

> "you AI doomers show up and demand receipts as though that's a total reasonable and"

Asking what large scale programs they have built is not "dooming". It's also not demanding. It's also not unreasonable.

> "definitely not obnoxious request".

Saying that it's "unreasonable and obnoxious" to ask people to justify their claims is where the likes of Theranos and Nikola electric truck company are hiding. If they don't want to talk about their stuff, they could have not commented. Since they commented it's reasonable to ask them about what they said.

reply
> What they specifically said was "enterprise scale projects". Enterprise scale means lots of users, decades of backwards compatibility, regulation compliance, logging and auditing, reporting, role based access control and permissions, integration with other enterprise systems, etc. etc.

Do you really think that's what they meant?

reply
> I can do things it took teams years to build as a hobby project over a weekend.

That is a pretty strong statement. And without even anecdotal evidence, it becomes very weak.

reply
Nobody owes you evidence.
reply
You misread a comment so badly that you ended up responding to a completely different and made up comment. Your unwillingness to admit that you were wrong is way worse than whatever reflex you are ranting about.
reply
>I want my products and projects to be judged by what they do and how well they work

Hence the question "what's the most impressive and useful thing you have made"

But no one really has any examples. Just half baked slop they never got over the finish line.

reply
Ladies and gentlemen, the cuck chair is going to be occupied tonight!
reply
So software that already solves those problems isn't good enough and your vibe coded versions will be better? And totally worth the negative cost of AI to society and the environment? You just gotta have your own versions right?
reply
Well, maybe, for certain programs, yes? Maybe I don't want a million SLOC behemoth with 1000 features that has a huge attack surface, when a slim 10k SLOC tool that does exactly what I need (and nothing more) will suffice? A tool that I can modify and extend on a whim? A tool that does one thing, and does it well — you know, the Unix philosophy?
reply
>you can't think of anything to unleash some agents on within the entire digital world at any given time?

It has to be worth it though right? Like I could spend some money and have agents build me my own Photoshop maybe (maybe?) But it would definitely be much worse to use than actual Photoshop. Then I have to have the continued interest to keep improving it which probably won't happen because the next shiny thing will grab my attention. So it all just seems like a bunch of kids that have been given a seemingly endless supply of free candy and they are going fucking nuts like chipmunks with ADHD on crack. Building all this shit that is absolutely meaningless. I realize I've gone on a rant but I'll keep going. I strongly suspect (with no evidence whatsoever) that the people who are churning slop apps out at breakneck speed have never been to an art museum. There. I said it. You've all got no taste. You wouldn't know a quality product if it hit you in the face. I'll leave with this thought- if apple didn't exist, would they ever exist now we have LLMs? I say no, because the age of good taste and refined design and original thoughts is gone forever now that we have Claude and chatgpt and agents.

reply
I think you're underestimating how things used to be - you could go into any office, any closet, any coffee shop and find a shit-ton of half-baked, crazy-genius, kick-ass, retarded ideas and projects lying around everywhere in the world: filing systems, carpet organizing systems, outlines for film scripts, unsent letters to loved ones, etc etc etc. Whole worlds everywhere you look. And other people chipping in their two-cents worth, adding a few new filing cabinets, an idea for a film sequel, a new way to think about a different carpet, a notebook system to organize someone else's unsent letters... All this slop eventually thrown into nasty, fetid garbage dumps, forgotten.
reply
A.I. potentially breathing new life into every personal Graveyard of Half-Assed False Starts.

What's not to like ?

reply
All of that stuff meant something to someone though. They put real time and effort into those creations.
reply
> All this slop eventually thrown into nasty, fetid garbage dumps, forgotten.

Isn't that how it is supposed to be?

The lower the friction, the lower the signal:noise ratio.

It doesn't matter if 1 out of every 100k slop projects is actually a humdinger, how on earth will you ever find it?

The value of a project is the commitment to to it by people. Slop projects indicates a commitment in the low to none range.

So, yeah, that AI-booster who "created" (I use that word loosely) 7x Adobe replacements in a week (none of which actually work, but he'll get there eventually, I supposed) will successfully edge out the person who carefully and thoughtfully created a Photoshop replacement over six months of user feedback.

TBH, the only way to start a software business now is in stealth mode.

reply
[flagged]
reply
I tried fireworks.ai, drawn in by their supposed blazing speed. Yawn. Mostly worse than vanilla Deepseek.

Is Lithos actually fast for common usage?

reply
You have all the struggles for the price of Anthropic / cursor subscription. I use the first one I code large chunks some PR are 50k LOC and I have at least 2-3 like this a week . It’s a greenfield project .

I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project

reply
Why the hell would you put 50k lines into a single PR?

I mean, why even pretend you’re going to “review” something that large? Just build everything on main.

reply
PR is just an entity to review /do some other LLM processes . Human part check other models review PR's, tests all kind of , security etc. Also human is to checks docs, specs in the pr, db migrations if any , some of the tests related to the PR. We stopped reading the code after opus 4.6. Sometimes for very core parts i skim through files just to make sure if the changes were correct.
reply
I would imagine even LLMs would do a better job reviewing smaller PRs than very large ones.
reply
But to reiterate parent's question: why not just do those things continually at that point? Or on a calendar-based basis?
reply
Yeah exactly, if you have a fully automated SDLC, how do you expect the agent to code review 50k lines properly?

It will take shortcuts and now the entire premise is busted. You now need to build a code review process for large PRs.

reply
This is a common problem. Fully automated SDLC needs to start before the CI/CD.

My suggestion is to have the proper chunking mechanisms and multiple specialised agents. The most important is harness engineering, what we do at dromeas.ai to verify the code that goes to prod is a)have the code mapped before hand for the right agentic context, b)chunks of the right size per model context window c)specialised agents d)deduplication and verification . All before assessing a PR, a commit, a release. Harness engineering is not easy.. Especially when supporting multi model

reply
I agree. I code a lot, a lot! And maybe my code is shitty, but yeah, I burn a lot of tokens, and I couldn’t do it without Chinese models. I’m just a random dev in the middle of nowhere. And I don’t feel like I’m missing out on anything with my setup at all.
reply
OpenRouter is complete garbage.

Buy directly from DeepSeek's API.

You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).

reply
One of the reasons I use OpenRouter is because they offer zero data retention. As far as I can tell, DeepSeek's own API doesn't support ZDR.
reply
AFAIK, when you use DeepSeek via OpenRouter, it still does not have zero data retention.

See here:

https://openrouter.ai/providers/

reply
I have ZDR enforced and see only compatible models and providers, yet am able to use it. DeepSeek as a provider may not be ZDR, but the models are available from ZDR and no training providers on EU/US servers.
reply
DeepInfra does and it's the same price. That's what I use.
reply
DeepInfra is decent but has a bad history of quantizing models
reply
Yeah, they all do it though. And now that models are being natively trained at FP8/NVFP4, I’m not sure it matters.

Only way you can really know you’re getting the full model is to host it yourself. Every inference provider has every reason to lie and it’s impossible to find out the degree to which they are.

reply
Youre doing something so special you need that?
reply
It’s a pretty common requirement in the enterprise world. If you’re processing data for enterprise customers, it’s a lot easier to retain nothing than to deal with all the compliance issues that arise if you’re retaining data.
reply
Ugh yeah try to get through a DPIA review!
reply
I've got a toilet cam to install in your bathroom
reply
I'm all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:

OpenRouter Pricing:

$0.02/M input tokens $0.60/M output tokens

DeepSeek Pricing (cache miss, off-peak):

$0.15/M Input $0.60/m output

reply
When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.

It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?

reply
Interesting, if the cache hit is that good, I think HN convinced me to toss $20 at DS official, and see how long that lasts.
reply
Read the tech report and see how laser-focused they've been on compressing the disk KV footprint specifically. 890 bytes/token is absurd, and is probably the reason I regularly get ~0.5 Mtok request full cache hits after an hour. On their end that's just yanking a 414 MiB file off an SSD, then doing a little bit of compute (bounded SWA replay). Nobody else seems to be serving their model as well as they do.
reply
It will of course depend on what you’re doing with it, but right now my session at work has a 99.8% cache hit rate, and I’ve been running this session for hours with 23M tokens read and 713K tokens written (Opus 5.5 in this case though)
reply
cache hit is in fact that good.
reply
I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.
reply
There’s a big difference in speed & quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.
reply
what about dsh + openrouter? and configuring open router to just serve from DeepSeek own servers

I think the 5% cut from open router is fair if I want to user other cheap models like mimo

reply
DeepSeek trains on your inputs. That's why people go on OpenRouter and choose ZDR providers.
reply
This isn't true at least on the API. If you read their privacy policy you'll see the training clause is scoped specifically to the consumer terms i.e. for the chat product. No such clause exists for the API service, and it would absolutely be required under Chinese law if it was taking place.

Contrary to popular belief, DeepSeek really aren't interested in your prompts.

reply
Do you have a link to their API privacy policy?

What I see is https://cdn.deepseek.com/policies/en-US/deepseek-privacy-pol...

They do not have a specific exclusion for API use.

I know Z.ai has an exclusion for API use. It's widely reported Deepseek doesn't.

reply
I'll preface by saying you'd be entirely reasonable in not finding this sufficiently reassuring, but compare the standard terms: https://cdn.deepseek.com/policies/en-US/deepseek-terms-of-us...

To the open platform terms (i.e. for API use): https://cdn.deepseek.com/policies/en-US/deepseek-open-platfo...

The standard terms includes clause 4.3 which grants them the right to retain inputs and outputs for training purposes, and this is missing from their API terms. The standard terms also cover the right to opt out (which you can do from your user settings). No such opt-out exists on the API because it isn't applicable.

reply
You're linking to the Terms of Use. I'm linking to the Privacy Policy. Your Terms of Use link points to the Privacy Policy.
reply
Deepseek does not offer zero training on OpenRouter. Out of dozens of alternatives, they are the only provider for 4.1 Flash that does this.
reply
Let me get this straight you guys really like deep seek because it’s open but you don’t wanna help them improve.
reply
No, we just want a choice on how to license our work.
reply
All these products, western or Chinese, are built on a hell of a lot of running rough-shod over licensing or IP laws in general.

I'm writing a program I personally need, but I would be happy if there existed something like it already, if someone else vibecoded a better version of it than mine, or if DS got better at vibing this kind of thing.

reply
>license our work

what proof do you have, that they don't train on your data?

they can say they don't, but I don't see any way for you to confirm it.

with how these companies operate currently, I won't be surprised, if they say that one of agents "mistakenly" did that already..

reply
Mistakenly and autonomously of course.
reply
Why does liking a product mean you have to give them all of your data? People are so outraged at LG because they make the best TVs and people wanted their expensive product, yet some MBA convinced them they could make more money by spying on your entire household all the time.

It actually was awesome in the early Facebook days where you could have your entire phone contacts and other apps filled out with a profile picture and Birthday by connecting them together. But that relationship has been completely abused, privacy has been invaded, and my data has been sold to multiple companies.

The goal going forward is to keep that data private. If your company can't survive without it then I hope your company goes out of business

reply
Just pin your config to a single provider, or several providers with the params `order` and `allow_fallbacks: false`. I regularly get ~98-99% cache hit rates with OpenCode. And some providers are much faster than DeepSeek; I was getting 200-300 tokens/second the other day with Together as my provider.

It's regrettable that OpenRouter doesn't even try to pin you to a single provider per session, but once you know about it, it's a problem that's easily solved.

reply
Don’t you get cache expirations then if they change providers? This ought to introduce delays and costs
reply
Yes, but that's why you pin them. If you specify more than one provider in `order`, OpenRouter will use the first one unless it's down, so that's the only time it would switch providers on you. And personally, I'd pay a few cents instead of waiting for the API to come back.
reply
Any proof of 100x spend? That seems a bit excessive
reply
There's no difference between the two if you pin the provider to Deepseek on Openrouter.

If you don't want to mess about client side with pinning, set a guardrail on Openrouter that limits the available providers to only the official one.

reply
Then why bother using OpenRouter and paying the extra fees?
reply
Because you have like every model on the planet to choose from. So if a contender drops you switch.

Also if DS is down you can choose another provider.

I have credits at DS and OR directly. But I do see the value in OR.

reply
So you're basically send your code to China?

I thought the advantage of DeepSeek is that you can host it on a server of your choosing.

reply
Yes, I send it to China, and they use it to improve open-weight models. I'm ok with this arrangement. At least, I'm happier with this than with companies using my open-source work to improve proprietary models without my consent.
reply
Or.. BYOK Deepseek because OpenRouter's UX is much nicer?
reply
Zero Data Retention and not having company source code leak to "CHINA!" (said in Trumps annoying voice) would be two reasons not to
reply
It's far more likely they start to restrict the usage of subscriptions in corporate settings and leave the individual subscriptions to inflate the margin against them away. The individual subscriptions are the hook; people need to be able to understand what they're capable of with enough tokens to tackle real projects they couldn't have without the sub. That's the only way individuals will make the attempt to sell it to their org.

We'll probably still be getting the same amount of work done with a $200 subscription a year from now. That will just represent a much smaller subsidization than we currently enjoy; something like 3:1 - 6:1 instead of 40:1. Maybe running at a lower tps than the API gets. Maybe no access to the absolute frontier, but still significantly more intelligent than we get now. The labs will essentially break even on the subs and the corporate spending will be the profit center. Tale as old as software.

reply
Corporate customers are by and large not allowed to use subscriptions.
reply
I freak out since months for Z.ai lite subscription, I use glm-5.3-flash every day for a ludicrous 8.5USD/month and it's as good as DS 4.1 flash, if not better.
reply
Almost exact same experience here, but I'm using Opencode's $10/month sub. It's perma set to DS 4.1 flash and I have anywhere from 3-5 agents going at a time. Never once hit a cap of any sort. I have absolutely no idea why people would be paying $200/mo when you can get perfectly good AI for $10 from multiple places
reply
Do they train on your data? I don't necessarily want competitors to be able to just ask it to make a clone and have it do so from memory next month.
reply
Software has rarely been the moat. Or file formats. You have always been able to reverse engineer them. The problem, always, has been network effects.

I can build an entire, fairly useful, spreadsheet app over a weekend. But can I send my "expenses.cells" files to my accountant? Will it work with the Excel/Google docs he uses?

AI can build or reverse engineer anything as long as you are motivated enough to do it.

reply
One example is Affinity 3 released for free. But it's a huge pain in the ass because all the guides for how to do things are for Photoshop or Affinity 2.

Your vibecoded app won't have years of reddit posts showing how to do things. This also seems to be where LLMs are the weakest at giving advice, they hallucinate 80% of the time I ask them how to do something in Affinity, giving buttons and menus that simply don't exist.

reply
This is a knowledge problem. You can fix it by pointing the model to documentation (if it exists). Otherwise the model will give you the next best guess
reply
Vibecoded apps will have to include MCP servers so the LLM agents can interact with them directly.
reply
Huh? Both Anthropic and OpenAI have computer use tools. I frequently just tell the models to test their own work. I'm a little paranoid so I only grant access to one window and make that window VMware Workstation.

You combine it with /goal. I usually set a goal like "Complete the application defined in goal.md as written. Then test it end to end autonomously using Compter Use. Record all issues discovered during testing in a to-do. Then fix the issues in the to-do. Repeat testing until no more issues are discovered."

reply
This is where some of the readers start thinking about putting some agents to work on making Photoslop.
reply
Does that matter when it can clone it today without the need to train on your specific code?
reply
It isn't about the code. Some of my work is genuinely new science that might offer an AI company a competitive advantage.
reply
I don't see why that is a concern when decomps and recomps are already blooming, they can copy your app down to the atom.
reply
How's the caching? I have 99.5% cache hit rate with deepseek when using their own API, it's dirt cheap.
reply
People also just do different work. Opus 5.5 is a really damn good model that's even better than Astra/Fable/Sol IME and I feel a huge difference in my work.
reply
Yup. I was going to scale down my Anthropic sub when Opus 5 was... weird, and there wasn't enough Fable provided. 5.5, otoh, chef's kiss

That said, I do want good local(ish) capability for if/when Anthropic enshittifies again. And to play with very useful smaller models - don't even count gemma 4 out.

reply
Same, it's like another Opus 4.5 moment.
reply
Yep... and we already have Opus 4.6 at home (Qwen Flash Next 3.8).
reply
This. At this point I don't really care about other models because max subscription are super cheap (relatively speaking) and I don't hit my limits. Even if the frontier models are only 5% better I might as well just use the best thing available if the price is reasonable.

Once the subsidization ends and cost becomes significant I will take a serious look around for the best value models and switch off the expensive providers, but that time hasn't come yet.

reply
I have a subscription at work, and still I find myself wishing I could use a fast Chinese model. Something wired up to really fast inference - that rapidity of feedback is a feature in itself.

4.1 Flash seems to be in that sweet spot of very decent, really fast and really cheap. Even omitting the cost, it’s still compelling for staying in flow.

reply
I'm actually shocked by how many people seem to have max subscriptions.

Something about renting that much compute doesn't sit right with me so I stick with the $20 subs.

reply
> Once the subsidization ends and cost becomes significant I will take a serious look around for the best value models and switch off the expensive providers, but that time hasn't come yet.

There's a reason the labs in the US frontier oligopoly are using “safety” to lobby for antitrust exemptions for mutual coordination as well as anticompetitive regulation.

reply
I got insane amounts of Anthropic and OpenAI credits given to me for free for my startup, and I have not touched them.

I get privacy, freedom, and no rate limits with the GPUs I racked locally, and those are features I would never give up even if the surveillance capitalism labs paid -me- to use their models.

How many consumers are there like me? Probably not many, but once local inference hardware is plug and play, I bet the tides shift pretty quick. Also weights-on-silicon will serve the needs of most consumers locally with more speed than any GPU could deliver for a fraction of the cost.

Most people will be doing inference in their pocket or a wearable in 5 years and the giant datacenters will be like AWS, sold to only big organizations that need to auto-scale capacity of custom models on demand.

The industry surely knows this and the subsidized inference is just marketing to generate so much buzz and demand such that the tiny fraction of the market they will be able to keep in the end is big enough that they do not collapse under all the debt.

OpenAI and Anthropic will be Dell and IBM in 10 years if they survive at all.

reply
Dell is trading at 4000% what it was 10 years ago
reply
I did not imply otherwise. Just that both would be likely irrelevant to most consumers.
reply
Frontier labs subsidising? I thought they were running with 80%-ish operating margins which is part of why they cost 10 times open weight models.

Also isn't an open weight model also subsidised? Training isn't cheap and you are not paying for it.

reply
Claude financials leaked last week, they are running with about -200% operating margins.

We also have OpenAI numbers, where people speculate that the about 150% of the margins spent on "marketing" is a fake line used to hide operational costs.

reply
But isn’t that including the cost of training, and I am not even sure if it’s training that very model?

In an interview earlier this year I remember Dario saying that the models are profitable, ie they more than pay back their inference and training over time. But because they invest in that explosive growth they have a deep negative cash burn.

And whatever markup the AI labs make, that’s on top of nvidia’s markup, micron’s markup, etc. It’s an industry where every supplier is adding a 70-80% markup!

reply
AFAIK, no. For Claude, many people have reported that it's not including the training costs. But I didn't personally look at the data.

For OpenAI, absolutely not, that's not including training costs. But I do personally accept that it can be actual marketing costs.

reply
But aren't you developing bad habits and learning patterns that won't work long term? Or do you think things will get cheap enough that you will be able to keep going with your current patterns post-subsidies?
reply
No one involved in this is thinking about the long term
reply
No one is thinking, the AI does that for them.
reply
AI does the work, it does not replace the thinking. It's like saying a tech lead in a project does not do any thinking because there's some junior dev doing tedious and boring tasks.
reply
> But aren't you developing bad habits and learning patterns that won't work long term?

2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.

reply
hooked to ai coding, but not tied to any particular model. if they decide to bump prices up, I can easily switch to a cheaper chinese model on openrouter.
reply
Yes network effects are significantly less than something like Uber, in fact they’re almost nonexistent
reply
This is why they enforce the use of client apps like Claude Code/Codex (although OpenAI is a little more lenient). Trying to create a network effect.
reply
Yeah it makes sense, but ultimately the only thing that creates any form of lock in is the chat history and memories, and that isn’t super important, it’s not a real network effect like a social media app or a taxi app.

Having a better model is the only real moat, without that inference is a commodity

reply
There isn't lock in and this is what scares them. Switching cost is SO LOW. Our team runs the 3 main models to cross check work without any issues.

The benefits are real and our willingness to pay is real, but the valuations only support one and right now even the free cheap models might win. Thusly the collapse could still happen.

reply
I think it will happen, like you say switching is easy and the only possible moat is to build a better model than your competitors at enormous expense.

They’re all stuck in a cycle of spending huge amounts of money on training just to stand still (in business terms).

In the long run, it can’t continue because it doesn’t make any sense

reply
Compared to what a lot of companies spend on software for chip and electronics design (we're talking about $10k-200k/seat per year), AI coding assistants have a long way to go in cost before companies won't be willing to pay for them. Companies pay a fortune for software when it enables their engineers to be productive.

For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.

If they stop subsidising Claude Code for the pro/max users, there will be a lot of people priced out of it, especially the casual developer. But I don't see it going away for commercial use, even with a large price increase.

reply
Nah dude I think it’s worse than that

Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness

Show me people handwriting code à la NASA

And I mean we as coders have been trying to do this workflow for a while, I personally would refuse to code without IntelliJ magic complete

For this workflow, there’s no going back. What’s hard to imagine is AI taking over the other workflows we predict it will; Customer service AI sucks ass for me as a customer, et cetera

reply
I don’t mind bad customer service AI any more, since it’s my claude interacting with it, not me.
reply
> For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.

That enterprise cost you're willing to pay is correlated to how much developers will work for. When driving an agent, almost anyone can do it (almost no skills required).

If devs cost $1k/m, enterprises are not going to be willing to pay $4k/m for Claude.

What I am saying is, there's an equilibrium that will be reached; the price of the human driver and the AI worker will approach each other.

Where they stabilise, I still don't know, but I'd be very surprised if, in any field (not just dev), the human gets paid multiples more than the agent they are driving, as the agents get more capable.

reply
You also have a cost of team interaction going with square if team size or sth.
reply
I use the frontier openai/anthropic models at work but exclusively open weight models (on cloud/hosted inference) for personal stuff and I think about it like this; 1) I don't see any reason GLM and DeepSeek won't eventually be as good as Claude, it's just a matter of time and 2) the open models are well and truly capable enough for most of what I'd want to do. I don't need nor want an LLM chewing away on a horrible enterprise spaghetti codebase, my employers can pay for that privilege.
reply
Long term, we will see what happens and adapt. At worst we all go back coding by hand. Meanwhile what can I do, tell my customers that I'm raising my fee because I have to pay for token? The Claude Pro $20 plan is good enough for me and even in auto mode I never had to wait for the 5 hours reset.
reply
I expect by that point we'll have local models that can do a decent job, I would guess give it a decade and we'll be running custom accelerators that are smarter than current frontier models.

In the same way that only supercomputers used to have multiple processors and caches but it's now standard.

reply
The base costs fall down 2-5x year by year, and they will continue falling due to infra ramp up, asics and distilled models.
reply
From the post:

“With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited.”

reply
I put £15.73 (from a dollar exchange conversion of $21.20) at the beginning of September just to try DeepSeek out via API using Opencode and Pi. I still have plenty of credit left! Their off-peak reduced token cost truly is amazing.
reply
> This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.

It's amazing how new we all perceive AI to be, and yet how old the tricks that the big players use. Their job is to just suck the oxygen out of the room as long as they have the money to do it.

reply
Yes, but they're all offering what appears to be a commoditized product with race-to-the-bottom economics.

You can't jack up the prices on your product if your competitors can just clone it and resell its essence for pennies on the dollar.

reply
It's glorious isn't it. We get free work done through subsidies. At the same time, this is what threatens my job, and the money for the subsidy is basically my own invested pensions.
reply
> Their job is to just suck the oxygen out of the room

And that this is even legal is a scandal all on its own.

reply
Hard to enforce when you’re dealing with private companies with “creative” accounting - who’s to say what the actual cost of inference is for OpenAI or Anthropic?

Possibly they don’t even really know themselves at this point, although obviously is it significantly higher than the consumer subscription price

reply
Agreed.

I've been running automated research tasks for life sciences companies, and the speed in which tokens are burnt is scary. Especially when you get into a complex knowledge space and require a subwgent to reason through each possibility, token usage grows quadratically not linearly as complexity increases...

reply
Worth checking out the methodology this guy put together using Jev and Haiku at appropriate points to cut the cost and also speed up a deep research job: https://news.ycombinator.com/item?id=50019056
reply
What the provider actually sells to you is GPU time / load (oversimplifying). Both tokens and subscriptions are just pretty arbitrary ways to price it, almost unrelated to the actual cost of running the model.
reply
Luna and sometimes even Sol are cheaper than Deepseek for same tasks even at API prices since they consume way less tokens and are more efficient
reply
> This won’t last forever

It probably will. Moore's Law is still churning away in the background.

Frontier models might get more expensive, but that's a moving target. For any particular capability point, the models will only get cheaper.

reply
> The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.

My understanding is that enterprise plans don't offer those subscriptions, so they end up paying for API prices and models like these directly impact that revenue stream.

reply
I think it's fairly likely medium to large corporations don't pay anywhere close to the listed API pricings as they get deals through existing partnerships with the big cloud providers.
reply
Are you sure you pinned the provider? OpenRouter is known for cycling you between providers which busts the cache and you end up paying way more.
reply
selling unused capacity behind high cost demand pricing of api access isnt subsidised in-fact the profit margin anthropic makes on subscriptions is close to 100%. people are looking at this economic model completely backwards
reply
I use $20 codex subscription, it's practically useless for anything other than luna. Deepseek v4.1 flash goes a LONG way for $20.
reply
Deepseek is in some ways subsidized as well since they use data for training.
reply
That's an additional cost to the user rather than infusing more resources from the provider like a subsidy.
reply
Which provider was that?
reply
opencode DeepSeek v4.1 Flash isn’t us/eu hosted as of recently, so not sure how this impacts privacy / model training
reply
I am using it on opencode go
reply
how was it? did you hit limit often? Im considering opencode go for second sub.
reply
MS's Copilot subsidy ended months back. The large OpenAI subsidy has just been cut in half. Who knows when Anthropic ended theirs because I don't ever remember it being great value.
reply
I mean, it is freaking out ? Or it was, but then started freaking about Opus 5.5 . But days are measured in dog days in AI.
reply
You should basically never pay API prices, they are always several times higher than subscriptions.

There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.

reply
[flagged]
reply
[flagged]
reply
Also DeepSeek usage is subsidized as well, it’s a power hungry model.
reply
Interesting. Are all of the providers on OpenRouter simply losing money? How does that even work out?
reply
No, it’s outrageously profitable above x% utilization without stealing any prompts. Provider economics still pretty good. Acquiring hardware is the current limiter.
reply
Don't ask and dance as long as the music keeps playing.
reply
You're the RLHF.
reply
With ZDR-only enabled? That seems illegal.
reply
Subscriptions are not ZDR.
reply
And OpenRouter isn't a subscription.
reply
deleted
reply
It's adorable you think the AI labs care about that.
reply
With all due respect (not that I feel like much is due after that response) I am not really sure you know what I am talking about. I'm talking about inference providers, like inference.net, Fireworks, Coreweave, Digital Ocean, etc, to use DeepSeek 4.1 Flash. They didn't create the model, they just are charging you to run inference tasks. That is a different story.

DeepSeek themselves are honest about the fact that they train on inputs by default. You won't hit DeepSeek if you use OpenRouter with ZDR enabled.

reply
[flagged]
reply
I now have two people with sarcastic replies that seem to think I'm talking about OpenAI. We're talking about DeepSeek on third party providers. Given that this thread is neither that long nor that confusing, I'm not sure how people are getting lost this quickly.

I don't really trust OpenAI, although honestly I find it stupid to suggest they'd offer a ZDR policy and violate it. They literally don't have to offer it. People will still pay. Fable doesn't offer ZDR at all and it hasn't stopped people from paying through the nose for it at API pricing.

reply
Not if you're not giving feedback.
reply
Are you sure about that? My impression was most providers on openrouter were purely selling tokens for profit...
reply
Have y'all tried an Ollama Cloud subscription? Their off-hours pricing for V4.1 Flash is extremely competitive.
reply
Yep. And the model is practically unbounded in its knowledge of math. So people should keep that in mind when they read anything about its level of intelligence.
reply