I do think its important long term to not be reliant on these companies as you don't have control over the system prompts, the thinking tokens, and once the subsidization stops or the company is public they will be required to start making money and thus raise prices.
But models may get more intelligent and cheaper once that time comes so it may be a non issue.
The system prompt is injected into your context on the server side.
here is a collection of public system prompts from claude code: https://github.com/Piebald-AI/claude-code-system-prompts
...the absolute state of the token maximizers.
I'm assuming this is a sarcastic response to the person who said "640K [RAM] ought to be enough for anybody!"
Because, as somebody who was around when we had 640 K RAM, people certainly weren't happy with that amount.
Although as soon as I wrote 'foreseeable future' I came to the realization that this is far far far less far out than it used to be. Which might mean I simply agree with you.
Checked yesterday, for that day alone I had used $168 worth on my $20 subscription in Claude Code. I still had plenty of weekly use left. Seems like subscriptions are discounted at a 1:10 rate?
EDIT: That was a delightfully thorough analysis. Nice to see Anthropic taking the value crown, only because Opus 5.5 is such a joy to use.
Because a big part of the price difference is due to Anthropic's models being more expensive that OpenAIs AND using more tokens for the same tasks (at least according to artificialanalysis.ai).
I use a personal Claude account for personal projects and can let rabl run for an hour and barely make a dent into my usage
On the enterprise I have to be a lot more careful or I can burn through 2k in a week
The excuse they give is the guarantees you get with enterprise plans that they won’t look at your data
This is why there are so many comments confused that you spent that much money on Deepseek. Those who got on one of the discounted plans could use very large numbers of tokens for trivial prices.
There is also a strange double standard for accounting for open models. People will look at OpenAI or Anthropic and say that we need to consider all of their training costs and employee compensation when thinking about the cost to serve their models, but when the models are released as open weights those costs are ignored. So in that way, the Deepseek models are heavily subsidized as well, with the possibility of them being served by a different company that paid nothing to develop them.
> Quality was ok, seems slightly above Luna quality perhaps?
I agree that it’s about in line with what you can get from Luna or Haiku, but I give the edge to Luna and Haiku when it comes to tasks that require world knowledge. It feels like their training sets were just cleaner.
Luna and Haiku are also close to free with a subscription plan, and they’re even very cheap at API rates.
So the amazing thing about Deepseek Flash is that you can almost kind of get that level of performance from an open weight model. It’s not as amazing when you start comparing it for how we really use smaller models from frontier labs on subscription plans.
You could then of course argument that oAI/ant/G's models were also heavily subsidized by using (scraping) the bulk of humanity's global knowledge for free while paying nothing for it and trying to privatize it.
That accounting doesn't make sense when looking at future models, but if you're evaluating the current state it makes sense to ignore training costs for companies that don't pay them.
You can also get subscriptions for the open models, which are typically much cheaper. It's not crazy popular because the open model crowd switches often, but they do exist if you want a firehose of tokens.
> Deepseek also came with heavily subsidized plans, at least at first.
Do you have a source for that?I got a source straight from the hoses mouth, which claims that DeepSeek-R1 was highly profitable:
> If all tokens were billed at DeepSeek-R1’s pricing (*), the total daily revenue would be $562,027, with a cost profit margin of 545%.
https://github.com/deepseek-ai/open-infra-index/blob/main/20...With DeepSeek-V4.1-Flash, the memory requirements for the KV cache have been reduced by over 50x and FLOPS by over 5x, so it is even cheaper: https://arxiv.org/pdf/2609.19969
Your link is about a different model.
The talk about subsidies is confusing because for OpenAI and Anthropic people usually include their salaries, training costs, and everything else. When the topic switches to open weight models we pretend the models appeared out of the ether at zero cost, and the only cost is running the servers.
Has anyone seen the neo-cloud profit margin on deepseek? I wonder if the deepseek served API prices account for the training cost? Because they don't / have not raised the money to fund their future operation / training- and they depend on that cashflow for now? That would suggest a really high profit margin on inference only.
OpenCode is building their own inference and they've independently stated that DeepSeek's old (cheaper) pricing is achievable without subsidies.
And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.
I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.
People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.
Those usage patterns don't correlate to output.
We ended up having to hire a full time employee to fix the performance of client built reports.
People run a stupid amount of expensive queries that end up costing way too much because they're asking Claude the wrong query.
Not to mention people running wrong queries, using the result as gospel, and then the result has to be sent to a data analyst to be reverse-engineered so the numbers make sense.
Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.
As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.
But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...
It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.
Which is an issue when you need to get department managers to manage their budgets around the amounts of credits their employees spend. The more of a black box it is, the more governance and corporate bullshit you have to deal with.
The copilot part of it runs "unlimited" on the license. Except it's not unlimited, and this is even more of a blackbox because you can't see any sort of spending and the limit is listed as "extensive use".
On your point about "monitoring", personally I feel this is toxic corporate IT culture, enabled and perhaps pushed by the likes of Microsoft with all their tools, which they of course make money off. People have cellphones with cameras making most points in this area moot.
An alternative used in other big corporates is to set budgets, with tiered authorisation approvals for higher limits. The users and their managers can justify why and what they're doing that they need the additional tokens. This also encourages more efficient use of tokens on other work. More efficient use is sometimes counterintuitive. Laissez-faire generally works best.
Cowork requires user approvals for high risk actions such as emailing.
The better pattern is to let it code the app and then you can use the app to target your data. So you only pay for it once, plus it's deterministic. But yeah, it requires setting up an environment, etc. It becomes "maintenance".
I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.
Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.
Eg I've used Sashiko locally for Linux kernel code reviews before sending out my contributions out to the world. Sashiko is a great system, but it can burn through tokens like there's no tomorrow.
https://github.com/sashiko-dev/sashiko and https://sashiko.dev/
I wish I could observe how some of us are using these tools.
I still struggle to spend $50 in tokens per month, and I exclusively use prepaid API tokens. This is in support of personal projects and two clients. There are billing cycles where I might spend upward of $400, but this is maybe once a year. This is offset by months like August wherein I spent $12 in tokens.
The other advantage with prepaid is that it handles the other direction much better. I don't even know what a quota limit feels like. Being blocked for hours is way more expensive to me and my clients than even $500/m. Losing an entire business day over this wouldn't work out.
I've managed to convince some others to try the same thing. $200/m flat fee is a pretty extreme constant expense if you can be more clever on average.
I think a lot of people are getting pushed around by FOMO effects into spending money on pointless subsidized tokens and have (valid) fears that if they don't maintain the same apparent economic leverage as their peers that they will be left behind. This isn't actually the case, much like lines of code are a really poor indicator for the quality or productivity over a codebase.
Each of these has a file it listens to in ~/tmp/<name>.io and whenever one needs something from the other, they message each other via that file. Tasks can bounce back and forth as issues are resolved and tested. At the same time, I keep each busy with a list of tasks
I can fairly easily run out of my $200/month subs every week, if I let Fable be the default model. With Opus it’s less likely. If and when I do, I just have an alias ‘claude.ds’ which fires up Deepseek instead, and burns through far less money, though I don’t think it’s as good at solving problems, just MHO.
Why isn't this just one coherent agent loop with subtools/agents as appropriate? If these tasks are related in some way, having a single context would probably make it go much better.
The freewheeling messaging part is where the token bloat is coming from. I suspect that for some of us this is actually the point. I think it's a mostly form of entertainment to do things this way. The next logical step from Factorio gameplay.
Parallel agents remind me a lot about multi core compute. It's incredibly easy to take a single core product and make it run much worse across a lot of cores.
I also find that it's easy to spend like that, but also easy not to with little impact on productivity. At the current moment I'm stuck with rider + copilot (not ideal), but e.g. using GPT 6.1 luna is really, really cheap, and lots of tasks are quickly and decently dealt with even at lower reasoning levels, (added bonus of having low latency). And that model is so cheap, I can't see a hitting 1500$ at api prices realistically - not even close. But it also depends on the harness and codebase.
I don't have the mental capacity to do a lot of context switching between active work streams, so I'm not doing stuff like leaving a big agent workflow running while doing other things.
All through claude code.
I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.
If your tasks are write-heavy, find a provider with cheaper output.
If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.
It’s PERFECT!
Lots of guys buying articles on TechCrunch saying they’ll build this, he’s bootstrapped
you can't think of anything to unleash some agents on within the entire digital world at any given time?
you motivate your own personal work only via gauging its' usefulness to others and your own prospects?
sheesh.
built anything for the sake of building yet?
God I hope so.
And even then, "enterprise" was often a dirty word in these circles. The over-engineered would-be-swiss-army-knife vendor that was mediocre-for-everyone but excellent for nobody.
I have built many tools in the recent past for myself. None of them need to be "enterprise scale." Most of them would be worse for it because the agent output suffers when the pile gets deeper and it's just adding more piles on top.
There's a recurring pattern on HN where any time someone talks about knocking dozens of personal projects off their list - things that almost certainly would never have actually been addressed in the finite span of a normal life, the way things go - and you AI doomers show up and demand receipts as though that's a total reasonable and definitely not obnoxious request.
It's like if you tell someone that you love your partner and they demand to sit in the cuck chair or else you're obviously lying. I keep hoping people will move past this "prove that you're actually productive" reflex, but it just keeps happening in basically every AI thread.
In reality there are many reasons not to list out projects that you've worked on with LLMs, and while "none of your damn business" is always going to be at the top, the simple truth is that I want my products and projects to be judged by what they do and how well they work, not by how they were made.
Nobody was talking about that very believable use case. What they specifically said was "enterprise scale projects". Enterprise scale means lots of users, decades of backwards compatibility, regulation compliance, logging and auditing, reporting, role based access control and permissions, integration with other enterprise systems, etc. etc.
> "you AI doomers show up and demand receipts as though that's a total reasonable and"
Asking what large scale programs they have built is not "dooming". It's also not demanding. It's also not unreasonable.
> "definitely not obnoxious request".
Saying that it's "unreasonable and obnoxious" to ask people to justify their claims is where the likes of Theranos and Nikola electric truck company are hiding. If they don't want to talk about their stuff, they could have not commented. Since they commented it's reasonable to ask them about what they said.
Do you really think that's what they meant?
That is a pretty strong statement. And without even anecdotal evidence, it becomes very weak.
Hence the question "what's the most impressive and useful thing you have made"
But no one really has any examples. Just half baked slop they never got over the finish line.
It has to be worth it though right? Like I could spend some money and have agents build me my own Photoshop maybe (maybe?) But it would definitely be much worse to use than actual Photoshop. Then I have to have the continued interest to keep improving it which probably won't happen because the next shiny thing will grab my attention. So it all just seems like a bunch of kids that have been given a seemingly endless supply of free candy and they are going fucking nuts like chipmunks with ADHD on crack. Building all this shit that is absolutely meaningless. I realize I've gone on a rant but I'll keep going. I strongly suspect (with no evidence whatsoever) that the people who are churning slop apps out at breakneck speed have never been to an art museum. There. I said it. You've all got no taste. You wouldn't know a quality product if it hit you in the face. I'll leave with this thought- if apple didn't exist, would they ever exist now we have LLMs? I say no, because the age of good taste and refined design and original thoughts is gone forever now that we have Claude and chatgpt and agents.
What's not to like ?
Isn't that how it is supposed to be?
The lower the friction, the lower the signal:noise ratio.
It doesn't matter if 1 out of every 100k slop projects is actually a humdinger, how on earth will you ever find it?
The value of a project is the commitment to to it by people. Slop projects indicates a commitment in the low to none range.
So, yeah, that AI-booster who "created" (I use that word loosely) 7x Adobe replacements in a week (none of which actually work, but he'll get there eventually, I supposed) will successfully edge out the person who carefully and thoughtfully created a Photoshop replacement over six months of user feedback.
TBH, the only way to start a software business now is in stealth mode.
Is Lithos actually fast for common usage?
I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project
I mean, why even pretend you’re going to “review” something that large? Just build everything on main.
It will take shortcuts and now the entire premise is busted. You now need to build a code review process for large PRs.
My suggestion is to have the proper chunking mechanisms and multiple specialised agents. The most important is harness engineering, what we do at dromeas.ai to verify the code that goes to prod is a)have the code mapped before hand for the right agentic context, b)chunks of the right size per model context window c)specialised agents d)deduplication and verification . All before assessing a PR, a commit, a release. Harness engineering is not easy.. Especially when supporting multi model
Buy directly from DeepSeek's API.
You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).
See here:
Only way you can really know you’re getting the full model is to host it yourself. Every inference provider has every reason to lie and it’s impossible to find out the degree to which they are.
OpenRouter Pricing:
$0.02/M input tokens $0.60/M output tokens
DeepSeek Pricing (cache miss, off-peak):
$0.15/M Input $0.60/m output
It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?
I think the 5% cut from open router is fair if I want to user other cheap models like mimo
Contrary to popular belief, DeepSeek really aren't interested in your prompts.
What I see is https://cdn.deepseek.com/policies/en-US/deepseek-privacy-pol...
They do not have a specific exclusion for API use.
I know Z.ai has an exclusion for API use. It's widely reported Deepseek doesn't.
To the open platform terms (i.e. for API use): https://cdn.deepseek.com/policies/en-US/deepseek-open-platfo...
The standard terms includes clause 4.3 which grants them the right to retain inputs and outputs for training purposes, and this is missing from their API terms. The standard terms also cover the right to opt out (which you can do from your user settings). No such opt-out exists on the API because it isn't applicable.
I'm writing a program I personally need, but I would be happy if there existed something like it already, if someone else vibecoded a better version of it than mine, or if DS got better at vibing this kind of thing.
what proof do you have, that they don't train on your data?
they can say they don't, but I don't see any way for you to confirm it.
with how these companies operate currently, I won't be surprised, if they say that one of agents "mistakenly" did that already..
It actually was awesome in the early Facebook days where you could have your entire phone contacts and other apps filled out with a profile picture and Birthday by connecting them together. But that relationship has been completely abused, privacy has been invaded, and my data has been sold to multiple companies.
The goal going forward is to keep that data private. If your company can't survive without it then I hope your company goes out of business
It's regrettable that OpenRouter doesn't even try to pin you to a single provider per session, but once you know about it, it's a problem that's easily solved.
If you don't want to mess about client side with pinning, set a guardrail on Openrouter that limits the available providers to only the official one.
Also if DS is down you can choose another provider.
I have credits at DS and OR directly. But I do see the value in OR.
I thought the advantage of DeepSeek is that you can host it on a server of your choosing.
We'll probably still be getting the same amount of work done with a $200 subscription a year from now. That will just represent a much smaller subsidization than we currently enjoy; something like 3:1 - 6:1 instead of 40:1. Maybe running at a lower tps than the API gets. Maybe no access to the absolute frontier, but still significantly more intelligent than we get now. The labs will essentially break even on the subs and the corporate spending will be the profit center. Tale as old as software.
I can build an entire, fairly useful, spreadsheet app over a weekend. But can I send my "expenses.cells" files to my accountant? Will it work with the Excel/Google docs he uses?
AI can build or reverse engineer anything as long as you are motivated enough to do it.
Your vibecoded app won't have years of reddit posts showing how to do things. This also seems to be where LLMs are the weakest at giving advice, they hallucinate 80% of the time I ask them how to do something in Affinity, giving buttons and menus that simply don't exist.
You combine it with /goal. I usually set a goal like "Complete the application defined in goal.md as written. Then test it end to end autonomously using Compter Use. Record all issues discovered during testing in a to-do. Then fix the issues in the to-do. Repeat testing until no more issues are discovered."
That said, I do want good local(ish) capability for if/when Anthropic enshittifies again. And to play with very useful smaller models - don't even count gemma 4 out.
Once the subsidization ends and cost becomes significant I will take a serious look around for the best value models and switch off the expensive providers, but that time hasn't come yet.
4.1 Flash seems to be in that sweet spot of very decent, really fast and really cheap. Even omitting the cost, it’s still compelling for staying in flow.
Something about renting that much compute doesn't sit right with me so I stick with the $20 subs.
There's a reason the labs in the US frontier oligopoly are using “safety” to lobby for antitrust exemptions for mutual coordination as well as anticompetitive regulation.
I get privacy, freedom, and no rate limits with the GPUs I racked locally, and those are features I would never give up even if the surveillance capitalism labs paid -me- to use their models.
How many consumers are there like me? Probably not many, but once local inference hardware is plug and play, I bet the tides shift pretty quick. Also weights-on-silicon will serve the needs of most consumers locally with more speed than any GPU could deliver for a fraction of the cost.
Most people will be doing inference in their pocket or a wearable in 5 years and the giant datacenters will be like AWS, sold to only big organizations that need to auto-scale capacity of custom models on demand.
The industry surely knows this and the subsidized inference is just marketing to generate so much buzz and demand such that the tiny fraction of the market they will be able to keep in the end is big enough that they do not collapse under all the debt.
OpenAI and Anthropic will be Dell and IBM in 10 years if they survive at all.
Also isn't an open weight model also subsidised? Training isn't cheap and you are not paying for it.
We also have OpenAI numbers, where people speculate that the about 150% of the margins spent on "marketing" is a fake line used to hide operational costs.
In an interview earlier this year I remember Dario saying that the models are profitable, ie they more than pay back their inference and training over time. But because they invest in that explosive growth they have a deep negative cash burn.
And whatever markup the AI labs make, that’s on top of nvidia’s markup, micron’s markup, etc. It’s an industry where every supplier is adding a 70-80% markup!
For OpenAI, absolutely not, that's not including training costs. But I do personally accept that it can be actual marketing costs.
2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.
Having a better model is the only real moat, without that inference is a commodity
The benefits are real and our willingness to pay is real, but the valuations only support one and right now even the free cheap models might win. Thusly the collapse could still happen.
They’re all stuck in a cycle of spending huge amounts of money on training just to stand still (in business terms).
In the long run, it can’t continue because it doesn’t make any sense
For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.
If they stop subsidising Claude Code for the pro/max users, there will be a lot of people priced out of it, especially the casual developer. But I don't see it going away for commercial use, even with a large price increase.
Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness
Show me people handwriting code à la NASA
And I mean we as coders have been trying to do this workflow for a while, I personally would refuse to code without IntelliJ magic complete
For this workflow, there’s no going back. What’s hard to imagine is AI taking over the other workflows we predict it will; Customer service AI sucks ass for me as a customer, et cetera
That enterprise cost you're willing to pay is correlated to how much developers will work for. When driving an agent, almost anyone can do it (almost no skills required).
If devs cost $1k/m, enterprises are not going to be willing to pay $4k/m for Claude.
What I am saying is, there's an equilibrium that will be reached; the price of the human driver and the AI worker will approach each other.
Where they stabilise, I still don't know, but I'd be very surprised if, in any field (not just dev), the human gets paid multiples more than the agent they are driving, as the agents get more capable.
In the same way that only supercomputers used to have multiple processors and caches but it's now standard.
“With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited.”
It's amazing how new we all perceive AI to be, and yet how old the tricks that the big players use. Their job is to just suck the oxygen out of the room as long as they have the money to do it.
You can't jack up the prices on your product if your competitors can just clone it and resell its essence for pennies on the dollar.
And that this is even legal is a scandal all on its own.
Possibly they don’t even really know themselves at this point, although obviously is it significantly higher than the consumer subscription price
I've been running automated research tasks for life sciences companies, and the speed in which tokens are burnt is scary. Especially when you get into a complex knowledge space and require a subwgent to reason through each possibility, token usage grows quadratically not linearly as complexity increases...
It probably will. Moore's Law is still churning away in the background.
Frontier models might get more expensive, but that's a moving target. For any particular capability point, the models will only get cheaper.
My understanding is that enterprise plans don't offer those subscriptions, so they end up paying for API prices and models like these directly impact that revenue stream.
https://cortecs.ai/detailedServerlessView/deepseek-v4.1-flas...
There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.
DeepSeek themselves are honest about the fact that they train on inputs by default. You won't hit DeepSeek if you use OpenRouter with ZDR enabled.
I don't really trust OpenAI, although honestly I find it stupid to suggest they'd offer a ZDR policy and violate it. They literally don't have to offer it. People will still pay. Fable doesn't offer ZDR at all and it hasn't stopped people from paying through the nose for it at API pricing.