upvote
Hint: The new models are really good at burning tokens.

I've had to use it a bit for work, and it's been remarkable watching the degradation in performance with the default suggested current models (Opus 5 as a prime example) vs the models that got them huge attention a year ago (Opus 4.6)

If you give 4.6 a spec, or existing code to implement a feature in, it will ask some pointed questions if there's something unclear in the spec, and then produce a plan and move to implement it.

5 will freak out at even a basic task, ask itself if it's own assumptions or your instructions are correct, proceed to re-assess it's own plan, and it's instructions 3-4 times, and then maybe produce code after burning several hundred thousand tokens (and quite a bit of time) analyzing existing code and thoroughly sweeping it for irrelevant problems both to the task it was given and the spec it came up with.

It's quite bizarre to me how well advertised the benchmarks and anecdotes from people one shotting MVP browser games are, compared to the experience of everyone I know that's had to actually use it to accomplish even a relatively basic task.

reply
I managed to lose around $300 in credits I had saved for some emergency /fast sessions the following way: switch to Fable. Work on the design. Downgrade to Opus for the build. If any of other parallel Opus session has /fast enabled it seems to enable it for the newly spawned session by default. Before I knew it, the $300 was gone. I think the bug is now solved, but it was rather unpleasant. I dont ever remember bugs that would drain my wallet - with claude code its just another Tuesday. Still love it.
reply
Claude code is just pool quality. They don't make how this thing will behave clear to the user, or give control. They fail at anything that needs an abstraction or model, not just APIs and shell scripts glued together. And "just ask AI" seems to be the default fix.

That vibe coding they brag about as if it was a good thing, it shows.

Take their notation for describing permissions. The docs are not comprehensive, and in practice it doesn't quite work how they describe it.

Or their management of sub-agents. I once lost a sub-agent, it finished and disappeared from UI. Apparently, you can't bring it back yourself: you have to ask the parent agent to do it for you. But the parent was Fable, and I ran out of credits, so I was locked out of using my opus sub-agent because of it.

Or an even more grotesque example: when you paste your claude API token to authorize, it covers characters with *. But it seems like an LLM has hallucinated a limit of API key length and the tail of your key stays visible.

reply
What amazes me is how, for a vibe coded product where all they have to do is use their AI to fix things ... NOTHING EVER GETS FIXED!

I've probably gone to file 20 bugs. In all 20 cases there wasn't just one issue already filed for it: there were several, each which had a bunch of upvotes. And in all 20 cases ... every. last. one. ... Anthropic closed the ticket with no comment.

IF YOU ARE GOING TO HAVE A SHITTY VIBE CODED PRODUCT, AT LEAST USE YOUR SHITTY AI TO FIX THE SHITTY PROBLEMS!

reply
[dead]
reply
so many ridiculous "how the fuck did this get through basic QA?" issues with Claude Code.

I can't believe how many critical bugs fall through.

My favourite one is the bug where Plan mode can execute destructive commands inadvertently.

Then all these get closed with `Closing for now — inactive for too long. Please open a new issue if this is still relevant.`. Awesome.

reply
> I can't believe how many critical bugs fall through.

Almost like CC is 100% vibe coded.

reply
I dont love it.

Opus 5 is just a token burner.

I use fable plan and spawn opus 4.8 workflows which seems to work alright.

reply
I suspect it must depend on how one manages their codebase - wrt to docs, ADRs, and general guardrails.

For me it is not great for design work - Fable is way better, and 4.8 was conservative and thus better (Opus 5 seems to jump to conclusions far more eagerly). But for overnight builds, where I give it 8hrs worth of work on LLDs created by Fable - its great. Where Opus 4.8 would often lose the plot and stop for questions clearly answered in the LLD - Opus 5 does manage to complete. Since it launched, I don't remember it ever disappointing me with builds. But designs? Boy, is this thing explosively stupid sometimes.

reply
Opus 5 loves to stop working "for safety reasons" and shuts down the session! I avoid it at all costs now. Opus 4.8 has been my default as well.
reply
deleted
reply
And at that cost they're still not profitable. It's going to be a bumpy road ahead...
reply
I thought they are making a profit on API pricing? A quick Google shows somewhere between 50-70% margins on API inference.
reply
API pricing is almost definitely profitable, but at this point I assume it's a small minority of their inference traffic compared to subscription usage, and unlikely to make up for the rest of their expenses on its own.
reply
Meanwhile I can do all that and more with reasonix harness for Deepseek with a cache hit rate of 99%. And that's with unsubsidized American providers like cloudflare or Digital Ocean
reply
I think profitability is a matter of accounting. Inference is where money is made, but training is where money is spent. We keep getting new models every few months, but frankly the old models are still quite usable. I suspect labs will soon start specializing in expert models per use case so they can increase the lifespan of individual models, and change the profitability per model.
reply
That's not the only reason to go to expert models. The more different domains you try to stuff in there, the more parameters the model needs to keep things coherent and not overload tokens in a way that induces errors. For example, if a model trained only on biology text sees "sonic hedgehog" there's no ambiguity, and this compounds for all the things that are "overloaded," in the training corpus, which turns out to be quite a bit.
reply
People keep saying this but from what we’ve seen, Anthropic models are marginally profitable and earn back their costs over their lifetime. The company is burning money building the next versions and other ventures (e.g. verticals), but the models themselves have been profitable.
reply
They're EBITDA profitable, not GAAP profitable.
reply
What’s the blast radius of this bubble popping? It’s all private investment still right?
reply
Two thirds of most of the DC builds are not compute. So it's a CRE play the last leg holding up that mess.
reply
It’s crazy how different the credit cost and subscription cost are.

With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.

reply
VCs are footing the bill for that $200 subscription.
reply
They have something like 80% gross margins, are at a $100B/yr ARR, and are growing at 10x per year... If that keeps up, they're going to be doing more revenue than Google in a year ($400B ARR, 20% per year growth)
reply
How can you sanely project the last 12 months forward? We have seen a huge uptick in usage. Last summer AI was a toy to most devs, now every enterprise developer I talked to uses it every day. Coding agent providers are surely going to hit market saturation in the near future.
reply
deleted
reply
At last, a valid usecase for VCs.
reply
Its tough to go from max account at home and pay per usage enterprise account at work with heavy usage limits... but the limits are there because pricing is insane. Feel like I'm in the $5 Uber rides phase at home.
reply
The Chinese models are the public transport in the uber analogy. Once the price the goes up catch the bus!
reply
Lol perfect analogy. I'm still paying for claude because the quality is unmatched.
reply
You plugged in a space heater on a roofless house.

There is some element of responsibility on the user to guide and monitor the model/harness and not let it rip to burn tokens.

reply
Anthropic is the new AWS.

Amazon's first principle is the Customer Obsession. Making customers happy.

Fun bit is that the human psychology rates personal looking fixes better than having no issues at all.

For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment".

Meanwhile, any other (small) cloud. Simple, no weird charges. Even _most_ of network egress is free. But, no reason to call support or feel "extraordinary". Comes out as "meh" against Amazon's "top tier" support model...

reply
I’ve never gotten a refund from Athropic.
reply
Anthropic's constant changing of its mind leads to instability which leads to unhappy customers
reply
> wayyy overpriced

Maybe they consider that hiring a person to do it would have cost at least as much and taken much more time, so paying them is a bargain.

reply
Yeah, but now we can hire the Chinese instead for 1/100th the cost. It's an even better deal.

Plus we get to own, keep, run, do whatever with the model. We don't feel trapped. Moreover, it's something we can truly build on top of and own our own destiny.

Anthropic and OpenAI are the new Oracle (Oracle pre-AI; Oracle is even worse now). Expensive, feels like dealing with a lawyer, and not at all open. They just became infinitely less cool than they were a month ago.

The whole of our industry is going to migrate to open weights. We're smart enough to know this is the better deal and technical enough to be able to pull it off.

The only thing that might save these OpenAI and Anthropic in the near-term is an abundance of enterprise contracts negotiated with non-tech companies. They'll soak consulting firms and F500 companies for "AI" integrations.

reply
> the new Oracle

I think that's exactly what they are going for - enterprise and government customers.

reply
You should just spend those towards a cursor subscription.
reply