upvote
What bothers me about this whole AI tokenomics situation is the lack of transparency. OpenAI and Anthropic have to perhaps be the most opaque companies in existence wrt their offerings. There's like a thousand variables that they can change on the backend at the push of a button which can wildly swing API spends within the same model (partly also due to the non-deterministic nature of LxMs, but still), and there's no objective way to measure them other than vibes.

When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.

reply
And it's all measured in "intelligence", a completely meaningless term. For coding i'd be much more interested in how much context actually works, what the complexity of algorithms it can understand and create is, for what languages. How much it manages to follow existing structures or that is just adds ad-hoc machinery to pass the test, etc etc.

A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.

reply
It's still insane that they stopped showing you all the tokens you pay for. They could inflate the billed reasoning token amount by a lot before it would raise any eyebrows.
reply
Internet Ad business has been like that for a long time with bot click & Co.
reply
Ain't that the truth.

In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D

reply
Totally hand-wavy and non-objective ... just like how employees are billing their employers.
reply
It's all Gacha for business.
reply
Yeah, you might get away with a singular 'wrt' with some consternation, but two?
reply
> Subagents are like trading derivatives. You can lose as much as you want.

Excellent pithy warning.

reply
Sadly this is true - for individual folks on the lower end of the spend spectrum.

But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.

reply
Afaik there is just pay-as-use with Deepseek
reply
Why use subagents at all
reply
Preserve context in the lead chat - let the subagents fill up their own contexts then only return the necessary information.
reply
Why not just have an agent that can branch its context?
reply
Forking conversations have been a thing for a long time, and it's not the same thing (e.g. you may have 50% of your context used up at fork time that is carried into "both agents" afterwards)
reply
deleted
reply
Why hire a junior developer if you have a perfectly competent senior developer already on your team?
reply
You can do a code review on a "less capable" model that costs less, and the key model gets its output / summary, then you can have that model build a plan, and feed it to cheaper models. It's a more efficient approach than just running everything through Opus, and now that Sonnet is a lot better I'll probably use them more frequently, one thing to note is don't ask it to spin up endless subagents, I'd cap it to 2 or 3 at a time, otherwise, yeah you'll hit your limit extremely quickly.
reply
Code reviews also work better in sub agents because the reviewer agent didn’t write the code being reviewed.
reply
Because two agents are faster than one.
reply
The models get dumb as context fills. Subagents allow them to accomplish a task with minimal context rot. You can also use cheaper models for subagent tasks
reply
Time is money. Parallelism is very helpful optimising one to get the other.
reply
Money is money too. Increasing contexts costs non zero money, even with cache hits. Also context rot is a problem that subagents help with
reply
This truism is intuitive to everyone but always funny to me how everyone never has any time, needs to save time, needs to hire staff workers for every mundane job and robots can't come soon enough… all so we can binge watch Game of Thrones and 90 day Fiancé.

And watch 10 hours of football on Sunday for our DraftKings bets.

reply
Money is also money, which anyone who makes heavy use of parallel subagents will quickly learn.
reply
> Time is money. Parallelism is very helpful optimising one to get the other.

Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.

There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.

reply