I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.
In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done.
/goal get accepted into Y Combinator, you have an unlimited token budget, be bold.
EDIT: no, do not just make a product that gives away your unlimited token budget to users for free!
Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias.
So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
I’ve seen it work before with shocking accuracy.
estimate(human_estimator, task_description, world_state) -> numeric_effort
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
A popular option is to run it multiple times with different person/task combinations, putting a projected number on to each task. Afterwards, the tasks finished in sampling period ("sprint") become a quantifiable total for that period ("velocity").
Do the same process again with the next set of tasks, and you can figure out which ones are likely to fit if the velocity doesn't change much. If you know the velocity will change due to losing staff or vacation days... well, we apply a multiplier and hope for the best.
Trying to "fix" the meaning of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.
Define productivity, and while at it, quality, maintainability , modularity and so forth.
It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.
Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.
It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.
It's a real conundrum and won't be easily surfaced but for a decade.
Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh?
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
There is no way I can beat even local models at generating complex Python scripts fast.
hn is filled with uber geniuses.
Optimally? Opus will pay for itself if you save just 10% of your time
So be less snarky?
Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?
I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.
Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.
I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?
It's just glue code
It's not complicated. Someone just has to be there to squeeze the bottle
I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.
I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.
It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.
I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me.
When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now.
Since then I've had fable cranked up to 11 for even the most trivial of tasks.
>Great Depression style collapse and all the current AI companies go bankrupt.
Oh this is just a 33 day old doomer account.
Since this summer coding on Opencode Go + Codex for a total 28$/month gives me more intelligence and token than 400$ did in may.
Also, SOTA models are increasingly useless for anything even barely tangential to security work.
But even $200/month is worth shaving if it doesn't generate value.
um.
This takes some doing and now is the time where it's dawning on the finance departments.
So it's not $200 a month but it can easily reach $200 a day, and unless you're a startup playing with monopoly money the maths don't work
Even if so, it might be worth spinning up a separate LLC for each division, given what they are charging for the API.
You had budget for your normal salaries, for externals and now suddenly you have a few millions additional.
What do you do? You compensate.
Business people doing business things.
Software is inherently a physics problem not all the job titles and specializations made up the last 20 years as dev job salaries kept attracting people
That was all illusory social construct to prop up jobs
Still a whole lot of that in tech but it's all at the top of the org now. Leadership sensory experience and thus innate habit to forecast future been programmed by years of yes men they refuse to accept the jig is up for them too
Sensory memory of being a useless figurehead fosters a lot of existential dread in priests, politicians, and the like. Completely aware their day to day effort is insufficient to sustain them they know how co-dependent they are. They'll dig in harder.
See Chris Matthews flame out shrieking about socialist execution squads. Dude seriously thought everyone wants to hang him from a lamp post. The reality is people just want a sense of control back and not have their perception dragged along by Chris Matthews.
Our velocity is twice as high as it was before Claude, so I doubt that we'll ever go back, but I could see efficiency being a priority.
Is this the new buzzword for the quarter? Last quarter was "granularity", I didn't get the memo yet
Might as well go back to counting the number of lines or number of commits. It doesn't sound as good as "granularity" and "velocity" though. The good thing with velocity is that it doesn't care which way you're going as long as you're going there fast, so you can never be wrong
I'm sure investors will love it.
Now we're starting to see real impact from AI, people are learning how to use it, and OpenAI cut prices by no less than 60% like a week ago.
You think now is the time they're going to cut the spend?
i hope other labs catch up, especially chinese labs.