I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.
People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.
Those usage patterns don't correlate to output.
We ended up having to hire a full time employee to fix the performance of client built reports.
People run a stupid amount of expensive queries that end up costing way too much because they're asking Claude the wrong query.
Not to mention people running wrong queries, using the result as gospel, and then the result has to be sent to a data analyst to be reverse-engineered so the numbers make sense.
Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.
As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.
But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...
It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.
Which is an issue when you need to get department managers to manage their budgets around the amounts of credits their employees spend. The more of a black box it is, the more governance and corporate bullshit you have to deal with.
The copilot part of it runs "unlimited" on the license. Except it's not unlimited, and this is even more of a blackbox because you can't see any sort of spending and the limit is listed as "extensive use".
On your point about "monitoring", personally I feel this is toxic corporate IT culture, enabled and perhaps pushed by the likes of Microsoft with all their tools, which they of course make money off. People have cellphones with cameras making most points in this area moot.
An alternative used in other big corporates is to set budgets, with tiered authorisation approvals for higher limits. The users and their managers can justify why and what they're doing that they need the additional tokens. This also encourages more efficient use of tokens on other work. More efficient use is sometimes counterintuitive. Laissez-faire generally works best.
Cowork requires user approvals for high risk actions such as emailing.
The better pattern is to let it code the app and then you can use the app to target your data. So you only pay for it once, plus it's deterministic. But yeah, it requires setting up an environment, etc. It becomes "maintenance".
I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.
Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.
Eg I've used Sashiko locally for Linux kernel code reviews before sending out my contributions out to the world. Sashiko is a great system, but it can burn through tokens like there's no tomorrow.
https://github.com/sashiko-dev/sashiko and https://sashiko.dev/