Yes, this is my production corpus, across 65 usage days, two subscription accounts, three machines, 25 project groups, and 213 sessions, across 43,261 invocations and 7,583 turns. Use your own data if you want to prove or refute what was seen in my corpus.
The "ups and downs in the chart" were not plotted with sub-daily resolution. Specifically, the two-month temporal chart uses a 3.5-day Gaussian bandwidth which is meant to reduce short-term noise while retaining broader changes.
Additionally, a separate episodic analysis identified multi-day changes in delivered thinking. And those episodes were predictive of held-out work. The more projects pulled into an ensemble, the more predictive they were of delivered thinking tokens for held-out projects during the episode.
The point of the benchmarks is to establish a baseline for what thinking-token counts one should expect from specific effort levels using published numbers, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.
So you don't have to provide a generous interpretation of my workload if you don't want to. Remove all of the zero-thinking token responses, redistribute those samples across the distribution, and then tell me if it magically shifts right and starts delivering anything close to published numbers. Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.
If you would like to denigrate a month of my time as vibe-slop, that is your prerogative. You can even be dismissive of my workload, if you want, even if my background should tell you otherwise. But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.
I don't think you understand. What you posted is highly dependent on your corpus. I can't "refute" anything because it's not available and it's the major variable in the experiment.
> The point of the benchmarks to establish a baseline for what thinking-token counts one should expect from specific effort levels using published, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.
I think you're missing something from that first sentence, but I assume you're talking about the comparison to ARC-AGI-2 published thinking tokens?
It should be blindingly obvious that you do not want your thinking token counts to be as high as a benchmark that was designed to push LLMs to their limit.
> Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.
What point are you even trying to make?
Again, you don't want invocations to be burning 16K thinking tokens except for rare problems that 1) must be solved in one step and 2) are designed to be entirely self-contained thinking in that step.
You're trying to compare development work to a benchmark that encapsulates complex thinking into a single step.
Coding work is iterative and works in incremental steps: It runs commands, reads more files, checks the web. Thinking tokens should be low for your turns.
ARC-AGI problems have an input and an output. They look like this: https://arcprize.org/tasks/b5ca7ac4 They have more thinking tokens because that's the entire state. They get one output and it's constrained.
> But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.
It is fair to discuss a published analysis. Saying that only people who bring their own month of equivalent analysis (which conveniently would take another month to produce) are allowed to critique it is just a cheap trick to shut people down.
If you post big claims, they are open to analysis and review by others
1/ if the point of setting xhigh or max effort on a "reasoning model" is -not- to have additional reasoning tokens applied to the problem, then I fear we are all using AI wrong 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.
You keep moving the goalposts. The core flaws in your analysis are that you assumed the variation was 100% server side and didn't admit that the inputs were random and different every day, and that you tried to compare to ARC-AGI-2 as a benchmark.
Comparing ARC-AGI-2 thinking tokens to agentic coding thinking tokens is as misleading as it gets, because these are completely different use cases.
Please stop and think about this for one second. Do you really want Anthropic to spend 30K thinking tokens on every input, just because that's what ARC-AGI-2 problems required? What would your inference bill look like if this was the case?
The premise of your ARC-AGI-2 comparison is broken.
> 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.
Ironic to accuse someone of cheap tricks as you pull out an ad hominem insult after someone explains the flaws in your reasoning.
My points stand: You can't claim this is a chart of Anthropic changing the server when you were feeding it random input every day and plotting the output as if the line should be flat. You can't compare agentic coding to single-turn ARC-AGI-2 problems.