1/ if the point of setting xhigh or max effort on a "reasoning model" is -not- to have additional reasoning tokens applied to the problem, then I fear we are all using AI wrong 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.
You keep moving the goalposts. The core flaws in your analysis are that you assumed the variation was 100% server side and didn't admit that the inputs were random and different every day, and that you tried to compare to ARC-AGI-2 as a benchmark.
Comparing ARC-AGI-2 thinking tokens to agentic coding thinking tokens is as misleading as it gets, because these are completely different use cases.
Please stop and think about this for one second. Do you really want Anthropic to spend 30K thinking tokens on every input, just because that's what ARC-AGI-2 problems required? What would your inference bill look like if this was the case?
The premise of your ARC-AGI-2 comparison is broken.
> 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.
Ironic to accuse someone of cheap tricks as you pull out an ad hominem insult after someone explains the flaws in your reasoning.
My points stand: You can't claim this is a chart of Anthropic changing the server when you were feeding it random input every day and plotting the output as if the line should be flat. You can't compare agentic coding to single-turn ARC-AGI-2 problems.