upvote
Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal
reply
Yes, but… more thinking tokens also means longer solution generation time. That said, v4 Flash is a fast model. I use it all the time because it’s very smart for the price. But it is verbose sometimes.
reply
The thinking trace was (preview) frustrating to read, I think I'd prefer a summary view of it at this point.
reply
It’s not outdated at all to use tokens to estimate performance, it’s directly related.
reply
But why should I care? If my metrics are speed and cost? How many tokens it takes as a user is arbitrary to some extent.
reply
> inefficient

That depends. Is it also more reliable?

If two books, one big one slim, prove the same thesis, what I would be interested in is the quality of the content, not the size. There can be a measure of efficiency in "have you really thought it through", but it is clearly complex - it requires measuring how solid the reasoning is.

reply
Roughly equivalent to Gemini 3.6 Flash in capabilities at 1/20th the price...

Mind you, until the recent price cuts to Luna - Gemini 3.6 Flash wasn't even egregiously priced (but oh how things change in just 1 week).

reply
you are correct. but in my experience -- not benchmarks -- g flash 3.6 is SO BAD for coding. I'm using all vendors all day and gemini is the worst by far. I built my own semi-deterministic orchestrator for coding agents.
reply