upvote
Yes, but… more thinking tokens also means longer solution generation time. That said, v4 Flash is a fast model. I use it all the time because it’s very smart for the price. But it is verbose sometimes.
reply
The thinking trace was (preview) frustrating to read, I think I'd prefer a summary view of it at this point.
reply
It’s not outdated at all to use tokens to estimate performance, it’s directly related.
reply
But why should I care? If my metrics are speed and cost? How many tokens it takes as a user is arbitrary to some extent.
reply