upvote
That always seems to happen upon release of a new model. People look at the benchmarks, which look good, because the model was almost certainly optimised for that. Then they start actually using it, and after a week or two we get the real assessment that is almost always less impressive than the initial reactions.
reply
Interesting, I generally use High for all models, even non Anthropic ones.

This was one of the few times I tried Max, and the problem was the code (it listed as a "hole") was directly adjacent to the problem area, and not especially complex.

Kind of like looking at a washing machine and telling the customer to be careful because the inlet pipe will pump water into an empty box.

reply
I believe the reasoning for not using anything higher than medium is that Opus 5 has a tendency to overthink at higher effort levels.

So I guess the washing machine analogy is that you've somehow added too much detergent because the new formula is 3x as strong, and your clothes are very clean, but the fibres have also degraded leaving your clothes a bit threadbare.

reply
GPT5.6-Sol is brilliant on Max in my experience.
reply