Both evals and Human pairwise tests for our use case are giving Haiku 4.5 first place in pretty much all tests.
No we'll try understand if we need to change our prompts to match performance ...
edit: maybe this will help: https://platform.claude.com/docs/en/build-with-claude/prompt...
Anecdotally we ran sonnet 4.6 for our more complex stuff and sonnet 5 was a LOT worse. 5.5 seems to have fixed it and we cut over our customer workloads. It’s strange, really.