I gave the task to codex first, sol 6 xhigh. it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.
Opus 5.5 high took the same prompt with no back and forth, it just went off and one-shotted a tool that takes pixel-perfect screenshots of exactly what my app looks like.
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.