I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following maximizing models seem to make worse free-form agents, but they're really all that some domains need.
I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.
It wasn't "better" it was better at kissing your ass which matches what a lot of people want in a partner.
For example, every day people teach teenagers how to drive and with only dozens of hours of practice, they are on the road.