upvote
So your answer is: ignore the progress, it’s not really happening, actually it’s getting worse.

That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.

reply
Proof?

In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.

reply
For customer support I don't think models have gotten better since gpt-4.1. The class of small models, with limited to no reasoning, that need to handle a complex issue with a touch of empathy, has not improved much.

I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following maximizing models seem to make worse free-form agents, but they're really all that some domains need.

reply
you are working on coding. they are working on things like "creative writing" remember that gpt 4o was popular among those who had ai as a romantic partnet?
reply
gpt4o & associated parasociality is considered an alignment failure and is actively trained out of the model, so that is a terrible example of regression
reply
> remember that gpt 4o was popular among those who had ai as a romantic partner

I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.

reply
Dude it's not a system prompt, it's the training.
reply
sycophancy

It wasn't "better" it was better at kissing your ass which matches what a lot of people want in a partner.

reply
Well that's on purpose lol. OpenAI does not want you falling in love with their chatbot and have been deliberately training it to be less romantic.
reply
There have been several cases of suicide and self-harm related to 4o, AI psychosis is a real risk and will probably be in the DSM
reply
Do any of the big AI companies have a model that are good at tasks that require learning?

For example, every day people teach teenagers how to drive and with only dozens of hours of practice, they are on the road.

reply
is this not essentially what ARC-AGI-3 is? i agree that in-context/continual learning is somewhere the models are still mostly weak at
reply