When people use words like “chatbot” to talk about their AI use and can’t really name basic things like the model they were using, it reveals a wide chasm between the people who are deep into understanding how to use these things and those who poke at easy options every few months and then look for ways to confirm their priors instead of trying to understand how other people are getting value.
The bigger the pile of stuff you have the tool create unsupervised, the more incomplete/inaccurate the overall picture of it is gonna be. That doesn't matter for most hobbyist/helper-tool/game-port/etc type of tasks.
I get a ton of value out of the tools and yet part of that value has come from being hyper-aware of the shape of the limitations, which changes much less and more slowly than the specific of the limitations.
"How'd you notice that one!?" - well, it helps to be the only person on the team who took the time to interrogate the e2e interaction of all the code...
The median piece of OSS software was already pretty rough once you had to depend on it in production at scale to make money. The median piece of the new AI-generated wave of software, also pretty rough.
The bet that OpenAI/Anthropic is making is that they can automate the interactions generally/globally enough to work around this and eliminate the need for humans in the loop. But it's a tricky problem because it's precisely where the permutations and special cases get nasty and much harder to deal with than more broadly-applicable cases of codegen/test-gen/mash-it-until-the-tests-pass-and-the-code-review-angent-is-happy.
If you talk to free chatgpt, sure.
But if you launch Opus 5.5 in claude code - for some reason Opus is the smartest for me, even for "general" type of questions, when launched from Claude Code. It has harness/prompts optimized for correctness, while more "customer-grade" products, even when used with top model tend to optimize for latency. And I tend to give "context" - e.g. I have "personal wiki" + skills for lots of those things.
I don't recall i last few weeks where Opus 5.5 in CC made an obvious lie, made an unreasonable assumption, or made a mistake that smart human couldn't make. E.g. if it "failed" it failed in a way that non-genius human could fail in similar way.
I tried to have it write a synopsis website for the different types of dance patterns in Dance Dance Revolution, with guides on their execution, explanations, and provided song sources. I provided all of the source articles and search websites needed for it to be able to do that, and gave it a sample set I wanted it to utilize and build off of.
It has bits and pieces of information wrong in 90% of the sections it created. Poor song choices, missing song choices, incorrect explanations, incorrect step-patterns, the list goes on.
If you're not an expert in a given field for which you're using it, it's going to look like a genius. If you are, it's untrustworthy and borderline useless.