upvote
Almost every time I use an AI chatbot it eventually turns out it's lying to me. I would not trust them with any of the things you mention.
reply
This is the kind of gap in shared understanding that is widening.

When people use words like “chatbot” to talk about their AI use and can’t really name basic things like the model they were using, it reveals a wide chasm between the people who are deep into understanding how to use these things and those who poke at easy options every few months and then look for ways to confirm their priors instead of trying to understand how other people are getting value.

reply
The incentives seem to be lined up to push users to the frontier of "how much can the harness around the model do and claim without most people pushing back on the claims."

The bigger the pile of stuff you have the tool create unsupervised, the more incomplete/inaccurate the overall picture of it is gonna be. That doesn't matter for most hobbyist/helper-tool/game-port/etc type of tasks.

I get a ton of value out of the tools and yet part of that value has come from being hyper-aware of the shape of the limitations, which changes much less and more slowly than the specific of the limitations.

"How'd you notice that one!?" - well, it helps to be the only person on the team who took the time to interrogate the e2e interaction of all the code...

The median piece of OSS software was already pretty rough once you had to depend on it in production at scale to make money. The median piece of the new AI-generated wave of software, also pretty rough.

The bet that OpenAI/Anthropic is making is that they can automate the interactions generally/globally enough to work around this and eliminate the need for humans in the loop. But it's a tricky problem because it's precisely where the permutations and special cases get nasty and much harder to deal with than more broadly-applicable cases of codegen/test-gen/mash-it-until-the-tests-pass-and-the-code-review-angent-is-happy.

reply
using skills with opus 5.5 and agents for each individual task that have shared knowledge of all of your records does not lead to a significant increase in trustworthiness
reply
I have ChatGPT Pro and it sources every single thing. For my Asia trip right now it's been helping me every single day to figure out optimized routes on what to do and where to go (if I feel like it, not to say I don't wander around too).
reply
That sounds terrible. I like to live like a human being
reply
When is the last time you used an AI chatbot, and do you only use the free versions?
reply
Which model? Super weird
reply
Every time I use Google search AI for anything more complicated than a search query I feel that way. A paid model from pretty much anyone is much better.
reply
> Almost every time I use an AI chatbot it eventually turns out it's lying to me. I would not trust them with any of the things you mention.

If you talk to free chatgpt, sure.

But if you launch Opus 5.5 in claude code - for some reason Opus is the smartest for me, even for "general" type of questions, when launched from Claude Code. It has harness/prompts optimized for correctness, while more "customer-grade" products, even when used with top model tend to optimize for latency. And I tend to give "context" - e.g. I have "personal wiki" + skills for lots of those things.

I don't recall i last few weeks where Opus 5.5 in CC made an obvious lie, made an unreasonable assumption, or made a mistake that smart human couldn't make. E.g. if it "failed" it failed in a way that non-genius human could fail in similar way.

reply
opus 5.5 uses agents and skills in chat sessions as well and maintains internal context that it tries to distill from what you provide

I tried to have it write a synopsis website for the different types of dance patterns in Dance Dance Revolution, with guides on their execution, explanations, and provided song sources. I provided all of the source articles and search websites needed for it to be able to do that, and gave it a sample set I wanted it to utilize and build off of.

It has bits and pieces of information wrong in 90% of the sections it created. Poor song choices, missing song choices, incorrect explanations, incorrect step-patterns, the list goes on.

If you're not an expert in a given field for which you're using it, it's going to look like a genius. If you are, it's untrustworthy and borderline useless.

reply
[dead]
reply