upvote
I think you're underestimating both their reliability for standard problems and the usefulness of that level of reliability.
reply
This is a good point. Opus does some silly shenanigans sometimes but then catches it later. It’s still an order of magnitude faster at getting to a working system than I am, for ones I don’t know.

It’s really a dream for setting up a homelab

reply
> out of 100 random questions I might think to ask, it’s likely to say something wrong or stupid a handful of times at least.

What are some examples?

reply
There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience
reply
I wanted some examples they actually experienced. Because I use these things daily and haven’t seen a hallucination in a long long time.
reply
I saw a hallucination just this afternoon about a spurious ca cert error. Definitely happens less often, but I do need to correct it occasionally. Maybe once a week so it still requires vigilance.
reply
Search a terminal with Claude Code for things like, “I got it wrong twice. I should look up the documentation instead of guessing.”

Does it about once a day, that I notice.

reply