upvote
It kind of does, mathematically. It doesn't label confidence. If 0.0001% of answers is a lie, without knowing which parts are a lie exactly, you cannot trust any of them.

If you need to independently verify every fact, why not just gather facts yourself in the first place.

Let's say, a mathematical concept of lie. I still use them every day, of course.

reply
"why not just gather facts yourself" vs "i use them every day", the duality of man. But yes, I appreciate that LLM output is theoretically completely untrustworthy- but when in practice I observe that it's around 90% accurate, I have to rely on my own internal calibration for how useful it is (depends on type, nature of task ofc)
reply
As they say, the less you know, the better LLM output is. :D
reply
> Your one example doesn't make all of LLMs a lie.

It's not a lie though, because the truth isn't guaranteed by the mechanism that generates the answer.

reply
deleted
reply
Okay, but the anecdote states that every model repeated the pseudo-factoid about Foobar square, not just the 4 GB open source model equivalent of a tabloid.
reply
I think the key phrase here is, "an obscure small town." There may only be a single mention of this place, hence the only one on which a response can be based. This says more about the user's understanding of LLMs than it does about LLMs.
reply
This says more about the user's understanding of LLMs than it does about LLMs.

"Tell me everything you know about (obscure small town), (state). Only what's unique to (town), not commonly-known facts" is an excellent way to test for hallucinatory tendencies in a new model, in my experience. Likely the best I've found.

Quality of results is almost linearly proportional to the size of the model in many cases. The largest models like K3 and GLM 5.3 will either confine their responses to known true facts about the town and its surroundings, or admit they don't have enough information to answer. Smaller ones will reliably make up hilarious or downright-strange things.

Another good test is https://whatever.scalzi.com/2025/12/13/ai-a-dedicated-fact-f... , which still works on the newest models. Of the open-weight models available, only Kimi K3 will consistently admit it has no idea who Scalzi's novel is dedicated to. The rest still make up random stuff and present it confidently.

TL,DR: progress is possible, and it has been made, but it's happening slower than many people think.

reply
Without disclosing what you were prompting for, it's impossible to evaluate your claim.
reply
Correct. We can either accept the claim or disregard it. The comment I replied to opted to accept it and then committed a fallacy, hence my response.
reply
by what mechanism do you believe LLMs verify truth? they are amoral token generators
reply