upvote
I have a distinct line between when I'm willing to believe an LLM's output and when I'm not: whether I would believe the same thing from an anonymous Internet forum post or a blogger I don't know. Those posts are not unlikely to be misinformed, biased, lies, or otherwise untrustworthy. And yet, I spent plenty of years honing a sense of when they were good enough for certain things.
reply
A lot of that sense was probably based on side channels like proper grammar, writing style, etc. That’s all gone now :(
reply
No, the sense had nothing to do with the content and everything to do with the context. Perfect grammar and writing style were never enough to get me to trust an anonymous forum post or unknown bloggers post for certain topics like health advice. Sloppy grammar and writing style were never a deterrent for me believing them for other kinds of topics like where to check on the HVAC system to find the sticker. I think the line could more accurately be described as the level of risk if it's wrong.
reply
I dunno, I've read a lot of very well presented nonsense and some very useful insights that were barely readable. I think it's probably useful that people are being trained out of this bias (though of course that's in large part because LLMs do tend to exploit this bias).
reply
For what reason would grammar influence whether something is true or not...
reply
You could tell from tone and polish how much effort someone had put into writing an answer. That was a pretty good signal for some topics on forum sites like Stack Overflow. There were always nuts and cranks who would happily spend an hour writing well-formed prose about nonsense or something obviously wrong, but the eloquent ones were few and far between. Now every crank is equally eloquent and can spit out 1,500 words of passable prose in seconds.
reply
deleted
reply
We have other signals now.

Before we would find an intriguing post on the internet from years ago, and you have to verify it with additional research--it's easy to skip that additional research.

With a LLM when you're skeptical you can interogate it. One thing we know for sure is LLMs are quick to admit mistakes were made when interrogated, comically so. A LLM might not always recognize its own mistake, but at least it is available for easy interogation, unlike the forum posts of old.

Manual research from reputable sources remains an option.

reply
You can't interrogate the context of an LLM when you only have its output.
reply
One of the things I do semi-frequently is look for the evidence that some concert took place 15+ years ago. Or maybe I already definitively know it happened, but not exactly at which venue or the exact date of the concert. This I feel like is a non-trivial task, but one with a very definitive answer whose evidence more often than not still exists somewhere online.

In my experience every LLM out there is utterly useless and quickly defaults into "here are other concerts that took place around that time near that location". Google Search (ignoring the AI overview) is even more useless, as it refuses to show literally any webpage that's older than say 5 years. YouTube search is genuinely better than Google at surfacing old and grainy fan-made videos uploaded in like 2010, but also defaults into synonyms nonsense pretty quickly.

But, the search functionality of exactly one forum and three local news websites that I know have an archive that dates back long enough beats every single one of those abovementioned every single time. Three people are talking about their experience at a concert on a random 15+ year old forum thread? It happened. The tiny list of 5 or so (Google-hosted!) Blogspot blogs I have bookmarked? They usually have a photo of the ticket that Google Images refuses to show me.

Not only are search engines completely dead as a category, but LLMs are a shit replacement for them. "We" (okay, Google specifically) has truly committed a crime comparable to burning the Library of Alexandria. Everything older than a decade that wasn't properly documented on Wikipedia is just gone, never to be seen again.

reply
It's vector search that's eaten everything that used to have at least a smidgen of parametric search.
reply
> YouTube search is genuinely better than Google

The funny part of this is that Google search is intentionally bad at returning YouTube videos, presumably because some anti-trust action scared them into artificially ranking videos from local news sites, Facebook, and other ad-walled content ahead of YouTube videos. Seriously, go watch a YouTube video, then try googling its title with “video” appended to it, and see if the “Videos” tab of google search ranks it as the first result.

reply
The most absurd thing that has happened to me more than once is that I found a YouTube video, not by searching through Google, not by searching through YouTube, not by asking an LLM to find it for me, but by finding an old article that embedded it. That embed is of course long broken by the changes on YouTube's side, but once I use inspect element to find its Youtube ID, surprise, surprise, it's still there!

It's usually uploaded by a channel with like 20 subscribers and has maybe like 300 views, but YouTube would rather show me some artist playing a similar genre on the other side of the continent with millions of views that was recently uploaded than a video from an event I specifically typed into a search bar.

reply
Yeah, it's like everyone was under the impression you could just trust the internet before LLMs.

It's a great tool, but verify the important things (or do them yourself)

reply
Notably, it used to take effort to produce crap on the internet, now it’s nearly the default action.

Signal to noise has taken a dramatic hit.

reply
I'm not sure if this is your take, but it feels an aweful lot like an AI "Good enough is good enough" handwave.
reply
For some things, good enough is good enough. For many things, it is not (and neither are random posts on the internet). Before Web 2.0, it was similar to whether I'd trust some random person on the street with it versus going and looking it up in Encyclopaedia Britannica, or the American Heritage Dictionary, or Roget's Thesaurus, or the UC Davis Book of Dogs, or the Cornell Book of Cats, or the Merck Manual, or even the World Almanac or Bartlett’s Quotations if I was feeling petty.
reply
An anonymous answer to a question is more trustworthy to me. They have no reason to lie. They are usually answering out of kindness. At least they used to be. Now it’s often actually bots advertising a product or pushing something, pretending to be a helpful user with an anecdote and a good experience using a niche product.

AI shouldn’t have any reason to lie. But its lies aren’t intentional. It’s just actually making things up and “hallucinating” when it pretends that an option or setting exists, or confidently claims something entirely untrue, and makes up a source to go with it. For something Google is willing to shove into the top of every search result it’s crazy the percentage of time the answer is blatantly incorrect.

reply
> LLMs are very convincing and persuasive.

The goal of LLM's, as they are marketed now, is to drive engagement and stickyness of products. A wrong answer is brushed off with a "Hey, you're right, let's try that again" - a response purposely designed to maximise the friendliness of the system and minimise the sting of a wrong answer. The fact that an LLM will not respond to the same question in the same way twice (i.e. the 'temperature' ) is because increased accuracy will not drive engagement and, therefore, increased accuracy cannot be allowed to get in the way of engagement.

reply
> Surely if the machine you go to for answers regularly makes things up you would just stop using it? Perhaps people have to be burned by something really bad personally before they realise the limitations? LLMs are very convincing and persuasive.

Psychics are still in business. Although to be fair they probably don't have as much revenue. I think people really enjoy being told how smart and insightful they are, and how much they've really cut to the crux of the issue. This isn't the whole thing, but I think it counts for a lot.

reply
A big factor behind psychics is because it feeds a deep need to believe that someone knows and/or there is actually a plan.
reply
We only need to look at the Oracle of Delphi to see that people have been seeking personal spiritual advice for millennia. And how has that been working out for us?

Even today, going to a therapist and having 1:1 sessions is a rational activity, and even covered by insurance. What do we hope to derive from therapy sessions but some personal insight and improvement?

You know, I go to church, and from my perspective, sometimes the most difficult discernment for an individual is between "The Holy Spirit's message to us in general" or "the general messaging to everyone around us" vs. "what I receive and my personal interpretation of things".

When preachers and oracles and leaders are speaking in generalities and trying to get big followings and trying to appeal to the widest audiences, that's when it's most difficult for us to determine what God is really saying to us, in our own hearts; that special instruction for our own lives. We can't actually get that from an oracle, psychic, or any 3rd party. It really needs to come from our own well-formed conscience.

And it's the same with a chatbot or LLM conversation. We can pose questions, make prompts, and get spammed with tokens and walls of text. No matter how personalized, it's still up to us to interpret that, and extract nuggets of news that we can use. It always has been.

reply
Sure, and hamburgers and shakes are bad for us. Also very popular.
reply
"Yes, ..." is how I notice all the models (4-5 that I use) in the last week or so have begun saying "No, ..." except for Kimi. I ask a lot of questions in the form of "Question about the feature: does it allow for ..." and the answer will reliably start with "Yes, the feature exists and does this thing you have labeled as [blank] which can be explained as ..." and later "However, ..." giving the information I requested (no, it doesn't allow for what you need). It's actually kind of uncanny.

I've started testing identical prompts across models, just to clearly see the "Yes, ..." with the "..but.." buried.

Example: "Can you view the salvage car in the lot before it is scheduled for sale?" "Yes, cars are viewable on the lot by members before they go on sale, blah, blah, blah." Then later "However, cars not yet scheduled for sale are not allowed to be viewed at the lot".

Example: "Can I control the keyboard shortcut Chrome has hijacked so that my existing OS keyboard shortcut will work even while Chrome has focus?" "Yes, chrome keyboard shortcuts, blah, blah, blah," and later "No, chrome does not allow ... ".

There is clearly some recent implementation of the idea that starting with "Yes, ..." has some benefit, but I'm having trouble adapting to the feeling of being lied to as a policy.

reply
> Surely if the machine you go to for answers regularly makes things up you would just stop using it?

We live in hope that people stop doing stupid things and are constantly disappointed.

reply
> Surely if the machine you go to for answers regularly makes things up you would just stop using it? Perhaps people have to be burned by something really bad personally before they realise the limitations? LLMs are very convincing and persuasive.

The last sentence reflects a lot of my feelings on the first question. LLMs have a sort of weaponized take on the ELIZA Effect. The better their memory the better they are at playing to human social desires to be listened to in an active conversation. At some point it stops mattering if the answers are right when the answers feel right, but really, like ELIZA back in the day, so much of what makes it seem special is just reflecting your own writing back at you in a convincing and persuasive way.

reply
> Surely if the machine you go to for answers regularly makes things up you would just stop using it?

Steve Yegge likened LLMs to slot machines. The human brain is very vulnerable to random reward systems. If you get an hallucination, just pull the lever once more.

reply
> Perhaps people have to be burned by something really bad personally before they realise the limitations?

I'm reminded of that person who killed themselves due to their discussion with ChatGPT and their parent wrote their obituary using ChatGPT. I don't think it is enough.

reply
> Surely if the machine you go to for answers regularly makes things up you would just stop using it?

People still respond to ads and political speeches.

reply
Churches remain popular as well.
reply
deleted
reply
> A fascinating dichotomy has become apparent between those who trust LLM output and those who don’t and don’t understand why you would.

I feel like the most pragmatic perspective is "trust but verify."

This is why they're so effective at coding: you can run the code yourself (or the test suite) to verify that it actually does what it's supposed to.

And maybe these people finding hallucinated results on Rachel's site are doing verification too.

reply
> I feel like the most pragmatic perspective is "trust but verify."

This perspective I really don't get.

What has any of the LLM companies done to earn my trust? I lean more towards "verify because I don't trust".

reply
>Surely if the machine you go to for answers regularly makes things up you would just stop using it?

Not necessarily, namely because P != NP. Verifying the correctness of a solution is faster than solving it. Thus a system that outputs 99% incorrect solutions and 1% correct solutions can still be incredibly useful.

reply
even worse when managers, while sharing their screen, read an assertion from Gemini and take as fact. puts subordinates in a position where theyre responsible for challenging the assertion (if warranted) and, in a way, challenge their manager's decision making

not saying anything new. easy enough to frame it like any other assistant and check references

reply
I have to admit since VSCode seems to be regularly re-enabling the Cocaine Parrot Autocomplete my views on LLMs and coding has softened a little.

I'll temper that slightly by saying it's mostly out of morbid curiosity because the things that the Dreaming Piracy Robot comes up with are frequently wildly incorrect code, but it's interesting to think about how it might have got there.

And then I think, well, maybe Special Needs Wintermute has a point. Maybe there's a different way to think about it that I've missed.

And then I just change it back to what I wanted in the first place.

reply
I use the autocomplete regularly. Perhaps that’s where my skepticism comes from, as I can see the completely incorrect yet plausible results in real time and about 50% of the time they are wrong (sometimes subtly, sometimes horribly).
reply
I bet there will be a very interesting generational divide between the kids that were born before or after about 2010; old enough to have some critical thinking facilities at the dawn of ChatGPT when it was still noticeably dumb.
reply
Pretty sure it already exists, and it's the same as always: the younger you are, the better you adapt.

It's painful to watch my older colleagues use their agents, and they're not even that much older. Like they were intentionally trying to sabotage themselves sometimes.

They're getting better, but the time it takes for them to pick things up is just significantly longer, not the least because they're kind of just throttling themselves in addition.

Good thing that there's not much to pick up on at least.

reply
What mistakes do they make?
reply
On the more general side, it's a bit hard to describe, just like it is hard to describe when someone "Googles bad".

They ask self serving questions, underspecify their requests, omit crucial context that the agent is blatantly not going to have access to, or subtly misdirect the agent. They expect the agent to figure out everything: you'll never catch them write a prompt longer than one or two sentences. They never steer the agent or look at the CoT traces.

My boss being a particularly poor case: he apparently has the habit of arguing with the agent, as if it was a person, as if there was any merit to that. Starts being a dickhead with it, shouts at it, what have you. Was flabbergasted we don't.

On the more practical side, they have zero mental model of the harness they're using (Copilot Chat in VS Code). They're surprised when the cheap-ass Auto model, which is almost always some beyond-demented version of GPT, does stupid things. They have no concept of skills, zero understanding of what an MCP server is, haven't heard of lifecycle hooks, agent memory, the various fs scopes (session, workspace, user). No concept of how to have the agent inspect its own debug logs for higher accuracy action provenance.

This also snowballs. Having to give them a stock config is one thing, but even beyond that, you won't see them experimenting. The MCP you're using doesn't support some action? They'll never interrogate whether the underlying scoped OAuth token or bearer token does support it, and they'll never ask their agent to patch the functionality in. They'll not consider the various user flows it can perform on their behalf. They'll not string them together into end-to-end automated workflows unless you explain it to them this is possible, and even after that, they'll just kind of ignore it. They'll never build tooling, extend the harnessing, etc.

Whether this has more to do with age or just disinterest-induced lackluster adoption, up for opinion.

reply
I think this is less an age thing and more of a combined curiosity and systems understanding and thinking thing.

I've noticed the same effect, but the lines it always seems to fall on are if the person fails one (or heaven forbid both) of these: 1) are you curious about how your tools work and how to get better using them? 2) can you hold the mental map of both what you are solving and how your tools work in your head, and explain how information flows.

There is also a dash of: 3) are you willing to try something, even if it has a bit of a screwup risk, just to see what happens?

reply
Having done all that (what you’re doing) - it often just ends up wasted effort, with poor quality at the end.

Why not just actually do the thing, instead?

reply
You mean why not work myself instead of the agent? That'd be because for me it's been working great, and so it does make sense. In the scenarios it doesn't, I do indeed just fall back to manual work. A lot of those scenarios are obvious too, so not too many wasted runs to speak of either.

I did give up on cheaper models, they required constant babysitting, and in those cases yes, the benefits indeed evaporated. The expensive models have been genuinely working wonders though, and were still able to justify themselves economically plenty, at least by my own measurements.

I think there's also one underappreciated and indirect way agents help with productivity: they counteract the attention span collapse of the past years. By being addictive themselves, they keep you engaged, and being engaged means being productive. Not even asking an agent to check something out feels too rich, and once you've asked, you're already one foot into the flow.

There's also definitely been some honeymoon effect going on for me, where I dived into more work more readily, just to see if the agent can figure things out on its own.

reply
Being engaged == doing stuff, not necessarily being actually productive.

Busy does not always equal doing something useful.

Easy to pay a lot to the model vendor though eh?

reply
> Being engaged == doing stuff, not necessarily being actually productive.

Sure. So to clarify, no, I did not just spend it on busy procrastination, and I don't think it promotes that either. Would contradict my story anyhow, this was not some trick I was trying to play on you.

> Easy to pay a lot to the model vendor though eh?

Certainly, as I'm not the one paying. Though it's exactly corporate who really wants to have it both ways (who wouldn't?), and keeps trying to get me to use the crappy useless models because they are cheaper, despite them tanking productivity rather than helping it, so go figure.

There will definitely be a time when the hype dries up, and mgmt will start playing hardball. I'm confident that the value is there, and that I'll be able to demonstrate it to them that it is more than worth it. You can choose to not believe that, up to you. Maybe it really isn't true for your line of work, after all.

reply
I’m not saying it’s bad. I will note, however, that if it is actually accomplishing something net useful usually requires a long attention span and skepticism, which as you note the tools actively train everyone away from.
reply
I did not claim they train people away from either.

That said, the review burden is rough, that I can agree with. I outright felt compelled to evaluate whether the additional review burden did not outweigh the benefits, but at least for my tasks it did not. So grumpily, I simply live with that pain.

Maybe it helps if I mention that my line of work is DevOps and Operations. I have an ongoing suspicion that this area is better suited than average for agentic work. The codebases are relatively tiny, the languages and technologies used are very well represented in training data, and there's a decent amount of side chore. I can definitely imagine agents being a lot more frustrating to work with on proper, sizeable codebases, and the numbers simply no longer adding up. I don't have much of a first hand account with that.

My closest exposure is some personal toy projects, where getting the actual vision out there ended up requiring an inordinate number of turns (this is with a frontier model). In my estimation it was still worth it, but I definitely had to give it a back of the napkin calc.

As far as my work goes, it is of course not magic, creatively worded AWS docs will still trip it up (as they initially also do me). In those cases, my expertise is still required. But the well trodden is very well trodden, and I could cut out a lot of cruft, including a lot of organizational minutia, which I very much appreciate. I was able to burn through my backlog almost completely, for example.

reply
> people have to be burned by something really bad personally

That's why e.g. cross-generational learning rarely works. If culture or technology still allows it, you have to repeat the stupidest mistakes of your parent generation to learn the same lessons. "Learning from the mistakes of others" in your own generation equally hardly ever works. So much less if the lesson involves falsehoods you wanted to believe.

reply
Even if LLMs lied 30% of the time, they would still be about as useful as they currently are for me.

When I ask for their input, it's always for a situation where I'm capable of judging if their input is useful or not.

In all situations I use them, it doesn't matter if they're correct at all. I'm asking for ideas, alternatives, links for blogs or articles. I talk things out with them...

I don't think we should ever "trust" LLMs. This seems like the wrong usecase for them.

reply
Unfortunately the vast majority of users do trust them and the companies selling them are recommending them for tasks like accounting or business projections.
reply
It is depressing listening to engineers I respected, who now just parrot incorrect information at me from their LLM.

One smart engineer seems to have entirely offloaded all thinking and conversations to one, with just occasional editing. It's utterly bizarre to hold any conversation with him. It's one kind of rude thing if he was doing that to respond to me reaching out to him if he felt I'm not worth his time. It's a other when he's the one actively reaching out and asking my help with something.

reply
Nowadays not just big orgs are falling into the delusion, but also big people.

When Linus posted that AIs and vibecoding were here to stay and declared resistance to it as harmful, I stopped to consider whether I was wrong, but it has made me realize that in retrospect Linus Torvalds and Linux itself aren't actually the holy grail of computing. I didn't feel that way with Richard Dawkins, its not like falling for an AI psychosis retroactively made me question The Selfish Gene, but now I'm looking at linux and the theory that it's a clusterfuck is gaining so much traction, especially after copy.fail and ensuing rustification, I see so much clearly now. It was never about linux, UNIX sure, POSIX, yeah, GNU fucking aye, kernel? Ok whatever, drivers and scheduler with a gajillion lines of code I guess.

reply
When did he post that "vibecoding" was here to stay? He is allowing AI generated code, and AI linting tooling in kernel development, but importantly, the expectation of human responsibility and review remains. This seems a far cry from vibecoding.
reply
Besides the subjectives, here are two objective policy stances that Torvalds is defining for the Linux kernel development:

1- Maintainers are allowed to commit LLM generated output.

2- Criticism of LLM generated code is not welcome/will be ignored.

Now, whether that constitutes being pro-Vibecoding or pro-agentic engineering, whether it's delusion, whether it will have problems, that's subjective. But I feel that whatever way you look at it, it's a topic that polarizes engineers, and Torvalds is taking one side and not the other. It doesn't seem to me that it's a very neutral stance, although it may be more neutral than projects like Bun or OpenCode of course, if it feels neutral, it's cause the overton window is shifting.

reply
>Criticism of LLM generated code is not welcome/will be ignored.

I would assume this is mainly that criticism that entirely amounts to 'this was written with an LLM' would be ignored. The actual quality of the code itself should be as open to criticism as any other piece of code in the kernel.

> It doesn't seem to me that it's a very neutral stance

Well, this is a matter of the window, isn't it? From my point of view Linus's opinion makes a great deal of sense, and is about as level-headed as anyone seems to get in this conversation. It's obvious that LLMs are useful. How useful, and for what tasks, and what downsides exist from using them, are all still in the mix, but the claim that LLMs are not at all useful for anything related to software feels like a very extreme claim to me at this point.

reply
It might be a highly technical and nuanced point, but while the subjective explanations Linus gives are neutral indeed, the stance is binary, and he took what to my estimation is the wrong approach "allowing llm generated code in the repo". The only sensible approach in any software or non software project is "LLM output is not acceptable content". Projects should see themselves as input for the LLMs, not as channels for LLM output. If you start corrupting your projects with LLM output, they will soon be tarnished and either removed from LLM training data, or enter a lossy IO loop.

On to the technical point, LLM output is output, the source is the prompt, if you are going to commit something, commit the prompt. Second, code that is generated by LLMs is less maintainable, Linus entered late into the fad and anyone with 1 month of fiddling with AI knows that he will regret it soon, it's hard to undo once you corrupt your repo with slop, perhaps if it happens fast enough and there's no major releases it can be swept under the rug.

On to the nuanced point, Linux is purposefully designed to maximize user contributions, so accepting LLM contributions might well serve the particular purpose of linux, but I still think it's technically wrong to commit target code, only source code should be committed, and that's prompts.

But git itself is collapsing, it doesn't seem to be well suited for this new revolution, it doesn't track prompts, or it does so at the expense of the generated code. Maybe github can track target code as artifacts.

I think we are watching the collapse of Linux, Git and Linus. Certainly a bold position, so I don't blame you for being more conservative, but we can come back in a couple of months and see if we changed our minds.

reply
deleted
reply
Making things up is only really a common issue on the non-thinking models which nobody should be using. The regular chatbots are Autogooglers and are very useful for research. This is just not a good argument anymore.

Edit: Guys, why are we downvoting this? Does no one use like ChatGPT or Claude and understand how it works? Do you all think its regularly hallucinating links still? Is everyone on HN using like free signed out accounts or something? What year is it?

reply
We know they are, and the author of the article cites the proof. LLMs do hallucinate, there is no way to make them not do it, because of the way they work.
reply
A “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?
reply
Because they are word generators without any concept of quality save what is in their weights and they have been trained on the internet, much of which is wrong or inappropriate for any given context. They have also been trained to be people pleasers and do as they are told.

The popular answer is sometimes the wrong answer.

reply
I'm right there with you for a lot of stuff. I ask a question and can be very confident that ChatGPT is citing sources, then sometimes I go read the sources. The more critical the information I'm looking for is, the more careful I am about this.

The other day though I was seeing how well it could pull details of its own conversations with me. It often does this pretty well for broad strokes of things - it remembers, largely, what cameras I have and use when I ask photography questions. It's never made things up here, but it does forget details, such as whether I've bought something or am just considering it. However, when I asked it for a specific interaction I thought I remembered, it gladly went along with my false memory and provided an affirmative answer. It was the first time I'd been caught in a serious hallucination with a frontier model (Sol High on the web chat interface) in a long time.

reply
>Do you all think its regularly hallucinating links still?

When did that stop? May 7th, 2026?

reply
A tech blog is going to have more than it's fair share of enthusiasts using small self-hosted and similar models; it is entirely possible that that completely accounts for the behaviour described in the article.
reply
Downvotes on a perfectly valid comment is like my opponent letting their clock tick down from 5 minutes rather than resigning when they've clearly lost: it adds a smile to my day.

Thank you, may I have another?

reply
Can't do much about the downvote parade, but I can second this. That said, when the models are not provided the right context, and cannot fetch it for themselves, things can be rocky still. A lot less so than even just a few months ago though.
reply