> You can make any modern LLM explain its reasoning
You can make any modern LLM create a plausible, self-consistent explanation that looks like reasoning, but it's not "the reasoning it used to arrive at that answer".
We often make a decision based on a gut feeling, and then backfill a logical reason supporting our feeling, without even realizing we're doing it -- rationalization.
When you ask people who are rationalising poor behaviour about the scenario, but it is someone else doing it, they may arrive at a better answer. Can we use multiple LLMs to achieve self criticism and critical thinking?
To your point, I agree that nominally there should be a way to give conceptual names to paths of weights, and when answering a question, notice which weights were and were not applied and retrospect on that.
That's not what reasoning traces as they currently exist are, though.
Like asking a human "how did you catch that fast ball coming at you?"
What makes you so sure your own brain doesn't work the same way?
Quite ironic given the topic. It seems that the author’s model indeed contained too much knowledge about old Gemini releases, and did not do enough tool calling.
>Why? If a model's weights claim that Bart Simpson became President in 2020, why does this fact suddenly become uncheckable?
Because in one case you have a source you can use to validate the fact, and in the other you don't. Though, as you explain earlier in your comment, the premise is misguided/hallucinated.
>The internet is full of wrong information and I cannot magically edit it to make it all correct, so this doesn't help me.
My favorite RAG experience was asking Bart (or whatever they were calling Gemini back then) an answer to a question I knew.
It gave me the opposite of the truth (as was common with LLMs at the time).
But weirdly, it had cited sources for this "fact."
I checked the sources. Two of them, both AI SEO slop.
In this moment, andai was enlightened...