upvote
How much experience do you have with LLMs exactly? It would be consistent with my experience if Claude stuck in a line of python that just emits a JSON literal with no justification, potentially buried in a large program where an untrained person might not notice it. I don't even trust them if the output consists of structured data paired with source images from the PDF, because I've experienced LLMs fabricating the source rectangles to match the output. I only use tools like this by asking for programs, because as you note LLMs are good at that, and the verification process consists of tool calls to legitimate PDF manipulation tools so I have some confidence everything is above board. Even then I only do this for hobbies, not anything that matters.
reply
Lawyer here. I used to trust Claude as hallucinations are near non-existent now. However for large volume tasks such as due diligence exercises, they still happen. We also tried Legora's tabular review, there were also numerous halucinated provisions in our due diligence exercise.
reply
Junior associates hallucinate too...
reply
And when they do, you can train them or fire them, and they learn not to do it.

LLMs change not a whit, and there's no one to take responsibility for the failure (and thus no way to fix it).

As the new variation on the old theme has it, "A computer can never be held accountable, and so very many people are trying to get them make management decisions."

reply
You can’t train people to never make a mistake, particularly when doing highly repetitive work like this. You must build your systems to account for that regardless.
reply
You can scold juniors and they will learn. You can't scold Claude.
reply
Surely the rate of improvement in new LLM models is the equivalent mechanism?
reply
You can scold Claude. Just doesn't make a difference.
reply
... until you hit Claude's risable "model welfare" protection.
reply
That's a recipe for disaster in my experience. I tried it (with Claude) on a simple tabular bank statement PDF, and it transposed two amounts, placinh each against the other's description. And the bot assured me the result was cotrect. The chance of a human checker catching such corruption is low.
reply
Interesting. Did that PDF have a text layer or did you ask Claude to OCR it? If the latter I'm not surprised at all.
reply