upvote
Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along.
reply
One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.

Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.

reply
Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.
reply
... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.
reply
> ChatGPT will say there are two Ds in "your mom" and one D in "uranus"

… Isn't it possible that it understands the innuendo and is going along with making the joke?

reply
Why is this getting downvoted? Is it not a reasonable question? I was wondering the same thing. Both sound like jokes to me. If the LLM is trained on text, including internet comments, how is this outlandish?
reply
How many LLM users have anything in their prompt against "going along with jokes"? I'd guess not many.

What a wonderful new world.

reply