upvote
deleted
reply
They're not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions.
reply
One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.

Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.

reply
Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.
reply
... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.
reply
we already fixed it with reasoning
reply