If we would know that, there would be no need for interpretability research.
I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?
> maths are internally coherent and entirely theoretical
Nope. This kind of wish-washy thinking is not what we mean by understanding.
This is like saying we don‘t know how a car works because a car can beat the best human athletes in 100 meter dash.