Anyway, I think what they're saying is that if you train a model purely on generating language, it'll lack many of the things about human brains (and the human experience) that result in the language they generate. The process can matter more than the result for language (specifically for creativity and art, too), so it follows that a model trained purely on the output is going to be missing something more fundamental, even when it does produce coherent language. This matches up with my LLM experience so far. That's not to say anything about their value or utility, just that they're not the same.