upvote
This really doesn't match my experience. I can ask an LLM to modify engineering plans using vague natural language prompts and it will find the right place in the plan from the description and then make appropriate modifications, which necessarily requires doing spacial reasoning.
reply
Or it’s just taking common examples from training and applying those copied heuristics to your problem? Doesn’t mean it’s actually reasoning about the space and how to solve the problem. It’s the equivalent of a student writing an answer they saw somewhere else without understanding “why”.
reply
Why does CoT significantly improve their performance? Most of what they are applying they learned in post-training by solving similar problems themselves. This isn't about regurgitating pre-trained knowledged.

Advancements in Math and coding are because RLVR at massive scale is so cheap.

reply
Looking at the reasoning traces it sure seems like it's reasoning. It internally debates which of the possibly matching parts of the input are the one described by me in the prompt and picks the right one based on sound reasoning.
reply
I mean, it's a text predictor. When you say <BEGIN_REASONING>, you'll get reasoning-like output next, whether or not the model is capable of reasoning.
reply
It's not just text that appears at first blush to resemble reasoning, it's actual sound reasoning. And it can chain it for hours at a time without breaking down.
reply
I think the point here is that LeCun was arguing that training on pure text would not grant spatial understanding. I believe most models are trained on spatial data as well, so you are both right.
reply
Is that what he meant? He works on models with an explicitly spatial internal representation, whereas I was using a standard LLM that edited the provided plan by using a bajillion python calls to inspect small regions of the image at a time and then generate edits.
reply
I've seen recent AIs make detailed and technically impressive 3D models. You might argue "they're not doing spatial reasoning, they're making measurements with code and doing math to configure relative positions". Fine, but at a certain point that becomes functionally indistinguishable from spatial reasoning.
reply
I don't know - real-world tests leave me unconvinced: https://youtu.be/ENWVpqtOdRI?t=867
reply
> when it can't be derived from the training data

This sounds like a goalpost on wheels. Can you define clearly where your stake in the ground is?

reply
we have benchmarks proving it can do spatial reasoning.
reply
reply
I’m not sure what you are trying to say
reply
When put to the test in real-world environment, the capabilities don't look as impressive as benchmarks and synthetic tests might indicate. So doubts about actual spatial reasoning capabilities remain.
reply
I see. Gary Marcus said that AI won’t be able to make a coffee in any arbitrary home kitchen.

I think it’s a good test and I think LLMs will reach it in 3 years. Current benchmarks maybe slightly incorrect.

I’m happy to make a 4:1 bet in my favour that I’m correct about the kitchen bet.

reply
I cant make coffee in an arbitary house kitchen. People tend to put stuff anywhere but where at look for them...
reply
Your prompter need merely to say “keep going” each time you report that you haven’t found the grounds yet.
reply
the statement was "can't do"

"doing poorly" is still doing

reply
Careful, you'll have them accelerating their goal posts up to the speed of light if you keep questioning them. That much kinetic energy is dangerous.

Of course this is the same reason LeCun holds very little sway with his words for me, they seem to be terrible predictors of the future.

reply
Aaah, the old benchmarks maxxing argument, having precise and clear definition of what "spatial reasoning" is, what, and most importantly WHY, the benchmarks of choice are would settle this debate, otherweise let's not delve into it.
reply