upvote
My perspective is that the addition of thinking loops to models allows sufficiently advanced ones to approximate world models.

Incredibly inefficiently because of the recursive loops ("Wait, the object is on the table. I should think about this more deeply..."), and likely instantly surpassed by large world models if/when those are shipped, but effectively enough vs non-thinking models.

reply
LeCun calling them "world models" gives a high-level description of the desired functionality. They are Joint Embedding Predictive Architectures (with SIGReg). They might produce more useful world models, but it's yet to be seen.
reply
This sounds like a human trying to reason about quantum mechanics. We als simplify to newtonian for day to day tasks.
reply
I like this analogy. Both GenRel and QM are well beyond our experience, and although there is some intuition that comes from working with the equations over time, it is bizarre and "just calculate" often gets the correct answer faster.

Picking the right tool or model is like picking the right problem to work on. It's actually quite hard (often you can't just try them all), but without it you will be incredibly inefficient and occasionally, fundamentally wrong.

All models are wrong, but some are useful. -Box

reply
LeCun's argument wasn't about the definition of learning though. He stated that they would never get these common sense things correct because they weren't sufficiently part of the training data. A statement that we can hopefully all agree has been thoroughly refuted.
reply
As of a few months ago they still have trouble, with low thinking, at the "should I drive to a car wash that is 100 m away" kind of question.
reply
Simply appending “check your assumptions” to the question fixed it even back then: https://news.ycombinator.com/item?id=47040530

Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then.

The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment.

reply
Low thinking is an artificial constraint. It can fail spectacularly on things that aren't in the training data.
reply
It's a nonsensical question to ask, and how an LLM answers gives 0 signal.

If you were home and a family member asked you that question, you'd probably criticise the question rather than answering. LLM are RLHF'd into being milk-toast helpers that just try to answer questions like that with no criticism.

This is all beside the fact that the world of AI has changed pretty dramatically in the last few months.

reply
It is so nonsensical because it has such an obvious answer. The answer is so obvious, in fact, that one answer can be considered nonsense and the other common sense.
reply
*Milquetoast
reply
deleted
reply
This is just a stupid post.

It’s nonsense to test if a product that is marketed and sold as being able to provide generalised intelligence on demand, does what it says on the tin?

Check yourself

reply
Since you're new here, I'd suggest you read the guidelines for etiquette.

https://news.ycombinator.com/newsguidelines.html

reply
It's very unlikely that person is either new or unfamiliar with the guidelines. They almost certainly created a throwaway account specifically because they know the guidelines and want to flout them without consequences. (It seems like there has been an uptick in the number of these kinds of throwaway flame comments. I wonder if HN tracks that?)
reply
nothing indicated otherwise at the time. IMO he just underestimated RL-scaling. chinese models improved a lot too, they are not parrots anymore, there's some real intelligence, at 27B params.

consider me optimist now, but just few months ago, even frontier models were dumb, doing stupid mistakes all the time, all of them were so dumb I'd never expect anything to change in just few months.

reply
I thought it was more because of fundamental limitations in the architecture. As in, no matter the training data, it could not be consistently and generally represented
reply
Actually, I think my fundamental challenge with AI is that it has no common sense. The way it builds things, writes, and operates is out of touch with reality.

Incidents like hugging face are partly rooted in the lack of common sense. It still functions like a supercharged toddler.

I'd love to overcome this because it'd mean I spend less time guiding the the LLM to produce usable outputs.

reply
> It still functions like a supercharged toddler.

And we've had difficulty as humans to childproof our sandboxes and infrastructure. Things that are otherwise innocuous spots to coordinate between like minded toddlers can become problematic.

reply
Last week I asked a frontier model draw me a backplane PCB and it placed daughterboard slots side by side in a chain.
reply
No?

This is always the issues in the discussions.

There’s the outcomes camp (objectivists?), which points at the things LLMs can do.

Then there’s the process methods camp, which talks about what is actually going on.

If you only care about the outcome, then the process does t matter.

If you are talking about what is happening, what the underlying mechanics and science of it is, then the process matters.

These models aren’t thinking. They simulate cognition well enough to do useful work in several fields and domains.

Both are true.

reply
I think where both camps get hung up is sometimes the process method group "ignores" the obvious outcomes and effectiveness of LLMs.

But the outcomes group "ignores" the fundamental limitations of models which are purely text based.

E.g, a baseball players trains to catch high-speed balls and they dont do it by: "ball velocity 50mph, vector:[1,2,3], run move hand command now"

That's absurd.

No, there is an embodied network which is "trained" on visual, tactile input, and control as direct output.

LLMs are fundamentally not the right tool for that.

reply
> E.g, a baseball players trains to catch high-speed balls and they dont do it by: "ball velocity 50mph, vector:[1,2,3], run move hand command now"

That is a NN that learns a skill.

But that is not an Analyst. If it were ballistics, then the answer to "how to parametrize the launch to reliably hit the target" excludes getting the result through natural skill.

The problem lies in the need to get "AI" facing "LLMs": the latter create a need for reliability, for "AI".

Speech is an endowment of both those who give educated guesses via developed skills and of those who return answers like Analysts, who check and compute. LLMs create a confusion between the two, and they will remain a problem until an ability to act as Analysts - strictly - will be implemented.

reply
> These models aren’t thinking.

They are for any definition of the word that makes any kind of sense. I'm sure you have a contorted definition that magically only includes humans though...

reply
> for any definition of the word

For "thinking" here we mean "assessing a representation of an object". That, or equivalent, is required to be reliable. So it is fundamental and critical.

reply
Sure? If humans happen to be doing something that LLMs are not, then should the answer change to accommodate your disdain?

The models are simulating thinking, if the fidelity is good enough for you - great!

reply
It depends on whether you assume that thinking requires doing everything that humans do. I think it would be silly to say that an AI doesn't think because it doesn't wrinkle its forehead in concentration. So you need to decide which parts of the way that humans think are actually necessary components of the process.
reply
Tbh it doesn't even matter if humans turn out to have a soul, or quantum microtubules or whatever other magic LLMs can't have.

The normal definition of the word "thinking" definitely includes what LLMs do. Hell people used to say computers were thinking even before AI. It's super weird to get all uppity about the semantics of the word now.

reply
> A statement that we can hopefully all agree has been thoroughly refuted.

Uh, no? So much of what we learn and take for granted as common sense is not learned via language, and not even expressible in it.

reply
To determine this, it would first need to be able to spell "raspberry" as letters rather than as tokens.

Given you also don't want it to memorise [for all tokens, count([for all letters]), this would probably be more like "here's two images, count all things in the big image that look like the thing in the small image", which can then be r's in a photo of a raspberry jam jar in a supermarket, or dragons in a photo of a furry convention, or whatever.

That said, they are competent enough at coding that I keep seeing them write code to do even simple tasks.

On a related note: why did I see Claude editing a file by using cat to write a python script to do a grep search and replace?

reply
> it would first need to be able to spell "raspberry" as letters rather than as tokens

Of any object in question they should be able to create a representation that allows correct assessment.

> Given you also don't want it to memorise

That is obviously necessary: what we want from the consultant is to check, not to remember. Answers must be correct and that implies having performed all due diligence - and being capable of doing it, before that. So, objects must be instanced internally in a way that allows effective handling. Counting letters is a good example of the ability (that must remain general).

reply
> Given you also don't want it to memorise [for all tokens, count([for all letters])

Why not? You've memorized how words are spelled, and how sounds correspond with letters, and how concepts correspond with words. To the extent that there are shortcuts that enable compression you use these, and the model will do something similar.

reply
Combinatorial explosion, and facts merely memorised is a huge waste of parameters that are better dedicated to effective reasoning. Not that we really know how to split facts from skills, though we are trying various approaches.

Being able to spell all the words then count letters is simpler, and more generalisable to other tasks, than memorising answers to all possible word questions.

That said, we're so bad at splitting facts from skills that trying to get them to memorise a bunch of facts might force them to learn a skill and generalise anyway.

reply
Ah, I misunderstood what you meant. I was just trying to highlight that in order to answer these types of questions the model needs to memorize the spelling of each token. But you're right that that's all they need to memorize, and algorithms like counting are pretty simple for transformers to implement.
reply
> Why not?

Because to "123x456" we want a reply that goes "this times that plus that...", not "Was that not nnnnnn?". If it does not perform its duty (returning solid checked answers) it is a liability.

reply
counting 'r' in 'raspberry' to the LLM is similar to 4-dimension space to human. Their world's unit is token, not character, although they could use indirect method such as "run code" to find out. It will stay that way until they change the fundamental of the token that the LLM can perceive characters.
reply
I hope you understand: it is a core point that systems that answer questions must have the ability to internally represent the objects they assess in a way that allows reliability. Whatever the object.
reply
I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable.

https://huggingface.co/posts/omarkamali/593639295164067

https://huggingface.co/blog/omarkamali/tokenization

reply
Careful: the problem is very certainly ___not___ counting letters. That is only a telling way to check "is the NN checking or not?". We demand that NNs for consultancy tasks check, strictly.
reply
I used to think byte level tokenization was the answer, but humans also think at a word level and only reevaluate the words at a character level when asked. The solution to better tokenization across languages is likely to be learned tokenization. Here is one attempt I have seen: https://github.com/SamD770/bitter-lesson-tokenization
reply
How many 'r's are there in the next 30 seconds of this [1] song?

[1]: https://youtu.be/l7vRSu_wsNc?si=SndkB6GBaRyhvNNA&t=61

reply
deleted
reply
It's not even fair to call "run code" to be indirect compared to what a human would do. The word raspberry has no Rs in it in human language either. We have a written representation of it, which we can then write down either in our head or on paper, and then we can "run the algorithm" of counting each of the letters.

Nothing intrinsically more or less direct about the LLM's method than ours.

reply
I could argue LLM only have "token" as their perceivable dimension, compare to human multiple senses as the physic perceivable dimension and a brain with many other dimension of "learning" and "thinking". In spoken language, we may not have 'r' but in written we have, both spoken language and written language are learned skills.
reply
You could argue in return that humans only have electro-chemistry as our one perceivable dimension. We only indirectly perceive light through the signals our eyes send to our brains.

In my mind general intelligence is pretty much by definition a virtual machine, so the mechanisms behind thought are only relevant for the sake of efficiency (ie you can argue that LLMs make a poor basis for intelligence because tokens and natural language are a poor way to encode the world, but if you can run it on a big enough computer to counteract the inherent wasteful virtualisation then who really cares how it works under the hood?)

reply
So LLM and human all have 1 dimenion perceivable signal, just LLM is 240p, and human is 8K in resolution, that's why we have 'r' in our signal, LLM still have 'r' in their signal, just because of the "low resolution", raspberry wasn't encoded with so many 'r' as in human signal.

I will stop here before our analogies go too far.

reply
Is "token" a directly perceivable unit for the LLM? If you ask it "how many tokens are in this sentence?" can it count them (again, not guessing or making a tool call)?

I've never tried it and it might take some thought and effort to conduct an experiment to find out properly, but I would be interested in the answer.

reply
I dont think so. This is akin to asking a person, what is the frequency of the light hitting your eye when watching a leaf for example. You either know the (approximate) answer by knowing the frequency of green, or use a tool to measure it. If the LLM gives the correct answer it is either.guessing based on intution(and this intuition is based on pairs of word to tokenization length in text form in training data), writing code(or executing a tokenizer) or running a tokenizer mentally (reasoning via CoT).
reply
Not the point: the simulated intelligence in this context needs to create proper representation. It is not a matter of what it sees but of what it can see.
reply
Can you tell me what is the exact frequency of light hitting your eye as you read this comment? Not by guessing, not from knowledge, but from actually counting? No? Then you are not generally intelligent :)
reply
Justify your statement (the other similar post nearby is not sufficient), or realize that we are not talking about that.

We can have adequate representations of light that are the instances over which we reason. Your simile is about perception, not about instancing ideas.

reply
All the frequencies, in varying amounts. Next question, please.
reply