I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.
Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.
For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).
They were never designed for machines or non-human animals.
Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.
It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.
"Please accept my resignation, I don't want to belong to any club that would have me as a member".
(Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)
IQ has precise definition. It's what the tests measure. Intellect has only fuzzy handwavey definition. Correlation between IQ and intellect is about as loose as the definition of the intellect. The way people make the correlation even looser is by defining intellect in even more fuzzy and nebulous manner.
Let LLM control a physical robot to perform tasks that average human can do.
Case in point: you had hundreds of millions of JPEGs to vacuum up and bitmap image generation is amazing. But if you ask them to recreate the same scene as vector art, they will struggle to generate a decent SVG. Like, kindergarten-style pelicans on bicycles are the state of the art. It should generalize seamlessly, but somehow, doesn't?
I think it will happen, just like self-driving cars are happening, but it will probably be a slow process.
Also, the skill of the human opponents matters. You'd want to test it against people who have practiced playing the game. Otherwise, it's like the difference between building a chess bot that can win against random undergrads who don't normally play, versus winning against grandmasters. And it's not like there's a pool of skilled human players of the imitation game.
No, it is just because they have difficulties at the bench.
> how gives we don't measure them by
We'd measure them by all the tests available. Not all test are usable in all circumstances.
I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments for good which is one way we define intelligence.
On the other hand saying "the gifted person is highly intelligent/smart just not good at some things" really diminishes the other tasks, because they really are very difficult tasks but are not measured by an IQ test.
Nobody says that the IQ test would "measure intelligence". We know it does test a form of it.
I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations
[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...
It spells out a form of intelligence - some can and some cannot.
Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.
LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.
So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.
This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs
Not fully relevant: timing is crucial in all-pass tests, not crucial in pass-or-fail tests. I.e.: first of all, they have to be able to reach the goal, and that is already an achievement. Then - and in parallel - the problem solving must also be optimized for efficiency. But "solving" and "efficiency" are non coincident dimensions.
Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.