From https://epoch.ai/latest/announcing-frontiermath-erdos
> Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours
> Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.
Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.
"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.
Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.
It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.
There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems
Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're past that now). New models will exhibit something novel by adding solutions. And the matter of solution is interesting as well (contradiction versus a positive proof).
I would venture to predict it'll take years to decades to get to 0 open problems. But I'd be very happy to have this comment look foolish in retrospect as models continue to improve
I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.
Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.
For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).
They were never designed for machines or non-human animals.
Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.
It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.
"Please accept my resignation, I don't want to belong to any club that would have me as a member".
(Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)
IQ has precise definition. It's what the tests measure. Intellect has only fuzzy handwavey definition. Correlation between IQ and intellect is about as loose as the definition of the intellect. The way people make the correlation even looser is by defining intellect in even more fuzzy and nebulous manner.
Let LLM control a physical robot to perform tasks that average human can do.
Case in point: you had hundreds of millions of JPEGs to vacuum up and bitmap image generation is amazing. But if you ask them to recreate the same scene as vector art, they will struggle to generate a decent SVG. Like, kindergarten-style pelicans on bicycles are the state of the art. It should generalize seamlessly, but somehow, doesn't?
I think it will happen, just like self-driving cars are happening, but it will probably be a slow process.
Also, the skill of the human opponents matters. You'd want to test it against people who have practiced playing the game. Otherwise, it's like the difference between building a chess bot that can win against random undergrads who don't normally play, versus winning against grandmasters. And it's not like there's a pool of skilled human players of the imitation game.
No, it is just because they have difficulties at the bench.
> how gives we don't measure them by
We'd measure them by all the tests available. Not all test are usable in all circumstances.
I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments for good which is one way we define intelligence.
On the other hand saying "the gifted person is highly intelligent/smart just not good at some things" really diminishes the other tasks, because they really are very difficult tasks but are not measured by an IQ test.
Nobody says that the IQ test would "measure intelligence". We know it does test a form of it.
I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations
[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...
It spells out a form of intelligence - some can and some cannot.
Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.
LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.
So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.
This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs
Not fully relevant: timing is crucial in all-pass tests, not crucial in pass-or-fail tests. I.e.: first of all, they have to be able to reach the goal, and that is already an achievement. Then - and in parallel - the problem solving must also be optimized for efficiency. But "solving" and "efficiency" are non coincident dimensions.
Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."
Well I dont know about all of you, but I am celebrating meat based humans...
I wouldn’t put it past a company like OpenAI with a long history of lying and being deceptive to record the tests and benchmaxx ARC. They have trillions of dollars of incentive to cheat any way they can.
none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"
I’ve ignored it thinking it would go away, but it keeps coming up.
I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.
Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.
But you’d be no nearer to solving consciousness.
Given this thought trajectory - what is AGI supposed to be?
Intelligence, and hence AGI, doesn’t require consciousness or emotions or sentience.
What human intelligence do you mean? Genius? Professional? Educated? Random person? “Dumb” person?
They are still different concepts of course, but I imagine that once one is achieved, the others aren’t far off.
I’m just trying to understand the implications of the current frontier model capabilities.
"highly autonomous systems that outperform humans at most economically valuable work"
https://time.com/article/2026/08/26/openai-sam-altman-interv...
We cannot in a declarative sense define what is economically valuable work even now let alone into the future.
People take what they can get for pay. Very few individuals can demand a wage. The value of employment is obviously designed around that, not some arbitrary definition of “valuable”.
Of course an AI will accept $0/hr, it doesn’t mean it does the job.
Anyone who could accurately define the value of work would be wildly successful without having to try.
That is not a useful definition for me unfortunately.
Do you want to take a stab at defining/quantifying it? I'm s afraid anything specific you can come up with will also be useless.
How can you conflate the two.
> solving consciousness
We are very much not interested in that. We just need a proper problem solver.
For example, how do you know that “feeling pain” is not a functional prerequisite for a task. And that consciousness is a prerequisite for feeling pain
I wonder at what point consciousness is necessary… that is, if you can have anything like that without it.
To the point that solving consciousness (and combining it with intelligence) is what gives you the autonomous, recursive, self-improving thing otherwise it can only drive in the dark and make big mistakes.
To your point I think - it’s why we don’t see too many non-conscious advanced biology (it rarely survives against those with it).
Examples:
- predict a coinflip: easy to verify, hard to learn
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
but we both know these examples go against the spirit of my point
also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye
Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)
Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.
Alignment++
There is no incentive for OpenAI to subsidize is you if no one reads /reports on your benchmark . They are only going to fund a few that are currently popular .
Community acceptance doesn’t automatically mean the best , it is combination of some level of technical quality and the ability of the promoter to socially influence or get support of influencers .
Prediction:
We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.
TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".
(I have not gone down the rabbit hole to understand how they achieve that 24% number)
If government don't step up and regulate outsourcing like 100% tax, big problems are coming up.
IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!
It's already happening :)
The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).