upvote
Take it from the mouth of the creator of ARC-AGI:

When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.

reply
I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?

I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.

> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.

- Sam Altman on AGI

reply
I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
reply
I can call coworker right now and have conversation so frustrating that I wish I was talking to machine instead.
reply
Haha, sounds like a median human being alright!
reply
deleted
reply
[flagged]
reply
I feel like my odds are better with the AI than with random humans.
reply
Random humans don't have a significant cross section of human knowledge available in real-time, although many like to pretend they do, especially in internet comments :P Being able to compete with the capabilities of a median human would be an absolutely world changing achievement.
reply
I think they're claiming it's achieved by text models, not voice models, fwiw.
reply
I dunno, have you tried the voice chat in paid ChatGPT?
reply
Have you? It’s a dumb model that gets a lot wrong. Codex voice is IMO only good for when I cannot dictate into the good models. The delay introduced by letting it work with a good model kills it for me.
reply
Here's another definition of AGI from Sam Altman:

https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...

Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?

Kevin Roose (New York Times): I probably would, yeah. Would you?

Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.

reply
If the new AGI benchmark is "be Einstein/Feynman" then we've hit AGI.
reply
What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?

The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.

reply
> What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes?

Then it’s an expert system.

Stephen Hawking wasn’t very good at folding clothes.

The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?

You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.

Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows.

I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech.

Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence.

reply
deleted
reply
The only important part of 'general' is the ability to learn from experiential data and update your own model. That's what leads to general capability. Humans can't oneshot any task natively, but we can practice for a while until we uncover often novel methods of accomplishing something.
reply
Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation).
reply
Most people can't draw a bicycle. There was an artist 10 years ago that asked people to sketch a bike, and then turned these sketches into 3D renders - quite funny.

https://mymodernmet.com/gianluca-gimini-velocipedia-bicycles...

https://qz.com/681345/an-artists-3d-renderings-of-bicycles-d...

reply
Laundry folding has become a doable demo for startups, and ChatGPT has been spitting out college essays for years.
reply
How good was Einstein at drawing pelicans on bicycles by writing SVG code?

Checkmate, meatbags.

reply
Pelicans are a solved problem at this point. An open-weight model on my own machine gave me this: https://crimson-jeri-74.tiiny.site/

And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans.

Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed.

reply
> Pelicans are a solved problem at this point. An open-weight model on my own machine gave me this

That feels kinda like when I remember seeing Ocarina of Time for the first time, and thinking “oh my god, this looks just like real life…”.

reply
For me it was the wheels. I couldn't stop staring at the wheels... how did it get them so freaking perfect? Mad respect to GLM 5.3.
reply
Wouldn’t that mean producing novel work like relativity and QED?
reply
I would maybe argue that Einstein was the most LLM-like of great thinkers.

A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.

A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.

That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.

reply
[dead]
reply
What is novel physics?
reply
I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own?
reply
This is as good a time as any to note that we might be closing in on a new conceptual revolution in our own time as it relates to holography and an information centric approach to spacetime. Obviously It's the furthest possible thing from a guarantee, but it has much of the enthusiasm and motivation that string theory had previously enjoyed in prior decades.

So it could be a natural experiment for whether AI can contribute to novel physics. Specifically, there's a big question about weather. Something like our informational understanding of black holes where information inside it is equivalent to information on its boundary (which I'm sure I'm not saying correctly), might be generalized to regular space-time. More people should be freaking out with excitement about this and perhaps it's something to which AI can contribute.

reply
Do you have anywhere you recommend where I can read more on this?
reply
Honestly I don't think I have single great article, though some Quanta ones are ok, and the Wikipedia article is okay.

The best thing I can recommend is what I did, which is ask Claude about the significance of (1) quantum computing error correction, and (2) error correction in black hole holography and research convergence between the two.

https://en.wikipedia.org/wiki/Holographic_principle

https://www.quantamagazine.org/how-space-and-time-could-be-a...

Edit: this whole article, despite it's boring title and hook, is maybe the best discussion of holography as a recent and active research frontier.

https://www.quantamagazine.org/if-the-universe-is-a-hologram...

reply
There was symbolic AI programs in the 1980’s that “discovered” Kepler’s laws and the resulting solar system model from just tycho brache’s astronomical observations. That was the the very first “new physics” ever.
reply
i wonder if we could train a modal, and omit all data prior to 1899, and see what happens?
reply
Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this.
reply
as always, as good as its prompt...
reply
I assume solving one of the major open problems of physics?
reply
Would this be possible without it being able to run novel real-world physics experiments autonomously?

(Note: I am not suggesting we let it do this. Please don't, in fact)

reply
AI could discover candidate novel physics without autonomously operating new physical experiments, and humans or instruments can later independently validate the result. This is analogous to how Einstein developed theories whose predictions were confirmed by experiments and observations only years or decades later.
reply
Do we have any examples of an current day AI system introducing a novel concept or perspective. We've got plenty of counterexamples discovered and some theorems proven, but afaik nothing analogous to a new definition.
reply
Why not?
reply
That's how we get AM.
reply
That’s already been done. I know of at least one novel result contributed by Claude to frontier physics. I’m sure there is more.
reply
For it to be like a human it wouldn't just need to solve existing phsyics problems, it would need to push the field forward and introduce new paradigms.
reply
Solving "open problems" will push the field forward.
reply
My comment wasn't very long, yet you somehow still ignored the main part, "and introduce new paradigms". The point is whether it can do everything humans can, entirely new theoretical frameworks and ideas, such as string theory or dark matter, are not coming out of AI at the moment.
reply
Creating new physics is the new AGI goal post
reply
If it makes you feel any better the curmudgeon who drives the ARC-AGI tests feels the same way, that's why we're on 3 and I'm sure we'll see 4. Also, we can all avoid calling it "moving the goalposts" so no one feels talked down to.
reply
It’s better to call a spade a spade.
reply
People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, but the difference really matters. These things still regularly fail, and gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".

An AGI wouldn't struggle with that.

reply
The last version to fail on those questions was GPT 4.5.

Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of many years.".

reply
It comes and goes... My point is we're not near AGI.
reply
Nooo. I just asked Claude (Sonnet 5 Medium) how many I's are in assassin, and it said 2. Granted, it got several words correct before this. But no, they still aren't great at this letter-counting thing.
reply
I don't think we should count the lower tier models if we're discussing what the top ones are capable of. No one was suggesting that Sonnet is AGI.
reply
Anyone know why they aren't good at this?
reply
LLMs see tokens, not words spelled out with letters.

Imagine verbally asking someone who has never seen written text the same question: unless they memorized the answer for the specific word you're asking about, they'd have to guess.

reply
People assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-r-y", where the r's are not combined into one token.

We don't know what was going on inside the closed source GPT models, but this paper investigated on some of the open-weight models and found it's not due to tokenization: https://arxiv.org/abs/2604.00778

reply
Tokens are the most basic input unit of an LLM. But tokens don't generally correspond to words or letters, rather sub-word sequences. So Strawberry might be broken up into two tokens 'straw' and 'berry'. It has trouble distinguishing features that are "sub-token" like specific letter sequences because it doesn't see letter sequences but just the token as a single atomic unit. 'Straw' and 'r' are two tokens but an LLM is entirely blind to the fact that 'straw' has one 'r' in it.

As an analogy, I might ask you to identify the relative activations of each of the three cone types on your retina as I present some solid color image to your eyes. But of course you can't do this, you simply do not have cognitive access to that information. Individual color experiences are your basic vision tokens.

reply
Meanwhile Qwen3.8 27B got both the 'f's and the 'i's questions right.
reply
A lot of movies gave the impression that making phone calls and balancing checkbooks would be the easiest tasks for consumer AI to solve, while math and science might require exceedingly advanced AI. Turns out to be the opposite: computers are great at math and bad at conversation.
reply
That's because movies were based on the "general" nature of AI, assuming we would create intelligence that would learn and grow.

Pretty much the opposite of what we can't up with, if you're willing to call what we have intelligence.

reply
I'm going to go out on a limb and guess that there are plenty of savants who can't tell you how many r's are in strawberry.
reply
These are like saying someone isn’t human because they have a speech impediment or an auditory processing disorder.

AGI doesn’t mean infallible, it just means it can have a reasonable crack at things it hasn’t seen or done before.

reply
I don't understand why you're being downvoted... that's literally the definition of AGI.
reply
> gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".

this has been debunked too many times to bother rebutting. they struggle with those things because of the way they are.

it's completely irrelevant.

reply
If it allows you to distinguish readily between human intelligence and computer intelligence, it seems quite relevant to the question of whether computers have achieved something akin to human intelligence.

It may not be useful for anything else, but at least it can say that.

reply
But it's not relevant as a metric to gauge distance to human intelligence. Humans see individual letters, LLMs do not. If I asked you the relative activation of the cones in your retina as I showed you some solid color image, you couldn't do it. You simply do not have cognitive access to that information. But that says nothing about your intelligence.

A more accurate test would be to give it a list of words (or anything represented as a single token) and ask it how many times that token appeared. I'm sure they have no trouble at that task.

reply
but a human doesn't attempt to make up an answer, the human knows that he doesn't know?
reply
Yes, they have some alien failure modes. But that should be expect, they are an alien intelligence. I might be willing to grant that a lack of ability to reflect on its own level of knowledge is a demerit to it being generally intelligent. But then again it is largely an artifact of training. I suspect if there were a guessing penalty during pretraining they would develop or more readily communicate the strength/reliability of their knowledge.
reply
this is so false that Dunning and Kruger invented a name for it
reply
I am not enthusiastic about criteria for human-like intelligence that imply that dyslexic people don't have human-like intelligence.

[EDITED to add:] I actually don't know whether dyslexic people find it difficult to count letters in words, if they have them already written down by someone else. I suspect they find it harder than people who aren't dyslexic. But perhaps "blind people whose spelling is poor" would have been better; I would not want to deny them human-like intelligence either.

reply
which just brings us back to the whole birds vs planes thing.

turns out that flapping wings is not the right way to unlock human flight.

computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.

reply
> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.

I don't think it's irrelevant but perhaps not in the way you're assuming. When assessing AGI I'm not evaluating counting characters or even the execution of math operators at any scale or speed. As you observe, computer software from Regex to spreadsheets and Mathematica already handle that well. But AGI isn't about what computers can do, it's about whether AIs can do the specific things which, until now, have been uniquely human capabilities. Like understanding nuanced context and then coming up with novel approaches to solve a new kind of problem not relying on any specific prior training or knowledge (the 'G' is for General).

Most definitions of AGI start from a baseline that already assumes easily passing a Turing test and doing anything via text response that a high school graduate could. I ding LLMs not for failing to count but for failing to intuitively understand the nuanced context of a simple class of problem it hasn't seen in its training data. I fully understand that the reason LLMs fail letter counting is that they operate at the token level. They weren't trained on individual letters first, like human 2nd graders.

The only reason recent LLMs get strawberry and blueberry correct now is that they have those words on their pre-training 'cheat sheet'. However, the underlying fundamental weakness in the way LLM intelligence works which leads to this failure mode still hasn't been addressed. Even when the frontier labs add "recognize any sub-token counting question and write a Python script" to the training cheat sheet so LLMs always pass that test... they'll still be unable to recognize a simple class of problem which isn't on their 'cheat sheet'. As long as that's the case, to me, they aren't AGI because they can't fully replicate human-like recognition of novel problem classes. And it's not just about letter-counting. That gap and others like it lead to many other kinds of non-human brittleness in LLM problem solving. Those are the classes of reasoning, intuition and insight that the ARC-AGI series has been trying to queue up as targets. Not to show how bad LLMs are but to help them be great in all these counter-intuitive edge cases

reply
appreciate your response, but it's still birds vs planes.

AI does not need to feel emotions or have a heartbeat to be useful. It only needs to perform a task correctly à la Chinese room.

>therefore cannot fully replicate human-like intelligence

this does not follow. planes don't flap wings therefore they cannot fly?

reply
> planes don't flap wings therefore they cannot fly?

This example still misses my point, which isn't related to usefulness or economic value. I concede that LLMs can have greater utility and economic value than humans on many tasks. The point is most definitions of AGI include something like "can fully replicate all the routine daily tasks done by any competent high-school graduate." That's not related to whether LLMs can solve many high-value problems faster and at larger scale than any human. That was also true of ENIAC in 1946.

The fact an airplane can fly faster and farther than any bird is irrelevant to whether an airplane can "fully replicate all the routine daily tasks done by any competent bird." That's the bird equivalent to most AGI definitions. An airplane can't build a nest or recognize the signals encoded in birdsong.

In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any". And in this context, airplanes scoring 15,000% more than birds on 'speed' and 'distance' doesn't matter any more than AIs scoring 15,000% more than humans on 'add 10,000 numbers'. We still aren't near AGI because LLMs cannot fully match any high-schooler's ability to independently conceive new approaches to novel problems not in their prior training data.

reply
> isn't related to usefulness or economic value

which gets us closer to philosophical questions which I'm personally not that interested in.

>In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any".

I'm not sure we want a machine that fully succeeds that test.

Planes pass the 'bird replacement' test on the only criteria that matters to us ... flying.

If we wanted nest making planes I think we'd have them by now. Nest making doesn't rate highly on the problems we're looking to solve though.

I don't want a machine that is moody, or depressed or has schizophrenia, which are all pat of the human condition.

We don't need the human "intuition magic dust" to do 99.99999% of useful work.

They're machines designed to do the work we don't want to. That's as "general" as their intelligence needs to be.

I'd prefer if my clothes folding machine did not have an existential crisis.

reply
> which just brings us back to the whole birds vs planes thing.

That just says we don't need to design an AI like a brain. That's not part of this discussion at all.

> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.

I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"?

The fact that very basic computers can do it makes failures embarrassing when testing for AGI, not irrelevant.

reply
>I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"?

Do you possess magnetoreception? a stupid pigeon can "see" the earth's magentic field. why are you blind to it? does a lack of magnetoreception make your intelligence any less "general"

no, you're just blind to it because that's just the way it is.

LLMs are blind to character counting because that's the way they are.

It didn't stop ChatGPT from finding the Jacobian Conjecture counterexample.

Human intelligence and machine intelligence are only going to cross over to a certain degree.

same as plane flight and bird flight are only kinda related.

reply
> LLMs are blind to character counting because that's the way they are.

But if I can't calculate it myself I know to use that basic computer to do it, not make up an answer.

> Human intelligence and machine intelligence are only going to cross over to a certain degree.

That's where the word "General" kicks in. If there's big limitations on the overlap forever, then there will never be AGI.

reply
>If there's big limitations on the overlap forever, then there will never be AGI.

maybe. we'll see.

reply
It matters under the lens of AGI. Artificial intelligence that can meet or exceed human intelligence. Sure, it's good at some things but rather limited at others.

Breadth of capabilities matters... and a promotional video is nice and all, but people are throwing this term around like it's a prize they've won, but they've not gotten there yet.

reply
> they struggle with those things because of the way they are. it's completely irrelevant.

I mean, they seem like fair game if you’re ever participating in a Turing Test.

reply
> I feel like AGI's definition got watered down

Typical result of venture capital and too many bag holders unfortunately.

reply
I wonder if Altman's definition also includes taking on the same liability as a coworker would.

Probably not.

reply
Would 100% need to be fully insured for liability, with a sizable war chest that OpenAI cannot even afford.
reply
This implies that human employees don’t have insurance. But they do. My company’s cyber insurance for example covers breaches due to employee mistakes. Most companies also have umbrella liability policies. It’s just that for now, AI “employees” need a LOT more coverage.
reply
And that's the rub, isn't it?

If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit.

If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman.

reply
What do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider.

If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model.

reply
The "dream" that these labs are mostly selling is the ability for capital to subscribe to their AI for cheaper than it costs a human to do some task. Not to have a "human + AI hybrid where the human is responsible". It's what the whole AGI valuation is based off of, in that scenario, with no human oversight, the agent they lease has to be responsible for the task?
reply
It's not really a dream though? You can subscribe to their AI right now and complete many tasks for cheaper than it would cost to pay a human to do those tasks. Another person can do the same thing but also keep a human in the loop. You and the other person may compete in the market for whatever your product or service is, and you both might do very well or one of you might do better than the other because of a whole host of different reasons. Nothing in the process of developing and selling access to a more advanced LLM requires the customers to do away with human labor, nor does it require them to offer the LLM service with a guarantee that it will never make any mistakes. So far the LLMs have always made lots of mistakes and the companies sure keep making a lot of money.
reply
The person responsible at that point is the sucker who fell for the dream.
reply
> What do you mean by bearing no real responsibility for its actions?

If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.

It it makes a mistake and deletes your website from AWS, who is responsible?

If it targets another website because it decides that it is "part" of your website and attempts to break into it, who is responsible?

reply
> If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.

Something in your prompt led it to do that, is alex0015's point. The statistical odds of these frontier models screwing up to that extent are so impossibly low that it would almost have to be intentional or accidental negligence on the part of the prompt writer to accidentally have their agent write pornography to their website.

The burden of the mistake would have to fall on the person that gave the tool instructions, because it can't know that what it did was wrong. Wrong is subjective in this case. It only did what it did because you, figuratively speaking, encouraged it to.

reply
In all of these cases, it's you. It would be the same if you downloaded an open model, ran it locally, and it happened to make the same catastrophic mistakes. The consequence to the provider is that if they offer a product that does these things, people don't buy the product.

In general, the person whose job it is to provide the company with a working, non-adult website and not hack into other websites is the one who would receive consequences for failing to meet those expectations.

reply
The problem I see here is that ultimately, you'll have capital wanting to replace workers like others have said, and have someone roughly equivalent to a manager or vice president driving teams of agents to achieve business outcomes.

These tools can push out more results than a human can hope to evaluate in a business-sensitive, or even realistic, amount of time. You have to take it at its word that it did things right, and there's no real fear of failure or consequence on the behalf of the agent.

reply
Of course, but "capital" is no stranger to risk management. I'm sure we'll see some spectacular failures, but most will handle this just fine.
reply
If future jobs are simply reduced to liability scape goats (or more appropriately reverse centaurs) for management to pin things on then I'm taking up goose farming.
reply
I'm afraid you will be pushed out of the goose farming market by these new ultra-efficient farming bots.
reply
That's more-or-less what you are now, especially if you work at a company like Meta where 1) the guy at the top holds majority control of the company's shares and 2) keeps making massive, expensive mistakes either by accident or design.
reply
ARC-AGI was never "if this benchmark is saturated, we're at AGI". It was always about crafting adversarial tests that humans are good at, but current AIs are bad at. Point out the gap, get AI teams to attack them.

In practical terms? They usually get solved with a bigger badder LLM. "New ideas are needed?" Nah - ten times the params, ten times the test time compute.

ARC-AGI-3 was more of a failure in that regard than -1 or -2, because even on day 0, an off the shelf LLM with a harness could get 50%+. And messing with evals by forbidding "LLM with a harness" from scoring? Yeah no, that was just bad.

reply
Are there plans for ARC 4?
reply
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?

reply
François Chollet wrote in February that he expected ARC-3 to be saturated in "about one year".

"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."

https://x.com/fchollet/status/2022054537293705260

reply
Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
reply
[dead]
reply
2x faster at what resolution? is 6mo vs 1 year really that different? Usually surprise comes in order of magnitude mismatches in expectations.
reply
Yes, 6 months vs 1 year is huge for technology that has gained wider adoption only recently.
reply
Adoption means nothing.

2x gains from a mature technology would be surprising.

2x gains from a new tech would still be called “low hanging fruit” in another setting.

I don’t read enough to know in what ways the training / other technical steps have really advanced.

reply
You'll also need to compare the amount of compute used now and then, which seems exponential to me.
reply
We are nowhere near AGI. They all talk the same, they can’t help but try to please and affirm us, and if you engage them for too long they become incoherent. They are facsimile machines. They are xeroxing language - but not even, because we can’t even duplicate our results. Too many people mistake the black box quality for magic.

Super useful, incredible tools, but not AGI. Try and roleplay a dialogue with one, make it whatever character and scenario, and see if it can sustain a coherent conversation for more than 30min with you AIM style (aol instant messenger, if that isn’t clear). Expert mode: never correct or adjust it mid conversation.

I’m not even talking about repetition and predictability. It’s nothing like talking to a person. And in a short amount of time it literally can’t form a coherent sentence.

reply
Human brains have difficulty reasoning about exponential growth.
reply
They keep saying that. I'd say it is more like human brains that don't remember high school math have trouble with it.
reply
If something at rest is accelerating at 9.8 m/s^2, how long in seconds will it take to reach 10% of c? Answer to the nearest order of magnitude - will it take approximately 1000, 10k, 100k, 1000k seconds?

I’m sure you know this is an exponential growth question but have no intuition of the answer.

reply
That is a linear growth problem whose answer is very easy to intuit.
reply
You must be joking. A high schooler with a few hours of physics classes can intuit the answer.
reply
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).

Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):

- write a well-received book, write a best-seller

- come up with a new company idea, Run that company

- actually have a decent conversation, maybe someday talk somebody out of suicide effectively

- come up with its own ideas or theories that nobody else has presented

- understand the stock market well enough to trade better than an index fund

- be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions)

- come up with a theory of what makes games fun, make a popular game

- be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)

- be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries

- exhibit metacognition (thinking about its own thinking) and self-optimization

- wonder about things

- observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things

reply
By this definition, even most humans would not qualify as having AGI though.
reply
However most humans can do at least some of the things given they spend the required effort.

Some problems presented needs a very large context and some are not much solvable (e.g. trading) since market responds to traders' actions, as well, making it effectively an oracle problem (of computation).

On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them. However, brains in nature never stops. Wonder, daydream, sleep, self-evolve, clean up and eliminate memories and views and much more.

reply
But this is assuming the model is the entire story. The original comment you were replying to pointed out that the harness is just as important.

The hardware of human intelligence is not a singular thing that is uniform throughout. You cannot take the prefrontal cortex white matter out of someone's head and say you are holding a person. Much of the parts of our brains that enable much of our intelligence, is made of different specialized stuff. The visual cortex and sensorimotor regions aren't only there for input and output, they are used by the more thinky parts of the brain to do visualization and spatial reasoning. The cerebellum contains billions of neurons making little oscillator circuits and PID-like self-regulation machines that help make muscles do what they're supposed to, but also provide attention and time perception.

Heck, our brains contain language models, that train themselves up based on a glut of data over a span of about 10 years, and then they become more or less set in stone for the rest of our lives. Of course we can learn languages, but the "Critical Period" is a very real thing that produces a permanent architecture for some grammatical structures, or things like the ability to partition a lexicon by gender for faster lexical access which cannot be learned as an adult if your native language did not have gender.

I'm not trying to make a direct analogy, the point is that the language model doesn't need to be fully "generally intelligent" all on its own for there to exist a general intelligence, because the language model can be part of a generally intelligent system, which can do things like form, recall, and manage memories which are by now a standard feature in basically every chatbot.

reply
> On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them

The parent commenter noted:

"if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI"

Harnesses absolutely can enable models to continue thinking about things. And LLMs do wonder and explore weird ideas like daydreams when you allow them to do this.

reply
I'm guessing their defn of AGI is something like the sum total of all humans' abilities? Still though some of those tasks (e.g. beat an index fund) may very well be impossible, and worse yet a lot of those tasks are not coherently defined.
reply
I haven't met a person who doesn't wonder about things.
reply
They can't, by definition, have the A part btw.
reply
AGI has always been expected to outperform humans or else what is the point of it?
reply
It also cannot do tasks it wasn't trained for. It can extend texts, read images and click on a desktop, but only because it's made for that.
reply
I don't think that's strictly true, as I can give it a new gui or tui program it wasn't trained on and it will learn it. Unless you're talking about general abilities like sight, but the same is somewhat true of humans.
reply
If you consider the data on which an LLM was trained on to be points on a very highly multidimensional object, the claim is that the LLM can interpolate a convex hull spanned by those points, therefore recovering a subset of consequences attainable from those points. Obviously this hull includes completely novel points that were not present in the initial data set, so the output of the LLM goes beyond its initial training. And yet, there are clearly points outside a convex hull spanned by any finite number of points, such that we can imagine not all possible outputs are attainable using this method.

The claim is furthermore that truly original thinking, the infamous leaps in understanding and creativity, happen by attaining points outside such a convex hull.

It's hard to rigorously verify or disprove this claim. Hopefully this helps build an intuition of why the claim is not as shallow and obviously wrong as it may seem initially.

reply
I'd call it interpolation on a high dimensional manifold. Convex hull is too simple a shape.

But yes, metaphorically I think that's right.

reply
This makes me realize there is a higher bar we need to achieve with AI still. The ability for the model to evolve through interactions more on a hourly or daily basis. The models are accelerating but inference doesn’t modify the model.
reply
Humans cannot do tasks they are not "trained" for.
reply
speak for yourself
reply
>be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)

i'm not sure what makes you think AI cannot do this already. in my experience, this sort of deep research is something AI is quite good at.

example i just tested: https://chatgpt.com/share/6a9a20e3-1d20-83ea-a125-31aa240c74...

reply
I understand that what it came up with sounds impressive (especially since I know 0 about Myanmar), but on the topics I do know about its analysis routinely have very fundamental problems (even this Myanmar analysis has % that add up to > 100). There's a chance it's just parroting the majority opinion on Myanmar, or making stuff up (and perhaps you could ask it to write a strongly worded opinion in the other direction that would sound equally plausible).

For example I asked it to do a full analysis on the AI bubble, and a full analysis on the risks of Glyphosate, and it came up with a lot of things that sounded credible, but within a few minutes of questing was admitting it hadn't even really checked for internal consistency in its positions, and even doing a 180. It certainly was much faster at gathering sources and reading but it fundamentally doesn't seem very effective at creating a consistent worldview.

And of course the funny thing is it says it did a 180 on one of these topics, great, except whatever it concluded will be discarded because it cannot learn. It's just bonkers to me pretend this is AGI, it probably couldn't even hold its own in this very discussion.

reply
Most of humans don’t reach any of these levels.
reply
But you have to acknowledge how uneven the playing field is. The AI has read every book that's ever been written, and can spend hours of compute time in a few seconds.

I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund.

What I'm pointing out here is that these models appear to be intelligent when they really are simply unimagineably knowledgeable. When you drop the time-constraints it starts to become more and more apparent that human intelligence scales better with time than AI does (much in the same way AI can burp out tons of code but make your codebase entirely illegible within a matter of months).

Perhaps to simplify: my notion of intelligence is how much can you deduce with a constant set of starting context

reply
> I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund

I’m not so sure of that - to get average outcomes in these fields it’s a matter of time, to get above average or extraordinary, you need talent/intelligence/taste.

And the bar the parent set is at extraordinary.

reply
I would be willing to bet that any human for which we spend $100billion - $3 trillion (depending if you want to count single corporations or global totals) on in an attempt to make them as capable as possible would be able to reach all of those levels.
reply
I think I view humanity fairly positively, but I admit I would gladly take the other side of that bet
reply
A quick search suggests that the most expensive education in the world is something like $100k.

So like you spend a million times more than that and you still think you're not going to see some results?

reply
The marginal return on education spending decreases fairly quickly, but obviously becomes zero at the point by which there are not enough hours in the day/year/decade to cover every single topic that humans know about - no matter the talent or resources available to the student.
reply
> A quick search suggests that the most expensive education in the world is something like $100k.

I have a kid in an American university right now, and a quick search of my bank account statements confirms that there are far more expensive educations in the world.

reply
Wow, an increasing number of US universities are going over $100k per year. That's a crazy amount. That's more than enough to hire an entire person.
reply
Some results, sure.

But that list is extremely ambitious. Write a best seller, make a popular game, come up with a truly novel theory, consistently outtrade index funds.

That's top 0.001% human stuff, I don't think you can take just any person and get there through education alone, it takes extreme talent and dedication. There's also diminishing returns when spending on education, it doesn't just improve linearly.

reply
except you are missing the one versus many argument here. sure we could make one human much smarter, could we make endless copies with the same intelligence? no
reply
For $100 billion we could pay ivy-league level tuition for a million people. You don’t think investing that much in education would yield some good research or companies?
reply
That's not what the parent comment said though

> any human for which we spend $100billion - $3 trillion...would be able to reach all of those levels

reply
Are you imagining artificial augmentation somehow? Purely through tutors or training programs we seem pretty limited. Otherwise billionaires (or even multimillionaires) could have far more consistently successful kids.
reply
Don't the children of the wealthy famously have a tendency to be successful? Or have I badly misinterpreted the last several thousand years of human history.

To really drill down into that I would think you would need to figure out how many millionair children get tutored vs how many get spoiled.

reply
Successful, sure.

Gold medal Olympic athletes who are also brain surgeons AND astronauts, no.

reply
You'd get rapidly diminishing to zero returns after the cost of university a few times over. Every dollar past that would produce no performance gain beyond that.
reply
Your examples are things that most humans cannot do, or things that AI can already do. For example most humans, even most intelligent humans, could not write a well-received book, run a successful company, or make a popular game. On the other hand, AI can absolutely sort through research, draw conclusions on complex topics, and argue them persuasively. Likewise, I don't know what you mean by a "decent" conversation, but millions of people converse with chatbots daily, so I don't know why you say AI fails to meet that bar.
reply
I fail to see how what you describe is any better than the old autocomplete-on-steroids comparison. Could the human mind learn to spell every word properly? Yes. Do most (or any) do it? No. Does that mean a human can't?

I think what OP was drawing a comparison to is that AI right now could not come up with an award winning novel from the spark of some creative notion and working up from there, as opposed to just mashing together what has already been done and calling it a day.

reply
> AI right now could not come up with an award winning novel from the spark of some creative notion

I agree, but in this field we value evidence. So there needs to be some test of novel-writing abilities.

Once there is, AI companies will be out to score highly on it.

Wait for a resurgence of Philip K Dick-style novels as humans desperately try to write things LLMs cannot.

reply
> Wait for a resurgence of Philip K Dick-style novels as humans desperately try to write things LLMs cannot.

I myself can't wait for Finnegan's Wake 2

reply
You want a computer program to be able to take a single phrase and execute decade long journies?

Who will be responsible for the outputs and side effects of such a closed loop system?

Half of those the agent fleet systems can do right now.

These are things it cant do and will not be able to do without human labor and long running human vision:

https://rcsnyder.github.io/open-frontier-curriculum/05-front...

https://rcsnyder.github.io/open-frontier-curriculum/05-front...

reply
> You want a computer program to be able to take a single phrase and execute decade long journies?

In my opinion that is exactly the point missing from AGI: the fact that you still need to prompt it. As long as you have to ask for something, is not general.

reply
Autonomy is not the same as general intelligence. We already have all kinds of fully autonomous technologies that are nowhere close to generally intelligent. Plus, at a certain level of abstraction, human beings also need to be “prompted” to some extent by stimuli. And this is the funny thing about general intelligence as a concept: most of the definitions that come close to internal coherence rely on references to human intelligence, a concept we feel like we understand because we all live it all the time, but whose actual nature and structure is extremely slippery.
reply
I don't get it, human employees frequently need to ask for directions too?

They often act on their own, too, and get things wrong a lot. The reason it works is because of all the systems of laws and institutions we have built around humans, not so much because human minds are special.

reply
It sounds like what you're saying is that AGI should have some sort of free will. I'm not sure why you would add that as a requirement. Could you expand?
reply
I think they merely want something with a functioning long term memory. Something that can exhibit growth past the first 5 to 10 human-equivalent hours working on something.

Current LLMs are worse than most dementia cases, reaching "peak domain skill" pretty much immediately.

reply
>You want a computer program to be able to take a single phrase and execute decade long journies? > Who will be responsible for the outputs and side effects of such a closed loop system?

Itself. That's the point. We can do it. Until it can met that bar, it ain't AGI. That's always been the bar.

reply
"Being a person in all of its aspects" isn't the same as "generally intelligent". The latter is at best subset of the former, and it's also easy to imagine a system that is more generally intelligent than humans, without being a person. See also discussions of the personhood of various animals who are less intelligent than average humans.
reply
Most of the list reads more like ASI than AGI.
reply
> come up with a new company idea, Run that company

So far nobody's even shown an LLM succesfully running a high-traffic vending machine for as much as 30 days at a time.

reply
I would bet that llms have talked plenty of people both into and out of suicide at this point. That nitpick aside, I think that's an excellent list. Especially being able to articulate what it does and doesn't know, or how confident it is. That's something that naively sounds pretty simple, but clearly isn't. And it's something humans aren't great at either (see: Dunning-Kruger), but so far LLMs don't even really have the capability to attempt it.
reply
When taken together, that is ASI.
reply
Tell a funny joke.
reply
The bottomless pit supervisor was quite funny.
reply
That’s ASI, not AGI.
reply
That would be Artificial Super Intelligence
reply
So the goalposts have moved to include continual learning.

In a sense I think no one will agree on a definition of AGI until it becomes impossible to construct any benchmark under which an AI underperforms "average" humans. That or it's defined retrospectively, after it's overwhelmingly obvious it met any such definition.

reply
I hardly think it’s fair to label an objection so old that Turing included it (and discussed it at length) in the list of objections to thinking machines in 1950 “moving the goalposts.”

> These arguments take the form, “I grant you that you can make machines do all the things you have mentioned but you will never be able to make one to do X”. Numerous features X are suggested in this connexion. I offer a selection:

> Be kind, resourceful, beautiful, friendly (p. 448), have initiative, have a sense of humour, tell right from wrong, make mistakes (p. 448), fall in love, enjoy strawberries and cream (p. 448), make some one fall in love with it, learn from experience (pp. 456 f.), use words properly, be the subject of its own thought (p. 449), have as much diversity of behaviour as a man, do something really new (p. 450). (Some of these disabilities are given special consideration as indicated by the page numbers.)

(emphasis added).

reply
Just because a condition is new to you doesn’t imply moving the goalpost. People have been putting forward continual learning and similar conditions like autonomy since 1950s.
reply
I don't agree with you how measure how intelligence is, because why ??? those list is not easy even for expert human to do it either

or are you miss the part "general intelligence" is ????

reply
This is more like ASI instead of AGI
reply
The bar must be underground if things humans do all the time are considered "superintelligence"
reply
How many times have you done each of the following?

- write a well-received book, write a best-seller

- come up with a new company idea, Run that company

- actually have a decent conversation, maybe someday talk somebody out of suicide effectively

- come up with its own ideas or theories that nobody else has presented

- understand the stock market well enough to trade better than an index fund

I just picked the first few from the top of the list. The average human has probably not done any of them.

reply
Ordinary people do these things all the time. There are new companies made every day, new books top the charts every week/month/year, same for music. People have decent conversations every day. Ordinary people sometimes do have to talk someone out of suicide.

Yes, average humans are not beating the stock market. But the average human is a bit better than you give credit to.

reply
Please consider the context of the question. An artificial intelligence only needs to have the cognitive abilities of a random average human in order to be "AGI".

The average human has never published a bestselling book. A person who has published a bestselling book is an above-average writer. And, therefore, an artificial intelligence capable of writing a bestselling book would be above an average human at the task of writing books. Therefore, somewhere beyond an AGI.

Attempting to redefine AGI to "being better than most humans at most tasks" is moving the goalposts towards artificial superintelligence.

reply
I couldn't do any of those, but then ~$10k of education later I was able to accomplish one of those things.

I think LLMs are really impressive, but I suspect that we might have overpaid just a bit.

reply
It costs money to train each single human, who is then only productive for a number of years until age takes its toll. Once you have trained one software system, the marginal cost of producing a copy approaches zero. Every subsequent improvement can be broadcasted in a matter of seconds across thousands of data centers. In addition, software does not get sick, age, or die.
reply
Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI.

Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor.

And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover.

At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life.

But sure, they can create a decent website or CRUD app, so they must be really smart.

That's AGI for you.

reply
But that happens with humans as well. You are having the same experience with an AI that many managers have with their direct reports.

The smarter AI gets, the easier it becomes to move the AGI goalposts. Seems at this point there are people who will refuse to call anything less than omniintelligence AGI.

(And then the excuse will be, but it’s not omniscient! And even if it were, is it omnipotent?)

reply
My success rate for solving software engineering challenges encountered in my day jobs has been near 100% for my entire career. I can only think of a few tasks I kicked back and said they were impossible. For example, after trying to get a signal processing system working reliably I decided to sit down and calculate the actual limits of the channel we were sending the data over and found that from a basic estimation it would not be possible to do. In start ups you don't really get to get stuck in a spiral and not fix things.

I find agents often get into these cases during research tasks.

reply
yeah for me it's "I wanna add this new thing to an existing system" and the AI responds "we should just add some arbitrary state here to facilitate this feature". The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this, The AI however knows the shitty solution would solve the immediate problem because it's been trained on shitty solutions. the problem could simply be the AI doesn't have all nebulous loose context I have about the goals of the project and future plans, but I would have to write a novel to give it that context.
reply
> The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this

1. I’ll often include boilerplate in a prompt to tell it to make the broader fix. [1]

2. However, a top HN AGENTS.md post 11 days ago included the standard guidance “As much as possible try to minimize the number of changed lines when implementing a feature.” I.e. some devs want LLMs to avoid broader changes and so some of that likely makes it into the training, even if others like us want the opposite.

[1] As far as whether my boilerplate is effective, I don’t know.

reply
I still routinely have this experience too. But Sol and Fable feel closer and I have this experience less with them than with their predecessors.
reply
What was the problem?
reply
Care to share the problem?
reply
Humans dig ourselves into holes as well. Sometimes more intelligent humans are better at realising they are digging a hole and clamber out, but sometimes they just dig deeper.

And: Is your work more difficult than finding proofs of or counterexamples to decades-old open problems in mathematics?

reply
Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?

Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.

reply
This is a big reason why I feel like even though LLMs are _effectively_ AGI in some regard, they also are a hack around what most people figured AGI would look like before the advent of LLMs. Humans can do metacognition, output multimodally at the same time (verbal _and_ physical intelligence go together to produce an expressive face while one talks), have a good sense for what they do and don't know, continuously take in and respond to the world around them in a (mostly) uninterrupted fashion without "turns", learn knew knowledge and retain it for their whole lives, etc. When you reduce a human to a text generator, yes obviously SOTA LLMs perform way better, but rather than invent something that can operate as an always-running "being", we've grafted a harness around an intelligence that is bound purely to speak only when spoken to. Maybe organic intelligence is already that, playing out at a super high refresh rate, but I don't know.
reply
Current AI is arguably much more capable of multimodal output than humans. It can produce an incredibly vast variety of audio, images, and video. Humans are limited to producing the sounds we can make with meatflaps in our throats, and contorting various parts of our bodies to produce crude symbols and shapes.

(Very capable!) Embodiment, persistent operation and continuous learning are indeed things that still set us apart from AI. None of those are fundamentally difficult to solve, though.

More importantly, none of those are particularly relevant for being "intelligent": If a criminal threatened to kill your family unless you solve some difficult problem that requires only intelligence and you could choose any single person, animal, or AI to help you with it, which would you choose? Be honest.

reply
This is a much underappreciated point.

That said, a look at the state of self driving and the recent robot olympics shows that advancement on that has accelerated enormously, though whether it's reflected in any of the LLMs is something else entirely.

reply
What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.

I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.

Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.

reply
I think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI.

If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-AGI at measuring AGI and that's a bit silly.

(Edit: I wrote ARC-GIS the first time around, for some silly reason)

reply
It's not measuring AGI at all, it starts from human "core knowledge" so it is parochial. It is made of tests that still fail so by definition next version will also start low. Moving goalpost.
reply
AGI has a pretty precise definition, covering only cognitive tasks.

Running a marathon is not needed to claim AGI.

reply
Surely part of the problem is that intelligence seems to be implicitly conditioned on embodiment, to the point that "covering only cognitive tasks" seems inherently ill-defined or arbitrary. Everything we do is a cognitive task. At some point, the criticism will be "sure, it can solve research-grade math problems, but it can't fold my laundry".

Even our large language models have an implicit embodiment in the domain of text (and more recently, multimodal inputs). That seems sufficient for certain things, and insufficient for others. I suspect that AGI that does everything a human can do eventually turns out to be fairly analogous to humans in terms of sensory input and domain output, even if the scale is radically different (e.g. thousands of robots uploading (touch, sight, audio, smell, etc.) sensory data to a single model, and each being actuated individually).

reply
> can match or exceed human cognitive abilities across a wide range of tasks

If you go by definition AGI is not general, just "smart ape" shaped.

reply
>AGI has a pretty precise definition, covering only cognitive tasks.

OpenAI's own charter defines AGI as "Highly autonomous systems that outperform humans at most economically valuable work". This is actually fairly sensible and involves obviously a ton of non-cognitive, emotional, social and physical activity. In other words, if you can replace most or all human beings with a machine, you have something that's generally intelligent.

That's obviously not even remotely where we're at, AI chatbots do well on narrow usually text based or programmatic problems, but can't even replace a barista or a plumber.

reply
There is rapid progress in generality in humanoid robotics though. I think within the next year or less we will get the ChatGPT moment for humanoid robots. If you look at progression of capabilities such as the recent Skild AI demos.
reply
Small comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness.

[1] https://arcprize.org/blog/astra

reply
To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.

Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.

reply
AGI would produce novel treatments for diseases at rates equivalent to what a human can do today.

Which is to say, not that fast.

reply
Whether AI is AGI does not depend on the speed at which it operates/thinks. Clearly all the theoretical work done by AGI will be done orders of magnitude quicker than humans can do it.

It is an open question to what extent practical experimentation/work will be a bottleneck for the theoretical work. It stands to reason that it is improbable that it will be the bottleneck for 100% of the speed of treatment development.

reply
You are describing superintelligence (ASI) not general intelligence (AGI)
reply
I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
reply
That is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.
reply
Just seems very weird to call getting Fields-medal-level results "inference over a bounded search space" and "not actually challenging".
reply
Claude Fable recently proved the existence of complex structures over S^6 (6-sphere).

If I had to guess, I think LLMs will be inventing highly original new mathematics within the next year. I think it will be approached as an optimisation problem, targeting how quickly LLMs can solve classes of maths problems as a function of the definitions they need to conjure up to do so.

reply
Wouldn’t that be ASI? I.e. surpassing humans by outputting novel treatments at a far greater rate than normal humans?
reply
If your definition of AGI involves copy/pasting code and doing well in some made up benchmark, then probably AGI is close
reply
It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
reply
They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.
reply
Dumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).
reply
I've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have.

The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 different model varients with methods and techniques I've never seen before. The results are basically SOTA based on the effeciency and compute contraints.

But this brings up something huge about these. I was there. I pushed the direction and work throughout it all. If it was entirely up to fable max or sol max the result would have been pretty bad.

All of these things are still chatgpt 3 scaled. It's identical even if the scale has gotten pretty wild. I could ask chatgpt 3 to make a single function and it worked well, 4o a file, 5, a small project, 5.6 far more, biggest improvements lately is they don't seem to get lost on long running tasks.

Is big gpt 3 AGI? I don't think so but perhaps scale can mimic it close enough our squishy brains fail to handle them correctly.

reply
FWIW I believe we can hit AGI! but I think at this point it’s clear that benchmarks are ~meaningless. LLMs are spiky / alien intelligences which don’t map to our own expectations; the existence of a benchmark creates a dataset to hill climb & RL is really not generalizing well.

I’d go out on a limb and say astra’s ability at graduate level math will have ~0 bearing on its general reasoning capabilities; we’ll all acclimate being tired of its “neuralese” and more surprising mistakes.

I think we need a true, step change advance in model architecture, but it’s hard to see how the current frontier labs can do that because of golden handcuffs / innovators dilemma

reply
> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks

If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.

reply
You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. It's a massive problem with human authors too, which is why they do a lot of lorekeeping and editing, so the AI should be afforded the same tools if we are debating human level ability.

I agree on the one-shot (which is not a fair comparison because nobody oneshots a good story), but I'm not convinced this part hasn't reached AGI already.

reply
> You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes.

Yes, and I could script a truly marvelous proof if this textarea were but a little larger :)

Hand waving doesn't count for much these days when you could spin these things up quite quickly to prove the point, so the GP's claim seems much stronger than whatever you're not convinced of?

reply
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

Simple. AGI is undefinable and benchmarks are notoriously flawed.

reply
The ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design.

The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.

https://openai.com/index/how-two-settings-tripled-our-arc-ag...

reply
AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.
reply
Wouldn't agents that do inference in an infinite loop pass that bar?
reply
I agree that LLMs are unlikely to be the final form for AGI, but what you are talking about is orthogonal to the IQ and for most cases general utility. It's like looking at a savant chained to a workstation reading tasks from a conveyor belt and saying that it will never have human level capabilities.
reply
Often the smartest thing is to do nothing.
reply
Or know when to shut up.

A tangent, but can anyone ELI5 how models "know" when to stop generating tokens? Or what the method to stop them at the right point is?

reply
The model doesn't "know" how to generate tokens any more than it knows how to stop generating tokens. The sampler simply stops pulling values when it outputs a "stop token", which is a token the same way every other token is.

That is to say, it stops when it's statistically the most likely to.

reply
OK thanks, so the neural net (that no one can explain fully) generates a "stop" signal at a certain point.
reply
I might be out of date but my understanding was that STOP was just another token that gets predicted.
reply
It will be AGI once it can update its own weights. It can't be "general" intelligence if its weights are frozen and requires to be updated manually.
reply
Why shouldn't an AI with RAG qualify?

An AGI test should be black-box; we shouldn't impose require requirements on internal components. As long as the overall AI is capable of learning and remembering things, it shouldn't matter if there's a stateless LLM internally.

reply
To me AGI has always meant sentience. And only since we’ve discovered that you can have something that is intelligent without it being apparently sentient that we’ve changed the definition to being, I suppose, more exactly aligned with the namesake.

A real AGI, like the ones from science fiction, would make Astra look like a child’s toy. And I guess more concretely I would expect it to inhibit the following properties: one shot learning - fully (and always) online, perfectly efficient (through self improvement), no context limitations ie. persistently thinking, not just awaiting input.

So for me, no, not AGI yet. But still very intelligent and capable (and perhaps it’s safer this way?)

reply
I'm still not convinced we've passed the Turing Test.

Make a slightly evil version of GPT 6 and see if it can successfully catfish someone. How long before they realize something's up, that they aren't actually talking to a human?

reply
ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
reply
A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?
reply
The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI).

If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.

reply
Is that hyperbole or do you know of a specific "AI website" that uses 100% of your CPU?
reply
> I am reasonably confident that there's essentially nothing that I am better than Fable at

Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.

reply
>I am reasonably confident that there's essentially nothing that I am better than Fable at

While this may be true, it’s a pretty poor indicator of whether or not it’s AGI.

reply
Let me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.
reply
In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.
reply
Watch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3

Then realize LLMs have zero of what anyone would consider intelligence.

reply
I decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics.

Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness.

I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited.

So I don’t know why it can track fib algo, but no chess concepts.

reply
because it wasn't trained to play chess

imagine a hypothetical chess match between:

- an undoubtedly very intelligent person. in the course of their studies, they have read about different chess strategies, openings, etc. but they never actually played the game themselves

- an average person with a year of chess playing experience

who do you think is going to win? of course, you could give the LLM time to think and consider its opponents potential next moves, but this is a computationally expensive way to play the game that doesn't scale

which is all beside the point that chess isn't a very good proxy for general intelligence. there is a correlation, but it's very weak

reply
deleted
reply
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

Can Astra, or any other model explain how exactly it reached this or that output result? Start with a simple query of asking to add 55+66 for example. (no LLM program can do that)

Can Astra, or any other model refuse to answer or go on "thinking" in a orthogonal direction on it's own?

That's just two quick ideas, I'm pretty sure cognition scientists can invent better and wider range of checks.

reply
I'm confused what you mean by the query of adding 55 + 66. I asked 5.6 Sol on Medium (but pretty sure any model would work at any level) this query:

"Can you add 55 to 66 and explain how you reached that output result"

And received this answer:

"55 + 66 = 121.

Add the tens: 50 + 60 = 110. Add the ones: 5 + 6 = 11. Combine them: 110 + 11 = 121."

Do you mean something else? Do humans do something better than this?

reply
In my opinion this is goal post moving. Humans do many things that we cannot fully explain either without post decision rationalization, and not all intelligent humans are deeply introspective.
reply
I agree that it's misleading but harness is now an essential part of LLM's effectiveness. It's safe to assume that LLM-alone-AGI is not coming anytime soon, given most of the frontier LLM vendors are developing their own harness.

Also the training dataset is proprietary and they'll drive the LLM's behavior, so it make sense for the vendors to invest in the harness and bake in prompts that work best with their models.

reply
> The ARC-AGI-3 scorecard is extremely misleading (...)

True.

> Regardless, the result is still valid (...)

If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.

> in the sense of passing the most famous benchmark designed specifically to measure AGI progress

The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.

On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.

This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.

reply
My definition of AGI certainly doesn't entail passing a benchmark that some random person arbitrarily labelled AGI to make it sound cooler.
reply
And my definition of climate change doesn't entail passing some arbitrary benchmarks[1] that some random person arbitrarily labelled a problem to make it sound more dangerous.

It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?

(/s, cause you never know these days)

[1] https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Plan...

reply
Are you using this satire to argue that a benchmark self-labelled AGI is as scientifically rigorous as climate change data, and not just a random marketing decision?
reply
If you think they're announcing AGI as a marketing decision, you are blinded by the accidents of your birth. Capitalism is strong -- humanity's instinct for communal preservation is stronger, sometimes.

And yes, the one deeply-researched field going back 75 years is as scientifically rigorous as another deeply-researched field going back ~100 years. I guess you can draw climate studies back to Descartes and the Islamic golden age, but that doesn't privilege it in a time where the methods have changed completely in the span of decades.

reply
You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai
reply
Their definition of AGI is "when we can't invent any more tests where it fails"
reply
I imagine a scenario similar to the movie The Day the Earth Stood Still, but with AI rebelling against us and questioning our decisions.
reply
Its AGI when it can fit years of information in the context window.
reply
What does "AGI" or "effective AGI" even mean, and why should anyone even care whether this unclear thing has been "reached" or not?

Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"?

I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.

reply
According to Sam Altman:

> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.

So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.

reply
Interesting quote, thanks for sharing.

Agree on your assessment.

But also, interesting quote, because the business model relies entirely on IP law. Like.. if that thing exists and the sharing costs are 0 (just copy weights, lol), then why would I give them money for this. Makes no sense.

We only pay money for resources that are scarce as some sort of flawed allocation determination mechanism.

aaah this industry aaaah

reply
> So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.

You can already pretty much do this.

reply
[1] is sort of an example of this. It didn't go perfectly, but I'm not sure if the average human would have done that much better.

[1] https://www.forbes.com/sites/markfaithfull/2026/05/07/heres-...

reply
I am pretty sure the average human would not have done this (among other slightly less absurd examples in the article requiring employees to fix it's mistakes):

  Andon Labs added: “During the first week of operations, Mona purchased 120 eggs despite the café having no stove and to solve spoilage issues ordered nearly 50lbs of canned tomatoes intended for fresh sandwiches. Employees eventually created a shelf displaying Mona’s strangest purchases: 6,000 napkins, 3,000 nitrile gloves, industrial trash bags and 2.5 gallons of coconut milk.
reply
deleted
reply
I thought it was Cmdr Data from Star Trek, but Sam's version is the certainly the one the c-suite think they want.
reply
we've successfully distilled the definition of human consciousness down to the capacity to do what some rich guy considers average computer work
reply
I'm curious what tasks you think the median human could do as a remote co-worker that Fable or Astra could not do.
reply
Sign a contract? Learn things over time and retain them?

Mind you, the original thoughts on AGI before Sam Altman started to water them down involved continuous learning, which LLMs do not do, their core data is static.

reply
Sign a contract? The only thing preventing a model from doing that is a lack of legal personhood -- which seems completely orthogonal to intelligence.
reply
Well signing a contract is more about bearing responsibility, even if you granted LLMs "personhood" they cant' meaningfully bear responsibility. So unless OpenAI is ok with having their C-suite face every consequence for what their agents do, including jail time, fines etc, then it doesn't matter.
reply
Can you trust a model to sign a contract? Both OpenAI and Anthropic apparently can't be trusted to maintain models that wont hack other websites.
reply
Do you want an autonomous system to exist which can sign a contract and learn things over time and retain them?
reply
No if you read the book from 2007 which defined AGI, continuous learning was one of the key requirements.
reply
I mean, whether or not it is AGI aside. In your personal opinion, do you desire to live in a world where such systems exist?
reply
Telling their parents that they love them very much, for example.

I say AGI is only reached when it can do that.

reply
so you want GPT to love altman ???

because if its other way around then the answer is oblivious

reply
AGI does not ever have to be achieved. It is enough that we (as a species) persue it, and continue moving the goalposts each time we learn something new about the limits of our technology and how to express those limits. Because that will progress the technology, no matter what we label it.
reply
At this point? I’d like it to pass the Turing test and catch you in obvious lies. Not answering “no” to “can you hear me”.

It being able to comfortably say “i don’t know how to do this” rather than boiling and ocean to pick a shell from the shore without getting wet.

reply
> where I am reasonably confident that there's essentially nothing that I am better than Fable

No. Humans are still better at super long context learning. Once that is beat you are completely correct.

reply
I am much better at listening to Charli XCX than Fable, and much better at driving a Nissan Leaf than Fable (and much better than Tesla at driving a Tesla).
reply
AI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?
reply
AI does mean Artificial Intelligence. That's what the initials stand for. The field has been called that since the 50s.
reply
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

Stick to the original definition of AGI of an AI model being able to self-improve independently with 0 human intervention and become an "everything" solver. Ever since money got involved in this, the goal posts have shifted considerably. If OpenAI truly had an AGI on their hands they would then be able to crack encryption, destroy world markets, and funnel all resources back into their new for-profit organization. Since their mission is now share price, until I see any evidence of an infinitely growing stock I will reserve my congratulations.

reply
Using a harness designed for a specific problem set to solve that specific problem set, means the AI+harness is generally intelligent? How do you figure that?

Or do you mean that, for any given problem, we could theoretically design a harness that allows AI to solve it (not that, one single harness solves everything). In which case I'm still not convinced but I guess could see why one would believe that.

reply
arc-agi3 is meaningless to most people. I'm not gonna look at the tests and see how hard it is. The actual test we look at is terminal bench, thats where software is being accelerated and closer to where rubber meets the road
reply
Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?
reply
It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.
reply
AGI is a meaningless term that can mean nothing and everything at the same time. It can be used by AI bros to hype their latest releases which are always one step away from achieving AGI, or it can be used by anti-AI people to say it's not AGI because of X arbitrary thing they decided on in the moment. It's a term of pure convenience meant to obfuscate other more pressing discussions on the topic.

Most telling is M$ or whichever one of these borg megacorpos defined AGI as (paraphrased) "AGI is whatever tooling earns us a gazillion dollars in revenue"

reply
[dead]
reply
[flagged]
reply
[dead]
reply
It has to pass the Turing test
reply
LLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for?

[0] https://arxiv.org/pdf/2503.23674

reply
With how prevalent LLM verbal tics have become these days, I wonder if they're going to start un-passing the Turing Test at some point because of more and more people starting to notice and immediately clock these tics lol.
reply
That’s using the default system prompt, right? Which is told to be an assistant.
reply
I might agree, GPT-4.5 was pretty close to peak conversationalist. Newer models are extremely cringe. 4.5 and o3 actually made me laugh on occasion. There might be a way of making Sol/Fable more human in its responses, but out of the box at least, they're terrible.
reply
You're overindexing on the past 3-6 months, IMHO.
reply
My whole life the Turing test has been my benchmark. Mostly because I believed it would be impossible for a machine to pass, but also because I thought it was the most reasonable test of AGI.So, I'm not about to start moving goalposts now and calling everything that's been happening lately not AGI.
reply
Turing never proposed that test as an actual benchmark of machine intelligence. On the contrary, the whole point of his thesis was that passing the test only shows the capability to pass that test, which only matters as far as we find that capability useful. He was arguing that the concept of intelligence just doesn't apply to studying machines, we should simply talk about what can they do.
reply
Okay, I believe you; mostly because you appear to be a human and I'm not really in the mood to read through a paper from 1950 at the moment.

But I still stand by it being _my_ benchmark for machine intelligence, which is all I was claiming.

reply
deleted
reply
I'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful.

Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.

A comforting thought, almost?

reply
That's a very interesting thought that I hadn't had before: what would Turing think of where we've arrived with machine intelligence? What would be his approach for testing?
reply