Scoring well in a benchmark that's called AGI does not make an LLM AGI.
If so I'm hoping we can track them down and have them tell us if they think this is AGI.
Once we have 1000 tps, i am sure robots etc.. will also start working like magic.
If I can't give it an arbitrary task and have it solve that task eventually, it's not a general intelligence.
(obviously it might take years for me to get good enough at something, or if you set the "arbitrary" task as something ridiculous, but lets work in good faith here and think of something the average human could do after learning about it)
If we progress to the point where an LLM instance can meaningfully learn to get better at something overtime without retraining, then I will accept that is basically AGI. Right now, they still seem to be pretty boxed into their training, even if you can prompt them to act differently.
But if you’re asking when a model has a sustainable general intelligence, for me, it’s pretty easy…
When it makes financial sense to run it 24 hours a day.
It makes either position pointless to argue.
Aren't we way way past that already? QPS to any of the frontier models for a given point in time is most likely (far) greater than zero.
Directly - something can be useful without being AGI.
"Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"
In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”
"""
You (and the rest of the media and many industry figures) are conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).
So now humans is "super" intelligence? it's nice to move the upper bar so that more stuff can be called "just" intelligence.
general intelligence for beavers or a birch forest would be very different than general intelligence for humans...
Is it rapid skill acquisition? -> ARC benchmarks are saturated Is it breadth of knowledge? -> See many ... many benchmarks Is it ability to do hard tasks? -> see terminal-bench and released outputs.
We are at the point where the starting point for most tasks should be "send your agent to work on it."
So where do we draw the line in a way that doesn't move every 6 months?
1 year ago we viewed models as tools and agents were just kinda toying around, that we now think the bar is literally an anything to anything converter through one agent is wild.
• 97.6% on frontier math
• 95.9% on CAD
• 100% on ExploitBench
Nothing modest about it
They've released two videos:
Vision video:
https://www.youtube.com/watch?v=1QNsdr-Qx_I
(kinda reminds me of these retro videos about the future home: https://www.youtube.com/watch?v=rnbaehgxdp0) ((can't find the other one where someone controls the home computer with voice))
Vibe coding with it:
Don't be surprised to see other (or even the same) people declaring AGI again and again, as it becomes the best time to do so for different parties.
Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc
I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain.
It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it.
It's scary, TBH.
But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.
Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.
> Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.
It's not good to be confidently wrong the way you're being.
>The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
Perhaps. But only at that point, not leading up to that point.
It's kind of like setting up scaffolding to build something. You spend all of that time and money to build something just to tear it down in the end. But the point is that it's simply a cost to be able to build the actual thing you're building.
If these companies are able to achieve the results they're looking for, none of the investors involved are going to care that the datacenters and infrastructure they spent so much money.
The HN crowd I'm sure will still be unhappy calling it AGI because "it's not AGI unless its speech comes from the cerebral cortex region of the brain, otherwise it's just sparkling emoji" or something.
Those are all things that humanity is doing everyday. What we have is amazing, but it’s not that.
True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
Harnesses magnify and make the intelligence actionable, but we have not reached limits on raw intelligence yet, not even close.
One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.
"The harness improvements are the real sauce" is like a sincere "It's gotta be the shoes" take about Micheal Jordan.
(For the younger: that line was from a series of Nike ads where his skills were being explained)