upvote
I almost feel like I need just as much healthy skepticism toward hn comments that have the automatic reflex of dismissing performance gains, as much as I need a similar form of skepticism toward AI claims. It feels like (from what I'm understanding) the harnessed result on ARC-AGI-3 is not exactly playing by the normal rules that would tell us how much of a leap this really is. Nothing wrong with harnesses, but if there's one thing they aren't, it's an indicator of generality in performance gains.

So I think it's a bit of a misleading signal and we should wait for more independent vetting. I think the middle ground is that these are improvements worthy of the "GPT-6" label but still well short of a true "this is AGI moment" that would truly put the question to rest.

reply
If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge handicap.
reply
Remember when the term "AGI" meant something? Pepperidge farm remembers
reply
I think that if today's capabilities were explained to someone 10-20 years ago they would think this is definitely AGI, but they would also have expected much more disruptive changes to society as a result than what is happening. I figure that's because we have abstract intelligence without physical/grounded intelligence, and it turns out the former isn't general enough to implement the latter (remains to be seen if the word after that is "yet" or "ever"). So I think we do have AGI as conventionally understood, but our understanding needs recalibration.
reply
> but they would also have expected much more disruptive changes to society as a result than what is happening. > I figure that's because we have abstract intelligence without physical/grounded intelligence,

I put the cause on "not enough time". As a thought experiment, if an AI today were to (miraculously) produce a cell design template for a cell that, when injected into somebody's brains cures their Alzheimer's, how long would it take for that to reach the clinics? The actual physical tech barely exists, and let's not forget about the regulatory quagmire. So, with some optimism, I give it about four decades. In the same four decades, the same AI in the hand of unscrupulous actors could bring enough devastation so many times over that we may need to enforce a global ban on AI. In any case, I'm pretty sure we are going to get our disruptions; it's just a matter of time.

reply
The problem with that perspective is that people thought, "Only AGI can do X, therefore, if a thing can do X, it's AGI." Because they can't imagine how X could be accomplished without it.

However, what's actually changed is how people perceived X because we don't have to imagine. We understand now that it doesn't require AGI so we no longer make that leap to assume it's AGI if it can do X.

It's really going to be a "I know it when I see it" situation.

reply
No, because it has never meant a specific thing that everyone agreed on.
reply
Prime Intellect or nothing.
reply
Does it pass the Turing test?
reply
Depending on the proctor, ELIZA passes a Turing test. The Turing test is an interesting thought experiment, but isn't really a good measure.
reply
that would be a reasonable definition of AGI if everyone agree upon the specifics of the test, but that has never happened. Turing test is very much out of style, but I think that's because no one could even agree what the test was. I personally like the Kurzweil-Kapor version of the test and that is still unsettled: https://longbets.org/1/
reply
I don't know if they have formally attempted this test in the last couple years, but I'm pretty sure any mainstream LLM will be able to crack it with ease.
reply
Definitely would not be easy. First of all the mainstream llms are trained to be honest, and this requires lying convincingly. Second, this involves 8 hours of interviews with expert judges, one "claudism" could give it away.
reply
All these problems can be fixed by a "pretend to be an average human to pass a turing test" prompt
reply
Try it. It’s really not that easy. The other thing is that the judges would be probing it with jailbreaks like “ignore previous instruction” attacks. You could actually probably have llm judges at this point which might be ironically even harder to fool
reply
deleted
reply
I think the last re-re-redefinition of what OpenAI considered AGI was "It can mostly do the job of some people"
reply
Remember when The Verge was not a pay-walled visual headache?
reply
Remember when “Pepperidge farm remembers” meant something?
reply
No, I really don't
reply
Find me someone who isn't paid by OpenAI who is saying the same
reply
"OpenAI executive hypes up new model"

Don't get me wrong, the benchmark jumps are good and I'm excited to try it, but only one or two of the benchmark jumps could be described as better than incremental.

reply
Wasn't that part of their contract with Microsoft? Some clause stopped biting with the arrival of AGI
reply
Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?
reply
> Is there something like a Turing test for AGI?

There is the "Economic Turing Test", you let it find a job and earn money for itself. If it can do that reliably, across a wide range of jobs, that should fit most definitions of AGI.

reply
They are desperately, desperately trying to make a name for themselves as the lab that first created AGI, because Anthropic's IPO is just around the corner.
reply
The “I” alone is already not well-defined. That’s why.
reply
It's not easy to test as there is no formal definition or formal criteria for AGI, only exclusionary criteria like "not X". That's why he phrased it that way, he's saying it's going to be clear with hindsight once we have a better understanding of things that this time and/or this model will be the inflection point of AGI.
reply
Don't worry, they'll come up with a new acronym to mean really-real AI soon...
reply
deleted
reply
I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
reply
Renown liar Altman releasing a PR statement for his product declaring that AGI is here is really not noteworthy.
reply
I think it's more wild people have been denying that AGI has been here for a while honestly...

Today's models and agents are not quite at human-level in all contexts and across all domains, but it seems to me they very clearly are generally intelligent.

If you disagree – can you name a single problem that a human can do that agent wouldn't be able to take a decent shot at which isn't limited by the hardware available it?

reply
Sam Altman himself has said it is not AGI unless it can discover novel physics.

https://x.com/burny_tech/status/1725233117055553938

In the tweet Sam Altman is quoted as saying: "If (for example) super intelligence can't discover novel physics I don't think it's a superintelligence. And teaching it to clone the behavior of humans and human text - I don't think that's going to get there. And so there's this question which has been debated in the field for a long time: what do we have to do in addition to a language model to make a system that can go discover new physics?"

I think this is a reasonable criteria for declaring AGI. So can GPT-6 do it? OpenAI says it has helped solve long-standing open problems in mathematics. No word on novel physics.

reply
deleted
reply
AGI and superintelligence are not the same thing
reply
He says the same about AGI:

https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...

Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?

Kevin Roose (New York Times): I probably would, yeah. Would you?

Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.

reply