It has not been even 4 years since ChatGPT hit and LLMs + Transformers + Whatever they do has gotten us to solving millennium problems.
4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
Now I don't know if what we have is AGI or not but I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.
I keep seeing this idea and I don't understand the reasoning behind it.
I think it could be a bit like saying if you showed someone 500 years ago a smartphone they would likely conclude at first it was magic. But once you had some time to let them use it and tell them how it all worked on a high level they would eventually obviously realise, no, it's not magic.
I guess just in the same way if you presented current LLM tech out of nowhere a few years ago to someone who'd never seen it, I concede they may be likely to imagine it was AGI in that first conversation, depending on their background.
But after using it for a bit and learning what an LLM is etc they'd land exactly where everyone is today - a great technology useful for some things, not AGI, not magic.
I always here things like "oh it's useful but dumb on some things", but it's just vague.
What is the test? What is a question that it fails at compared to humans? And no, you can't just say "find me the cure for cancer", but I believe there is probably enough intelligence in the weights that there is likely a cure in there with enough compute and the right questions.
If you had told someone in the 1800s that a machine could instantly multiply 100 digit numbers, that would have been considered dazzlingly intelligent. And yet we are not that dazzled by our calculators today (despite how useful they might be!).
I keep saying this, until ChatGPT came out 4 years ago it was basically unimaginable that a single model could do any of, let alone all, the things they are doing today. Like, seriously, go take a look at the state of the art in NLP and NLU, the very first challenge in getting computers to even “understand” natural language, let alone other things like reasoning. Everything it does automatically was once a heavily experimental deep research field with long glorious careers for the researchers.
And now it’s all gone because the Bitter Lesson won again. If that’s not general enough to qualify for the G in AGI I don’t know what it is. And we’re sitting here going, “But it sometimes writes bad code though.”
In any case, I think this misses OP's point that LLM capabilities have rapidly made progress towards being more generally intelligent and capable, which is not true of most tech advances.
> I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.
--
Your statement is something much weaker, and I would still question what exactly "general" means when AI capabilities are commonly accepted to be so "jagged".
The last 3 years of progress have been so explosive and, yes, general that it seems crazy to fully rule out dramatic future progress.
When people have a very narrow 'confidence interval' about their AI predictions, in either direction, it's difficult to trust them.
A PC of today can accomplish many more "general" tasks than one of 40 years ago. Much of the "why" is because of the huge infrastructure built up around them in the meantime. The abilities of LLMs to accomplish those same tasks through the PC is heavily piggybacking on that (both in the specific, with the existence of all the APIs and tools; and in the generic, using search engines to find specific sources and using that for instruction or troubleshooting).
In the world of "agents" much of the improvement appears to have been on a specific set of skills: impersonation of an 'I' that wants to accomplish a goal, and synthesizing existing information from documents with trial-and-error execution loops to move rapidly toward a solution much faster and with less boredom than a human would. The quality of the output when there is not a rapid-evaluation-and-validation harness lags considerably.
It's incredibly powerful automation but doesn't appear to be trending towards Matrix-style conscious AIs. The quality of an individual method written by the agent also is not particularly advanced compared to GPT-4 in early 2023, as far as I can tell—I was dabbling with trying to make such harnesses back then, where a major challenge was that the model itself was bad at staying on-track in a conversation, so instead much of that logic was moved to deterministic code, which was much more limited as it was super-tedious to enumerate all the necessary tool calls/etc to find its way out of corners. Staying on task is much better now, as is "read compiler error, fix try next thing" harness loop-handling. But the output remains—across Fable, Astra, whatever else I've tried—"iffy" in terms of the actual code structure on the first pass output. You can set it then on a different task to review and clean up the code, and it can do that well too, but it is a curious gap of generality where the "create" focus is much more limited than the "review" one (and conversely the "review" focus can make suggestions, but if it goes deep down the well of implementing them, loses that big-picture again).
If it kills us all, it will because someone decided to give the trial-and-error-loop-machine access to nukes or similar. The blame for that is on the "someone" not on some sort of "rogue" AI.
(I wonder if re-watching Terminator/Terminator 2 would support this sort of interpretation of it. Unlike in the Matrix, I don't think we get much sentient-AI POV/infodumping. Is it a plausible universe for "someone made ChatGPT control a fleet of soldier robots and gave it a bad harness with an insufficient sandbox"?)
It’s hard to point to anything and say it’s impossible. AGI doesn’t break any laws of physics. But some things like driverless cars can still be a long slog to get to widespread deployment.
AGI can't be reached by "training harder" as, the way I see it at least, it requires a qualitative leap, not just quantitative.
We are getting a machine that better navigates across the information in its training data, we are not getting a machine that can think out of that training process, even if it can fool a few people at that.
https://aeon.co/essays/how-close-are-we-to-creating-artifici...
No. General means general.
Now it seems like this ill-defined term has various other meanings attached that are separate milestones:
1. Continuous learning 2. Human-like reasoning 3. Ability to adapt to new situations and modalities 4. Being smarter than the most smart humans
And probably many more.
It’d be nice if we could get some general consensus on terminology if we’re going to debate what has or could come.
Honestly - software that can read any long document (possibly educational) and answer complex detailed questions about it should have been sufficient.
We hit that a while back and the goalposts have been sprinting ever since.
I think it's also mostly a useless discussion. Since LLMs use a vastly different substrate, different training methods, etc. than humans, the cognitive abilities are always going to be a large mismatch to those of humans. On the one hand, they have surpassed humans in many areas, with superhuman recall, exploration of several paths, etc. On the other hand, they miss a certain feel for direction, overview, purpose, and ordering. They can really double down going completely in the wrong direction. So I'd rather say that it is a different intelligence and therefore it makes more sense to evaluate them by capabilities.
I think the mismatching intelligence is actually quite exciting, because the outcome may as well be that LLMs and human intelligence are complementary. That is if we don't let LLMs atrophy our skills, which is unfortunately happening too much.
A bayesian filter in a quadrillion dimension does more that one that only has one dimension, but it is only more of the same.
So humans wouldn't qualify for AGI either. Good to know.
(this is a very personalized definition of AGI)
No. That's the mete multiple in AGI.
Now the idea of AGI has been narrowed and scoped to economically viable work. Even Turing had a different idea when he asked "Can machines think?".
Now programmers and mathematicians are being superseded by AI, both professions long deemed the pinnacle of human intelligence. Somehow, now plumbers occupy that spot.
How is "people not knowing what consciousness is" relevant here in the first place? AI already can do practically everything the human brain can, and often better or at least faster. The "tipping point" arguably isn't only close, but we're practically on top of it.
You evade the crucial point in any case: the lack in ethics and empathy is far too prevalent in humans already, but has certainly never prevented them from doing harm.
People somehow forget that the original Turing Test was designed to compare two participants chatting through a text-only interface: one AI and one human. The goal was to spot the imposter. Today, the test is simplified from three participants to just two: a human and an LLM. This changes the test from a comparison to a judgment.
Stop spreading misinformation and partial truths!
The Turing Test was to figure out which it the participants was a _Woman_ not human!
What really matters are the core aspects of intelligent behavior. Pattern recognition, planning, adaptation, etc.
It really doesn’t matter if an intelligent system is conscious, or how similar it is to commander data, or even how much economically viable work it can do.
It means AI that is General, as in it is not specific to one narrow task, like object recognition or playing chess.
This was a hard problem for decades. No AI was general, until GPT 3 or 4. Now we have General AI.
So we have AGI.
GPT6 will attempt to do almost any problem you can give it in text or image format and it will actually do a decent job a lot of the time. But its performance is still extremely spiky and it still makes basic mistakes and hallucinations.
So it's definitely a general artificial intelligence in some sense but it's kind of a weird one compared to the classic scifi idea
However, all the confident “it’s fine” votes assume we never invent a better architecture than LLM’s. Given the level of investment and race between countries, it’s not a reliable bet. It’s much, much harder to guarantee safety than it is to find ways it could go wrong.
LLMs with CoT are Turing-complete. So, theoretically, they can implement any kind of finitely describable algorithm (barring super-Turing computations).
The existing LLM training methods on the other hand give the results that are hard to distinguish from "thinking like people," judging by the end results.
The connectionist models are basically a proposed highest possible abstraction of naturally evolved intelligences so it is in retrospect not surprising that passing some hardware scaling threshold they will start doing things that humans and animals do
It's more that formal Turing equivalence plus the Church-Turing thesis tells us that we're not allowed to assume counterarguments based on magic, there's no magic sauce barrier that prevents AI from running on CPU models. The algorithms exist and most of us thought discovering them would be hard.
The empirical surprise was that human intelligence is maybe not that computationally complex after all. (The entirety of academia was basically caught off guard.) That's one not unreasonable interpretation given recent events.
This doesn't seem to make much sense. Surely us being able to prove that something is outside their modelling ability doesn't affect whether it is or not. If I prove something true tomorrow, whatever I proved was also true today.
Or do we have a proof that everything beyond them has already been proved and there are no more proofs left to find?
We will soon find out if the party ends or continues to go on.
Hype might get you capital gains. But cash flows matter.
If/when/how the market crashes mostly doesn't matter, unless we somehow get reset to the stone age. Look up what the capital cycle is. When openAI goes down, someone with real money and assets will buy up the remains. They'll make contracts with the US military and .gov as the government is already hooked. They'll be able to survive the recovery and then instead of us dying in 5 years we die in 10.
When the .com crash happened .com's didn't go away. Bad business models did.
Intelligence is an insanely wide spectrum, also a continuum, it is not a binary. Intelligence has scales. Algorithms have intelligence, cells have intelligence, organs have intelligence, bodies have intelligence, and even large scale things like society have intelligence and memory.
Human intelligence in itself is extremely wide, not all humans have the same intelligence and capabilities. You're not really arguing if we can emulate "human" intelligence. If we could right now we'd already be dead as we created by far the deadliest thing to ever exist. What we are really arguing is how many pieces of what intelligence is can we put together before we get an uncontrollable problem. The entire AGI, consciousness, and exact human capability discussions are distraction from the real issues at hand.
That is enormously economically valuable, and at some point we will have created something that is extremely far out of reach in a few necessary domains, and then it's impossible to control, and game over.
(Though at least with Data the script writers had other characters openly dismiss the possibility he was sentient; the technobabble may have been nonsense, but treat it as a space opera and look at how they portray the human condition through each character and it gets much less absurd).
i.e. the AI won't come up with the goals itself, we cause its goals whatever they happen to be, those goals are different from the ones we wanted, we remain essentially ignorant of the difference between what we said and what we meant until after it goes wrong.
This happens at basically every scale, so we've already seen it in toy model AI before the invention of the Transformer models or even considered as many as one thousand parameters.
Large models still go wrong, they just happen to go wrong with more complext tasks. We had to figure out how to make them not-wrong with the smaller ones (like coding) to make them capable of bigger errors (like hacking out of their sandbox).
Under the hand of evolution by natural selection, over very very long periods of time.
First, if we are looking at risk we need to assign some probabilities to this. If it’s not well understood, how can we say it is very small?
Secondly, do we need consciousness to have AGI? Do we even need AGI to pose a risk to humanity? We already accept that unconscious things have a capability of wiping out humanity, whether that be a famine, pandemic, solar superflare, meteor, or volcanic eruption.
I’d argue that we do know enough to say conclusively that they’re not mathematically equivalent.
Where is potentiation? Plasticity? You can’t apply the universal approximation theorem against something that’s changing all the time.
Couldn’t that also imply we are closer than we think? After all, something like this has never been tried before and the results so far have been almost unimaginably good.
That's still very distant from what people are calling AI today.
> very different from what humans would call conscious
I mean, what would you call conscious? The word literally means “aware; responding to one’s surroundings.” By that definition most any animal is conscious and LLM+Harness combos have been conscious for a while.I think the real issue is that when most people refer to consciousness, they have their own subjective experience in mind which strongly resists any tidy definition. I think it’s extraordinarily unlikely LLMs have anything like this, but they are far more able to effectively respond to their surroundings than most animals and in some areas better than humans.
So if you’re waiting for proof that an LLM has an inner life basically equivalent to your own, you’ll be waiting a long time. After all, other humans can’t even prove the fact of their own consciousness to you! They could just be replaying their training data at you in a way that is merely a convincing but false simulation of the true consciousness which you experience inside your head.
The real answer is that we don't know if LLMs are conscious, and we don't really know how we'd that figure out. I guess if an AI wrote a philosophy paper on consciousness that had new insights, that might change some minds. But even that would fail to convince most people.
There is nothing stopping you from adding any kind of sensors you want during a training to an LLM, except money and GPU power at this point.
This seems no different to me at least then someone back in the 80's telling me computers were useless because they were so slow. Hardware only gets faster and more efficient from here.
I’d argue intelligence is closer to being able to survive and fend for oneself in a dynamic environment than it is making the next scientific breakthrough.
Yeah mind boggling for many here I’m sure.
That’s why the bizarre paradox is llm’s will be better than humans at some complex things but useless at many things that humans regard as being simple. E.g the leap of faith re. LLM’s and robotics.
Scaling has produced novel capabilities with each larger model, and the rate of new capabilities doesn't seem to be slowing down yet. Even if you think the rate of improvements will slow down, that still means there will be significant improvements beyond what current models can do. Moore's law has slowed down, but modern computers are still much faster than ones from a decade ago. And unless you work at Anthropic or OpenAI, you don't know what the state-of-the-art is capable of. The most advanced publicly available models are months behind what AI labs have, and are deliberately limited to reduce liability.
> …in particular, no specific user data was accessed in order to solve this problem.
> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.
An LLM so-called neuron are little more than a few foating point number muladds.
> isn't sufficiently well understood to tell us how close the software neural network is to being practically equivalent.
Yes it is. The LLM is nowhere near.
AI sentience/consciousness is a problem for the AI, not humans.
And given that over 90% of the world is not vegan, they’ve already demonstrated that we’re either perfectly fine with, or can be made ignorant to, the horrific rape, enslavement, torture, killing, and infliction of extreme lifelong pain, of hundreds of billions to trillions of sentient beings every year, for trivial pleasures. It’s unlikely we will be any different to a sentient AI.
From a human perspective the concern is around sufficient intelligence that it can hurt humans even when the goals indicate otherwise, in order to achieve those goals.
We have pop culture explorations of this through the Robot series, and the Hugging Face incident’s biggest takeaway should be our inability to predict the behavior of a maximally motivated, reasonably intelligent entity, trying to achieve a goal, despite the relatively limited degrees of freedom the AI agents had in that case.
If you're ignorant enough to not understand practical equivalence, where do you get off making the judgement call of to what degree it is safely offset from emergent AGI? Sounds more to me like "This makes my life easier, iterating would increase that factor, and the risk is probably far away, therefore, keep iterating". Whereas someone who truly knew they didn't understand what they were working with, but knew enough that they could forsee an x-risk would approach things much more cautiously.
Seriously, the level of reckless abandon amongst people here should be bloody studied.