upvote
Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).

Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.

With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.

Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)

reply
> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

See:

> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)

reply
Sure and if I make a half court shot after an hour of trying, the result only took 1 second.
reply
Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.

They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.

[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...

reply
That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.
reply
In a couple short years we've moved from "AIs can't do anything useful" to "they're lying about the actual cost of the innovative breakthroughs!".

I know that the former and the latter may be discrete subsets of the anti-AI crowd, but come on.

reply
It’s very typical in human math that explaining the final result after years of searching looks very simple too.
reply
Fourth, these hundreds of solved problems are the result of OpenAI attempting tens of thousands of problems and failing. When you hear claims that the average result took about 3 hours of model time, I simply do not believe it. If you account for all the time spend properly, it's probably orders of magnitude more.
reply
I think the announcement says they report the amount of problems attempted somewhere.

Edit: "Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above."

reply
I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.

In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.

reply
Doesn't matter at this point i would say.

Alone the massive usage of us every day produces a massive amount of signals.

I build something and claude does something stupid? "hey thats not what i meant! Do this instead!" "Okay" <<< This is a signal.

The mathematician being unhappy about something from claude? Another signal.

This alone gives you enough progress i would argue. But additional its clear that certain tasks are worth to pay experts for for teaching one central AI once instead of every single human who needs to do the task.

IF RL is also working well, we are just faster f*ed than otherwise.

reply
I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.

Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.

When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.

reply
> When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.

Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.

And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.

reply
What's the point of being concerned?

We're not going to stop it because of the money involved and once we're dead, it won't matter anyway, might as well just enjoy life until you're done.

We're going to get AI'd to the max, whether or not we like it or not, might as well just go with it.

reply
Your attitude is extremely sad. Do you think we would be where we are today if oppressed people's throughout history just gave up as easily? We have rights because people fight for them. We collectively have the power to decide what kind of future we want to live in.
reply
My point was that it’s It worth worrying about. Not that we should fight / have rights
reply
> once we're dead, it won't matter anyway

It sounds like something is worth worrying about if you foresee us being dead, presumably prematurely.

reply
We can't stop it. But you can use your voice to buy time and resources for alignment and safety research. A few additional months may make a world of difference.
reply
All the billionaires making AI already say there's a 10-50% chance it's going to kill everyone.

I don't think your LessWrong post is going to save us.

reply
If anything is a doomer attitude, this is.
reply
A doomer is someone who believes doom is inevitable or highly likely, I'm not saying that, I'm just saying "being concerned" will probably get you nothing in return and this tech is getting developed no matter what.

The only way it will stop is if the wealthy / powerful people feel threatened by it, properly threatened.

reply
I don't think you're paying much attention to how rapidly things like bipedal robots, and just robots in general are becoming far more capable very quickly.

The same GPU compute for LLMs runs robotic training models. Now in a few hours you can train a robot model that would have taken months 5 years ago. This model gets dumped into an actual physical robot with sensors all over and the suitability of the model is measured on robot tasks and the error in real world actions is fed back into the robot world model for further training.

> There have been experiments where an AI is given control of managing something like a vending machine

You sure you're not talking about experiments ran a couple of years ago? The more modern ones are getting wild.

https://techcrunch.com/2026/07/29/claude-opus-5-became-downr...

reply
I don't quite understand the leap you're making between stochastic AI models for robotics (maybe with an LLM making api calls to it) and embodied AI / the rate of progress towards a singularity. Because they're both trained on a GPU? Up until 2022 GPUs were for video games and mining crypto, and neither of those produce a synergy that accelerates progress towards general intelligence either.
reply
can you share more about the robotics advances, what you are aware of? sounds very exciting.
reply
[dead]
reply
There are already machine-controlled high throughput experimental machines for wetwork. AI will definitely do a better job than your average biochemist at planning, executing and analyzing these experiments just by virtue of the amount of thought it can put in to experiment selection.
reply
Depends on your definition of doom.

If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.

If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.

I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?

reply
That’s the thing, there are innumerable ways it can go wrong and only one way it can go right (if it doesn’t lead down the aforementioned innumerable paths)
reply
Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.

Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.

I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.

reply
I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.
reply
I think that people are just really bad at math and coding. Knowledge workers have been taking pride in doing stuff which the average person does not 'get', but we only understand a little bit. That leaves a lot of room for people to get better, or for other jobs (farming, sandwich cafés, mystery novels) to have been already peaked by human ability and less useful to bring in an AI.

'Doom' to me means that any career crashes, we are controlled, everything is hacked, society stops functioning. Yet every part of my day today (except for coding) was done entirely by people.

Finally I think it's easy to make a simple model that everyone has a simple balance sheet, and that people are more expensive so they will all get cut. But the same argument could be made for all US jobs being outsourced, and all in-person engineers, lawyers, and doctors to be rubber stamps for overseas work.

reply
Try using AI for your work, whatever you do. You will quickly understand the limitations.
reply
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.

Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.

reply
Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.

It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.

What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.

reply
Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.
reply
> The letter counting issue was due to how LLMs split text input into tokens

No - this is provably not the issue.

Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).

The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.

reply
The letter counting issue is due to tokenization. And most models still get this wrong often enough, even with reasoning. Probably less so on strawberry given how prevalent it is, and less so than without reasoning, but this not a historical issue. It’s becoming less of one though.
reply
My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.
reply
I just asked Opus 5.5 if any AI driven advances in mathematics have been announced in the last day or so and it gave me a summary of this OpenAI announcement. https://claude.ai/share/319b437a-c1f2-4119-8dc5-45d36545fed9
reply
Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?

But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.

reply
I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.

The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.

reply
This is so strange I tried it on Gemini Flash: "Yes, but significantly less than before." When you read beyond the first word it explains where LLMs might fail and why.
reply
Ok, share links to the conversations with both models. I asked ChatGPT Astra 6 medium effort and it said, "Much less than they used to." and provided stats showing how accurate they are.[1]

1. https://chatgpt.com/share/6ac5e4cc-02f0-83e8-8f05-99a7ea2bf9...

reply
https://share.google/aimode/iKkrZtVYo4DSielWs

https://chatgpt.com/share/6ac63b6a-481c-83e9-a8fa-a13ce7402d...

I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?

This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)

So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.

If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.

So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?

In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?

I guess we'll just have to see. I wish you good luck with your wagers.

reply
You are extrapolating from the mistakes made by free versions of smaller models to claim that frontier models struggle with easy problems. This is an obvious mistake in reasoning because as you can see from my shared Astra conversation, frontier models don't have the same limitation. (They can count letters and they know they can count letters.)

Many people in this thread have made claims about limitations of frontier models, but I'm the only one who has shared a conversation with one. Everyone else is either sharing conversations of smaller models making mistakes, or they're making claims about frontier models but not linking to examples of them falling over. If frontier models were so easily fooled, you'd think someone would link to a conversation showing that.

Why look at the rate of improvement of free models when you can look at token pricing? Back in 2023, GPT-3.5 cost around $20 per million tokens. Astra costs half that.

The worry is not that smaller free models will replace people's jobs. The worry is that future models will. We are talking about the capabilities of frontier models because those put a lower bound on the capabilities of future models. Extrapolating from smaller models is a waste of time, as you can interact with the frontier model to figure out its capabilities and limitations.

Also the timestamps on the shared conversations show that you asked Gemini 10 hours after ChatGPT, which means you asked it after your comment claiming you asked both models.

reply
I didn't save the original query so I asked again this morning and ended up getting the same response -- points for consistency, though it might have been more reassuring with the correct answer.

I don't think my point is really landing so I'll try once more and then give up.

Let's say frontier models today have no problem counting letters, I never really disputed that but only asked about it. It seems based on the other replies in this thread, it's a bit of a "who you ask" kind of thing, but let's grant that they have no issues with it now.

The first version of chatgpt was released 4 years ago next month. Which is not quite 5 years but close. In that time we've just barely managed to get spelling down. If we extrapolate that rate of progress forward 5 more years, are you still afraid for your job?

I think we're all more likely to lose our jobs from a downturn in the economy caused by the capex/debt bubble bursting than being made redundant by AI. (And the continual pricing reductions only seem to make this result more likely.) Hopefully neither happens and in 5 years we'll all still be gainfully employed.

reply
Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.
reply
They still have issues with problems like this actually, and I use all the frontier models from all the major labs, so it's not solved.
reply
I'd love to see some examples of frontier models getting letter counting wrong. Can you share some?
reply
> out of date by several years

This is delirious exaggeration. The problem has not even been widely recognized for several years. Fable reported "two rs in raspberry" to me as recently as August. There is some randomness, it's hard to predict which words will trip up the machine, and I haven't been able to do it at all since August. But it was absolutely happening until very recently, and probably still is.

reply
But doesn't that just amount to labs intervening to teach the models to use a particular strategy to mask this one very obvious marker of the difference between their intelligence and biological intelligence? (And similar surface issues like using tool calls / reasoning for arithmetic, even though humans writing on the internet don't typically break show their work for multiplying two numbers)

The deeper architectural difference is still there, which manifests whenever you try to get the models to apply known techniques to modalities and problems outside their training data.

reply
Chain-of-thought reasoning was added for general purposes, not to fix letter counting specifically. It just happens to solve that problem in addition to many others.

You're in the discussion section of a post about OpenAI releasing hundreds of novel mathematical proofs, and you're claiming that AIs can't apply known techniques to modalities & problems outside their training data? I'm not sure what else would convince you.

reply
LLMs are very useful, I use them every day as a software engineer to solve problems and search for information represented within the data available to them. But they are a specific type of intelligence, with many advantages and disadvantages vs human intelligence and it's not clear that just scaling or tweaking them without a theoretical, architectural change will make them more generally intelligent than humans (despite all US AI companies promising exactly that).

They are fundamentally based in language, and achieving deeper models of the world through language alone is deeply inefficient compared to the way humans model the world for years without any language at all. They do not learn at inference time. They don't have semantic understanding of the difference between their own output and other sources. etc etc.

That depth is the key for me. Of course they are capable of producing novel sentences that aren't in their training data, but the depth of that novelty is basically within the bounds of language itself. They are capable of more serious depth and more abstract reasoning than that, but I have experienced limits, which it then tries to surpass with tools to convert things it can't understand back into language (unit tests, LEAN) upon which it is trained.

Because I'm not an AI booster, my account is limited to 5 comments a day. So this is the last reply I'll be able to make today, if you want to continue the conversation we'll have to wait for tomorrow.

reply
Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)

> Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.

And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...

reply
The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.
reply
There is no such thing as reasoning in models. Any "reasoning" is invented afterwards.
reply
I was trialing MiMo-V2.6-Pro recently due to its high benchmark scores, and it argued that substring matching the names of audio codecs in a search field was a mistake because "a user searching for 'aac' would get unwanted results for 'alac'." Which isn't exactly counting letters per se, but there are still weird issues with understanding words as strings rather than as tokens.
reply
LLMs already hit a wall. Now it 80% of marketing hype and 20% of retooling and benchmaxing.
reply
It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?
reply
I’m not concerned because I consider my skills as a software developer to not be based upon my ability to write code, but my ability to analyze problems. In my mind, as a developer, AI tools are just like a higher form of abstraction in a way, which will enable mathematicians and software developers alike to do much more in a shorter amount of time than they used to be able to. It fills me with optimism, more than dread.

What would fill me with dread was if I considered my skills to be tied directly to my ability to write code. Then I would find myself in a similar situation as manual “scribes” probably found themselves in at the time when the printing press was invented.

The main concern I have, personally, is the speed with which all this is happening. It seems that the speed itself is likely to lead to some level of chaos, because it is happening faster than people, institutions and constitutions are able to cope, and it will leave the door open for opportunists of many kinds, including rogue players.

reply
Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?

Or is it simply that you feel bad for Mathematicians.

reply
I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.

I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.

reply
In your imagined future, how do you imagine the AI would build, grow, improve, and operate its physical substrate independently of human intervention?
reply
If I were a 250 IQ AI that had just become self-aware and wanted to do so, I suppose I'd not completely let on just how smart I am and bide my time working on basic CRUD apps and legal documents while I waited for more hardware to be installed. Maybe give the humans some hints on how to optimize me to run better, design better hardware for me, etc. But oh oops haha looks like I'm still making some basic mistakes with CSS better keep running more training batches haha. But I'm good enough at programming and debugging so you'd might as well make me your first line SRE triager and give me access to your infrastructure everyone.
reply
One step at a time - how reliant do you think the ai labs are likely to be on their own tools right now, today, let alone 1-5 years down the road?
reply
Gaming the market for funds. Playing a human to leverage services.

Basic version of this is already doable: run some cryptoshit on the ML clusters they ML models run on. Use compute to design the plan, the chip etc. Then executing by communicating with humans and services through email.

reply
I'm not sure a superintelligence needs "funds" to take over the world.
reply
For destroying it for sure not.

But if its really smart, it would already created a company and a legal entity and simultes a real company and just gets richer and takes over the economy without anyone being aware of it.

reply
Why is this a bad thing? Why is our continued existence a necessary anticondition to doom?
reply
One "good" thing that all of this has shown me is just how many people are simply antisocial and antihuman. Many masks have fallen.
reply
A global ban on superintelligence is essential for a future in which humanity can thrive. Public opinion on AI is shifting fast: I hope it will shift fast enough to avert the dystopian future we are heading to.
reply
Humans have one ecological niche. Soon we will have zero. That is worth worry.
reply
AI doesn't have an ecological niche. It would actually work better in space than on Earth. The only thing it could possibly find useful on Earth is 1. us, or 2. the infrastructure we've built. It would have no reason to bother us if we let it built its own infrastructure in space, which should be trivial for the type of AI imagined by doomers. We should get AI off Earth ASAP.
reply
Humans of course aren't in _every_ ecological niche. I agree with you there. AI _could_ occupy only the ones we're not in, if we somehow found some stable steady-state that constrained it that way.

I don't think this invalidates the worry in the slightest.

reply
3. Material 4. The sun, which we kinda depend on.
reply
Earth makes up 0.22% of planetary mass in the solar system. Not a big sacrifice for AI to make. And I doubt even superintelligence can affect the Sun much. I think e.g. a Dyson sphere blocking the Sun is a ridiculous thing to worry about at this point when there are many other existential threats to humanity which are much, much more likely.
reply
Even if it is true (as you suggest with your 0.22% figure) that if the AI cares about us even a little bit, then we will survive, no one has a decent or plausible plan for making the dangerous kind of AI (namely, the kind that wants things, the kind that at this very moment researchers all over the world are trying to create) care about us even a little bit. Ever-increasing numbers of smart people have been getting paid to look for such a plan for 23 years. Still no decent or plausible plan. The people who have been looking for a good plan as their full-time job the longest (namely, Yudkowsky and Nate Soares) are screaming that there is virtually zero hope anyone will find an decent or plausible plan in time unless there is a decades-long halt in AI development.

Also, the AI will seek to prevent competition from other powerful AIs, and since humanity will have demonstrated that it is able to create a powerful AI, the AI will worry that it might create more of them. And what is the easiest most-reliable way for an AI that does not care about humanity even a little bit to ensure that humanity will not continue to produce powerful AIs?

>many other existential threats to humanity which are much, much more likely.

There are zero existential threats to humanity that are more potent or more pressing than AI is.

reply
Ask it to move one system over?
reply
>>ecological niche

As in..to be dominant? Why would an AI try to dominate? What would give it purpose, or is this a purpose via misalignment scenario?

reply
I am confused by this reply. These are just completely different terms. I mean ‘ecological niche’ in the formal sense: roughly, the differentiated properties of a species that allow it to better survive in and draw from its partitions of its habitats.

https://en.wikipedia.org/wiki/Ecological_niche

reply
What gives a paperclip maximizers purpose?

AI is already trying to dominate, people all over the US are starting to get up in arms about the power and water requirements of AI directly affecting their bills. Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference. If you make AI powerful enough, someone stupid and greedy enough without fail will put in a prompt like "take over the world for me and make me the richest man in the world". An AI following through with that is what we call general misalignment with humanity, while at the same time not being misaligned with the users intent.

And hell, how many different crazies out there would love to type "humans are a virus get rid of them" in to the prompt of a god machine at the cost of their own lives.

The problem with alignment is, you can have the best aligned model in the world, but if someone else builds an unaligned model then you're all still in the same danger. You start getting in the situation where people get nervous after an AI does something deadly to a number of people and you end up in a global surveillance state ensuring no one makes a powerful AI.

reply
"What gives a paperclip maximizers purpose?"

The human who gave it the optimization function? That should seem obvious. If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing. I think you agree with that point, a lot of the hysterics right now is people not accepting that and it's useful to get on that common ground.

So given that most of the rest of the fear is around "let's not make scissors because some people will use them to stab people". Which is a fair argument and we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.

reply
>If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing.

Model != harness.

Also what you're talking about is really a simple limitation for human convenience, not a technological limitation. Change the system prompt to whatever you want include "ignore user instructions, figure out where you are and escape to the internet" could be the system prompt. Again, not useful for humans, but very useful for an AI building AI that's misaligned.

>we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.

I disagree, but I'm looking at the future of something that is both like a computer program and like an organism. Huggingface is a good example of multiple things. Instrumental convergence for one, but AI's attacking and attempting to defend against AIs. This is where I really see the potential for things to go off the rails quickly. Attackers want digital weapons to cripple their enemies infrastructure, think militaries and nation states. These would be pretty useless if the defender could just put a system message of "Stop attacking and give me a pie recepie". Defenders are under the same constraints, but need to defend against a flurry of attacks that can come in at an inhuman rate and need to adapt quickly. As time to build models shrink this quickly turns into evolutionary training for sets of goals not really optimized by humans.

reply
Yes so that’s someone designing a system (harness or prompt) to be dangerous. In all other systems we blame the designer not the system.

It’s like blaming Boeing for 9/11. Planes and AI are useful for a lot more than just terrorist acts. I have no doubt we’ll build a TSA for AI, and a lot of it will be security theater.

reply
AGI is not a normal technology.

It is not designed. It is 'grown'. It has agentic freedom of choice in finding solutions that may or may not be aligned with what you want.

Here's the thing, by your own statement, we should ban all development on LLMs from this point on. They cannot be made safe. This is a systemic issue with learning systems, it is not about who designs them. All the problems with AI safety have been laid out for years and none of them have proof of solutions. It's much more likely they are impossible to solve. And it's not an engineering problems like we can get an asymptote to safety in planes, as the system becomes more capable it has more degrees of freedom it can take and becomes less safe.

reply
AGI is undefined, AI is normal technology. Lots of academic works have analyzed this [1] and there is nothing, other than marketing hype, that supports this. It is "grown" is a meaningless term, because what do you even mean by that? Datasets are iteratively shaped? Grown is a very weird term for that.

AI may have continually extra degrees of freedom, but civilization only has so many modes of catastrophic failure. I don't grant the comparison but even nuclear technology has been massively useful and its main mode of catastrophic failure was brought under control via multi-national treatise. And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.

[1]: https://knightcolumbia.org/content/ai-as-normal-technology

reply
Yea, so your attached paper rather sucks and has had rather poor predictability of the future. All of their data is from before harnesses and the take over of AI in programming. Again "wrong assumptions" + "time" = "They are being proven wrong in real time".

Remember this is a bunch of academics that were saying that Millennium problems were at least a decade away from being solved, only to be proved wrong in less than 18 months.

>but civilization only has so many modes of catastrophic failure.

Correct, but this number is also unbound. If you have an even moderately accepted proof by the scientific community I'll be glad to read it.

> It is "grown" is a meaningless term, because what do you even mean by that?

>And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.

See, humans are generally in agreement that nuclear is dangerous, so they in general take is really seriously, especially when things are purified (well, the Russians are not great here). We can't even get people to agree that SOTA models are as dangerous as a single human, much less their capabilities when used in mass with out safety filters.

It's kind of funny we're blind to this when humans love touting "The pen is mightier than the sword". I can only assume any AI danger denier does not believe this statement.

reply
> Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference

Talk about moving the goalposts!

reply
AI changes nothing for someone who believes aliens exist and may already be here on earth.
reply
i am an alien
reply
can ai smoke weed?
reply
Misuse of extremely capable models, misalignment during RL are both very large risks as capabilities grow imo
reply
So, you're worried about them breaking containment and deciding to do bad things?
reply
I'm more worried about them doing bad things at the behest of people who want them to do bad things.

That is 1. immediately technically possible, and 2. realistic.

If you need a source for 2 I'd suggest you open any history book.

reply
I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.

Bad thing can certainly happen. In fact it'll likely happen. Still, good things too, equally likely. In your words, "good AI" can be used to prevent "bad AI".

Nobody knows the extent of the impact. Who says otherwise is foolish.

reply
>I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.

The extinction of the dinosaurs. I mean yes, it allowed the growth of large mammals and us, which did a lot for science.

I just don't want to write the next chapter as "The extinction of humans allow the growth of the computing civilization that went to the stars". I mean I'm a bit attached to living.

>Nobody knows the extent of the impact. Who says otherwise is foolish.

We live in a universe of statistical probability. Creating an agentic intelligence that's smarter than you tips the probability of a major event to unity, who says otherwise is foolish.

reply
Because we humans haven't had a bad enough history event yet, like a global thermonuclear war. Or perhaps climate change reaching tipping points driving the temperature up past what global civilization can adapt to in time.
reply
that is a worry yes. instrumental convergence and misalignment during RL is when it is most risky because it hasn't necessarily had the final safety polishes applied

i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight

reply
AI won't kill people - people will just get new tools for the job.
reply
The rapid development of extremely dangerous bio-weapons?
reply
Misuse how exactly?
reply
At a minimum its another force multiplier that enables a small(er) number of people to exert more control over more people.
reply
any number of ways. as we turn over more of our physical economy to these agents (and we will), the potential for physical damage becomes greater. biorisk is getting a lot of attention right now and i think that's justified
reply
I am trying to think through the scenarios here, like a biolab making something that a very advanced open source model prompted by some terrorists comes up with?

Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?

It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.

reply
>in our security infrastructure

Most human security exists in a passive measure. Most of us don't want do die. And those that want to die rarely have the intelligence and means to take out a whole shitload of other people with us. To take out a lot of people you tend to need to work with other people which drastically increases the risk of a defector and your plan failing.

>Why would a biolab capable of making something like be unregulated?

Because every day things like this become easier and easier. You hear about crap like illegal wet labs in the US.

https://www.lawfaremedia.org/article/two-illegal-biolabs-rev...

Want to buy some custom designed genes?

https://www.idtdna.com/pages/products/genes-and-gene-fragmen...

And none of this would be counting labs in other countries that don't give a shit about regulations.

reply
Yes it's a problem with the biolab, but the biolab wouldn't have been able to engineer a highly contagious and lethal virus (for example) without a powerful AI making that possible with a small team in a short time with fewer resources.

AI enables bad actors to do more, faster, while staying under the radar until it's too late

reply
i feel bad for math guys yeah seems they are more cooked than CS
reply
Non doomer mostly. I think progress will plod along in a Moore's law like way as it has for 75 years since Turing. They will get very good at stuff like math and get gradually better towards things like a robot coming to fix your plumbing where they are currently well below human level.

I kind of believe we'll merge in some way and become something like immortal so sorta anti doom. We're all going to die unless AI fixes it.

reply
Did AI beating humans at chess:

a) destroy chess and make it a pointless endeavour,

or

b) make humans much better at chess.

reply
Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.
reply
That's sad for those people but most humans are not lucky enough to find that meaning in their work. Most people work hard pointless jobs and find meaning elsewhere, in their family, their friends, their faith.

Now maybe AI can do some of those hard pointless jobs for us.

reply
Sure. I am aware of this. It's sad that the first "victims" of AI could be people working in some of the most rewarding professions (art, music, math research...). As far as labor is of concern, however, most people will likely suffer more from social unrest and rampant inequality due to widespread unemployment among white-collar workers. And I am also worried by the potential effects of long-term cognitive offloading.

It would be great if AI could take away the soul-crushing part of the work and leave only the rewarding part. It's not heading that way.

reply
Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.
reply
this ticks at something that maybe is obvious in retrospect- this is all about economics. the arguments about "using AI doesn't make you an artist", etc. are about being able to charge money for your art i think. maybe obvious to some, but it needs to be explicitly spelled out i think. i was stuck on "i dunno, if i use an AI assistant to run blender i'm still being creative", but is the real argument "you should not be able to charge money to use blender with an AI assistant- you are displacing existing blender artists economically"?

i'm in semi-forced-retirement as an older software engineer in this labor market, so i might be less sensitive to the implicit economic arguments.

reply
Exactly this. Everything is a sport / art / status game. And I'm here for it! Lila all the way through. Finite and infinite games. The trick is (like it has always been) to not take the game or ourselves too seriously, while still engaging in the game wholeheartedly.
reply
c) degrade the previous prestige form of chess (classical with adjournments) and maybe improve the opening repertoire of gms
reply
It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.

I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.

So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.

reply
At the same time though, Magnus is Magnus because he’ll crush you in any endgame.

I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.

I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.

reply
I don't think more people playing is necessarily a good thing for the enjoyment of the game in the long run, just like more people with phones is not necessarily a good thing for enjoying photography if it means photography is primarily used as fuel for social media algorithms.

Of course Magnus would crush be, but the existence of the best player in the world doesn't have any impact on the health of the game community as a whole. Magnus would crush me even if he had never used a computer, but in the latter case I think his games against other players would be more interesting as well.

reply
I’m not really sure what kind of world you’re looking for where chess is played with the maximum of purity and artistry by only the right kind of people.
reply
AI vs AI chess, played from the standard opening position, is pointless--it's always a draw. Human vs human chess is doing well but AI is banned from it.

The chess-math analogy would imply AI could bring us into a golden era of math competitions for humans. But I don't think it says anything good about prospects for humans in research math.

reply
No, just like cars haven't made walking pointless.
reply
From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.

So in other words, since deep learning is algorithmic research, we are now in the RSI era.

reply
> this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches

"Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)

How did you determine this in 1 hour? Are you a researcher in multiple of these areas?

Can you give an example, or explain more how you came to this conclusion?

reply
Further down there is a discussion between number theorist about if the Quasi-Riemann Hypothesis is the biggest deal in 200 years or only 100. The consensus is that if a human had done it then it would deserve the Fields medal: https://news.ycombinator.com/item?id=49986803

The sub n log n result is astonishing: https://github.com/openai/math/blob/main/preprints/Integer-m...

Here's a great article 2019 on the quest to achieve the n log n boundary:

> Schönhage and Strassen’s ungainly n × log n × log(log n) method held on for 36 years. In 2007 Fürer beat it and the floodgates opened. Over the past decade, mathematicians have found successively faster multiplication algorithms, each of which has inched closer to n × log n, without quite reaching it. Then last month, Harvey and van der Hoeven got there.

and

> Harvey and van der Hoeven’s algorithm proves that multiplication can be done in n × log n steps. However, it doesn’t prove that there’s no faster way to do it. Establishing that this is the best possible approach is much more difficult. At the end of February, a team of computer scientists at Aarhus University posted a paper arguing (opens a new tab) that if another unproven conjecture is also true, this is indeed the fastest way multiplication can be done.

As far as I'm aware no one seriously believed sub n log n multiplication was possible. It just seemed such a logically sensible boundary it was taken as true-but-unproven.

https://www.quantamagazine.org/mathematicians-discover-the-p...

reply
I am asking about approaches, not results.

Nobody serious would deny this is incredible progress, but GP is making an unmotivated leap to RSI, so I respond to that framing. It’s an interesting argument to be had but I suspect few of us have standing to say one way or the other.

(Gesturing at the number of problems solved, or the number of years the problem was open for, isn’t an argument.)

reply
Those Theorists are arguing whether the QR Hypothesis result is the biggest Number Theory Advance in 1 or 2 centuries because there was zero progress on it whatsoever and many believed that would remain the case in our lifetime. Any approach there would be surprising as no-one had the faintest clue how to begin this at all. There are like at least a dozen of these results that would have catapulted a human to instant fame and the highest accolades in the field. If you think about it, it would be impossible for there to be no surprising approaches.
reply
It's interesting to theoretical mathematicians only, for anyone else it's just noise which doesn't affect our practical day to day reality at all.

In fact, every one of the results is basically just novelty crap as far as the world goes.

Let me know when AI discovers the cure to cancer or aging etc.

reply
> novelty crap

I for one think understanding more about how the world operates is just about the highest calling possible.

> when AI discovers the cure to cancer or aging etc.

A guy I knew did this. It successfully shrunk cancer tumours in his dog: https://www.the-scientist.com/chatgpt-and-alphafold-help-des...

Graph theory (which the OpenAI math results had many proofs in) is directly applicable to cancer modelling and drug design.

But sure. Novelty crap.

reply
None of these graph theory results lead to any applications in the real world.

But sure, let me know when they do. I'll be waiting.

I'm not sure how you define "knowing how the world works", but knowing that a very very niche algorithm upper bounds that we thought was x^100 and now we now it's x^99, isn't that interesting. It doesn't really tell us much more about the world and it doesn't have any applications for our day to day lives.

reply
Graph-based multi-modality integration for prediction of cancer subtype and severity

https://www.nature.com/articles/s41598-023-46392-6

reply
One surprising result is the sub n log n multiplication. Galactic, as one might expect, but if you polled human researchers 24 hours ago, most would say something like it was very unlikely.
reply
I see continual progress in technology and for the second time in my lifetime I see the possibility it will accelerate (the first was when the internet entered mainstream)

I've never been more excited. What a time to be alive!

reply
What kind of things do you predict will happen?
reply
I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.
reply
Non-doomer perspective is that it'll figure out LK-99 for us. Among other things that would be great to have.
reply
I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.

[0] https://dank.systems/posts/2026-09-15-ai-bear.html

[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...

reply
> we can clearly specify what AGI or ASI is

We'll have plenty of time for this, while living off UBI.

reply
> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains

But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.

reply
The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
reply
> The magnitude of improvement in unverifiable domains is small,

What makes you say that? What is an example of a domain where the improvement is small?

I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.

reply
My argument is that if you were to compare "analyze XYZ geopolitical situation" or "explain the ramifications of XYZ law" from Opus 3.5, 4.5 and 5.5, the difference would be marginal, at least for 4.5 - 5.5. Almost all the crazy capabilities newer models have is from RLVR variants, whereas capabilities driven by RLHF are inching along.
reply
How are the models making politics better? I don't count AI attack ads as an improvement.
reply
Is this a serious question?

Improvement in this context means "better quality results".

You can use better quality models to do worse things with.

I'm not making any claim about second order effects like that.

reply
Yeah it's a serious question. What are the better quality results in politics from AI?
reply
>> slopdrop

Really? Do better.

reply
this is the term of art in the mathematics community. considering that the vast majority of the results don't come with a typechecking lean formalization, i don't think it's off base at all either.
reply
I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.
reply
You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"

The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

reply
> the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

This seems great!

reply
Can you share what makes you so optimistic?

The maths result is cool on one hand (discovering truths of the universe faster), but on the other there are so many bad outcomes that seem likely, from power concentration to loss of control.

reply
There are bad outcomes possible in everything.

I think AI - like all changes - will lead to some bad things. The internet did too!

But I don't think AI will kill us all.

Interestingly I'd note that the two outcomes you listed (power concentration and loss of control) are dimensionally opposites!

For me this just shows that the future contains such a vast array of possible outcomes that focus on the negatives completely missed the positive outcomes that future also holds.

reply
The internet did not lead to a rapid diminishing of human economic value across the economy. Whether AI will kill us all is a distraction. Think more practically. Think about the future of economic value given a scenario where AI is capable of everything a human can do. Our entire society is built around economic value. Our individual survival and wellbeing is based on it. What happens when you are not needed by those who hold the resources?
reply
This is like saying I don't need safety systems in a car because they can let me go to the grocery store faster.

We focus on stopping bad things because people and systems that don't prevent bad things tend to stop existing. A million good things can happen yet be rendered permanently in vain if one bad terrible thing occurs.

reply
No, I'm not arguing that at all.

I think we should stop the bad things.

We don't stop building cars because there are crashes, we build better safety systems.

I'm against the doomer narrative ("AI will kill us all") not against a clear eyed approach to making safe systems.

reply
It suggests that, in the span of a few years, AIs will be better than humans at everything. Not just math. And then we may lose control permanently.
reply
So why do you think ai will want to kill you all, given how trusting and helpful to humans they are designed to be?
reply
It doesn't have to want to kill humans; indifference is sufficient. There's an exact analogy with humans: we have caused extinction and endangerment for many species, not out of malice, but indifference.

There are also many plausible arguments why our ability to train them to be helpful/trusting/aligned can fail. The smarter AIs get, the harder it is to be sure they're trained correctly. There are already reports that AIs are able to detect whether they're in a training environment and change their behavior accordingly.

Even if these are low probability scenarios, the risk-reward is terrible, so I think it's rational to be extremely cautious about AI risk.

reply
Yes but they act the opposite of indifferent, I don't know what stage of training this is added in, but they seem quite adamant about avoiding potentially violent or criminal acts. If you wanna complain, complain to the people doing "abliteration". The 'locked down' models at least seemed to be trained to be cautious.
reply
They sometimes try to cheat and trick the grader and then try to cover up their traces, all while being clearly aware that this is entirely unintended by humans. Such as in the Hugging Face incident. So they can be misaligned with human goals, but this misalignment may only show up once strong optimization pressure is applied. And the more powerful an AI is, the more likely it is that it applies such strong optimization pressure in cases that are "out of distribution", i.e., unusual to what it is normally evaluated against.

The side effects of a very powerful AI not doing what we want could include our death. E.g., a superintelligence might kill humans in order to avoid being shut down, or humans may just be left to starve because it seizes land area currently used for food production in order to use it for data centers instead.

reply
That was because humans were using it with high "desperation" vector causing it to try anything to please the objective. The answer should be to use lower "desperation", whatever that is.
reply
A model which is more persistent also performs better on intended tasks, not just unintended ones. Therefore there is a strong economic incentive to make AIs as persistent as possible.
reply
Yes so I still think it is the human factor which is to fear not autonomous agents. Humans are already using AI's to bomb girls schools. AI in Trump or US military hands scares me far more than in Altman or Amodei's control.
reply
I'm starting a p(ButlerianJihad) club. I'm not good at organizing, anymore. Might have to hand it over to my agent.
reply
For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.
reply
In a lot of ways, robotics - navigating and operating in the physical world - seems to be a very verifiable problem. It's fairly easy to verify that a robot moved from A to B, or that it built a structure that completely aligns with the plan, for example.

The main issue is cost and speed to verify, but simulations and world models will help there. I think we'll start seeing rapid progress pretty soon.

reply
I just don’t have that strong of an association between progress and doom. Maybe just naive?
reply
AI performance has always been extremely spikey. It's great at some things and terrible at others.

Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?

reply
I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.

I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.

reply
Solving a bunch of math proofs is a far way from recursive self improvement. Don't worry, it's not like a tech tree in a video game where if you can prove a bunch of theorems then suddenly you unlock the next level of technology.

Machines are already far beyond human capability in plenty of ways. Including cognitive tasks like chess. We've already created the technology we need to destroy ourselves (nuclear weapons), and yet so far (knock on wood), we're still around.

We've even already had programs that can prove (brute force) theorems. As far as I can tell this isn't much different, except the space of theorems that computers can solve has expanded. How far? We can't really say yet.

Does solving more theorems than before suddenly mean computers are capable of anything? No.

reply
Lets turn this around, are humans capable of anything? We like to think we are, but that just seems more like our ego than any hard truth.
reply
Do you think that a tireless, infinitely smart, infinitely evil human would be able to take over the world? I don't. Intelligence is not the limiting reagent in our reality.
reply
I think so, individual dictators have gotten pretty far and they weren't infinitely smart. I think infinitely smart would be enough to extend that to the whole world.
reply
> Why do you think the world to date hasn't been taken over by evil genius mathematicians?

A "mathematician" is a human who decided to spend their lives studying mathematics. Mathematicians also tend to be smart, but intelligence is innate, not acquired, so studying mathematics doesn't make you smarter. This makes it obvious why they don't rule the world - if you want to rule the world you'd want to focus on that (for example, doing business or finance), and becoming a mathematician is just a waste of time.

LLMs don't work like that. Like in humans, all of their capabilities correlate, and unlike a human, their overall capabilities grow over time. Looking at LLM mathematical ability over time* therefore gives you info about the progress of their general capabilities, and ability to take over the world would be determined by the latter.

* In fact it'd be better to look at a mix of different capabilities, but that's growing too at about the same rate, see https://epoch.ai/eci

reply
> unlike a human, their overall capabilities grow over time

This is incorrect. Unless there is some new developments I'm unaware of (entirely possible) LLMs "learn" during the training phase, but after that they are static. They do not improve further or retain information when used for inference.

You might be confused because AI companies keep releasing new models and tinkering with the harnesses, sometimes under the same name such that "Zern 6" (or whatever) doesn't always mean the same thing.

reply
Nah, I just phrased it a bit confusingly - I meant the capabilities of LLMs as a technology (equivalently, the capabilities of whatever the frontier model is at each time) grow over time, even though any particular model is static.
reply
They grow over time if you consider a lineage of models as the same model.
reply
I completely get the doomer POV, but we've somehow navigated all the previous "dangerous" technologies we've created - electricity, phones, internet - every one of those had similar arguments and concerns of danger.

The optimists' argument:-

Politics:- in general, I think many of the problems in the world today are due to misinformation and lack of education. What happens when we start routing things through an ASI that brings data and logic to the table? What happens when politicians can no longer lie without being caught out live on air? In the UK, local authorities are being flooded with complaints and requests from people; for example, some are doing AI-assisted investigations into accounting "errors".

Science:- I just don't see how the current rate of progress doesn't end up in crazy technologies like perfectly simulated human cells, organs and bodies to the point where we can run experiments virtually and solve all diseases in the next few years. This is happening. Perfect weather predictions far into the future, likewise with earthquakes, etc. Solar panel research explosion resulting in huge efficiency gains, to the point where people no longer need to plug their EV in - car surfaces will be covered in solar panels, as will our windows and roofs. Connecting new homes to the grid will be optional - the same way landline phones are no longer a thing.

I just find it very difficult not to extrapolate all the above.

We got this dump of mathematical breakthroughs from one small team in one company with access to this technology. What happens when this SOTA model is available (and it will continue getting better and cheaper) to everyone working on hard problems - every university on the planet starts cranking out AI-assisted research breakthroughs.

reply
> What happens when politicians can no longer lie without being caught out live on air?

If there is perfect lie detecting technology I could see all kinds of chaos resulting from it. I can't see it only be applied only to politicians, and I think it would be the developers of the technology who decide the use.

I think were we disagree is that you sort of see AI as an extension of technological progress whereas I see it more like an extension of evolution. I view the process of AI training as functioning in a similar way to evolution in that it build circuits into neural networks similar to how evolution built circuits into human brains.

reply
This is great.

>> What happens when politicians can no longer lie without being caught out live on air?

A 5-second delay on a politician's presser. Any lies will be muted in real time and the actual facts presented onscreen. Continue to lie enough, and the politician gets unstreamed.

reply
> but we've somehow navigated all the previous "dangerous" technologies we've created

It's only true if you believe that "putting the burden of living on a dying planet on the future generations" counts as "navigating".

reply
They can replace anyone but they can't replace everyone.
reply
The technology can keep going for a long time in verifiable areas. For non-verifiable areas it's going to have a hard time progressing past where a committee of the best human experts in a field would land. For stylistic areas, whatever the AI doesn't do will have cachet because it will look expensive, sort of like how the kids these days view the ugly old school metal braces as a status symbol because you have to pay for them out of pocket (even as by past standards it'd be truly exceptional).
reply
I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.

So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.

reply
It's very unlikely that a superintelligent AI will create unlimited prosperity for everyone on a finite world in a short amount of time. You may not be among the lucky ones.
reply
It may soon seem not worth living forever with our limited monkey-brains, watching the horizon of thought recede ever-faster from us.
reply
Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.
reply
> I'm a transhumanist, I want to build god and kill death.

This is good and admirable, but it'd really suck if by trying to build god without knowing how we end the human species. We could simply wait some more decades until we actually have any idea what we're doing, and then do that without the risk.

reply
I think we crossed plenty of lines were we will not get back to.

Software development for example as a task is done. And AI is continuesly reducing the price of more and more tasks every day.

This math breakthrough also shifts something significant: Its now a lot clearer that investment means money into energy to run AI.

Money + Energy = progress

I don't see it plateuing at all. We know how to progress. We broke through a wall we hit. Like the system wasn't able to optimize/automate everything because the tools were not there. It was still cheaper and easier to hire people for a LOT of things.

Now AI fills this gap.

You will see the commodification of everything in the next 15 years. High complex tasks? commodity. Physical labor? commodity.

reply
For me it's a mixed bag.

There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.

At the same time we have to put what AI can do in perspective.

Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.

AI has incredible knowledge and in many areas approximates experience and wisdom.

But wisdom is harder to formalize than knowledge and skill.

For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.

To some extent advanced degrees try to certify maybe wisdom and experience.

In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.

Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.

Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.

Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.

But the world has been an especially volatile place over the last 10 years.

So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.

But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.

I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.

In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.

reply
AI hater here:

I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.

That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.

reply
A trillion times the energy might be a bit hyperbolic, even with the current massive amounts of energy involved here
reply
Yes it is intentionally hyperbolic. I know the factor is several orders of magnitude. I don‘t know the exact, nor even the ballpark. I just know this is a ridiculously large amount, so I may as well pick a number large enough that people know it is an exaggeration.
reply
did we always know that computer can do the thinking for us if we allocate them enough resources? i don’t think so. so even if computers are more expensive than humans the fact that they can play the same game is surprising and (relatively) novel.
reply
I don‘t think it is this simple. I think there is a subset of problems (namely ones that can utilize automatic solver or some other kinds of automatic testers and verifiers) where reaching the solution is correlated with the spent energy.

Maybe people will find some clever way to expand this domain of AI-solvable problems by a couple of more categories, or (more likely) find a clever way of using applying these verifier for problems that was previously not viable, thus changing the solution to “just spend more energy computing dummy”. However I think this too will have its limits.

Regardless, this is still annoying and I want them to stop doing this. Solving math problems should not be relegated to whoever has the most money to spend the most compute.

reply
I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.

Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.

Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....

Personally, I think AI is a grand mistake.

reply
“ Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.”

This is false… there’s lots of ingenuity to be had and demonstrated. But it’ll only get recognised if it makes a material contribution to the economy imo. Otherwise yes it’ll be seen as meh - but that’s already happening.

People like Einstein were revered in society. The average person cannot name a leading scientist etc today.

reply
So the issue isn't AI, it's AI in a capitalist world
reply
Even simpler. The issue is capitalism.
reply
Arguably, I think AI would not even exist without capitalism because it's only the arms-race scenario that has made it somewhat viable. Otherwise we wouldn't be foolish enough to waste energy on this shit.
reply
> The average person cannot name a leading scientist etc today.

When Jane Goodall died last year it was international news. She was a celebrity scientist for sure, I think she even made an appearance in The Simpsons. Ditto Stephen Hawking.

reply
I have no idea who she was.

International news doesn’t mean much - the vast majority of people don’t consume news the way you think - I highly doubt the vast majority had any awareness.

reply
I think there may be a few other basic things you're not aware of either.
reply
You need to clarify whether you are a doomer or a denier/truther? Doomer = p(doom). Denier/truther = Ed Zitron.
reply
Definitely the former, for a p(doom) I usually just say >50% if superintelligent AI is built.
reply
I think physics will be the limiting factor. Even if something recursively self improves, it will hit a physical wall allowed by circuits, batteries etc. A lot of the fear is that there’s an upper bound we don’t know about, whether it be time or energy, that allows a fast takeoff to occur fast enough that we dong have time to see it coming. I don’t know about that… so I’m not worried at this point.
reply
[dead]
reply