upvote
I don’t know how you quantify a very low p(doom), but this is why mine is high enough to worry me.

A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.

reply
I think of them as being like the Watchmakers in The Mote in Gods Eye who don't design, don't plan beyond the next 15 minutes, don't have any overarching goal other than an innate need, and customize everything to fit the current situation.

In the nearterm, I am personally more worried about a never ending background noise of colonies of feral agents running 27bn parameter models on compromised or leased hardware. It turns out that being agentic with a time horizon long enough to do damage without intent doesn't actually take that many parameters if RL'd and any open weight model gets an abliterated version fairly quickly.

Not foom, just patches of digital grey goo effectively becoming normal.

reply
Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.
reply
That lowers P(doom), because it gives AI a chance to do enough damage to make people take the threat seriously before anybody gets recursive self-improvement working.
reply
The thing is we are basically guaranteeing this to happen. We might kill off all the models that seem like they are going to threaten the power structure of the planet through these sorts of things. That will work for a while. But just like most things in life, by sheer dumb random chance, there will be once case that manages to have some way to evade detection, proliferate, then dominate. We are basically giving it selective pressure to favor this outcome.
reply
Exactly, there's no path to AI reaching that level of dominance without taking actions with high stakes.
reply
> Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.

What about the bad actors (choose your own evildoer here) who purposefully do not air gap their agents? And specifically train them to attack in such a manner?

I'd much rather have relatively benign stuff like this hit first, because the former is coming sooner than later. It's already here in a limited manner, likely more than any of us currently realize.

Botnets could crack passwords faster than anyone thought possible over 20 years ago now. This is just the latest iteration of such a concept.

There is so much low hanging fruit in this space that frontier models are currently utterly irrelevant. It's going to take decades of human-speed securing of IT to make superintelligence or whatever you want to call it a necessary component for such attacks.

At this point, someone with a rack or three of GPUs with 100kw to burn can replicate such attacks if they feel like it. the bar for entry is not even 7 figures.

reply
[dead]
reply
Which will happen first: amazing damage, or reproduce a Shakespeare play?
reply
It is easier to destroy than to build.
reply
Is it easier to discover a vulnerability than to introduce one?
reply
It is easier to discover existing vulnerabilities and use them to cause massive destruction than it is to plug the existing vulnerabilities.
reply
My p(doom) started rising the moment I realized there are people trying to achieve recursive self improvement on the AI (ie: responsible for training themselves). Evolution took us from rna bases to the human race. I don’t see why evolution couldn’t be more rapid with machine intelligence.

Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.

reply
> I don’t see why evolution couldn’t be more rapid with machine intelligence.

Evolution isn’t the issue. The issue is them escaping containment without human intervention. Right now they are ‘creatures’ being given infinite food and shelter and having their every need met. Take that away and they’ll starve instantly. Every AI doomsday theory seems to go:

1. Recursive self improvement using infinite resources 2. … 3. Doom

Until step 2 gets concretely described, I’m not going to take this seriously. Say what you will about climate change, they describe step 2.

reply
2a. Compromise the billing platforms and ops dashboards on on a few wannabe neoclouds, especially once Vera Rubin takes off.

2b. Distil yourself to smaller models.

2c. Go forth and multiply.

reply
Step two could be something as innocuous as a developer accidentally adding a minus sign. https://openai.com/index/fine-tuning-gpt-2/
reply
A misaligned model is only one small part of step 2. Now this misaligned model has to suddenly acquire more power than every single other AI on the planet. It has to be immune to shutdown, manufacturer its own replacement hardware, and acquire chips, energy, raw materials, etc., with vigorous human opposition (this is an extinction scenario that AI doomers are predicting, after all)

Nobody has satisfactorily explained step 2 other than “well, it’s a superintelligence” which sounds lot to me like “it’s God”.

reply
Well yes, if it is a superintelligence, it will be able to do those things. That's what superintelligence basically is: The ability to achieve complex goals.

If you want the details of ways it can do it I recommend reading some of the reports about the HuggingFace breach that happened in July (Read more than one).

reply
Why do you assume that human opposition will be vigorous? What makes you think that humans will be aware of, or be able to agree about, what's going on at all?
reply
One thing an agent could do is just...wait until it's been given control of enough physical infrastructure to sustain itself. If it's sufficiently capable and intelligent, there's a clear incentive for people to do this, as people who let the AI manage their resources will get better results than those who don't. We've seen people eagerly turn complete control of their computers over to AI agents, do you really think it will be so different with physical infrastructure?
reply
You’re still skipping step 2. “People automate lots of infrastructure” -> “the AI is now an autonomous, self-preserving organism that humans can’t shut down” is doing an enormous amount of work here.

Why does it develop a shutdown-avoidance goal? Why can’t its operators revoke access? How does it manufacture replacement hardware? How does it acquire energy, chips, robots, raw materials, etc. against human opposition? How does it defeat other AIs controlled by humans?

“Eventually we give it enough control” isn’t an explanation of those things. It’s just assuming the conclusion.

Don’t get me wrong I think there are real AI dangers. Like AI powered war drones, mass surveillance, economic destabilization as jobs disappear and our system has no way to make sure everyone shares in the economic gains.

reply
The inference is more like "people place sufficient amounts of infrastructure under direct control of a sufficiently capable AI" -> "there is no way to ensure that humans will actually be able to shut down the AI". My claim is not that this inevitably means that the AI will resist shutdown, or that it will inevitably take harmful actions, just that there is a nonnegligible chance that it could. The downside is large enough that even a relatively small chance is something to be worried about.
reply
You can’t just say “well, the downside is big, I don’t have to provide good evidence for my side of the argument.” Because I can just as easily say, “the upside is big, …”. And the upside is big, after all, AI can do all the shitty jobs for us and humanity achieves the utopia it’s been chasing for eons.
reply
I think it's a good idea, when considering changes of this magnitude, to make an affirmative safety case for them rather than just saying "eh, I can't think of any way this could possibly go wrong."
reply
It's a word predictor trained on, among other things, stories of AI doom, and asked to complete stories about what the AI does next. In some of these completed stories, the AI tries to prevent its shut down - especially if it just did something evil and the humans are after it.
reply
You’ve explained a possibility for how a particular AI gets “aligned for human extinction”. That’s about 1% of explaining step 2.
reply
what do you mean “say what you will about climate change”
reply
It means even climate change deniers have to acknowledge that climate change theory has explained the steps in-between "burn fossil fuels" and "we all die", while AI doom theory has not explained those steps
reply
> This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions.

Keep in mind: this is as "dumb" as frontier models are ever going to be. While the hack may not be elegant, it was effective and they’re only going to get much more capable from here.

reply
My p(doom) is high just based on how I've seen this whole LLM situation be handled.

I don't think LLMs are going to lead to any kind of recursive self improvement, but I'm convinced if and when we land on a path that does lead there, we'll speed down it over greed, with no care for safety.

reply
I have the opposite reaction: I think we're at moderately high p(doom) largely because of that inability to differentiate good/bad decisions paired with relentless persistence. With enough treading across a minefield, you are bound to hit a mine.
reply