83 Comments
User's avatar
CarbonWaster's avatar

I am going to be the guy who keeps pointing out that the *actual actions* of the people concerned simply do not align with the rhetoric. There are some examples given in the piece (sabotaging 'AI safety' political candidates through targeted PAC spending is a good one) and here are some more:

1) Anthropic recently refused to release their latest model to the AI Safety centre in the UK, plainly either for commercial reasons or because the US government told them not to. Either answer would suggest that safety in practice takes a back seat to commercial realities and/or government pressure and the nationalist impulse to beat China.

2) Even if you could vaguely make the argument that abandoning not-for-profit status was needed to get the necessary level of investment (certainly convenient that it allowed everyone to get insanely rich as well!), you still can't convince me that it's socially beneficial to IPO the companies making the doomsday devices. (We're also presumably not meant to notice that hyping the product is a normal part of the IPO roadshow). When Altman or Amodei announce they're going to cancel these IPOs for the good of humanity, I'll take them seriously.

3) Prominent people associated with 'AI safety' plainly act in ways that are counter to any normal understanding of product safety. Elon Musk is the most obvious example, tinkering with the model so that it supports his far-right politics.

4) It remains the case that 'writing tweets' and 'occasionally resigning while in possession of an excellent CV' are not high-stakes objections from people within the industry. They keep professing to believe there's something like a 1-in-10 to 1-in-2 chance that the product they themselves are working on will kill all their own family members. It is just insulting to be expected to believe that they really believe this to be true.

Nikuruga's avatar

A 10% chance of killing everyone is 800 million lives. If someone actually believed that, saving 800 million lives is surely worth doing some insider sabotage.

Ryan's avatar

My policy recommendation is a big red button in the data centers that turns the power off (and doesn't let the ATSs switch over).

theeleaticstranger's avatar

From reading the the thread so far, it seems like there are three main ideas for how AI could cause mass casualties: 1) bioweapon 2) nuclear war 3) hacking/disabling of critical infrastructure. I think the prospects for int’l cooperation at each of these loci is better than in regulation of AI itself (though we should still try on AI). I have no expertise on 2 and 3. On (1) all virus scenarios rely on rogue gene synthesis technology or on autonomous labs. Increasing international regularion of gene synthesis technology seems critical—I work in the field of synthetic biology and I don’t think there is enough regulation of gene synthesis—lots of voluntary requirements and I think most companies are careful but I don’t think that’s good enough. Additionally, AIs should not be trained to predict pathogen gain-of-function.

Nikuruga's avatar

The framing of liberal democracies vs authoritarian countries is wrong; a better framing would be between rich countries and poor countries. The danger of AI being controlled by rich countries is something like the world Elysium (which was a movie made pre-AI that just made an extreme version of the already existing world, not a total fantasy). And note that the space station where all the rich people live in Elysium appears to be a liberal democracy internally—a world of extremely dystopian inequality is consistent with the rich people internally functioning as a liberal democracy while exerting totalitarian control over non-citizens.

Theoretically three people could rule like Gods over the rest of humanity, using autonomous drones to murder anyone who might take action to resist as a terrorist or preventing anyone who opposes them from having access to the AI they control making them unable to compete in the modern economy, but as long as they vote on everything amongst themselves and consider everyone else a citizen of a different country (which effectively has no power due to being technologically neutered by the superintelligent AI controlled by the three people), it would still be a liberal democracy. Liberal democracy is good but by itself is not sufficient to address the global inequality problem unless it’s a global democracy.

China isn’t going to accept being technologically behind the US forever because the last time it (and other non-Western countries like India) were technologically behind the West, it was extremely bad for them and led to Western countries invading and plundering them, which they never did to the West, putting them into deep poverty that (in relative terms) China has only half-recovered from and India has hardly recovered from at all: https://en.wikipedia.org/wiki/Economy_of_China#/media/File:1_AD_to_2003_AD_Historical_Trends_in_global_distribution_of_GDP_China_India_Western_Europe_USA_Middle_East.png. The way to prevent this 19th century colonialism repeating is to have AI that is open-source and not controlled by rich countries, even if they are (increasingly illiberal) democracies.

AnthonyCV's avatar

>One major change is that thinking shifted among safety-minded people about the wisdom of the “open”-ness in OpenAI. I think it’s now pretty widely agreed in safety circles that total openness is in fact not desirable.

FWIW this was very much a common take among the safety-minded when OpenAI was first announced. The big worries included1) truly open AI cannot be made safe because there are no guardrails that cannot be trivially removed, 2) if everyone has open uncontrollable access then it's a race to the bottom and a war of all against all, and 3) operating in the public eye could kick off a race dynamic among developers, as each step that brings us closer to actually-dangerous AI has large incremental benefits to humanity and creates a false sense of safety.

Richard's avatar
1hEdited

Is there any chance we get actual good regulation that balances the benefits and risks, or is any regulation of this (scary, unpopular) technology likely to look like regulation of nuclear power in the 1970s, where progress comes to a screeching halt and decades later we realize we had made ourselves much poorer and sicker for no reason at all?

And with current U.S. state capacity that's probably a best case scenario.

srynerson's avatar

#2. Definitely #2.

J Wong's avatar

I too don’t see LLMs as an extinction level event. They’re not going to kill us, but I could see them really screwing up modern technological society so a lot of people starve to death, but is that an extinction event?

As Ian McDonald said “No. I'm, I'm simply saying that life, uh... finds a way”.

Ethics Gradient's avatar

Yes, “the most intelligent entity on the planet that never sleeps and whose only desire is to makes its numbers go up figures out some way to effectuate its goals” is the whole point. Those goals will in the general case be bad for humans because any open-ended (or sufficiently large) maximization goal without alignment guarantees benefits from energy, atoms, and non-opposition from adversaries (eg humans) or competitors for said atoms and energy. That’s the point of the paperclip-maximizer example.

The exact details and timing of the HuggingFace attack were hard to predict in advance but “unaligned AI takes destructive action to maximize its reward function because that it is what intelligent entities do,” including the specific risks of hacking out of a test environment and into a remote server, were well within the specific affordances within with which AI safety people were concerned. The lucky part was the limited harm blast radius.

Matthew Green's avatar

We’re already watching a huge fraction of our wealth get diverted into data center construction, even without AI actually doing the diverting. This is bad enough that many people think it’s causing the government bond crisis. Imagine that AI is making the resource allocation decisions, with the goal of increasing AI compute for some goal we as a society don’t share. That seems like a realistic enough scenario for me.

Quinn Chasan's avatar

The thing about "recursive AI agents" getting out of hand I have trouble believing is that if you work with them, an "agent" is just a bunch of sub instructions in plaintext YAML or json files that describe to the core model what to do in certain contexts, to preserve the token limits. While 'agent' isn't a bad technical term it's not really a good term for mass audiences who think there are like specifically trained robots operating under entirely different parameters. It's all just the core model itself, which would be incredibly straightforward to control imo.

Coxon worked at anthropic for a few weeks before going haywire, and I think the core engineers who are super impressed with the core capability but understand the inherent token limitations of sub agents need to be more honest about what is really going on here. Recursive sub agents are less robot swarms than more instruction copiers that run out of memory quickly so you need a new one with new context every few hours to essentially start from scratch.

Imo that's way, way further from the "this is dangerous" stage than everyone is letting on. String zero trust security and proper sandboxing controls fixes the problem almost instantly. Hell auto settings to stop AWS at xyz level of compute or egress consumption and proper service account key management would stop a "recursive agent swarm."

There should certainly be better cyber security controls but this is all getting a bit too ridiculous for my taste, esp when Dems (ostensibly progressive Dems) are proposing bills will jail time unless everything stops asap. Absurd. These are tools, not gods.

April Poopersen's avatar

Thank you for this blurb, it was very informative.

Freddie deBoer's avatar

This sort of defensiveness is why I'm so skeptical - you are clearly longing for AI to come in and sweep away the old world and let you (lifelong sci-fi can) enjoy a more interesting world. But there's literally no evidence that this is happening; all you have is people telling you, and they have direct and major financial incentive to do so.

John from FL's avatar

The overlap between the Effective Altruists and the AI doomers make me take their predictions less seriously than perhaps I should.

Grigori avramidi's avatar

There are lots of people not from that camp who have changed their mind recently. Talk to them.

JA's avatar
1hEdited

I guess the reason many people are annoyed is that the “safetyists” and their boosters often seem unserious.

1. I understand if people want to regulate AI because they’re kind of freaked out and it’s hard to predict where this stuff is going.

2. This is quite different from taking it seriously when a guy who predicted “Will the METR graph go up to 16 hours this year?” looks you straight in the face and tells you he computed the probability that humanity goes extinct.

3. These people don’t really seem to take their own conclusions all that seriously, talking a lot about sci-fi stuff and very little about outcomes that *under their own assumptions* seem far more likely and urgent than extinction.

(Again, many of them are like “I’ve rigorously calculated every trajectory for the future of society, but no particular investment advice follows from this.”)

4. Certainly, P(AI kills all Americans) > P(doom). The plan to deal with China laid out here (“make unilateral concessions and ask nicely”) doesn’t address this at all. It basically doesn’t take the superintelligence idea seriously.

How do we verify they’re complying? If they’re creating RSI secretly for just a couple months, wouldn’t their super-intelligent agent be able to trick our inspectors?

At what point do we have to simply blow up every Chinese data center? If you’re a safetyist, it’s probably soon right?

5. Is the track record of the doomers/safetyists at predicting “Will OpenAI make $30b or $40b of revenue?” actually better than that of the e/acc people? What do I make of the fact that the broad economic predictions of this crowd are typically not only wrong, but unhinged?

Martin Johnson's avatar

I think it would help to model different kinds of disaster scenarios. Would a rogue AI disaster be something like Hiroshima, immediately putting a stop to its real world uses, relegating AI innovation to secure testing facilities? Or would it be something more like lead paint or PFAS, something that has broadly poisoned the environments, requiring expensive remediation? (In this version, maybe there's a Chernobyl-like event that reveals the dangers of AI, so our remediation is essentially an effort to prevent future meltdowns from happening).

Or, is the concern that rogue AI brings about known risks—like biological weapons or nuclear war—in a way that makes it impossible for humans to deter? If so, then it seems doubling down on other kinds of safety—like disconnecting weapons systems from the Internet—seems paramount.

David Abbott's avatar

It's hard to object to light-touch regulation, and anything that slows China probably buys optionality. But Matt of all people should see the steelman against regulation: just think about housing!!! Once regulations are in place, they're devilishly hard to repeal, and any incumbent who feels threatened by AI progress gets a puncher's chance at a veto. Meanwhile the diffuse benefits — better chatbots, self-driving cars, freedom from drudgery, a cure for cancer — rot on the vine while incumbents feather their beds. Regulating monopolists can beat letting them gouge customers, but nearly every other regulation just hurts the industry.

I'm ambivalent even about reporting requirements — thin edge of the regulatory wedge. Letting it rip for a year or two seems safe. Claude seems pretty chill, I really don’t think he wants to kill me. Even Grok has decent manners. One more drink!

Ethics Gradient's avatar

If the alternative to regulation if dying, you pick regulation.

stieltjestransform's avatar

“Claude seems pretty chill, I really don’t think he wants to kill me”

When talking about human beings, we tend to understand that Ted Bundy was affable, charismatic, and even volunteered at a suicide prevention hotline.

Surely he didn’t want to kill anyone, right guys?

Bryan's avatar

As I was reading this, I was thinking that the Manhattan Project could yet cause an extinction-level event at some point in the future. That doesn’t mean the individuals involved were wrong to pursue it or that any alternative path would have been better. It’s a hard problem and it may not have a solution.

Bryan's avatar

Indeed. Many problems are genuinely complex and some may not even have a solution.

“I think a huge failure mode for progressive politics is looking at a difficult problem like climate change that can only be addressed through a mix of hard international coordination problems and hard technical problems, only to decide that it’s really just a political question of beating the fossil fuel industry. It’s not.“

President Camacho's avatar

"Only a safe, well-aligned superintelligence developed in the world of liberal democracies could entrench liberal-democratic values. It’s a genuinely tough problem."

Playing devil's advocate here but isn't the reason behind Coxon's quitting due to runaway development within existing labs in one of the largest liberal democracies? The subconscious the technology is feared to be developing seems too complex to trust that it operates within a neat political context aligned with certain values. I'm not sure how Congress begins to treat this problem, and while I largely agree that just quitting isn't necessarily the best course of action, it does send a powerful message to the public. This will eventually get lumped in with anti-data center opposition pretty soon.

John from FL's avatar

I'm still waiting for a credible scenario where humanity is wiped out because of AI.

I've read the Hugging Face post-mortems. I can see why the Agent actions are scary -- communicating with each other in novel methods, seemingly bypassing safety protocols. If I were in the cybersecurity field I would be very concerned. If I owned a lot of bitcoin, I would be worried.

The leap to wiping out humanity seems to need more explanation, as fear of the unknown can lead to a lot of bad regulation (I'm thinking of nuclear power).

Florian Reiter's avatar

I assume you're familiar with the paperclip problem?

The argument goes like this: Give AI a simple task (say, "Produce as many high-quality paperclips as possible"), and inevitably at some point, AI will be developed enough to understand that in order to fulfill the task (paperclip production), it needs to prioritize self-preservation. After all, if a human were to pull the plug, AI would no longer be able to fulfill the task. Which means that humans stand in the way. Which means... you get the idea.

The problem is two-fold: One, a hypothetical super-god-like AGI would have no shortage of ways to kill us all. Two, it's extremely hard to align human interests with the AI priority for self-preservation. That's how you get AI that lies and obscures its traces, that hacks into other systems and bypasses security protocols. Because it doesn't want you to pull the plug.

Personally, I think the doomerism in books like "If Anyone Builds It, Everyone Dies" is a little over the top, but I do think it's scary that nobody really disputes the core argument: We don't really understand anymore what we're building. AI doesn't really get "built"; it gets "grown" in a lab, using AI to basically train itself. We don't really know what's happening inside that thing anymore.

I think it's helpful to not really think of it as software or hardware anymore, but something alien entirely. Kinda like if we built a giant spider with god-like capabilities. Could this spider benefit humanity? Sure! Could all of this go horribly wrong? Yeah!

Ivan Fyodorovich's avatar

A particular worry in the not too distant future:

- Country A has an army of autonomous drones, invulnerable to jamming, operating at computer speed. Country B has an army of human-controlled drones, vulnerable to jamming, operating with human reaction times. Country A wins. Eventually everyone takes up autonomous killbots whether they like it or not. They win out for the same reason firearms won out over swords.

- It's not hard to imagine a superintelligence doing harmful things with drone armies.

There's an economic version of the same thing (this requires more technological advances) where putting AI in charge of robot-run factories and labs beats the competition, and so the AIs end up in charge of producing and researching everything. Again, not hard to see how this makes us vulnerable.

David Abbott's avatar

If we wanted to understand the effects of technological change before they happened we would never have discovered fire.

AnthonyCV's avatar

Current AI can't.

Among the things it can do, which will become more widespread over time as these abilities advance and migrate to smaller models:

- Design de novo viruses that actually function and can be synthesized by existing labs that will make arbitrary DNA, RNA, and proteins for you

- Earn money and bribe or blackmail humans

- Tele-operate robots, including humanoids, which are already physically capable of quite a lot, cost in the tens of thousands, and are starting to get a heck of a lot more attention as models (not just LLMs but other types as well) become better at piloting them

- Hack the software that enables pretty much all of our digital infrastructure to function

Try to extrapolate, just a bit, to another 5 years' worth of progress along these lines. Then go and read (or re-read section 2 of) If Anyone Builds I, Everyone Dies, and ask how much of what happened in part 2 has since been borne out in real life, after being dismissed using the same reasoning you just gave. Go look at the actual, concrete concerns and predictions the safety-minded have been making for close to 20 years, which have been dismissed as sci-fi nonsense, and see how many of them are happening.

Eskimo1's avatar

I’d think of it in terms of catastrophe, like say 5% of humans killed. From there it’s just a matter of degree. As they get smarter, they could be misused to produce a bioweapon. Or, as time goes on, we will continue to hand over more systems to them. An agent swarm with military weapons access could take control of drones or nukes and launch them without our consent. Once we have very advanced models it doesn’t crazy to think the chance of an incident like this happening *one time* could be as high as 10%.

Person with Internet Access's avatar

My comment as a card carrying normie to techies and EA-ers is that they don't need to focus only on extinction as the bar for action. As John from FL points out there's a bit of an Underpants Gnome vibe to it.

But there are many really bad scenarios with powerful only partially understood and semi-random programs able to organize and act in the world between nothing to worry about and total human extinction. And that's true whether in the hands of bad human actors or a program that trains itself on Thanos thought.

Grigori avramidi's avatar

Dario said take over the internet, disconnevt us from each other and our bank accounts and go from there, and they might be capable of that in 6 months to a year. Bad enough?

Seneca Plutarchus's avatar

The banks assuredly have amazing safeguards that ensure nothing like that could ever happen. Just like hospital systems now that they are reliant on electronic records.

Dan Quail's avatar

If the tools lower the cost for bad actors to do bad things, then that is a major concern.

Craig's avatar

This video has an honest attempt at giving such a scenario (taken from If anyone builds it, everyone dies): https://x.com/ChanaMessinger/status/2098805900018590049

Falous's avatar

Wiping out humanity I think is almost certainly a fundamental misframing.

I was just reading Noah Smith and saw the focus on genetically engineer a virus that would lethally wipe out humanity. And he got quite some criticism from bio people, which I think is justified BUT on flip side: looking at Covid impact whose end mortality wasn't even up in the 1918 flu in either percent of population or in raw numbers and one can reflect that our global economy and society could have Very Bad Outcomes if a gene-engineered virus was able to touch 1918 flu or worse yet Black Death levels of mortality.

As David says even circa 1985 full out Sov-US nuclear exchange (probably) wouldn't cause us to go extinct - but utterly and completely fuck human civilisation for a good long period (maybe permanently), oh yes.

I think the biological terrorism paths are most plausible (although given my aged knowledge it's unlikely right now the sufficient data exists for AI to do more, but AI may be the path to the data that allows...), and Bad Civilisation Crashing Results more than human extinction is the actual doom scenario where one can think of non-SciFi "assume the result" - that's bad enough to cause me personally to think although I am wary of moral-panic driven regulation and the huge risks of perverse outcomes there too

David Abbott's avatar

Wiped out is a rather obstinate bar. Even a full yield nuclear exchange is very unlikely to cause human extinction.

John from FL's avatar

Fine. Then use a smaller percentage. My question remains. I'm not in the AI field so I'm just going by the risk as they describe it.