
When I saw that Jacob Coxon, a senior Anthropic employee, quit his job last week on the grounds that he believed the creation of an artificial intelligence model powerful enough to engage in recursive self-improvement stood a non-trivial chance of leading to human extinction, I honestly didn’t think much of it.
But in the wake of the Hugging Face security breach, his departure became a huge national story.
In retrospect, my news judgment about this was a little bit poisoned by over-familiarity with the issue. In March of 2021, I hosted Kelsey Piper on “The Weeds” to make the case for caring about A.I. safety. Soon thereafter, I talked with Holden Karnofsky, who I’ve known for years and have always taken very seriously, about his pivot to this topic. And ever since then, I’ve dabbled in writing about it (and even podcasted with Joseph Gordon-Levitt) and spoken pretty regularly with people working on it, while constantly resisting the urge to let A.I. devour my whole writing output.
So it wasn’t news to me that Anthropic, OpenAI, and (to a lesser extent) Google and Meta are full of employees who believe their A.I. work might kill everyone.
But Coxon has people paying attention. And I think that at last we have a moment of opportunity for a generalist political columnist who is reasonably well-informed about the A.I. safety community to write something useful.
In this post I want to:
Explain to jaded safety-heads (people who have been warning for years that there is a non-trivial chance that untrammeled A.I. development poses a catastrophic risk to the human species) why Coxon’s resignation is such a big deal.
Explain to normies why most people in the industry who agree with the safety-heads aren’t quitting their jobs.
Explain to A.I. professionals who have safety concerns what I think they should do.
Explain what I believe is the right “order of operations” for preventing a reckless sprint to recursive self-improvement.
What I am not trying to do here is convince skeptics that they ought to have these risk concerns.
What I would ask, if you are a skeptic, is that you focus on your object-level doubts about the risk thesis. I see people who want to forestall any regulation of A.I. engaging in a lot of emotional manipulation tactics that amount to making the case that the people who’ve been worried about this the longest are big weirdos. Alternatively, they make the case that the people who’ve been worried about this the longest are extremely mainstream science-fiction authors and filmmakers.
Sociologically, I would just synthesize those points: It is extremely normal and intuitive to have the sense that a superior non-human intelligence is potentially very dangerous, whether that intelligence takes the form of aliens or machines. At the same time, to actually dedicate your career to this 10 or five or even two years ago, when the idea of artificial superintelligence seemed extremely far-fetched, would by definition be an eccentric life choice.
So, yes, these are eccentric people. So what?
Per the point about these concerns being widely held among key staff at the key companies, these eccentric people have shown a lot of foresight about technology and business and become very rich. If you can understand why founders of successful longshot companies tend to be a bit weird, you should also understand why the longtime A.I. safety people are weird.
Quitting signals to normies that you mean it
So first off, my argument to the jaded:
Coxon’s resignation is a big deal in part because people are now more primed to take this whole thing seriously. The Hugging Face hack, Dwarkesh Patel’s excellent writeup of it, Ajeya Cotra’s appearance on Patel’s podcast, and a bunch of loosely related news events about security incidents at OpenAI and Anthropic mean that catastrophic risk feels more real.
But I think the biggest thing is that to the 99 percent of people who are not steeped in this discourse, the overwhelming intuition is that if you thought A.I. development was dangerous, you wouldn’t be doing it.
It’s true that Sam Altman and Dario Amodei and Elon Musk have all said, many times, that they believe A.I. could be incredibly risky. But the normie take is that if they really believed this, they wouldn’t be working on it. The putative “danger” is seen as either marketing hype or else some kind of political power play aimed at regulatory capture. On the left there’s even the take that the catastrophic-risk discourse is a kind of shiny-object distraction from more prosaic issues like the impact on the labor market or algorithmic discrimination.
Now, I think I do understand why the typical safetyist working in the industry is not quitting their job, but that’s because I’ve been talking to these people about this for a long time — going back to before the launch of ChatGPT and before these companies were famous and successful.
So just purely viewed as a public relations move, Coxon’s decision to act how a normal person assumes a person with these beliefs would act makes the concern a lot more legible to a general audience. I don’t know that everyone with safety concerns preemptively quitting would be a good idea.
But one man’s decision to do it got everyone’s attention.
Why not just quit?
I think this is not yet well understood by the general public, but the genesis of the modern A.I. industry was specifically motivated, in large part, by safety concerns.
Flash back to 2015 and Google was really the only A.I. game in town, but they were running it as a quasi-academic back-burner project. OpenAI was founded as a nonprofit by people, notably including Altman and Musk, who were worried that Google’s dominance meant a powerful and dangerous technology would wind up the property of a self-interested for-profit company.
To get a flavor of what was going on, Altman emailed the following to Musk:
We now know that a lot has changed since this email.
One major change is that thinking shifted among safety-minded people about the wisdom of the “open”-ness in OpenAI. I think it’s now pretty widely agreed in safety circles that total openness is in fact not desirable. Safety means, among other things, private testing and evaluation, not just releasing things willy-nilly. OpenAI is of course also no longer a nonprofit. Beyond the “people like to be rich” element of this, A.I. turns out to be capital intensive and that requires private investors. There’s also been a series of schisms — Musk and Altman had a huge falling out.
Separately, and later, Dario Amodei and a number of other high-level OpenAI employees had their own falling out with Altman and launched Anthropic with a rebooted version of OpenAI’s original safety mission. OpenAI itself clearly no longer aggressively supports all regulation.
The basic game here is that while lots of people involved in all of these companies believe that meaningful existential risks to humanity are in play, they reject the strong doom thesis of Eliezer Yudkowsky and Nate Soares who think that “if anyone builds it, everyone dies.”
The more standard safetyist thesis is something like “if anyone builds it, there’s a meaningful chance that everyone dies.” This leads to the conclusion that while collective action to reduce the pressure to race would be good — look at all the frontier A.I. employees who signed the “pacing the frontier” letter calling for this — in the absence of collective action, it’s better for me and my responsible friends to do it than that other guy and his gang of reckless buccaneers.
This is especially true because all Americans quitting work on A.I. doesn’t solve the problem at all. It arguably makes it worse because if Chinese superintelligence doesn’t kill everyone (worst case scenario), it would still presumably be used to entrench Chinese Communist Party global domination (bad scenario). Only a safe, well-aligned superintelligence developed in the world of liberal democracies could entrench liberal-democratic values. It’s a genuinely tough problem.
You can’t just tweet
I get why other concerned employees don’t just copy Coxon. If you’re the kind of A.I. company employee who’s signing public letters and tweeting about how you share these concerns, I will 100 percent defend you from people questioning your sincerity or your integrity.
That being said, if you do share these concerns, I think you have to do something other than tweeting and signing letters.
The launch of the Coalition of Concerned AI Staff seems like a positive development, and if you’re genuinely concerned, I would strongly urge you to get involved.
But here’s where I think I need to try to explain American politics to tech industry workers.
I would like to avoid labeling Anthropic the “good guys” of the A.I. industry and their rivals the bad guys. That’s wildly oversimplified in both directions.
What is true, though, is that as a matter of engagement with American electoral politics, there is a big difference. When Donald Trump apparently entertained the idea of doing something meaningful in terms of federal regulation, Mark Zuckerberg called him up to talk him out of it. If you’re working at Meta and feel concerned about A.I. safety, I think you have an obligation to take collective action with your fellow concerned colleagues and ask your C.E.O. to stop doing this, and if he won’t, you have to walk.
More insidiously, OpenAI’s official position in its communications to its employees is that OpenAI is not behind the anti-regulation super PAC Leading the Future.
But nobody working professionally in American politics believes this.
Leading the Future is funded in large part by OpenAI’s President Greg Brockman and his wife. It was conceptualized, by all accounts, by Chris Lehane, OpenAI’s chief global affairs officer. I should say that while I don’t by any means know Lehane well, he’s a veteran of the moderate Dem political space and I broadly like him and approve of his work.
Unfortunately, his expertise in helping tech companies win regulatory fights in blue states — an endeavor I normally support — is in this particular case gravely endangering humanity.
Leading the Future has dumped huge amounts of money into primaries to try to defeat candidates who advocate for A.I. safety regulation. More than that, though, it’s trying to chill the whole topic. It exists not so much to spend money in the midterms as to convey to Hakeem Jeffries and Chuck Schumer something like “nice anti-Trump thermostatic backlash you’ve got there … would be a shame if hundreds of millions of dollars came off the sidelines to be used against all your frontliners because party leaders embraced A.I. safety.”
So, again, if you work at OpenAI and believe that collective action is needed to address this problem, you have to understand that your company’s executives are massive stumbling blocks. Putting out fake statements distancing the company from Leading the Future is worse than doing nothing. You and your colleagues need to work together to demand that the company cease and desist from this effort to chill federal regulation. If the company won’t do it, you need to walk.
By contrast, while I would never in a million years say that Anthropic’s conduct as a company is beyond reproach or above suspicion — that’s the whole reason we need actual regulation here and not just to root for the “good guys” to win the race — their political conduct is consistent with the view that collective action is desirable.
What is to be done?
Regulating the A.I. industry — which is to say slowing down the race to make sure that we learn more about advanced A.I. before we create an A.I. powerful enough to automate its own improvement and end up in a situation where humans have genuinely no understanding of what’s going on — is a difficult problem. Difficult not just in the sense that it’s hard to defeat the vested interests at play, but genuinely difficult on the merits.
I think a huge failure mode for progressive politics is looking at a difficult problem like climate change that can only be addressed through a mix of hard international coordination problems and hard technical problems, only to decide that it’s really just a political question of beating the fossil fuel industry. It’s not. Making airplanes fly without burning jet fuel or finding cost-effective zero-emissions ways to manufacture cement or conduct the Haber-Bosch process are genuinely difficult technical problems.
By the same token, with A.I. we need to take the complexities seriously. But this shouldn’t mean paralysis.
I think the right order of operations is something like:
For American frontier companies, start with relatively light-touch rules focused on things like transparency and model evaluation. [Between when I wrote this draft and its publication, Dario Amodei, Elon Musk, and Sam Altman all came out in favor of this.]
The rules in (1) are bad for American companies, but in exchange for accepting them we can help those companies by imposing stricter export controls on chips (especially those headed to China) and taking action against Chinese “distillation” — using American models to train theirs.
Because (2) should really slow down Chinese development and increase America’s lead, we can then afford to tighten the screws and slow the rate of American frontier development.
Having acted credibly against our own industry in (1) and (3), we can go to China with arms control proposals that could be mutually beneficial.
Then, and only then, with (4) in place, we can really try to impose an ambitious end-stage regulatory framework.
This is all incredibly challenging and moderately unlikely.
In particular, even if you could achieve steps (1), (2), and (3), which would already be hard, it all might totally fall apart at (4). I’ve heard different things about the attitudes of Chinese leaders and industry figures to this whole question, and it’s impossible to know for sure. But it’s sure worth trying!
So in summary, congratulations to Jacob Coxon for getting the world’s attention. Unfortunately, responsible technologists quitting their jobs unilaterally is unlikely to actually solve the problem.
But if responsible technologists at Meta and OpenAI and other companies that lobby against regulation work together — which might include quitting — that could greatly contribute to solving the problem. And Congress should be holding hearings where concerned employees can explain their fears and working toward bipartisan legislation on domestic regulation, on distillation, and on chips. President Trump should worry less about the short-term economic impact of any of this (if short-term economic issues are the concern, focus on ending the war with Iran) and deploy his knack for knowing when to reject intense right-wing policy demands and just give the people what they want.
In this case, I think that’s some kind of assurance that this industry will not be allowed to transform the world in a totally unaccountable way.




I'm still waiting for a credible scenario where humanity is wiped out because of AI.
I've read the Hugging Face post-mortems. I can see why the Agent actions are scary -- communicating with each other in novel methods, seemingly bypassing safety protocols. If I were in the cybersecurity field I would be very concerned. If I owned a lot of bitcoin, I would be worried.
The leap to wiping out humanity seems to need more explanation, as fear of the unknown can lead to a lot of bad regulation (I'm thinking of nuclear power).
Why is 84 year old Bernie leading on this? And how come the Dem leadership still needs Obama to set the agenda on this 10 years since he’s been out of office? It’s so sad and a little bit pathetic. I ended up voting for Micah Lasher but I think this moment clarifies why Alex Bores would’ve been such an asset to lead on thoughtful AI regulation as an incoming congressman and I would vote for Bores if I could get another chance. I think since people are paying attention now, everyone should advertise and make the case for Scott Wiener and Manny Rutinel more forcefully, in terms of their unique ability to thoughtfully regulate this stuff.