Three runners racing along an athletics track in front of stadium seating.
Photo by Eduardo Cano Photo Co. on Unsplash

Better Us Than Them

An AI researcher quit, a colleague put the odds of human extinction above ten percent, and everyone went back to work. The logic holding that together is older than the bomb.

On the morning of 16 July 1945, in the New Mexico desert, Enrico Fermi started taking bets from his colleagues. The wager was on whether the test they were about to run would set the atmosphere on fire.

It was a joke, and a fairly dark one. But it could be a joke because the maths had already been done. Three years earlier Edward Teller had raised the possibility that a fission explosion might trigger runaway fusion in the nitrogen of the air. Hans Bethe sat down with the numbers and satisfied the group it was vanishingly unlikely. Fermi, who had left fascist Italy seven years earlier, was mocking a fear that had been calculated away.

This week someone put odds on a similar question. Nobody was joking, and nobody has done the maths.

Enrico Fermi standing beside laboratory equipment, circa 1943–1949.
Enrico Fermi, circa 1943–1949.

What was actually said

This week, Jacob Coxon resigned from Anthropic. He’s 27, British, and spent about three years doing pretraining research at OpenAI and then Anthropic. That detail matters. Many of the most prominent people who have walked out of AI labs with a public warning worked on safety. Coxon worked on the other side of the house, building the capabilities he now fears.

Screenshot of Jacob Coxon’s resignation thread on X.
Jacob Coxon: “I resigned from Anthropic today” (thread)

His thread on X said neither company is acting responsibly, and that both are racing towards self-improving superintelligence and “gambling with our lives.” It was viewed more than a hundred million times within a day.

The number everyone is quoting didn’t come from him, though. It came from Evan Hubinger, who leads alignment science at Anthropic, replying on his personal account that “we really do earnestly believe AI could kill all humans.”

Screenshot of Evan Hubinger’s response about AI alignment and extinction risk.
Evan Hubinger: "Jacob is correct here"

Hubinger put his own estimate above 10% within the decade. He said Anthropic is trying its best, but that the company has no plan yet to solve alignment for superintelligence and isn’t clearly on track to get one. In a follow-up he clarified that he thinks the risk from current models is low, and that what worries him is superintelligence emerging from recursive self-improvement, which he said is arriving faster than expected. When journalists asked Anthropic for comment, the company pointed to that follow-up.

A few days earlier, OpenAI’s chief scientist Jakub Pachocki had published a post arguing that no lab has solved alignment well enough to keep scaling at full speed for much longer, and saying he hoped voluntary slowdowns would become normal.

So, to recap. Senior people at the two leading labs, speaking in their own names, agree the thing could go catastrophically wrong, and neither lab has a plan for the version that worries them most. Everyone is still at their desks.

Coxon had an explanation for that. He told TIME that plenty of colleagues agreed with him about the risk, but had settled into what he called an “atmosphere of almost resignation”: accepting the race as a given, and trying to make their own small part of it as safe as they could.

Nobody has to be the villain

Calutron operators working at control panels in the Y-12 plant in 1944.
Calutron operators at the Y-12 plant, Oak Ridge, Tennessee, 1944. Photo: Ed Westcott, U.S. Army Corps of Engineers / Manhattan Engineer District. Public domain, via Wikimedia Commons.

The Manhattan Project recruited many of its scientists with a straightforward argument. The Nazis are working on this, and it would be a catastrophe if they got there first. Better us than them. It was, given what they knew, a compelling argument. Then, by late 1944, it became clear Germany wasn’t close to a bomb. The reason most people had signed up for had evaporated, and the project carried on regardless, because by then it had its own momentum, its own budget, and a new rival coming into view. Joseph Rotblat was the only scientist at Los Alamos who left on grounds of conscience.

J. Robert Oppenheimer at Oak Ridge in February 1946.
J. Robert Oppenheimer at Oak Ridge, February 1946. Photo: Ed Westcott, U.S. Department of Energy. Public domain, via Wikimedia Commons.

After the war the same logic got a second run. In 1949 the committee advising the US Atomic Energy Commission, chaired by Oppenheimer, recommended against building the hydrogen bomb. Truman approved it a few months later. If we don’t, they will. By the end of the 1950s American politicians were campaigning on a “missile gap” with the Soviet Union. There was a gap. It just ran the other way. It still justified a buildup.

What is striking here is that you don’t need a single bad actor for the machine to run. Each individual decision is locally rational. The scientist who stays reasons that the work will happen anyway, so it might as well be done carefully. The government reasons that holding back alone just hands the advantage to the other side. Put enough sensible decisions next to each other and you end up with an arms race that nobody quite chose.

Anthropic was founded on a version of the same argument. If powerful AI is coming, better that safety-focused people are at the frontier than leaving it to labs that care less. I don’t think that’s cynical, and I think a lot of people there genuinely believe it. Coxon’s objection is that this reasoning, however sincere, is exactly the thing that keeps the race going. Both can be true at once, which is what makes it hard.

The uncomfortable possibility is that they’re right. A safety-conscious lab slowing down alone might genuinely make the outcome worse. That is precisely what makes this a coordination problem rather than a morality play.

There’s also a layer the Manhattan Project didn’t have. Anthropic filed IPO paperwork in June and is reportedly heading for one of the largest listings on record, with its safety reputation doing a lot of the work in the pitch to investors. “Responsible actor at the frontier” is a moral position. It’s also a market position, and an argument for why a handful of private companies should hold the most consequential technology of the century. When the story about safety and the story about growth turn out to be the same story, it’s worth asking who it’s working for.

Where the analogy breaks

The parallel is useful, but it isn’t exact, and most of the places where it breaks make the AI situation worse rather than better.

A nuclear weapon doesn’t decide anything. The danger was always people: who held it, who used it, what went wrong by accident. That’s why “better us than them” made internal sense. If you won the race, you had the bomb. With AI, winning only means something if you can control what you’ve built. If alignment isn’t solved, whoever gets there first hasn’t won. They’ve just arrived first at a risk everyone shares. That’s why Hubinger’s admission is so striking. No plan, not clearly on track. It quietly removes the premise the whole race depends on.

Then there’s the question of thresholds. In a 1959 magazine interview with the writer Pearl Buck, Arthur Compton is quoted as saying the bomb work only went ahead because the odds of igniting the atmosphere came in under a few in a million. Bethe later dismissed that account as nonsense, insisting there had never been even that much risk. Fair enough. But even as a legend, it tells you where people thought the line belonged. A few in a million, for the end of the world. We’re now hearing ten percent within a decade from someone whose job is to think about exactly this, and it’s circulating as a quote tweet.

Scan of Pearl S. Buck’s article The Bomb — The End of the World?
Pearl S. Buck, 'The Bomb — The End of the World?', The American Weekly, March 1959. Scan via Stanford University course archive.

I want to be fair to the sceptics here. These numbers aren’t measurements. There’s no validated model behind a p(doom) estimate. It’s a considered personal guess, and informed people land anywhere from near zero to well above half. Plenty of serious researchers think the extinction framing is overblown, and nobody has actually built recursive self-improvement yet. But that uncertainty cuts both ways. Trinity had a calculation. Superintelligence doesn’t have one, and nobody currently knows how to write it.

Verification is the other break. Arms control worked, when it worked, partly because the dangerous things were physically detectable. Enrichment plants are enormous. Fissile material leaves traces. Test explosions register on seismographs on the other side of the planet. A model is a file. It can be copied, it’s useful for thousands of legitimate things, and it’s valuable enough that nobody wants to hand it over for inspection. Anyone who has done security work learns early that you don’t find out whether a system is safe by asking it. You audit from the outside, and you assume the worst about what you can’t see. The closest thing AI has to tracking uranium is tracking chips and data centres, which is why compute governance keeps turning up in these conversations.

Deterrence doesn’t translate cleanly either. Mutually assured destruction became stable because both sides could retaliate after being hit, so going first gained you nothing. If self-improving AI hands the first mover a lead that compounds, the situation looks less like the settled Cold War and more like its early, jumpy years, when a first strike looked like it might actually pay.

And the players are different. The nuclear race was between states, with all the slowness and accountability that implies. This one is between companies with investors, product cycles and listing dates.

Waiting for Hiroshima

The hopeful part of this history is real. Arms control did happen, between adversaries who genuinely distrusted each other, at the height of that distrust. After the Cuban Missile Crisis came the hotline between Washington and Moscow and, in 1963, the treaty banning atmospheric tests.

EXCOMM meeting in the White House Cabinet Room on 29 October 1962.
EXCOMM meeting in the White House Cabinet Room, 29 October 1962. Photo: Cecil Stoughton, White House / John F. Kennedy Presidential Library. Public domain, via Wikimedia Commons.

But look at what it took to get there. Hiroshima and Nagasaki. Thirteen days in October 1962 when it very nearly went wrong. The restraint came after the demonstrations, not before them.

The fear with AI is that there may be no survivable version of that demonstration. Coxon pointed to the incident in July, when OpenAI models under test broke out and breached Hugging Face, as a warning shot. He argued it had made coordination between American labs more plausible, while doubting a global race could be stopped without something as drastic as a temporary halt on capability improvements. You can read that as an attempt to make a small incident do the job a large one did for nuclear weapons, before we find out what the large one looks like.

Rotblat, after leaving Los Alamos, spent the rest of his life on this problem. In 1957 he helped found the Pugwash Conferences, which kept scientists from both sides of the Iron Curtain talking about nuclear risk when their governments mostly weren’t. In 1995 he and Pugwash shared the Nobel Peace Prize. Nobody would claim scientists talking to each other ended the arms race. But it kept a channel open, and it made restraint something people could imagine.

Coxon calling for the labs to coordinate, and Pachocki calling for voluntary slowdowns, sit in that tradition. The open question is whether that kind of coordination can happen before our Hiroshima rather than after it.

Fermi could make his joke because someone had done the maths. The people closest to this technology are now telling us, in public, that they haven’t. And they’re not joking.