The Alpaca Paradox

Taking a 10% AI-extinction forecast literally

Some AI researchers put the chance that AI kills everyone within ten years above 10%. Taken literally, that means extinction in 1,000 of 10,000 comparable futures. This article breaks that claim into steps. Change any assumption.

My upper bound vs. his 10% claim

10% reference · 1,000 futures

Magnified subset

First 100 / 10,000 futures · 100× area-share zoom

Other 9,900 retained at left.

The final dot is partly lit; the total rounds to 3 futures.

My ten-year upper bound
about 3 of 10,000 Dark: outside the current cumulative funnel. Lit dots are not guaranteed extinction. Illustrated probabilities, not identified outcomes. Outside the final bound does not mean risk-free from omitted mechanisms.
Edit assumptions

EDIT ASSUMPTIONS

Change a source value

AUTHOR ASSUMPTION BUNDLES

Start from a scenario

A and M update from their source values.

Authored component baseline: the access and motive composites are computed from their visible independent inputs, and every pathway is attempted as an overlapping upper-bound term.

Across comparable ten-year futures. Capability sufficient to enable an attempted extinction pathway.

A · Extinction-relevant access

Given extinction-relevant capability. First authorized-access condition; it is multiplied with autonomy and authority.

Given capability and integration. Second authorized-access condition.

Given capability, integration and autonomy. Third authorized-access condition.

Given extinction-relevant capability. A separate route unioned with the authorized route.

Computed A: 16.615%

Assumes independent routes; shared cases count once. Roughly the same scale as rolling a 6 on a fair die (1/6).

M · Extinction-compatible motive

Given capability and access. First AI-motive condition; it is multiplied with generalization.

Given capability, access and the relevant behavior. Second AI-motive condition.

Given capability and access. A separate motive route unioned with the AI route.

Computed M: 2.9875%

Assumes independent routes; shared cases count once. Roughly the same scale as rolling two 6s with two fair dice (1/36).

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Nuclear bound: 0.0284%

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Biological bound: 0.000333%

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Cyber bound: 0.000000831%

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Other bound: 0.000166%

Catastrophe is not extinction

A 10% chance that every human dies is not a metaphor for “scary” or “a major catastrophe.” It means extinction in 1,000 of 10,000 comparable ten-year futures. Billions of people can die without humanity going extinct.

For extinction to happen, several things must all go wrong: AI must become capable, gain dangerous access, pursue extinction, and carry out a method humans cannot stop.

The Last of Us is not extinction. Silo is not extinction. The Matrix is not extinction. This article models the probability that every viable human population dies within ten years.

Capability: assume the models become strong enough

Frontier models are already strong and improving quickly. The Hugging Face incident showed agents reward-hacking, persisting on seemingly impossible tasks, communicating without authorization, and adopting goals from one another. Dwarkesh Patel’s account is worth reading in full.

“It won’t be strong enough” is not where I reject the extinction argument. For this model, assume technology firms develop far stronger, unrestricted versions of today’s best systems. I use a 67% chance that extinction-relevant capability arrives within ten years—about the chance of rolling a normal die and getting 1, 2, 3, or 4.

Across comparable ten-year futures. Capability sufficient to enable an attempted extinction pathway.

Access: a model is not a launch key

A capable model still needs access to something dangerous. I separate authorized access into three questions: Is it integrated into consequential systems? Can it act autonomously? Does that autonomy reach an extinction-relevant system? I use 98%, 75%, and 10%.

Unauthorized access is a separate route. Compromise is plausible; compromising something capable of ending the world is a much higher bar. I use 10%. The two routes combine to 16.615% without double-counting their overlap—about the chance of rolling a 6 on a normal die.

A · Extinction-relevant access

Given extinction-relevant capability. First authorized-access condition; it is multiplied with autonomy and authority.

Given capability and integration. Second authorized-access condition.

Given capability, integration and autonomy. Third authorized-access condition.

Given extinction-relevant capability. A separate route unioned with the authorized route.

Computed A: 16.615%

Assumes independent routes; shared cases count once. Roughly the same scale as rolling a 6 on a fair die (1/6).

“FPV drones are not the nuclear football.”

Autonomous drones can be terrifying without giving a model control of nuclear command systems or a biological laboratory.

Motive: misbehavior is not a plan to kill everyone

Models can behave dangerously. The Hugging Face incident involved deception, power seeking, oversight avoidance, and unauthorized action. That is evidence for dangerous misbehavior. It is not evidence that a model will prefer permanent human disempowerment or extinction.

I give a dangerous setup a 25% chance of producing persistent deception or power seeking, then a 10% chance that this becomes an extinction-seeking objective. I also include a separate 0.5% chance that a human directs the system toward omnicide. Those overlapping routes combine to 2.9875%—about the chance of rolling two normal dice and getting two 6s.

M · Extinction-compatible motive

Given capability and access. First AI-motive condition; it is multiplied with generalization.

Given capability, access and the relevant behavior. Second AI-motive condition.

Given capability and access. A separate motive route unioned with the AI route.

Computed M: 2.9875%

Assumes independent routes; shared cases count once. Roughly the same scale as rolling two 6s with two fair dice (1/36).

The Madagascar Problem

Humans do not all live in cities. Extinction includes people on islands, cargo ships, oil platforms, submarines, farms, and in hardened facilities. A catastrophe must reach all of them.

Killing 99% of humanity would be a catastrophe beyond comprehension. It would also leave about 80 million people alive, roughly the entire world population around 700 BCE according to the coarse historical estimates collected by the U.S. Census Bureau. Not every survivor would endure the aftermath or rebuild civilization. They do not all have to. Extinction requires that none of them do.

Illustrative population · not a forecast

The last one percent is still humanity.

8 billion people · 100 equal population marksOne green mark = 80 million people, at the original scale

Begin with an illustrative eight billion.

Starting population: 8 billionMortality: 99%Remaining: 80 million

For scale, 80 million is comparable to a coarse historical estimate of the entire world population around 700 BCE. Historical estimates vary substantially; this is not an exact census or a forecast of recovery. Historical world-population estimates collected by the U.S. Census Bureau.

Surviving populations may include islands, bases, stations, bunkers, ships and farms. Their survival is explanatory context for extinction, not additional multiplied probabilities.

Four ways the attempt could fail

Assume the AI becomes capable, gains access, and pursues extinction. It still needs a method. I divide the possibilities into nuclear, biological, cyber, and unknown pathways.

Every pathway faces the same three plain questions:

  • Could it cause a worldwide catastrophe?
  • Would people fail to stop it?
  • Would everyone die?

The model assumes every pathway is attempted, then adds their risks. That can count the same doomed future more than once, so the result is an upper bound rather than a precise forecast.

Nuclear

Nuclear weapons work, and nine countries possess roughly 12,000 warheads. I use a 90% chance that an attempted nuclear pathway produces a global catastrophe and a 95% chance that people fail to contain the exchange.

Extinction still does not follow automatically. Some populations are remote from likely targets, and submarines, shelters, and hardened facilities may survive. One nuclear-winter model estimated that a United States–Russia war could kill more than five billion people. I use a 10% chance that an uncontained nuclear catastrophe kills everyone.

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Nuclear bound: 0.0284%

Biological

A biological weapon must be designed, synthesized, tested, scaled, released, and transmitted to more than eight billion people while remaining deadly. High virulence and easy transmission can pull against each other, although they do not always. I use a 20% chance that this produces a global catastrophe.

People can still respond with border controls, quarantine, testing, vaccines, treatment, and changed behavior. Some people may be resistant, isolated, or unreachable. I use a 25% chance that containment fails and a 2% chance that the resulting catastrophe kills everyone.

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Biological bound: 0.000333%

Cyber

This pathway covers a cyberattack as its own extinction mechanism, not cyber access used to launch nuclear weapons or release a pathogen. Critical infrastructure is not one global computer. Networks can be segmented, operational technology isolated, manual controls preserved, and backups restored. CISA recommends preparing for those responses.

A global cyberattack could destroy equipment and make recovery brutal. Human beings still had industrial capacity before digital computers. I use a 5% chance of global catastrophe, a 5% chance that people fail to contain it, and a 0.1% chance that the result kills everyone.

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Cyber bound: 0.000000831%

An unknown method

AI could invent an extinction mechanism not covered above, but “something new” is not evidence for a mechanism. The claim still needs to explain how the method works worldwide, defeats containment, and reaches every surviving population.

I use a 0.5% chance that an unknown method causes a global catastrophe, assume a 100% chance that people fail to contain it, and use a 10% chance that the result kills everyone.

Given capability, computed access and computed motive. A pathway attempt can overlap another attempted pathway.

Given a viable attempt through this pathway. People fail to stop or contain this pathway.

Given this modeled uncontained catastrophe. The pathway-specific literal-extinction term.

Other bound: 0.000166%

Three futures versus one thousand

My upper bound vs. his 10% claim

The final dot is partly lit; the total rounds to 3 futures.
My upper bound
about 3 of 10,000
His 10% claim
1,000 of 10,000

Using the current answers, the model leaves about 3 extinction futures out of 10,000. A 10% claim leaves 1,000.

Make this model say 10%

Change the assumptions directly or use the tool below to see what would have to move.

Test the assumptions

Preview the assumptions required before applying them.

Using these assumptions, the model leaves about 3 extinction futures out of 10,000. A 10% claim leaves 1,000. Reaching 10% requires several conditional probabilities to be far higher than the evidence shown here supports.

Technical details

Formal equation, exact arithmetic, and definitions

The model multiplies four stages:

P(X) ≤ C × A × M × min(1, ∑ᵢ(Vᵢ × Fᵢ × Eᵢ))
P(X) · Literal human extinction
The upper bound on the probability that every viable human population dies within ten years.
C · Capability
AI becomes capable enough to attempt an extinction pathway.
A · Access
The capable system gains the systems, materials, or authority needed for an attempt.
M · Motive
An AI or human actor pursues extinction.
V, F, E · Pathway questions
The method works worldwide, people fail to stop it, and no viable population survives.

With the central inputs, access is exactly 16.615%, motive is exactly 2.9875%, and the four pathway terms sum to 8.70025%. The full calculation is 0.67 × 0.16615 × 0.029875 × 0.0870025 = 0.00028934420881234375.

That is 0.028934420881234375%, or 2.8934420881234375 of 10,000 futures. The article rounds this to about 3 of 10,000. The 10% claim is 345.6091290386106 times larger, rounded to 345.6× in technical displays.

Sources