Current Affairs explainer · 11 September 2026 · Ethics + S&T coverage of the Anthropic resignation wave
- Why did Jacob Coxon resign from Anthropic?
- The corroboration that made it a global story
- The week it belonged to
- Ethics-classroom angles (GS-4)
- The >10% number is not a fringe estimate
- Inside Anthropic’s safety model — and the gap insiders allege
- The ethics toolkit, expanded (GS-4)
- Frequently asked questions
- Who is Jacob Coxon?
- What did Anthropic’s alignment lead say?
- Is a 10% extinction estimate from AI mainstream?
- Revision card
- Sources
The news in one line: Jacob Coxon, a 27-year-old AI-safety researcher at Anthropic, resigned on 9 September 2026, warning of a “reckless race toward superintelligence” — and the company’s own alignment lead, Evan Hubinger, publicly agreed that AI killing all humans is a real possibility he rates at “>10% within the next decade.”
Why did Jacob Coxon resign from Anthropic?
- Coxon quit over concerns that Anthropic and its competitors are not taking safety seriously enough amid the race to superintelligence.
- His resignation message to colleagues: without action, superintelligent AI carries “a risk of causing human extinction.”
- His bluntest line: “The people building AI earnestly believe that it could kill us all by the end of [the decade].”
The corroboration that made it a global story
Anthropic alignment lead Evan Hubinger posted on X: “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” When the engineers inside the race say the quiet part loudly, the story moves from fringe to front page — Forbes, WSJ, The Washington Post, NBC, PBS and the BBC all ran it the same cycle.
The week it belonged to
The resignation landed amid a cluster of AI-risk headlines: OpenAI leaders calling the current phase “scary, uncontrollable”; Anthropic’s “When AI builds itself” research on recursive self-improvement; and OpenAI’s disputed Navier–Stokes claim. Together they sketch the 2026 AI debate: capability announcements and extinction warnings arriving in the same breath.
Ethics-classroom angles (GS-4)
- Whistleblowing vs loyalty: resignation as moral witness — was speaking out inside enough, or was public exit justified?
- Consequentialism vs precaution: how to weigh a >10% catastrophic tail against certain near-term benefits.
- Epistemic responsibility: engineers estimating extinction probabilities in public — expertise, uncertainty, and rhetoric.
The >10% number is not a fringe estimate
What makes the Hubinger quote newsworthy is that it tracks what AI-risk surveys have shown for years. Large surveys of machine-learning researchers have repeatedly returned median estimates in the 5–10% range for catastrophic outcomes from advanced AI; the 2023 CAIS statement — “Mitigating the risk of extinction from AI should be a global priority” — was signed in one line by the field’s founders (Hinton, Bengio) and its chief executives (Altman, Hassabis, Amodei). The resignation converts a statistical debate into a workplace story: someone who ran the safety work concluded the practice was not matching the papers.
Inside Anthropic’s safety model — and the gap insiders allege
Anthropic built its brand on structured safety: the Responsible Scaling Policy with AI Safety Levels (ASLs) that gate how capable a model may be trained based on demonstrated risks, plus interpretability and alignment research teams. Coxon’s charge, and the support it drew, is that policy has not kept pace with practice — that launch pressure (competitors shipping frontier models on schedule) compresses the evaluations the policy promises. For answers, frame it as a case study in organizational ethics: formal frameworks exist, incentives undercut them, and the whistleblowing channel (resignation + public statement) becomes the residual check.
The ethics toolkit, expanded (GS-4)
- Whistle-blowing vs loyalty: test with the public-interest condition — was internal escalation exhausted before the public exit?
- Precautionary principle vs proactionary racing: irreversibility and magnitude of harm justify action under uncertainty; the counter-position is that racing yields the safety knowledge fastest.
- Moral agency of engineers: professional codes (ACM/IEEE analogues) place responsibility on builders, not only firms or states.
- Precedents to cite: Geoffrey Hinton leaving Google (2023) to speak freely; the Manhattan Project’s scientists’ debates — the classic parallel for “builders warning about the build”.
One exam-smart closing line: the story’s significance is not the probability estimate — it is that the estimate comes from the people closest to the machine, which is either the best-informed warning in technology history or an indictment of the culture that produced it. Both readings belong in the answer.
Frequently asked questions
Who is Jacob Coxon?
A 27-year-old AI-safety (alignment) researcher at Anthropic who resigned on 9 September 2026, warning in his exit message that the industry is in a “reckless race toward superintelligence” with human-extinction risk left unaddressed.
What did Anthropic’s alignment lead say?
Evan Hubinger publicly endorsed the warning on X: “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade” — the corroboration that turned a resignation into a global story.
Is a 10% extinction estimate from AI mainstream?
Within the field it is a minority-adjacent but not fringe view: surveys of ML researchers have repeatedly returned median catastrophic-risk estimates in the 5–10% range, and the 2023 CAIS statement on AI extinction risk was signed by the field’s founders and chief executives.
Revision card
- Who: Jacob Coxon, AI-safety researcher, Anthropic; resigned 9 Sept 2026.
- Warning: “reckless race toward superintelligence”; extinction risk.
- Backing: Evan Hubinger (alignment lead): “>10% within the next decade.”
- Context: same week as OpenAI “scary phase” remarks and Anthropic RSI research.
Sources
- NBC News — resignation
- Washington Post — reckless race
- Forbes — alignment lead warning
- PBS — dangers of AI development
Quick revision
- Coxon quit over concerns that Anthropic and its competitors are not taking safety seriously enough amid the race to superintelligence.
- His resignation message to colleagues: without action, superintelligent AI carries “a risk of causing human extinction.”
- His bluntest line: “The people building AI earnestly believe that it could kill us all by the end of [the decade].”
- Whistleblowing vs loyalty: resignation as moral witness — was speaking out inside enough, or was public exit justified?
- Consequentialism vs precaution: how to weigh a >10% catastrophic tail against certain near-term benefits.
- Epistemic responsibility: engineers estimating extinction probabilities in public — expertise, uncertainty, and rhetoric.
Have a doubt on this topic?




