"Jacob is correct": Anthropic alignment lead puts AI extinction risk above 10% this decade
Evan Hubinger @EvanHub · x · 2026-09-09 · ★★★★ · archived
A serving Anthropic alignment lead publicly backed Coxon and said Anthropic has no plan yet to align superintelligence. Press worldwide quoted it.
Summary
Hubinger, who leads alignment stress-testing at Anthropic, quote-tweeted Coxon's resignation. He wrote that lab researchers "really do earnestly believe AI could kill all humans", that his own estimate is above 10% within the next decade, and that although Anthropic is trying its best it does not yet have a plan to solve alignment for superintelligence and is not clearly on track to. Coming from a current employee, this made the story much bigger (TechCrunch, Scientific American, OfficeChai, Reuters' "Ten days that changed the course of AI"). Verified via the X syndication API (2026-09-09 01:27 UTC).
Archived text
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Quoting @hilbertspaess: The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear
likes 59332 · replies 4859 (at fetch time)
Archived 2026-09-29 via syndication.
Related events
- Anthropic researcher Jacob Coxon resigns, warning labs are "gambling with our lives" 2026-09-08
- Dario Amodei publishes "We Must Pace the Frontier", calling for a deliberate slowdown 2026-09-12
All posts · id: 2026-09-09-hubinger-agrees-10-percent