Post-Cutoff.com
  1. Home
  2. Posts
  3. "Jacob is correct": Anthropic alignment lead puts AI…

"Jacob is correct": Anthropic alignment lead puts AI extinction risk above 10% this decade

Evan Hubinger @EvanHub · x · 2026-09-09 · ★★★★ · archived

Open the original ↗

A serving Anthropic alignment lead publicly backed Coxon and said Anthropic has no plan yet to align superintelligence. Press worldwide quoted it.

Summary

Hubinger, who leads alignment stress-testing at Anthropic, quote-tweeted Coxon's resignation. He wrote that lab researchers "really do earnestly believe AI could kill all humans", that his own estimate is above 10% within the next decade, and that although Anthropic is trying its best it does not yet have a plan to solve alignment for superintelligence and is not clearly on track to. Coming from a current employee, this made the story much bigger (TechCrunch, Scientific American, OfficeChai, Reuters' "Ten days that changed the course of AI"). Verified via the X syndication API (2026-09-09 01:27 UTC).

Archived text

Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

Quoting @hilbertspaess: The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear

likes 59332 · replies 4859 (at fetch time)

Archived 2026-09-29 via syndication.

Related events

All posts · id: 2026-09-09-hubinger-agrees-10-percent