Evan Hubinger warns about unresolved superintelligence alignment
An Anthropic alignment researcher gives a personal risk forecast and distinguishes future superintelligence from current models.
I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence
About these dates
Statement date displayed on X in the inspected browser; not converted to UTC.
Responding to Jacob Coxon, Anthropic alignment researcher Evan Hubinger put his personal probability of AI killing all humans above 10% within the next decade. He credited Anthropic’s efforts but said there was no plan to solve superintelligence alignment and progress was not clearly on track. This records a personal judgment, not a measured extinction probability or an official Anthropic assessment. Evan Hubinger ↗Evan Hubinger ↗
His follow-up separated that forecast from present models, which he considered low risk. His concern was future superintelligence reached through recursive self-improvement, rather than an assertion that current systems already posed that level of danger. Evan Hubinger ↗
Sources & attribution
- Commentary 8 Sept 2026Personal warning about superintelligence alignment ↗
Evan Hubinger. Original source inspected. Personal views.
- Commentary 8 Sept 2026Clarification on current and future models ↗
Evan Hubinger. Original source inspected. Personal views.
- Organizational disclosure Accessed 10 Sept 2026Professional profile ↗
Evan Hubinger. Professional role checked; date is the access date.