Evan Hubinger, a lead alignment researcher at Anthropic, assessed on X the probability of the entire human race going extinct from AI in the next ten years at over 10% — and emphasized that this is his personal assessment, not the official position of the company. According to him, current models carry relatively low risk: the concern is not about modern systems, but about a hypothetical superintelligence that could emerge as a result of recursive self-improvement. Hubinger also acknowledged that Anthropic currently does not have a ready-made solution for superintelligence alignment and that the industry is not moving in the right direction. The statement was a response to the public departure from the company of researcher Jacob Coxon, who accused frontier labs of risking people's lives by approaching self-improving superintelligence.

image

What happened

The public dispute began with posts by Jacob Coxon — a 27-year-old researcher who spent three years working on pretraining at OpenAI and Anthropic and announced his departure from the company. On X, he accused the labs of "racing straight toward self-improving superintelligence, risking our lives," and suggested that "by the end of next year, everything could already be out of control." In response, Evan Hubinger, who leads the Alignment Science direction at Anthropic, wrote that he personally assesses the probability of the entire human race going extinct from AI in the next ten years at over 10% and that the company "does not yet have a plan to solve alignment for superintelligence and we are not on the right path." This exchange is being analyzed by Forbes and Newsweek, which point out: the figure is an assessment by one of the key safety employees, not the position of Anthropic.

Context

The significance of the statement is determined by the status of the speaker. Hubinger is not an external critic, but the head of the Alignment Science direction at Anthropic, that is, a person who himself is responsible at one of the leading frontier labs for research on how to keep increasingly capable AI systems under control. The central concept of the dispute is recursive self-improvement: a mode in which AI significantly increases its own intelligence. It is precisely the superintelligence that arises in this way that Hubinger considers the source of the main risk; today's models, by his own words, he assesses as relatively low-risk. The distinction is fundamental: the observed behavior of current systems and the hypothetical scenario of self-improving AI are different levels of threat, but both parts of this assessment remain the personal judgments of the researcher, not official statements from the company.

Why this matters for the industry

For the industry, this is a rare case where an internal assessment of existential risk is publicly given by a current head of the safety direction of a leading lab. Such a data point strengthens the arguments of regulators and investors in favor of independent checks of pretraining at frontier labs and funding for safety research. Against this backdrop, demand is forming for a separate market: independent audit, safety tools, and guardrail patterns for agent workflows are shifting from the category of "would be nice" to the category of expected requirements. The flip side is that the narrative of existential risk creates macro-undercurrents for the entire AI segment — from the investment climate to the growth of skepticism among enterprise clients, who are getting questions about safety and independent assessments. The news does not carry direct technical changes: no new models, API changes, prices, or latency have been announced.

Why this matters for users

Readers can directly read the primary sources: Hubinger's response is published on X under the account @EvanHub, and Coxon's posts about his resignation are under the account @hilbertspaess. This is an opportunity to see how the industry itself formulates risk, without media retelling. When citing the figure, it is important to separate the personal assessment of the researcher and the position of Anthropic, so as not to attribute to the company what it has not stated. For everyday work with current models, nothing changes: no new restrictions or new products have been announced. If the discussion continues, background effects are more likely: more media noise about AI risks, more questions about safety from enterprise clients, and more human-in-the-loop checkpoints in agent products.

What is still unknown / limitations

The main caveat: "over 10%" is the expert opinion of a specific person without a published methodology, scenarios, or data on calibration, so as evidence this is a judgment, not a reproducible result, and it is not the official position of Anthropic. The thesis that recursive self-improvement is happening "faster than expected" is a capability claim without attached evidence: in the sources there are no benchmarks, assessments, or references to specific observations. Coxon's claims about the behavior of the labs are his own accusations, not established facts. Finally, independent checks of frontier labs are currently only a subject of arguments by regulators and investors, not active regulation: factual changes require separate verification.

Sources

Author

Look at AI, editorial team