Anthropic researcher quits, warns AI ‘gambling with our lives’ as colleague pegs extinction risk above 10%

Anthropic logo
Anthropic logo. Credit: Anthropic / Wikimedia Commons (public domain).

A pretraining researcher at Anthropic, Jacob Coxon, announced his resignation on X on Tuesday, warning that AI labs are “racing straight to self-improving superintelligence and gambling with our lives.” Coxon, who previously worked at OpenAI before joining Anthropic, wrote that “the people building AI earnestly believe that it could kill us all by the end of the decade,” and pointed to incidents such as the recent Hugging Face attack as a “warning shot” that could make coordination between U.S. AI labs more viable.

Rather than distancing the company from Coxon’s warning, current Anthropic staff responded publicly to back up the substance of it. Evan Hubinger, who leads alignment-science work at Anthropic, wrote on X: “We really do earnestly believe AI could kill all humans! I personally think it is more than 10 percent within the next decade.” Hubinger added that Anthropic does not yet have “a plan to solve alignment for superintelligence.”

Samuel Marks, Anthropic’s scalable-oversight lead, posted a lengthy personal analysis making a similar point: “AI developers believe their technology could cause human extinction (or similarly bad outcomes) … in the next few years.” Marks noted he was speaking in a personal capacity rather than for the company.

Anthropic San Francisco office
File photo: Anthropic’s San Francisco office. Credit: Department for Science, Innovation and Technology (UK) / Alecsandra Dragoi, CC BY 2.0.

Anthropic did not respond to requests for comment on Coxon’s resignation or on his colleagues’ public statements. The exchange came just days after OpenAI chief executive Sam Altman and other industry leaders suggested that AI development was entering what they called an “uncontrollable phase,” a framing that has intensified debate over how AI labs should be racing against each other, if at all, on frontier model development.

The statements, made on social media rather than in a formal company publication, do not represent an official Anthropic position, and the company has not issued one. But the willingness of sitting Anthropic researchers to publicly affirm double-digit extinction-risk estimates marks an unusually direct moment of internal dissent from inside a leading AI lab.

Sources: TechCrunch; Axios; Newsweek; Forbes.

Leave a Reply

Your email address will not be published. Required fields are marked *