On Tuesday evening, Anthropic pretraining researcher Jacob Coxon announced on X that he’d resigned. Within hours, two of his colleagues — still employed at the company — went public with variations of the same message, warning that the technical problem of aligning superintelligence remains unsolved even as the race to build it shows no signs of slowing down.
Coxon, 27, spent three years doing pretraining work at OpenAI and then Anthropic. He didn’t frame his departure as a protest against a single employer, but that both companies are moving toward self-improving superintelligence without adequate safeguards.
“The people building AI earnestly believe that it could kill us all by the end of the decade,”
“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote. “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately.”
Evan Hubinger, Anthropic’s Alignment Science Lead, responded directly to Coxon’s thread.
“Jacob is correct here — we really do earnestly believe AI could kill all humans,” Hubinger wrote. “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
“Jacob is correct here — we really do earnestly believe AI could kill all humans,”
Hubinger runs the team that stress-tests Anthropic’s own alignment techniques — probing for the ways they might fail before those failures show up in deployed models.
Source link







