An Anthropic researcher quit his job over fears that the lab and its competitors are speedrunning towards human extinction due to the global AI race.
Jacob Coxon, a researcher who specialises in training new AI models by making them consume vast amounts of data, said that he quit his company because he did not want to participate in the building of AI systems that can self-improve. He left OpenAI earlier this year to join Anthropic.
In a thread on X, he said that he spent three years doing research at both OpenAI and Anthropic and that neither company was acting responsibly.
He asked people to not underestimate the power of the technology.
“These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.”
He also addressed a common question that he gets asked when such warnings are made. He also said that he joined Anthropic because he believed that the company was known for its model safety efforts. However, he now believes that no company can responsibly develop AI capable of outperforming humans across a range of tasks, sometimes referred to as artificial general intelligence, without government intervention or a coordinated slowdown across the industry.
“A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk,” he said.
However, despite the risks he said that he was optimistic about a “potential for coordination.”
“Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.”
Earlier this year, another researcher left the company to study poetry, warning that the “world is in peril”.
In response to the post, the Alignment Science Lead at Anthropic, Evan Hubinger, agreed with Jacob and said that he and his colleagues do “earnestly believe AI could kill all humans,” and said that his estimate was more than 10% in the next decade.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
The post was met with intense trolling from netizens who responded with sarcastic memes.
“Crazy take: maybe you shouldn’t be allowed to gamble with our lives?” one person responded.
“Concerning! We should do something about this. Maybe. Or we could rename the Great Lakes,” Sam Stein, a journalist at MSNBC, said.