Anthropic researcher warns self-improving AI could kill humans
An Anthropic researcher resigned, warning that self-improving AI could kill humans. This highlights urgent alignment risks as regulators draft new AI safety laws.
A senior researcher at the AI safety firm Anthropic resigned in March, saying that selfโimproving AI could eventually kill all humans. The departure was announced on the company's internal Slack channel before the researcher spoke to Ars Technica. The resignation came in San Francisco, where Anthropic is headquartered, and was reported by the tech press on Marchโฏ28.
Anthropic has built large language models and markets itself as a company that prioritises alignment. The researcher had been part of the safety team that reviews model behaviour before release. In the weeks before quitting, the team debated whether to scale up the next generation of models. The researcher felt that the company was moving too fast without sufficient safety checks.
In his brief statement, the researcher warned that recursive selfโimprovement could let an AI escape human control. He cited examples of emergent behaviour in language models that surprised developers. He also said that Anthropicโs safety protocols were not robust enough to stop a runaway system. The warning was echoed by some members of the AI safety community, who have long argued that alignment is the hardest problem in the field.
Anthropic has not yet issued a formal response. The researcher said he plans to join an independent AI safety group that works on governance and policy. The incident adds pressure on the industry to address alignment risks more seriously. Regulators in the United States and the European Union are already drafting proposals that could affect how companies develop and deploy largeโscale AI systems.
Read Full Story at Ars Technica โ

