Researcher Jacob Cookson announced his resignation from Anthropic, warning that the race to develop a self-improving superintelligence could endanger human lives. He also criticized two companies he worked for over the past three years: OpenAI and Anthropic, according to Bloomberg.
Cookson wrote in a series of posts on the X platform: "Neither company is acting responsibly. They are racing headlong toward a self-improving superintelligence and gambling with our lives."
The resignation highlights a growing debate within the AI industry about the ability of safety measures to keep pace with the development of more robust and autonomous models, even at Anthropic, which has made responsible technology development a cornerstone of its identity. This comes at a time when calls are increasing for international coordination to slow the race if risks escalate.
The Wall Street Journal quoted Cookson, a 27-year-old British researcher, as saying he moved from OpenAI to Anthropic earlier this year because of the latter's reputation for model safety, but now believes that developing these systems responsibly requires government intervention or a coordinated slowdown among companies.
Cookson specializes in pre-training models, the stage in which they learn from massive amounts of data. Anthropic did not immediately comment on the resignation, according to the newspaper.
Is it necessary to trust Chinese AI for use?
Concerns about systems beyond human control: In his publications, Cookson warned that this technology could soon produce systems that surpass humans, possessing vast capabilities to hack and disrupt scientific advancements in any field, and acquire resources and influence. He argued that the decision to enter this phase shouldn't be made through internal discussions on Slack at a private company.
The Wall Street Journal quoted him as saying that he found Anthropic's safety efforts serious, but he fears that competition between companies, and with Chinese developers, will force compromises in this area. He indicated, according to the newspaper, that some of the more accelerated scenarios could, in his estimation, lead to things spiraling out of control by the end of 2027.
Chip stocks led Asian gains, supported by a new model from OpenAI.
These concerns center on the potential development of superintelligence that surpasses human capabilities across broad fields, and on what is known as iterative self-improvement: that is, AI systems becoming capable of designing and developing even more capable successor systems, potentially accelerating progress at a pace difficult for humans to monitor and control.
Anthropic itself discussed this possibility in a paper it published titled "When AI Builds Itself," explaining that systems developing their successors independently could bring significant benefits to science and healthcare, but could also increase the risk of losing human control. She emphasized that achieving complete self-improvement is not inevitable, and that humans still excel at determining research directions and selecting important problems.