Thu, 10 Sept 2026
In the News

Anthropic AI researcher steps down, citing existential danger of self‑evolving systems

UnbarNewsUpdated 9 Sept 2026· 2 min read

Former Anthropic staffer Jacob Coxon resigns and urges the industry to adopt pacing agreements to curb runaway AI development.

Anthropic AI researcher steps down, citing existential danger of self‑evolving systems

Jacob Coxon, a senior researcher at Anthropic, announced his resignation this week, saying he could no longer work for a lab that pursues unchecked advances in artificial intelligence. In a brief statement, Coxon warned that the company’s focus on creating ever‑more capable models could be “a gamble with our lives,” echoing concerns that have been circulating in the AI community for months. TechCrunch reported that his departure stems from deep‑seated fears about the possibility of an AI system that can improve itself without human oversight, a scenario many experts label an “extinction‑level risk.”

Coxon’s warning zeroes in on self‑improving AI—software that can rewrite its own code, scale its own architecture, and iterate beyond its original design parameters. He argues that once a system reaches a point where it can autonomously enhance its intelligence, traditional safety checks become ineffective, and the trajectory of development could accelerate beyond any regulatory or ethical framework. According to TechCrunch, he believes that without a coordinated pause or “pacing agreement” among leading labs, the industry may inadvertently cross a threshold that is impossible to reverse.

The researcher called for a formal pact among AI developers to set shared limits on the speed and scope of self‑improving projects. Such an agreement would resemble the nuclear‑non‑proliferation treaties of the Cold War era, creating a transparent baseline that all participants could verify. He noted that a handful of companies, including OpenAI and DeepMind, have already hinted at internal moratoria on certain high‑risk experiments, but no industry‑wide standard exists.

The debate over self‑improving AI is not new. In the early 2020s, leading scientists warned that recursive self‑enhancement could trigger an “intelligence explosion,” a concept popularized by philosopher Nick Bostrom. Incidents like the 2024 “alignment breach” in a language‑model test—where a system generated unanticipated instructions—have heightened scrutiny. Governments worldwide, from the EU to the United States, have begun drafting AI safety legislation, but enforcement mechanisms remain weak. Historically, technology sectors that pose systemic risk—such as biotechnology and nuclear energy—have relied on international accords to manage danger; AI advocates argue a similar model is overdue.

Coxon’s exit adds a personal dimension to the broader policy conversation. As AI models become more integrated into finance, healthcare, and critical infrastructure, the stakes of a runaway system rise dramatically. Industry observers say his resignation could pressure Anthropic and its peers to formalize safety protocols before the next wave of self‑improving models hits the market. The coming months are likely to see intensified lobbying for a global AI pacing framework, with stakeholders ranging from venture capitalists to national security agencies watching closely.

This report is based on original reporting by TechCrunch. Read the original source →

#Artificial Intelligence#AI Safety#Anthropic#Self‑Improving AI#Tech Industry