Anthropic Researcher Quits Over AI Safety Fears

Anthropic Researcher Quits Over AI Safety Fears

Jacob Coxon has resigned from Anthropic, warning that leading AI companies are racing toward self-improving superintelligence without adequately addressing the risks.

Coxon is an AI researcher who says he spent the last three years working on pretraining research at both OpenAI and Anthropic.

His resignation became public through a post on X that quickly attracted widespread attention across the AI community.

“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”

— Jacob Coxon, @hilbertspaess

Coxon’s warning is notable because it comes from someone who has worked inside two of the companies at the center of the current frontier AI race.

Jacob Coxon’s Original X Post

Coxon announced his resignation directly on X and then published a longer thread explaining why he decided to leave.

The opening post lays out his main concern: AI development is moving toward increasingly capable systems while the industry has not solved the fundamental safety problems associated with superintelligence.

Original source: Jacob Coxon’s post on X.

What Did Jacob Coxon Say About AI Labs?

Coxon argues that the problem is not simply that AI companies are developing powerful technology.

His concern is that companies may be approaching a point where AI systems can improve their own capabilities at a pace that humans cannot reliably understand or control.

He described this as a race toward self-improving superintelligence.

In another post from the same thread, Coxon wrote:

“Do not underestimate the power of this technology.”

He argued that future AI systems could become superhuman across multiple domains and potentially gain access to significant computing power, information, and resources.

These are Coxon’s assessments and warnings. They should not be interpreted as proof that such outcomes are inevitable.

Coxon Says AI Researchers Already Fear Catastrophic Outcomes

One of the strongest claims in Coxon’s thread concerns what AI researchers and executives allegedly believe privately.

He wrote:

“The people building AI earnestly believe that it could kill us all by the end of the decade.”

Coxon said he does not view these concerns as a marketing strategy.

Instead, he argued that some executives and senior researchers publicly use more cautious language while privately expressing much greater concern about advanced AI risks.

This is an important distinction for readers.

Coxon is reporting his own experience and interpretation of conversations inside AI labs. His statement is not an official forecast from OpenAI or Anthropic.

Why Would AI Companies Keep Building If They Fear the Risk?

This question sits at the center of Coxon’s argument.

If researchers genuinely believe advanced AI could create catastrophic risks, why continue developing increasingly powerful models?

Coxon’s answer is competition.

He argues that the AI labs are caught in a race where each company fears that slowing down could allow another company to move ahead.

According to his characterization, OpenAI and Anthropic approach the problem differently.

He argues that people at OpenAI have not fully internalized the civilizational stakes, while Anthropic understands the risks but remains under pressure to keep competing.

This creates what Coxon sees as a dangerous feedback loop.

One company accelerates because it fears another company will accelerate.

The result could be faster capability development without a comparable increase in confidence about safety.

What Is Self-Improving Superintelligence?

Self-improving superintelligence describes a hypothetical AI system capable of improving its own abilities beyond the normal process of human researchers training a model.

A highly capable AI could potentially assist with software development, research, model architecture, training methods, and other parts of AI development.

If future systems become capable enough to substantially improve the systems that replace them, researchers could face a new level of control and alignment challenges.

This possibility is often discussed in the context of recursive self-improvement.

The central concern is whether human oversight would remain effective as AI capabilities increase.

Anthropic Alignment Researcher Evan Hubinger Responds

The story gained further attention when Evan Hubinger, Anthropic’s Alignment Science Lead, publicly responded to Coxon’s post on X.

Hubinger agreed with the core warning and said:

“Jacob is correct here—we really do earnestly believe AI could kill all humans!”

He then added:

“I personally think it is >10% within the next decade.”

Hubinger also acknowledged a major limitation in current AI safety research. He said Anthropic is “trying its best,” but does not yet have a plan to solve alignment for superintelligence and is not clearly on track to do so.

Importantly, Hubinger’s greater-than-10% figure is his personal estimate, not an official Anthropic prediction that AI will cause human extinction.

Original source: Evan Hubinger’s response on X.

Independent confirmation: Forbes and Business Insider both reported Hubinger’s comments and identified him as an Anthropic alignment researcher.

Why Evan Hubinger’s Response Matters

Coxon’s resignation alone could have been viewed as one researcher expressing personal concerns about AI development.

Hubinger’s response adds another perspective from inside Anthropic’s AI safety research community.

It also highlights a difficult reality in AI safety research.

Researchers can recognize potentially extreme risks while still lacking a proven method for reliably aligning a future superintelligent system.

That gap between understanding the risk and having a solution is central to the current AI safety debate.

Does Anthropic Ignore AI Safety?

No.

Anthropic has built its public identity around AI safety and alignment research.

The company publishes research on model behavior, evaluations, security, alignment, and responsible development.

Coxon’s criticism is therefore more specific.

His argument is that acknowledging the risks is not enough if competitive pressure still pushes companies toward increasingly powerful systems.

In other words, the debate is not simply AI safety versus AI development.

It is about whether safety research, safeguards, evaluations, and governance can keep pace with rapidly increasing AI capabilities.

Could AI Really Become a Threat to Humanity?

There is currently no established evidence that today’s consumer AI systems are destined to cause human extinction.

The concern raised by Coxon and Hubinger is primarily about future systems that could become significantly more capable and autonomous.

Possible risks discussed by AI safety researchers include loss of human control, deceptive behavior, autonomous replication, cyber capabilities, misuse, and failures in alignment.

These remain areas of active research and debate.

Therefore, claims about AI killing humanity should be presented as risk scenarios or expert estimates, not established future events.

What Is the AI Race Really About?

The competition between frontier AI companies is no longer only about making chatbots better.

Companies are investing heavily in reasoning models, coding agents, autonomous systems, AI research tools, robotics, and increasingly general-purpose AI.

As these systems become more capable, the commercial value of reaching the next capability milestone can also increase.

That creates a difficult incentive structure.

If one company slows development while competitors continue, it may fear losing its technological position.

This is the competitive pressure Coxon says makes voluntary restraint difficult.

Coxon’s Call for a Different Approach

Coxon is not arguing that AI research should simply disappear.

His thread instead focuses on changing the conditions under which frontier AI development happens.

He urged researchers to seriously consider whether they should participate in increasingly powerful training runs before there is a rigorous understanding of what those systems are doing.

His message is particularly directed at people working directly on frontier AI systems.

The question he raises is straightforward:

Should researchers continue accelerating AI because “it’s happening anyway,” or should they demand stronger safety conditions first?

What We Know and What We Do Not Know

What is confirmed: Jacob Coxon announced his resignation from Anthropic and publicly criticized the direction of frontier AI development.

What Coxon says: He believes OpenAI and Anthropic are racing toward self-improving superintelligence without acting responsibly enough.

What Hubinger says: He publicly agreed with the seriousness of Coxon’s warning and gave his own greater-than-10% estimate for AI-caused human extinction within the next decade.

What remains uncertain: There is no established evidence that AI will cause human extinction by the end of the decade.

These distinctions matter because AI safety discussions can quickly turn into sensational headlines.

Why This Anthropic Resignation Is Important for AI

Jacob Coxon’s departure highlights a growing tension inside the frontier AI industry.

Researchers are trying to build systems with increasingly broad capabilities while also attempting to understand and control their behavior.

The difficult question is whether those two efforts are progressing at the same speed.

Coxon believes they are not.

His resignation has therefore reopened a much larger debate about AI governance, safety research, competitive pressure, and the development of superintelligence.

Frequently Asked Questions

Who is Jacob Coxon?

Jacob Coxon is an AI researcher who says he spent the last three years doing pretraining research at OpenAI and Anthropic.

Why did Jacob Coxon leave Anthropic?

Coxon said he resigned because of concerns about the rapid development of advanced AI and the race toward self-improving superintelligence.

What did Jacob Coxon say about Anthropic and OpenAI?

He said that neither company was acting responsibly and accused both labs of racing toward self-improving superintelligence.

What did Evan Hubinger say about AI risk?

Hubinger publicly agreed with Coxon’s concern and said he personally estimates a greater-than-10% chance of AI causing human extinction within the next decade.

Does Anthropic believe AI will kill humanity?

Anthropic has not issued a prediction that AI will kill humanity. Hubinger’s greater-than-10% figure is his personal estimate.

What is AI alignment?

AI alignment is the field focused on making advanced AI systems behave in ways that reliably match intended human goals, values, and safety requirements.

What is recursive self-improvement?

Recursive self-improvement refers to a scenario where an AI system helps improve its own capabilities, potentially creating increasingly capable successor systems.

Final Takeaway

Jacob Coxon’s resignation from Anthropic is more than another employee departure.

It has brought a serious AI safety disagreement into public view.

Coxon argues that frontier AI companies are moving too quickly toward self-improving superintelligence.

Anthropic alignment researcher Evan Hubinger has separately acknowledged the severity of the potential risk while saying that a reliable solution for superintelligence alignment does not yet exist.

None of this proves that AI will inevitably destroy humanity.

But it does highlight one of the biggest unresolved questions in artificial intelligence:

Can humans safely build systems that may eventually become more capable than the people who created them?

As the AI race continues, that question may become just as important as the race itself.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *






Join Our Newsletter

Get articles and updates delivered straight to your inbox regularly.

No spam ever. Unsubscribe anytime easily.