Jacob Coxon, a researcher who worked on AI pretraining at both OpenAI and Anthropic, has departed and publicly accused both organizations of deliberately accepting extinction-level risks. His departure comes alongside statements from Evan Hubinger, a safety researcher at Anthropic, who estimates the probability of a misaligned superintelligent AI killing humanity within ten years at above ten percent.
Coxon's exit marks an escalation in internal dissent within leading AI labs. Researchers at the frontier of AI development are increasingly vocal about what they view as inadequate safety precautions. Coxon's decision to leave and speak publicly suggests frustration with the pace and seriousness of safety work relative to the speed of capability advancement at both companies.
Hubinger's probability estimate carries weight because it comes from inside Anthropic, a company founded explicitly on AI safety principles. Anthropic was created in 2021 when Dario Amodei and several colleagues, including Hubinger, departed OpenAI to focus specifically on building safer AI systems. If a safety researcher at Anthropic privately believes ten-percent-plus extinction odds exist this decade, this reflects genuine concern within an organization structured around risk mitigation.
The timing matters. Both OpenAI and Anthropic are racing to scale AI systems that approach and potentially exceed human-level reasoning across many domains. Claude, Anthropic's flagship model, now handles complex analysis, coding, and reasoning tasks. GPT-4, OpenAI's primary system, operates at similar capability levels. The gap between current AI and systems powerful enough to pose existential risk may narrow faster than safety infrastructure can adapt.
Coxon's specific accusations that companies knowingly accept extinction risk suggest an active choice rather than negligence. This implies leaders understand the hazards but prioritize speed, capability gains, or competitive positioning. The accusation cuts deeper than standard safety disagreements. It alleges that senior people made conscious risk calculations that favor near-term progress over long-term survival odds.
Anthropic has invested in interpretability research and constitutional AI methods designed to make models more aligned with human values. These efforts acknowledge safety concerns. Yet Hubinger's public probability estimate suggests these measures, however sophisticated, don't close the gap to safe levels in his view.
The ten-percent figure represents a stark departure from reassurance narratives common in tech leadership. Most public statements from AI executives emphasize safety progress and controlled development. An inside researcher from a safety-focused company putting near-term extinction odds in double digits contradicts the "we have time to solve this" framing.
This moment reflects deepening fractures between researchers who believe current AI development trajectories include unacceptable risks and companies that continue scaling. Coxon's departure and Hubinger's public statements suggest the internal conversations about extinction risk are not producing agreement or action sufficient to satisfy safety-focused researchers.
The claims remain contested within AI circles. Many researchers dispute extinction risk models or believe current safeguards work better than Hubinger suggests. But when scientists embedded in the leading organizations building these systems go public with multi-digit extinction probabilities this decade, the debate shifts. These are not outside critics or doomsayers. These are people who watched the systems develop from the inside.
Whether other researchers will follow Coxon's path, or whether Anthropic and OpenAI will address these specific concerns with concrete changes, remains unclear. The pattern of departures and public accusations suggests the pressure is building.
