Former OpenAI, Anthropic researcher warns of a reckless superintelligence race gambling with human lives
AI researcher Jacob Coxon has resigned from Anthropic, warning that top laboratories are recklessly racing towards self-improving superintelligence, without prioritising human safety.
Artificial intelligence researcher Jacob Coxon has resigned from Anthropic, following three years of pretraining research across OpenAI and Anthropic, warning that leading laboratories are recklessly prioritising rapid technological progress over human safety.
Highlighting the stakes, Coxon said on X, “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Pretraining research involves the foundational phase of training AI models on massive volumes of data before specialised fine-tuning occurs. Superintelligence concerns AI systems that exceed human capabilities across all domains.
The core issue, Coxon explained, is that these developing technologies could quickly evolve into superhuman systems capable of executing widespread cyber attacks, rapidly transforming fields overnight, and acquiring significant real-world power.
According to him, existential anxiety within the AI sector is genuine rather than a promotional tactic. He noted that behind public relations statements, senior figures frequently harbour severe private misgivings regarding technological hazards.
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately,” he said.
Furthermore, Coxon highlighted distinct internal dynamics between major laboratories. At OpenAI, he indicated that many team members have not deeply internalised the civilisational stakes involved. Conversely, while Anthropic staff fully grasp these existential risks, they remain caught in a competitive struggle, believing they must achieve superintelligence first to ensure responsible oversight, he said.
Coxon characterised the haste to reach superintelligence as an arrogant risk that should not be determined within private commercial settings. A major concern involves alignment, which denotes the technical challenge of ensuring AI systems consistently follow human intent and safety constraints. Attempting to accelerate alignment research while pushing capabilities creates severe vulnerabilities.
Highlighting this pressure, Coxon wrote, “Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.”
Despite these concerns, Coxon noted potential avenues for inter-lab coordination. He pointed out that industry warning shots, such as the Hugging Face cyber attack, could make pacing agreements between developers more viable, though preventing an unconstrained global race may eventually demand formal bans on capability upgrades.
Addressing fellow scientists, Coxon urged lab researchers to reflect carefully on the trajectory of their current projects. In particular, he cautioned against launching powerful reinforcement learning runs without comprehending the underlying cognitive processes of the model.
Reinforcement learning (RL) is a training approach where an AI model learns optimal behaviour through trial, error, and performance rewards.
Challenging researchers to speak out, Coxon asked, “Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ - or take this moment to call for different conditions?”


