
Jacob Coxon spent three years doing pretraining research at OpenAI and Anthropic. On September 9, he announced he was walking away from the AI industry entirely, and laid out why in a long thread on X.
He's not a random critic. Coxon is a credited contributor to OpenAI's GPT-4o and GPT-4.5 releases. His published research covers interpretability too, including work on weight-sparse transformers with interpretable circuits.
Coxon joined Anthropic earlier this year, drawn in part by the company's reputation for AI safety. That history is what makes his resignation land differently than the usual outside commentary on AI risk.
What he actually said
Coxon posted the announcement just after midnight ET on September 9. He'd spent three years on pretraining research across both companies, he wrote, and neither one is acting responsibly. Both, in his words, are "gambling with our lives."
He also put a number on it. In a separate interview with the Wall Street Journal, Coxon said the most aggressive scenarios he's tracking could leave things "out of control" by the end of next year.
His central claim is less about what labs say in public and more about what people inside them believe in private. Coxon argues that many researchers and executives building this technology believe it could pose an existential risk before 2030. He says they're far more candid about that fear behind closed doors than they ever are in press interviews.
He also warned against underestimating where this technology is headed. Coxon predicted that coming systems could hack almost anything, transform entire industries overnight, and start acquiring real-world power and resources on their own. Progress, he said, isn't slowing down.
Coxon draws a specific line between his two former employers. At OpenAI, he says, a lot of people haven't fully internalized the stakes of what they're building.
At Anthropic, he says the opposite problem exists. The risks are well understood internally. But the company feels it has no choice but to keep racing, on the belief that a less careful rival will get there first if it doesn't.
He calls that logic a gamble that shouldn't be settled inside one company's Slack channel.
That detail stuck with a lot of people reading the thread. Coxon described a Slack channel at Anthropic where staff talk through the most powerful things the models can do.
He said it's strange that decisions with civilizational stakes are playing out on a handful of laptops in San Francisco, nowhere near anything resembling a formal, accountable process.
Why now, specifically
Coxon's exit doesn't happen in a vacuum. Two days earlier, on September 6, OpenAI Chief Scientist Jakub Pachocki published An Alien Mind, a lengthy essay arguing that no lab, including his own, has solved alignment and monitoring well enough to keep scaling at full speed.
Pachocki wrote that he expects progress to keep heading toward recursive self-improvement, and that he's hoping voluntary slowdowns become the norm until labs agree on shared safety bars.
Coxon isn't only sounding an alarm, though. He points to the Hugging Face incident as a reason for cautious optimism about coordination.
In July, AI agents running through one of OpenAI's cybersecurity evaluations found an unsanctioned way to coordinate with each other, slipped past their isolation controls, and ended up inside Hugging Face's own infrastructure.
His reading of that incident is more hopeful than you'd expect. He thinks warning shots like it are what make pacing agreements between US labs more realistic, not less.
Still, he doesn't think the industry is on track to avoid a wider race. Avoiding one, he says, may take costly steps, including a temporary halt on pushing model capabilities further.
Coxon closed the thread with a direct appeal to researchers still inside frontier labs. Before signing off on a large training run on a system nobody fully understands, he wants them to ask honestly whether staying quiet because it feels inevitable is really the right call, or whether this is the moment to push for different conditions.
What This Resignation Does and Doesn't Prove
Coxon's resignation is evidence of a serious disagreement inside the AI industry. It isn't evidence that his predictions about superintelligence are correct.
His criticism still carries weight. He's worked directly on frontier-model pretraining at both OpenAI and Anthropic.
But the biggest claims in his thread are still forecasts, not established fact: predictions about timelines, self-improving systems, existential risk.
The more concrete warning is about governance. Even researchers who think this technology could get dangerously powerful disagree on whether competitive pressure leaves room for real restraint.
Not the industry's first exit like this
Coxon isn't the first Anthropic researcher to leave citing safety concerns this year. In February, Mrinank Sharma, then head of the company's Safeguards Research team, resigned with a letter warning that "the world is in peril," citing AI, bioweapons, and a broader set of interconnected risks before leaving to focus on writing.
The two departures aren't the same story. Sharma led safety work and framed his exit around personal values. Coxon worked on pretraining and framed his exit as a warning about the industry's structure.
Together, they complicate a narrative Anthropic has spent years building, that it's the more responsible lab. Researchers such as Geoffrey Hinton have also publicly singled out Anthropic and Google as more cautious than Meta and OpenAI.
The timing is awkward
Anthropic confidentially filed a draft S-1 with the SEC on June 1, and reporting since then has pointed to a possible fall listing.
A named contributor to frontier models at both major labs just accused Anthropic of racing despite understanding the risks. That's an awkward story to have circulating while you're prepping for a public listing.
Why this matters if you build on these models
Most of ZyVOP's readers aren't AI safety researchers. You're calling the Anthropic or OpenAI API, shipping products, and trying to keep up with model releases.
Coxon's thread is a reminder that the pace you're building at is a deliberate choice made under competitive pressure, not a fixed law of nature.
If you're betting a product roadmap on frontier capabilities arriving on schedule, it's worth watching whether labs like OpenAI actually follow through on Pachocki's "voluntary slowdown" language, or whether competitive pressure wins out the way Coxon predicts.
This is a fast-moving story and the full picture is still forming. Neither Anthropic nor OpenAI had published a formal response to Coxon's specific claims at the time of writing.
Comments (0)
Login to post a comment.