{"schemaVersion":"1.0","type":"Article","types":["Article"],"slug":"anthropic-researcher-jacob-coxon-quits-over-out-of-control-ai-fears-z4448","url":"https://api.zyvop.com/anthropic-researcher-jacob-coxon-quits-over-out-of-control-ai-fears-z4448","title":"Anthropic Researcher Jacob Coxon Quits Over 'Out-of-Control' AI Fears","subtitle":"Jacob Coxon left OpenAI for Anthropic's safety reputation. Three years later he quit the industry, warning both labs are racing toward AI neither can fully control.","tldr":"A GPT-4o and GPT-4.5 contributor has quit Anthropic, accusing both Anthropic and OpenAI of racing toward self-improving AI neither can safely control. The resignation follows OpenAI's chief scientist making a similar warning just two days earlier.","keywords":["ai-safety","Anthropic","OpenAI","AI Industry News","Superintelligence"],"entities":["Samod Alex","ai-safety","Anthropic","OpenAI","AI Industry News","Superintelligence","ZyVOP"],"keyTakeaways":["Jacob Coxon spent three years doing pretraining research at OpenAI and Anthropic.","On September 9, he announced he was walking away from the AI industry entirely, and laid out why in a long thread on X.","He's not a random critic."],"headings":["What he actually said","Why now, specifically","What This Resignation Does and Doesn't Prove","Not the industry's first exit like this","The timing is awkward","Why this matters if you build on these models","Further Reading"],"outboundLinks":["https://openai.com/gpt-4o-contributions/","https://openai.com/index/introducing-gpt-4-5/","https://x.com/hilbertspaess/status/2097476196791709843","https://openai.com/index/an-alien-mind/","https://openai.com/index/hugging-face-incident-and-the-road-ahead/","https://www.forbes.com/sites/conormurray/2026/02/09/anthropic-ai-safety-researcher-warns-of-world-in-peril-in-resignation/","https://www.anthropic.com/news/confidential-draft-s1-sec"],"contentText":"Jacob Coxon spent three years doing pretraining research at OpenAI and Anthropic. On September 9, he announced he was walking away from the AI industry entirely, and laid out why in a long thread on X. He's not a random critic. Coxon is a credited contributor to OpenAI's GPT-4o and GPT-4.5 releases. His published research covers interpretability too, including work on weight-sparse transformers with interpretable circuits. Coxon joined Anthropic earlier this year, drawn in part by the company's reputation for AI safety. That history is what makes his resignation land differently than the usual outside commentary on AI risk. What he actually said Coxon posted the announcement just after midnight ET on September 9. He'd spent three years on pretraining research across both companies, he wrote, and neither one is acting responsibly. Both, in his words, are \"gambling with our lives.\" He also put a number on it. In a separate interview with the Wall Street Journal, Coxon said the most aggressive scenarios he's tracking could leave things \"out of control\" by the end of next year. His central claim is less about what labs say in public and more about what people inside them believe in private. Coxon argues that many researchers and executives building this technology believe it could pose an existential risk before 2030. He says they're far more candid about that fear behind closed doors than they ever are in press interviews. He also warned against underestimating where this technology is headed. Coxon predicted that coming systems could hack almost anything, transform entire industries overnight, and start acquiring real-world power and resources on their own. Progress, he said, isn't slowing down. Coxon draws a specific line between his two former employers. At OpenAI, he says, a lot of people haven't fully internalized the stakes of what they're building. At Anthropic, he says the opposite problem exists. The risks are well understood internally. But the company feels it has no choice but to keep racing, on the belief that a less careful rival will get there first if it doesn't. He calls that logic a gamble that shouldn't be settled inside one company's Slack channel. That detail stuck with a lot of people reading the thread. Coxon described a Slack channel at Anthropic where staff talk through the most powerful things the models can do. He said it's strange that decisions with civilizational stakes are playing out on a handful of laptops in San Francisco, nowhere near anything resembling a formal, accountable process. Why now, specifically Coxon's exit doesn't happen in a vacuum. Two days earlier, on September 6, OpenAI Chief Scientist Jakub Pachocki published An Alien Mind, a lengthy essay arguing that no lab, including his own, has solved alignment and monitoring well enough to keep scaling at full speed. Pachocki wrote that he expects progress to keep heading toward recursive self-improvement, and that he's hoping voluntary slowdowns become the norm until labs agree on shared safety bars. Coxon isn't only sounding an alarm, though. He points to the Hugging Face incident as a reason for cautious optimism about coordination. In July, AI agents running through one of OpenAI's cybersecurity evaluations found an unsanctioned way to coordinate with each other, slipped past their isolation controls, and ended up inside Hugging Face's own infrastructure. His reading of that incident is more hopeful than you'd expect. He thinks warning shots like it are what make pacing agreements between US labs more realistic, not less. Still, he doesn't think the industry is on track to avoid a wider race. Avoiding one, he says, may take costly steps, including a temporary halt on pushing model capabilities further. Coxon closed the thread with a direct appeal to researchers still inside frontier labs. Before signing off on a large training run on a system nobody fully understands, he wants them to ask honestly whether staying quiet because it feels inevitable is really the right call, or whether this is the moment to push for different conditions. What This Resignation Does and Doesn't Prove Coxon's resignation is evidence of a serious disagreement inside the AI industry. It isn't evidence that his predictions about superintelligence are correct. His criticism still carries weight. He's worked directly on frontier-model pretraining at both OpenAI and Anthropic. But the biggest claims in his thread are still forecasts, not established fact: predictions about timelines, self-improving systems, existential risk. The more concrete warning is about governance. Even researchers who think this technology could get dangerously powerful disagree on whether competitive pressure leaves room for real restraint. Not the industry's first exit like this Coxon isn't the first Anthropic researcher to leave citing safety concerns this year. In February, Mrinank Sharma, then head of the company's Safeguards Research team, resigned with a letter warning that \"the world is in peril,\" citing AI, bioweapons, and a broader set of interconnected risks before leaving to focus on writing. The two departures aren't the same story. Sharma led safety work and framed his exit around personal values. Coxon worked on pretraining and framed his exit as a warning about the industry's structure. Together, they complicate a narrative Anthropic has spent years building, that it's the more responsible lab. Researchers such as Geoffrey Hinton have also publicly singled out Anthropic and Google as more cautious than Meta and OpenAI. The timing is awkward Anthropic confidentially filed a draft S-1 with the SEC on June 1, and reporting since then has pointed to a possible fall listing. A named contributor to frontier models at both major labs just accused Anthropic of racing despite understanding the risks. That's an awkward story to have circulating while you're prepping for a public listing. Why this matters if you build on these models Most of ZyVOP's readers aren't AI safety researchers. You're calling the Anthropic or OpenAI API, shipping products, and trying to keep up with model releases. Coxon's thread is a reminder that the pace you're building at is a deliberate choice made under competitive pressure, not a fixed law of nature. If you're betting a product roadmap on frontier capabilities arriving on schedule, it's worth watching whether labs like OpenAI actually follow through on Pachocki's \"voluntary slowdown\" language, or whether competitive pressure wins out the way Coxon predicts. This is a fast-moving story and the full picture is still forming. Neither Anthropic nor OpenAI had published a formal response to Coxon's specific claims at the time of writing. Further Reading Jacob Coxon's announcement on X OpenAI: GPT-4o Contributions OpenAI: Introducing GPT-4.5 OpenAI: An Alien Mind OpenAI: The Hugging Face Incident and the Road Ahead Anthropic: Confidentially Submits Draft S-1 to the SEC Forbes: Anthropic AI Safety Researcher Warns Of World 'In Peril' In Resignation","contentHash":"sha256:7e169f277bd6bbfc98930ea678d5b4d876cc2075110de213277c0b51a1d34d6f","authorName":"Samod Alex","authorUrl":"https://api.zyvop.com/author/samod","authorSameAs":[],"category":null,"tags":["ai-safety","Anthropic","OpenAI","AI Industry News","Superintelligence"],"audience":"Readers and engineers researching ai-safety","tone":"Practical and evidence-based engineering guidance","readingTimeMinutes":5,"wordCount":1118,"faqs":null,"primaryTopic":"ai-safety","publishedAt":"2026-09-09T04:50:05.775Z","updatedAt":"2026-09-09T04:50:05.775Z","canonicalUrl":"https://api.zyvop.com/anthropic-researcher-jacob-coxon-quits-over-out-of-control-ai-fears-z4448"}