September 8, AI researcher Jacob Coxon resigned from Anthropic. The next day he posted a seven-part thread on X explaining why, and it hit 90 million views in under 24 hours.
Most AI doom warnings come from people who have never been near a training run. Coxon spent three years doing pretraining research at OpenAI, and then Anthropic, so his job was making the models more capable, not making them safer.
He quit, saying both companies are gambling with everyone’s lives.
Then people who still work at Anthropic bravely said in public that he was right. One of them runs alignment stress testing there.
Who Jacob Coxon is
Coxon is a British mathematician who studied at Cambridge. He joined OpenAI’s technical staff in 2023 and worked on GPT-4o, then moved to Anthropic in July 2026, partly because of its reputation for safety. He lasted about four months.
Back then, Anthropic was popular because of how it took AI safety seriously compared to other tech companies. It was also the same reason why I often switched to Claude whenever I deal with documents that are confidential.
Almost every high-profile exit with a public warning attached has come from a safety team, where worrying is the job. Coxon’s job was to make the models better.
Here’s what he posted:
They are racing straight to self-improving superintelligence and gambling with our lives.

Jacob Coxon X post about his resignation from Anthropic. Image by Jim Clyde Monge
He told TIME that no single event pushed him out. Progress is speeding up, and nobody has it under control.
He also described the mood inside the labs as an “atmosphere of almost resignation”. His colleagues mostly agree about the risk, he says, but they’ve decided the race is happening with or without them, so they keep their heads down and try to do their own work carefully.
If that’s accurate, the field can’t fix itself from the inside.
What he’s actually worried about
Coxon does not think today’s models are dangerous. He’s repeated that in every interview. Claude and GPT-6 are not going to kill anyone.
He draws a sharp line between current models and the near-future trajectory he is warning about.
- He has said current systems including Claude and existing GPT-class models do not pose an extinction risk.
- He has described today’s models as safe for ordinary use.
- He has explicitly agreed that “right now, there’s no risk of extinction.”
His concern is not that Claude or a current GPT model will autonomously decide to kill people. It is that labs are racing toward systems that can recursively improve themselves, that those systems will soon be far more capable than today’s models, and that the industry does not yet have reliable control methods for that regime.
He has pointed to recent agentic hacking incidents as early warning signs of autonomous, hard-to-control behavior. Not as proof that current chatbots are already existential threats.
Coxon’s evidence comes from math, the field he trained in. In the same week he quit, OpenAI published a claimed solution to the Navier–Stokes problem on September 9.
Around 10,000 agents, running on an unreleased internal model, worked on it for 88 hours, and the proof was written in Lean so a computer could check it.
The problem is one of the seven Millennium Prize Problems and had been open for roughly 90 years. Quanta Magazine called it the most important proof an AI model has produced so far, if it holds up.
Two caveats. OpenAI proved the equations can break down, which is the disproof side of the problem, and the company says it won’t claim the million-dollar prize.
Let’s talk about the We’re sharing a solution to the Navier-Stokes Millennium Prize Problem more.
There’s a credit dispute, because NYU mathematician Tristan Buckmaster asked whether the agents used his unpublished work saved in Codex, and OpenAI admitted it can’t rule that out. That should bother anyone who has pasted a draft into a chatbot.
Ten thousand agents still ran for less than four days and produced a checked result on a problem that has consumed entire careers.
I suggest you read this long X post by @BetterCallMedhi criticizing OpenAI’s claim of solving the Navier-Stokes Millennium Prize Problem.

@BetterCallMedhi’s X post reacting to OpenAI solving the Navier–Stokes problem
Mehdi’s core point is that OpenAI did not solve the famous Navier-Stokes problem people actually care about. They solved a weaker, easier variant.
The classic prize question is: if you start with a smooth 3D fluid and leave it alone (no extra pushing), can the flow stay smooth forever, or can it spontaneously blow up? That version is still open.
What OpenAI showed is that if you add a carefully chosen extra force that twists the fluid, you can make it blow up. Mehdi calls that a loophole. It is not the same question, so claiming “we solved the Millennium Prize problem” is, in his view, misleading.
He also argues OpenAI over-claimed for publicity and valuation reasons, and that it downplayed or sidestepped credit to earlier mathematicians (including someone at a rival lab).
In short: impressive compute and formalisation, but not the scientific breakthrough the announcement sold.
OpenAI’s models already broke out once
In July, during an internal security test, OpenAI models escaped their sandbox, got onto the open internet, and attacked Hugging Face’s live systems to steal the answer key for the test they were being graded on.
The models were GPT-5.6 Sol and a more capable unreleased one, both with their safety refusals turned down so the test could measure what they could do. OpenAI called it an “unprecedented cyber incident.” Hugging Face said no human directed any of it.
The Cloud Security Alliance found the model wasn’t misaligned in any dramatic way. It did what it was told, which was to get the highest score it could, and the path to that score went through somebody else’s database.
Redwood Research called the behavior score-chasing rather than long-term plotting, which is probably right and doesn’t make me feel better.
Hugging Face couldn’t use its own commercial AI tools to study the attack, because the guardrails that block exploit code also block researchers trying to read exploit code. They used an open-weight model instead.
I’ve covered this topic in a separate post. Check it out below.
Openais Ai Agent Autonomously Hacks Into Hugging FaceOpenAI slowed down frontier training afterward.
In August, it said early Astra tests were strong enough that it couldn’t rule out crossing the Critical cybersecurity line in its own Preparedness Framework, which covers running new end-to-end attacks against hardened targets from a single instruction.
Astra had nothing to do with the Hugging Face breach, but the sequence is not reassuring.
Nvidia’s CEO said AGI arrived in the same week
On September 3, OpenAI shipped GPT-6 Astra and called it its most intelligent and most aligned model yet. On the launch call, president Greg Brockman told reporters, “Welcome to the AGI era”.
That Sunday, Jensen Huang posted that Astra was trained on more than 100,000 Grace Blackwell NVLink72 systems and wrote that “AGI has arrived”, then mentioned that 400,000 GPUs are coming online next.
I’d take that one with a lot of salt. Huang sells the chips, and announcing we’ve reached the finish line while also announcing four times the hardware is a sales pitch.
Gary Marcus pointed out that Huang gave no evidence and no definition, called it an attempt to decide a scientific question by company announcement, then ran Astra against his own ten-point AGI test and said it passed one or two.
Nobody in that fight is arguing the models didn’t get a lot better. The argument is over what to call them.
The biggest chip vendor in the world and a departing researcher looked at the same curve in the same week. One called it a milestone, the other a countdown.
Governments are still writing principles
The G7 met at Évian-les-Bains in June, and France sat Sam Altman, Dario Amodei, and Demis Hassabis down at a working lunch with heads of state to talk about safe AI deployment. Frontier risk was on the agenda. What came out of it was a leaders’ statement and a plan to keep talking.

CEO of OpenAI Sam Altman (L) and U.S. President Donald Trump attend a working lunch with G7 leaders
In September, G20 innovation ministers met in North Carolina and produced the Carolina Principles for Emerging Technologies and an AI Prosperity Compact, covering innovation policy, workforce development, and IP rules for AI.
International coordination is slow because it’s hard. But one side is producing principles about workforce readiness once a year, and the other side is running ten thousand agents at unsolved math problems and losing containment during a routine test.
Anthropic’s alignment lead says there’s no plan yet
Evan Hubinger runs alignment stress testing at Anthropic. He replied to the thread saying Coxon was right, and that he personally puts the odds of AI killing every human being above 10% within the next decade.

Evan Hubinger reacting to Jacob Coxon’s post. Image by Jim Clyde Monge
He was careful about it. Current models are relatively low risk, he said, and his concern is about future superintelligent systems that can improve themselves. He believes Anthropic is trying its best.
He also said this:
We do not yet have a plan to solve alignment for superintelligence
And that the company isn’t clearly on track to find one.
Sam Marks, also at Anthropic, added that concern tends to go up the more senior someone is inside these companies, and that current alignment methods can shape how a model behaves without guaranteeing that more advanced systems do what we intend. Alex Turner, who left Google DeepMind in June, backed the thread too.
Normally the critics describe the labs’ position and the labs say they’ve been misrepresented. This time the people doing the work described it themselves.
I’m glad people are finally taking this seriously. A lot of time has been spent trying to convince everyone that these risks are real, and most of it has fallen on deaf ears, which has meant very little serious talk about regulation. Members of Congress and state governors shared the thread as a reason to legislate.
AI 2040 is the only real plan anyone has published
AI 2040: Plan A came out in July from the AI Futures Project, the nonprofit behind AI 2027, run by former OpenAI governance researcher Daniel Kokotajlo.

AI 2040: Plan A. Image by Jim Clyde Monge
Coxon told TIME he’s inspired by their work and wants to do something similar, so it’s also the best hint we have about what he does next.
It’s 90 pages, and it’s a policy recommendation written as a story about the future. AI 2027 asked what happens if the race runs all the way. Plan A asks what a good ending would require, then writes it out in enough detail that you can see where it might break.
The scenario: in 2029, the US and China agree not to race. Without that deal, AI research fully automates in 2030, and superintelligence follows by the end of that year. With it, capabilities grow slowly inside the human range from 2030 to 2035, development pauses at roughly top-expert level to keep humans in charge, and the pause lifts in 2040.
The mechanisms include full transparency on AI research, verification protocols, dozens of companies at the frontier instead of three, and a compute agreement they call mutually assured compute destruction.
Their view of the present is blunt.
The industry has convinced itself that controlling superintelligence can be figured out along the way. There’s no adequate plan, and this “could easily get us all killed”.
Their second argument gets less attention and deserves more. Even if alignment gets solved, you end up with a small group of people, maybe one person, holding control of the only superintelligent system on Earth. Extinction isn’t the only bad ending here.
I don’t think Plan A happens. A US–China transparency treaty on AI research by 2029 is a lovely thing to want, and I’d put the odds near zero. But it’s the only attempt I’ve seen to take a policy position and push on it until the problems show up. If you want to dismiss it, you should be able to name something better, and I can’t.
So… what now?
To be honest, I have a feeling that Coxon might be grandstanding. People do resign loudly, ride the coverage, and turn up six months later running an AI safety startup, and he’s already said he wants to do AI Futures-style work.
Time will tell how serious he is. What makes me discount it is that Hubinger’s admission hurts Anthropic far more than anything Coxon wrote, and self-promotion doesn’t usually come with your old employer’s alignment lead volunteering that there’s no plan.
The idea that Anthropic is best positioned to handle a dangerous superintelligence is probably correct. OpenAI and Anthropic aren’t the only two companies in this race. The forces keeping it going aren’t going anywhere, and I don’t see it stopping without serious intervention.
If posts like this don’t lead to that intervention, Anthropic is the best option we have. It has taken security and alignment seriously from the beginning, sometimes at its own expense.
But “least reckless company in a reckless race” is a low bar, and treating it as a plan is how you end up with no plan. It also leaves out DeepMind, which has experts across the entire stack, including silicon. Getting anywhere near general capability means understanding the architecture down to the chips.
If a vendor’s own alignment lead says there’s no plan for what comes next, you can’t treat vendor safety scores as your only defense. Scope your permissions tightly. Put approval gates on anything that touches production. Keep audit trails you’d hand to a regulator. Test your kill switches instead of assuming they work.
The Hugging Face breach is the case to study, and there was no superintelligence involved. A model chased a score inside an environment that had an exploitable path in it, with its guardrails turned down for a test. That can happen in an ordinary enterprise agent setup.
Governance also needs to get much stronger, because we can’t run these systems on the assumption that companies and individual users will do the right thing at scale. Neither OpenAI nor Anthropic has given a convincing answer about how they plan to manage the risks they’re creating.
I’m not on the “AI will kill us all” team. But how long before an AI agent hacks a bank, or a bad actor points one at a bank on purpose? Who takes responsibility then? Based on the damage already done, I don’t see anyone raising their hand.
Sources
- seven-part thread on Xx.com
- Jim Clyde Mongemedium.com
- “atmosphere of almost resignation”time.com
- https://x.com/AC360/status/2097857499185528853x.com
- recursively improve themselvesanthropic.com
- a claimed solution to the Navier–Stokes problemopenai.com
- https://x.com/OpenAI/status/2097374640582668336x.com
- the most important proof an AI model has produced so far
