AI Can Improve AI Safety. Why Did a Researcher Leave Anthropic, the Company Behind Claude?
recursive self-improvement
A repeated cycle where AI helps improve more capable AI.
alignment
Work that tries to keep an AI’s actions within human goals.
benchmark
A standard test used to compare an AI’s performance.
What happened
On September 9, 2026, Jacob Coxon announced that he had left Anthropic, the company behind Claude. He had spent three years doing pre-training research at OpenAI and Anthropic. Coxon said leading labs are racing toward self-improving superintelligence and taking a dangerous gamble with human lives.
Evan Hubinger, Anthropic’s alignment science lead, publicly supported the concern. Hubinger said his personal estimate gives AI more than a 10 percent chance of killing all humans within the next decade. This is a judgment, not a measured probability.
The background
Self-improving AI does not mean today’s chatbot suddenly becomes alive. It means an AI helps design, train, or test a stronger AI. A repeated loop is called recursive self-improvement. The concern is that each loop could move faster than human understanding and oversight.
That concern appeared alongside an Anthropic study published August 28. Claude ran an automated alignment workflow. It searched research literature, proposed methods, trained a target model, and tested the results.
What the study found
The study covered ten categories of alignment failure. Anthropic said the system improved every target benchmark without reducing measured general abilities. The best methods also worked on hidden benchmarks and on models up to 4.7 times larger.
Claude also outscored 28 human safety researchers in one comparison. The comparison was not fully equal. The human researchers could not iterate in the same way. The result supports a useful workflow, not the claim that human researchers are obsolete.
Why it matters
The result offers hope. AI could help safety teams keep pace with faster development. But the same kind of automation might speed up other AI research too. Safety work must stay ahead of capability growth, or human checks may fall behind.
The study did not create a rogue AI. It tested narrow safety tasks under controlled conditions. Coxon’s resignation warns about a possible future. It does not show that extinction is underway.
What remains unknown
The benchmarks are limited proxies for real-world behavior. They do not cover every dangerous ability. A method can preserve tested abilities while affecting abilities nobody measured. Anthropic also said it had not shown that gains survive extensive training on other tasks.
There is no established scientific method for checking the ten-percent estimate. Public posts also do not reveal the full safety plans of Anthropic or its competitors.
What to watch next
Watch for independent tests, wider benchmarks, reliable monitoring, and public response plans for models that evade oversight. Watch whether rival labs can agree on safer pacing.
Coxon suggested coordination and, in an extreme case, a temporary halt to capability improvements. That is a proposal, not an industry rule.
The central question is simple: can human safety checks improve as quickly as AI-assisted research?
Sources: Ars Technica, TechCrunch, Anthropic’s research note, and TechCrunch’s study report.
AI Can Improve AI Safety. Why Did a Researcher Leave Anthropic, the Company Behind Claude?
📰 Full story: AI Can Improve AI Safety. Why Did a Researcher Leave Anthropic, the Company Behind Claude?
Anthropic, the company behind Claude, tested AI that helps improve AI safety. A researcher left because he fears the race is moving too fast.
self-improving superintelligence
A very powerful AI that helps improve itself or later AI systems.
AI-assisted research
Research where AI helps people find and test ideas.
rogue AI
An AI that acts outside human control.
💡 The gist
- AI helped test ways to make other AI safer.
- A researcher left Anthropic because he fears the race.
- The study showed progress, not a guaranteed safe future.
Anthropic published a study on August 28. Claude acted like an automated research helper. It read papers, suggested methods, trained another model, and tested the results.
The study used ten safety tests for unsafe behavior. The system improved every tested category. The target models kept their measured general abilities. Some methods worked on tests the system had never seen. They also worked on models up to 4.7 times larger.
This could help safety researchers. They might find problems faster. That matters because AI labs are building more capable systems quickly.
But the study had limits. Its tests were small examples of real behavior. A good score does not prove complete safety. The researchers did not measure every ability. They also did not know whether the gains would last after more training.
On September 9, Jacob Coxon said he had left Anthropic. He had done pre-training research at OpenAI and Anthropic for three years. He feared labs were racing toward self-improving superintelligence. That means AI helping create stronger AI.
Evan Hubinger, an Anthropic safety leader, agreed publicly. He gave a personal estimate above ten percent within the next decade. He meant the chance that AI could kill all humans. This is a prediction, not a proven fact.
The warning concerns future systems. It does not show today’s chatbots planning an attack. The study also did not create a rogue AI.
The key question is speed. Can human safety checks keep up with AI-assisted research? People will watch independent reviews, wider tests, monitoring plans, and agreements between labs. Coxon suggested slowing capability improvements in an extreme case. That remains a proposal.
AI Helps AI Learn. Why Did a Researcher Worry?
📰 Full story: AI Can Improve AI Safety. Why Did a Researcher Leave Anthropic, the Company Behind Claude?
Anthropic, the company behind Claude, asked AI to help with safety checks.
Anthropic
The company that makes Claude.
safety check
A simple test for dangerous behavior.
AI is a computer helper that learns from examples.
Anthropic gave Claude a helper job. The helper read papers and tried safety ideas. It checked another AI. It did ten kinds of safety tests. The scores got better. The other skills did not get worse.
But ten tests are not the whole world. Some problems might stay hidden.
Jacob Coxon worked at Anthropic. He left the company. He worried that AI could help build stronger AI. Evan Hubinger, another Anthropic safety researcher, worried too. He guessed about the next ten years. He said the chance could be above ten percent. That is a guess, not a fact.
This story is about a possible future. It does not say people have been wiped out. People need safety checks that keep up.