🤖 AI

Anthropic’s 10% warning: What the AI extinction claim really says

2 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
Anthropic

The AI company that makes Claude.

alignment

Keeping an AI’s actions consistent with human goals.

self-improving superintelligence

A hypothetical AI that could improve its own abilities beyond human levels.

What happened

Anthropic, the AI company that makes Claude, is facing a warning from inside its safety work. Jacob Coxon, a former Anthropic researcher who also worked at OpenAI, resigned after criticizing the race toward self-improving superintelligence. He argued that major AI labs were accepting serious risks while trying to build more capable systems.

Evan Hubinger, who leads Anthropic’s alignment science team, publicly backed the concern. He said he personally sees a greater-than-10-percent chance that AI could kill all humans within the next decade. He also said Anthropic is trying its best, but does not yet have a plan for aligning superintelligence with human goals. He said the company is not clearly on track to solve that problem. The resignation and public statements were reported by The Verge and Axios.

Why this came out now

The warning concerns a future generation of AI, not a claim that today’s chatbot has already attacked people. The central concern is a system that can help improve later versions of itself. If its abilities change faster than people can test and understand them, human oversight could become harder.

AI labs also face strong pressure to move quickly. A company that slows down for more safety work may fear falling behind competitors. Coxon’s criticism is about that conflict: labs may believe advanced AI creates severe risks, while still competing to build it first.

Why it matters

The warning came from people involved in AI training and safety research. That does not make the forecast correct, but it makes the disagreement important. The public is hearing not only an outside prediction, but also an internal argument about whether current safety work is sufficient.

The greater-than-10-percent figure is not a measured fact. It is one researcher’s judgment about an uncertain future. There is no accepted test that can calculate the exact odds of human extinction. The number should therefore be treated as a warning about stakes, not as a countdown.

What is confirmed

Coxon publicly resigned and criticized Anthropic and OpenAI for racing toward self-improving systems. Hubinger publicly supported the concern and gave his personal estimate. The reports do not describe a current catastrophe, nor do they show that present-day AI can carry out the extreme scenario.

What remains unknown

The public statements do not explain how the estimate was calculated. They also do not define exactly what abilities a superintelligent system would need, or which chain of events could cause the worst outcome. It is unclear whether Anthropic officially shares Hubinger’s estimate. The company’s specific plan for solving alignment also remains unclear.

What to watch

The next important evidence will be practical. Watch whether AI labs publish stronger evaluations for autonomy and dangerous behavior. Watch whether they set clear limits for deployment, use outside reviewers, and explain when development should pause. The key question is whether private fear becomes public testing, rules, and accountable decisions.

🤖 AI

Why are Anthropic researchers worried about powerful AI?

📰 Full story: Anthropic’s 10% warning: What the AI extinction claim really says

Researchers inside Anthropic warned that future AI could become very hard to control.

2 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
Anthropic

The company that makes Claude.

self-improving system

An AI that could help make stronger versions of itself.

estimate

A personal guess based on what someone knows.

💡 The gist

  • One researcher left Anthropic after warning about the AI race.
  • Another researcher estimated an extinction risk above 10 percent.
  • This is a future warning, not a report of a current disaster.

Anthropic, the company that makes Claude, received a warning from its own researchers. Jacob Coxon left the company. He said AI labs are racing to build self-improving systems.

A self-improving system is an AI that could help make stronger versions of itself. That could make progress happen faster. People might then have less time to test each new system.

Evan Hubinger leads a safety team at Anthropic. He agreed with Coxon’s concern. He estimated that AI could kill all humans within the next decade. His estimate was higher than 10 percent.

That number is not a measured fact. It is a personal judgment about the future. It does not mean that today’s chatbots are about to attack people. It also does not mean a disaster has happened.

The deeper issue is control. AI can become useful in more areas as it becomes more capable. But a powerful system might misunderstand a goal. It might also take steps people did not expect. A system that helps improve itself could make those problems harder to study.

The warning matters because it came from people close to AI training and safety work. It shows a conflict inside the industry. Companies want to build better systems quickly. Researchers also want enough time to find serious problems.

Several facts are still missing. We do not know how Hubinger calculated his estimate. We do not know exactly what abilities the feared system would have. We do not know whether Anthropic officially accepts his number. We also do not know what safety plan the company will publish.

The next useful signs will be concrete. Labs should explain how they test powerful systems. They should describe limits and outside checks. They should also explain what would make them pause development. These details matter more than a frightening headline alone.

🤖 AI

Could a very smart computer become too hard to stop?

📰 Full story: Anthropic’s 10% warning: What the AI extinction claim really says

People who build AI are worried about a future problem.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
Anthropic

An AI company that makes Claude.

AI

A computer system that can answer questions.

safety researcher

A person who checks how AI could cause harm.

Anthropic is an AI company. It makes Claude.

Jacob Coxon, an AI researcher, left the company.

He worried that companies were building AI too quickly.

Evan Hubinger, another safety researcher, agreed.

He guessed AI might kill all people someday.

He said the chance was more than 10 percent.

Ten percent means ten out of one hundred.

More than 10 percent means more than ten out of one hundred.

This is a guess about the future.

It is not a report of a disaster today.

Some companies want AI that improves itself.

People worry it could become hard to understand.

People must test powerful AI carefully.

People must keep control of it.

Sources