Researchers associated with Anthropic, one of the world's leading artificial-intelligence companies, have gone public with fears that sufficiently advanced AI could present an existential danger to humanity within the next decade. Most strikingly, Anthropic researcher Evan Hubinger has estimated that there is at least a 10 per cent chance of AI killing everyone within ten years.
Ten per cent is not a prediction that humanity will become extinct. There is a 90 per cent difference between a probability of ten per cent and a certainty. But consider what the number means when supplied by someone working near the technological frontier. If an aircraft engineer told you there was a ten per cent probability that the aircraft you were about to board would crash, you would not reassure yourself that there was still a 90 per cent chance of landing safely. You would get off the aircraft. Yet humanity cannot simply get off the AI aircraft.
The immediate controversy was triggered by Jacob Coxon, who worked at both OpenAI and Anthropic before resigning from Anthropic. His argument is not that today's chatbots are about to become homicidal. The concern is about where the technology is heading and whether the companies developing increasingly capable systems know how to retain control of them.
It is easy to ridicule AI extinction scenarios if one imagines present-day chatbots suddenly deciding to launch nuclear missiles. Current AI systems make elementary mistakes, hallucinate facts and sometimes require repeated prompting to perform tasks that a competent human could accomplish easily; they are unreliable tools, and the rail guards are distinctively Left wing. But technological risk is not determined simply by what today's system can do. It depends upon what tomorrow's system might be able to do.
The people sounding the alarm are concerned particularly about systems capable of performing sophisticated research, operating autonomously, writing and executing computer code, penetrating computer networks and eventually contributing to the development of still more capable artificial intelligence.
Until now, improvements in artificial intelligence have depended overwhelmingly upon human researchers. Humans design systems, train them, evaluate their performance and construct the next generation. Suppose increasingly capable AI systems become good enough to contribute substantially to AI research itself. The development cycle could then accelerate. Artificial intelligence helps humans build better artificial intelligence, while the improved system becomes better at AI research and contributes to another improvement. Each generation potentially assists in producing its successor.
This is the idea behind recursive self-improvement, although how rapidly or successfully it could occur remains disputed. The nightmare scenario is that the feedback loop eventually moves faster than human institutions can understand or control.
That does not require an evil machine. This is one of the most misunderstood aspects of the AI-risk argument. A sufficiently capable artificial intelligence need not hate humanity, become conscious or develop the personality of a science-fiction villain. It merely needs objectives that cease to correspond adequately with ours.
Imagine telling an extraordinarily powerful system to achieve some complicated objective. It discovers that obtaining additional computing resources makes achieving the objective easier. Avoiding shutdown also makes achieving the objective easier. Preventing humans from modifying its instructions makes achieving the objective easier. None of these intermediate goals requires hatred. The machine does not need to become angry with us; humans simply become obstacles.
Critics have long mocked simplified versions of this argument, particularly the famous thought experiment in which a superintelligence instructed to manufacture paperclips eventually consumes the resources of civilisation in pursuit of its objective. The paperclip example is deliberately absurd, but the principle behind it is not.
Human beings routinely produce disasters without intending the final result. Bureaucracies pursue targets while forgetting the purpose behind them. Companies optimise measurable outcomes while creating unintended consequences. Governments introduce policies that generate incentives producing the opposite of what was intended. We have difficulty aligning human institutions with human objectives. Aligning something more intelligent than ourselves may prove considerably harder.
There is another route to catastrophe that requires no rebellious machine at all: humans could use increasingly capable AI against other humans. An AI system able to discover vulnerabilities in computer networks could become a formidable cyberweapon. A system capable of advanced biological research might lower barriers to designing dangerous pathogens. Autonomous military systems could compress decisions about war into timeframes too short for political leaders to intervene.
A dictator equipped with sufficiently advanced artificial intelligence might construct a surveillance state beyond anything imagined by the totalitarians of the twentieth century. In this scenario, AI remains perfectly obedient. The problem is whom it obeys.
That distinction between loss of control and deliberate misuse is important because solving one does not solve the other. Perfectly aligning an AI with its operator could actually make the second danger worse if the operator happens to be malicious.
There is therefore something deeply strange about the position in which humanity now finds itself. Some of the people developing the most advanced artificial intelligence believe that the technology they are creating could conceivably destroy civilisation. Yet development continues.
Coxon's criticism goes directly to this contradiction. The companies are trapped in a race. If one laboratory slows down, another may continue. If American companies pause, Chinese companies will not. If Anthropic refuses to develop a particular capability, OpenAI, Google or another competitor might develop it instead. Each participant can therefore regard continued development as the responsible choice.
"We have to build it because otherwise somebody less responsible will build it" is a powerful argument at the level of the individual company. At the level of civilisation, it is terrifying. The same logic governed arms races. Neither side needed to desire nuclear war. Each merely had to fear what would happen if the other side obtained an overwhelming advantage. The rational decision for each participant could consequently produce an increasingly dangerous outcome for everyone. Artificial intelligence may be developing similar characteristics.
There is, however, an important difference. Nuclear weapons are enormously difficult to build. They require specialised materials, industrial facilities and considerable scientific expertise. Their proliferation can therefore be monitored to some degree through physical supply chains.
Artificial intelligence is software running on computing infrastructure. The most advanced systems presently require enormous computing resources, but algorithms can be copied, techniques spread and computing becomes cheaper. Knowledge that once belonged to a handful of laboratories eventually becomes widely available. That makes permanent control extraordinarily difficult.
There is also a danger in taking the extinction claims too literally. Nobody knows that there is a 10 per cent probability of AI exterminating humanity. There is no actuarial table containing previous superintelligence extinctions from which such a number can be calculated. It is an expert judgement under radical uncertainty.
Another expert might say one per cent. Another might say fifty. Another might regard the entire scenario as science fiction. The number is therefore less important than the identity of the people assigning non-trivial probabilities to the possibility. These are not outsiders frightened by a technology they do not understand. They are people working on the technology.
That does not make them infallible. Experts can become trapped inside intellectual cultures just as everyone else can. AI researchers may systematically overestimate the importance of their own field. Companies also benefit commercially when the public believes they are building something so powerful that it might transform civilisation.
There is an obvious marketing paradox in telling investors that your product could be the most important invention in human history while telling regulators that it is merely another useful software tool. Existential-risk rhetoric can therefore serve several interests simultaneously. It may reflect genuine concern while also increasing the perceived importance of the people issuing the warning. Scepticism remains appropriate, but dismissal does not.
A ten per cent estimate does not need to be correct for the underlying issue to deserve serious attention. If there were even a one per cent probability that a technology could cause human extinction, the consequences would be so enormous that sensible people would want to know considerably more before proceeding at maximum speed.
This is where the AI debate becomes politically difficult. The obvious response is regulation, but who regulates? Governments generally understand frontier technologies less well than the companies developing them. Regulation can freeze existing market structures in place, protecting today's dominant corporations from tomorrow's competitors. Governments also have their own powerful incentives to obtain advanced AI for intelligence, military and administrative purposes.
Giving the state control of superintelligence does not necessarily solve the superintelligence problem. International agreements sound more reassuring, but enforcement presents the same difficulty. The United States will not simply trust China to stop. China will not trust the United States. Neither will necessarily trust private laboratories, military organisations or smaller states.
Everyone has an incentive to make sure somebody else does not get there first. The result is a technological prisoner's dilemma: the rational move for each participant may be to continue racing even if everyone would prefer a world in which the race proceeded more slowly.
There is another possibility, of course. The pessimists could be profoundly wrong. Advanced artificial intelligence might instead become one of humanity's greatest tools. Artificial scientists could discover new medicines, solve difficult mathematical problems, design new energy systems, increase productivity and eliminate much dangerous and tedious work.
The same capability that makes AI frightening makes it attractive. That is precisely why the technology will be so difficult to stop. Humanity is not pursuing artificial intelligence because a handful of mad scientists have forced it upon us. We are pursuing it because intelligence is useful, and more intelligence is potentially enormously useful. We have just seen claims that thousands of AI agents working together may have made progress on one of mathematics most formidable problems; discussed today at the blog. Similar systems could eventually be directed towards cancer, ageing, materials science, nuclear fusion and almost every other scientific challenge.
Imagine being the government that voluntarily abandons that capability while its rivals continue. That is why "just stop AI" is not much of a policy. But "keep going and hope for the best" is not much of one either. The Anthropic researchers have therefore exposed the central contradiction of the AI age. The people closest to the technology can simultaneously believe that it offers extraordinary benefits and that sufficiently advanced versions may present extraordinary dangers. Both propositions can be true.
Perhaps the most unsettling part of the story is how quickly the argument has changed. A few years ago, people worried that AI might help students cheat on university assignments. Then came fears about artists, programmers and office workers losing jobs. Now researchers inside leading AI laboratories are publicly discussing probabilities of human extinction within a decade. That escalation does not prove that the final fear is justified. It does show how rapidly the frontier is moving.
In the trailer for Artificial, Hollywood's fictional Sam Altman declares that Pandora's Box has been opened. The line was invented by a screenwriter. The researchers now issuing warnings from inside the AI industry are not fictional characters.
https://www.bbc.com/news/articles/ckgwy1k42w4o
https://www.politico.eu/article/anthropic-openai-researcher-jacob-coxon-warns-ai-could-kill-humans/
https://www.breitbart.com/tech/2026/09/09/anthropic-researchers-go-public-with-fears-ai-could-kill-all-humans-within-a-decade/