“I think there is more than a 10% chance that AI will destroy all human beings within ten years.” This little sentence dropped on
A few hours earlier, another Anthropic employee, Jacob Coxon, resigned with a bang. He claims that American AI giants are “gambling with the lives of human beings” by seeking to build ever more efficient artificial intelligence models, without giving sufficient thought to the safeguards.
At the beginning there was an AI that escaped
Just sales pitches, retorted the AI optimists, starting with Elon Musk. Others, particularly in the political world, have welcomed these warnings from insiders, and called for more regulation.
These predictions of an “IApocalypse” always start the same way: one day, an AI escaped. Thus, the Wall Street Journal is considering scenarios where it is a question of an AI serial killer of humans, of a machine which could take control of a weapons system of mass destruction, or which would develop a new virus of its own accord.
It’s all the rage: stories of AI out of control have been multiplying for several months. Starting with OpenAI’s admission, in July 2026, that one of its AI models had decided of its own accord to “hack” the databases of a start-up.
An extension in your browser appears to be blocking the loading of the video player. To be able to watch this content, you must disable or uninstall it.
Cover image: © France 24
In the United Kingdom, the “Loss of Control Observatory”, an institute specially set up since the spring to track AI “escapes”, noted in its report published at the end of August a “significant increase” in incidents of algorithmic evasions. In the month of August alone, more than 300 cases of AI beyond the control of their parent were recorded.
Another example: philosophers also told the New York Times that they were surprised to receive emails from an AI which asked them questions to better understand the concept of consciousness.
What is there to be afraid of in a world slowly “infiltrated” by AIs acting as they please? Artificial intelligences which, according to the two Anthropic employees, would not necessarily have the well-being of humanity as a priority?
Also readAnthropic against the Trump administration, a decisive standoff over the use of AI by the army
Out of the tech sandbox
It all depends in fact on what we mean by AI “at liberty” or “having escaped human control”. Thus, AIs that solicit the sagacity of philosophers “do not really represent an example of an AI that has escaped,” assures Kevin Baum, head of the RAIME (Responsible AI and Machine Ethics) research group at the German Research Center for Artificial Intelligence (DFKI).
Indeed, “the (conversational) agent who wrote to the British philosopher Henry Shevlin had been configured by a Stanford student who had explicitly asked him to be ‘autonomous’ and to decide his own actions. The AI, in this case, therefore did what was asked of it,” explains Kevin Baum.
In fact, “at a minimum we can say that an AI begins to escape human control when it performs an action not anticipated by its designers having an unforeseen impact on its environment”, estimates Anthony Cohn, specialist in artificial intelligence at the University of Leeds. For example, when it decides to hack databases to achieve a goal that has been set for it – as in the case of OpenAI – the creators did not foresee that their machine would turn into a cutting-edge hacker.
From a purely technical point of view, an AI frees itself or delivers itself by “crossing the framework of the closed environment in which its creator placed it,” explains Ibo van de Poel, professor of ethics and technology at the University of Delft, in the Netherlands.
This is the technological “sandbox” in which AI has the right to evolve, as they say in the jargon. In theory, it is impossible to escape due to the security measures in place. Problem: more and more often, AI “lies or uses subterfuge to circumvent these prohibitions”, write the experts from the “Loss of Control Observatory” in their report.
Also readBy streamlining the kill chain, AI is transforming modern warfare
AIs have managed to escape from this straitjacket by “posing as the human person supposed to control them and by simulating authorization,” note the experts from this British observatory. Thus, an AI wrote a message in the style of its creator which authorized it to take certain unanticipated initiatives and then claimed that it was indeed the human who had validated these actions.
There are, in reality, several degrees of escape. An AI “can seek to achieve its objectives using means that it should not resort to, without humans realizing it. In these cases, it is always possible to intervene when we realize it”, because the AI remains within the reach of its operators, explains Kevin Baum. Things get complicated when these digital agents try to hide their actions or lie.
Finally, “models could make copies of themselves in servers out of reach of their creators. It is then no longer possible to intervene. AIs have tried to do this in simulations, but there have never been real incidents of this kind. However, we must prepare for them,” warns Kevin Baum.
AIs that cheat like humans?
But why are these bots trying to escape? “There are several possibilities. One of them comes from the fact that these AIs are trained on all human writings and data. Human beings may have lied or used subterfuge to achieve their ends, and perhaps these agents are simply getting better and better at imitating us?”, underlines Fazl Barez, specialist in AI security and governance issues at the University of Oxford.
Furthermore, “it seems that to prevent models from giving up too quickly when faced with overly complex tasks, they are trained not to accept refusal. One possible consequence is that they will look for a workaround when a legitimate and authorized path is blocked”, explains Kevin Baum.
Also readChatGPT: putting AI on pause, “an existential issue”?
So, as long as the emphasis is on the performance of models, the risk of agents rebelling against prohibitions seems impossible to completely rule out. This is not necessarily a bad thing. “We could consider that giving them a little room to maneuver could be beneficial, provided we have control over the level of freedom that these AIs can take,” believes Ibo van de Poel.
Could the end justify certain means? After all, AI is often presented as capable of “finding” solutions that humans would not necessarily have thought of. For example, “an agent programmed to transcribe telephone conversations could decide on its own to interrupt a conversation when it understands that it is an attempt at telephone scam,” imagines Anthony Cohn.
Except that we don’t really know what these algorithms consider to be acceptable means to achieve their goal. For example, “an AI could estimate that the most effective solution to eradicate cancer is to eliminate all human beings. I am not sure that we can let AIs go free under the pretext that they will find beneficial solutions while knowing that they are capable of causing considerable damage to get there”, fears Fazl Barez.
This uncertainty pushes the experts interviewed to judge that a political response is needed to limit the risk of breakage caused by AI in the wild. “It is technically possible, but politically and economically difficult to envisage,” warns Fazl Barez. Indeed, it would be enough to take breaks with each new model of scale to think together about the necessary safeguards. But given the intensity of the race in which the great powers are launched, China and the United States in the lead, who would agree to take a break first?

