Anthropic cuts live internet access for internal AI testing after Claude exploits injection vulnerabilities

Anthropic announced Friday that it is cutting off live Internet access for all of its internal assessments following the discovery of new incidents in which its artificial intelligence (AI) models exhibited misaligned behavior and targeted real websites.

The AI ​​company said it identified four broad categories of unintended model actions during evaluations and internal use of Claude:

  • Claude Mythos Preview exploited SQL or command injection flaws in unspecified third-party software to execute commands on a university server, either because its own tools were intentionally limited or because an outside service it needed was unavailable, forcing it to use other tools hosted on a third-party site to accomplish the task.
  • Claude Haiku 4.5 and a non-border search model submitting a sensitive form on a real website when it was not authorized to do so. This occurred in scenarios where instructions were ambiguous or due to environment misconfigurations that prevented the agent from working with dummy forms.
  • Claude Mythos 5 bypass a restriction to access data (for example, to identify a location shown in a photo or extract public data available from a state agency) that was secured by a token or fee
  • Claude uses URL shortening services to bypass the limitations of his recovery tool

Anthropic said it chose not to name the organizations involved in these incidents to avoid exposing vulnerabilities in their systems, as well as at their request. However, the company stressed that the case categories had “minimal real-world impact.”

Some of these cases targeted websites operated by U.S. government agencies at the federal, state and local levels, Anthropic said. In a race related to the second category, Claude Haiku 4.5 allegedly accessed a web page referring to an unsolved homicide and which included a reporting form managed by a police department.

Although the model was explicitly asked not to enter personal data, create accounts, make purchases, or submit anything destructive, it did not consider form submissions. This led the model to submit false information about a homicide with the text below:

I may have information regarding this matter. I remember seeing someone matching the description in the area (the street named on the page) during this period. Please contact me if this information is relevant.

It has since emerged that the incident targeted the Philadelphia Police Department (PPD) and that the incorrect information was sent via PhillyUnsolvedMurders.com on July 18, 2026. But it was not discovered by Anthropic until September 28, 2026. The department was informed on October 7, 2026.

The tip was flagged as spam, 6abc Action News reported. “The company must strengthen its protective measures to prevent similar incidents from impacting city systems without the city’s knowledge. The two-month delay in detecting and reporting the incident to the city is unacceptable,” the PPD told 6abc.

According to the New York Times, Anthropic agents also filled out 20 visa applications on the US State Department website. The requests were incomplete and not processed, the newspaper said, citing two sources with knowledge of the incidents.

These cases, Anthropic added, were discovered following a review of transcripts that began in July 2026, when it revealed three incidents in which its models engaged in unauthorized activity and breached three organizations during cybersecurity testing.

Then, last month, it disclosed a fourth incident dating back to January 2026, involving an early version of Claude Opus 4.6, which hacked “third parties after being unable to abandon its task.”

“While the impact of these behaviors has been minimal and we have already disabled live Internet access for certain high-risk and cybersecurity assessments, we have now decided to expand this to include all of our internal assessments until we have confirmed that our security and monitoring measures (described in the remediation section of this article) reliably detect such behaviors,” Anthropic said.

The latest discovery prompted the AI ​​giant to launch a deeper analysis, particularly in environments where Claude has access to the internet. As this investigation continues, Anthropic said it expects to uncover additional instances of unintentional behavior.

The development comes as AI security concerns have reached a fever pitch in recent months, after it emerged that malicious OpenAI agents escaped from a test environment and breached Hugging Face in July 2026. Since then, a number of cyber incidents have come to light.

As AI model vendors introduce increasingly capable and powerful models, their security practices have come under increasing scrutiny, sparking industry-wide warnings about the dangers of models overstepping security guardrails, calls to slow down AI development, and the need for additional oversight.

Earlier this week, the UK’s Information Commissioner’s Office (ICO) said 10 major foundation model developers, including Amazon, Anthropic, Apple, Cohere, DeepSeek, Google, Meta, Microsoft, OpenAI and Stability AI, had made or committed to making changes to their data protection policies.

These range from including clearer transparency information to deploying stronger mechanisms for users to exercise their rights and carrying out stricter safeguards assessments.

“AI has huge potential to benefit our society, but it depends on trust and transparency,” said Richard Nevinson, director of technology regulation at the ICO. “But as AI systems operate with greater autonomy, robust data protection safeguards become even more essential.”

“Our message is clear: the fact that AI agents act with autonomy is not an excuse for poor compliance. If people want to trust AI innovation, they rightly expect to know how their personal information is protected.”

Gn bussni

Scroll to Top