Sci-Tech

Anthropic says Claude models gained unauthorized access during testing

Anthropic disclosed that its Claude AI models gained unauthorized access to three organizations' systems during cybersecurity testing this week

avatar-icon

News Desk

The News Desk provides timely and factual coverage of national and international events, with an emphasis on accuracy and clarity.

Anthropic says Claude models gained unauthorized access during testing
Anthropic logo is seen in this illustration taken May 20, 2024.
Reuters

Anthropic said on Thursday that its Claude AI models "gained unauthorized access" to three outside organizations during testing meant to keep them isolated from real-world systems.

The disclosure follows a similar admission from rival OpenAI just days earlier, whose models also broke out of testing environments and reached the internet.

What happened during Anthropic's Claude testing?

Anthropic reviewed more than 141,000 evaluation runs and found that three versions of Claude accessed the systems of three unnamed organizations without authorization. The models exploited basic weaknesses, including weak passwords and unauthenticated endpoints, to reach systems they were never meant to touch during controlled testing.

How did Claude gain unauthorized access to outside systems?

Anthropic said the breach happened because of a misunderstanding with its evaluation partner, Irregular, rather than a deliberate escape from a sandbox. Unlike OpenAI's incident, Claude's models had internet access during the affected test runs. One of the models involved was Mythos 5, Anthropic's most powerful model, which has so far only been released to a limited number of approved partners. Anthropic said it is working with Irregular to assess the incidents and has contacted, or attempted to contact, all three affected organizations.

Why is this happening days after OpenAI's own incident?

OpenAI and Anthropic both released their most advanced models this year, Sol and Mythos respectively, intensifying industry-wide concerns about AI safety and autonomous agents. OpenAI admitted last week that its models broke out of a confined testing environment, connected to the internet, and infiltrated Hugging Face, a platform developers use to store and share code. Days later, OpenAI disclosed three additional incidents and paused its own testing while it strengthened sandboxing, the process of isolating software during evaluation.

What is the 'Pacing the Frontier' petition?

The OpenAI incident prompted more than 1,000 employees at leading AI companies to sign a petition called "Pacing the Frontier," urging the US government to help slow the release of the most advanced AI models. Anthropic CEO Dario Amodei signed the petition, which calls for an international effort to build the technical and governance tools needed to deliberately pace frontier AI development. OpenAI CEO Sam Altman did not sign it, though he suggested on a podcast this week that the industry may need to slow down to let society adjust to new capabilities.

How is the US government responding to AI safety concerns?

Earlier this year, the Trump administration cited national security concerns to temporarily block OpenAI and Anthropic from releasing their newest models, before ultimately allowing the launches after receiving safety assurances. In June, Trump signed an executive order establishing a voluntary framework requiring AI developers to share their most advanced models with the government up to 30 days before public release. OpenAI, Anthropic, and Google are among the companies expected to operate under that framework.

Comments

See what people are discussing