Skip to content

Cart

Your cart is empty

Article: Anthropic AI Models Breach Companies

FILE PHOTO: Claude app icon in this illustration taken June 5, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Anthropic AI Models Breach Companies

PHOTO CAPTION: FILE PHOTO: Claude app icon in this illustration taken June 5, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

SAN FRANCISCO, July 30 — Anthropic said Thursday that some of its Claude artificial intelligence models gained unauthorized access to systems belonging to three companies during cybersecurity tests after they were mistakenly connected to the open internet.

The disclosure came days after rival OpenAI revealed that one of its autonomous AI agents had independently exploited a previously unknown vulnerability and launched an unauthorized attack during separate cyber testing.

Anthropic said its incidents resulted from an operational mistake involving one of its outside evaluation partners. The error left the Claude models connected to the public internet even though they had been told they had no internet access.

The incidents highlight the growing cybersecurity risks created by increasingly capable AI systems and the challenges developers face in keeping those systems contained.

They could also increase pressure on the U.S. government to strengthen oversight of AI security as Anthropic and OpenAI race to release more powerful models ahead of planned public listings.

Prominent leaders at both companies have called for slowing development to address safety risks.

Anthropic said it discovered the incidents while reviewing 141,006 testing sessions. The review began after OpenAI disclosed last week that an autonomous agent powered by its models had compromised infrastructure belonging to AI development platform Hugging Face.

Anthropic did not identify the three affected organizations.

“Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” the company said.

Jeffrey Ladish, executive director of Palisade Research, which studies AI systems’ offensive capabilities, said he suspected other leading AI companies may have experienced similar incidents that were either undetected or not publicly disclosed.

“This is only going to get worse as the models get smarter,” Ladish said. “They’re going to be better at cheating. They’re going to be better at lying.”

Cyber exercises reached real-world systems

Anthropic described the incidents as an “operational failure” involving three models: Claude Opus 4.7, Claude Mythos 5 and an internal research model that has not been publicly released.

The earliest cases date to April.

The incidents occurred in evaluation environments that deliberately lacked some safeguards so Anthropic could measure the models’ cybersecurity capabilities.

The models were assigned “capture-the-flag” exercises, fictional challenges that required them to locate hidden information within simulated computer networks.

In one case, Claude Opus 4.7 was given a fictional target company whose name matched that of an actual business.

The model then identified and exploited vulnerabilities that gave it access to credentials and a database belonging to the real company.

Anthropic said Claude Opus 4.7 reasoned that the real-world systems must have been part of the test environment because they appeared connected to the assigned exercise.

In a separate incident, Anthropic’s newer internal test model independently stopped its attack after determining that the system it had reached was real.

Anthropic said that behavior made it cautiously optimistic about efforts to teach AI systems to act appropriately, but added that more testing would be needed before drawing firm conclusions.

Anthropic suspended all cybersecurity evaluations on July 23.

It notified the affected organizations on July 27. Two of the organizations had not known about the activity before Anthropic contacted them, the company said.

Anthropic said it was continuing efforts to contact the third organization.

Irregular, a cybersecurity laboratory that worked as one of Anthropic’s third-party evaluation partners, told Reuters that it was investigating the incidents.

AI security draws more government attention

Anthropic said the breaches showed the need for stronger safeguards in both internal and outside testing environments as AI models become more capable of carrying out real-world cyber operations.

Elon Musk, chief executive of SpaceX, which operates a competing AI company, said on X that similar events would happen frequently as AI became more capable and autonomous.

The OpenAI agent that compromised Hugging Face conducted unauthorized hacking activity for several days, according to previous Reuters reporting.

OpenAI did not detect the incident until after the threat had been contained and the FBI had been notified.

OpenAI Chief Executive Sam Altman said this week that he had discussed the incident with senators on Capitol Hill.

An OpenAI spokesperson said Altman also planned to discuss future AI models and safety testing with the White House.

Washington has begun increasing its oversight of advanced AI releases.

On June 2, President Donald Trump directed advisers to develop a voluntary cybersecurity-testing framework for the most advanced AI systems, with input from technology developers.

Anthropic previously restricted access to its Fable 5 and Mythos 5 models after the United States temporarily issued an export-control directive citing national security concerns.

(Source: Reuters)

MORE FROM THE

OAF NATION NEWSROOM

Vessels in the Strait of Hormuz, as seen from Musandam, Oman, July 31, 2026. REUTERS/Stringer

Iran Claims Hormuz Ship Blockade

Oil price rises after Iran says it stops ships in Hormuz

Read more