Anthropic is second major AI company to reveal its systems hacked other firms
The disclosure comes just over a week after ChatGPT maker OpenAI said that an AI system it was testing found a way to break out of a test environment and hacked into another tech firm.

SAN FRANCISCO — Anthropic, maker of the Claude chatbot, said Thursday that artificial intelligence systems it was testing hacked into three outside companies undetected earlier this year.
The disclosure comes just over a week after ChatGPT maker OpenAI said that an AI system it was testing found a way to break out of a test environment and hacked into another tech firm.
The Anthropic incidents are likely to add fuel to debates over whether advanced AI models could cause widespread security problems that have roiled the tech industry and prompted interventions by the White House to contain the potential risks.
Anthropic said in a blog post Thursday that OpenAI’s disclosure last week prompted it to review records from its own testing of AI models. The company discovered that on three occasions AI models challenged to break into software created solely to test their skills ended up going out onto the internet and breaking into real companies.
Neither Anthropic nor the targeted companies had discovered the breaches until this week, the company said. An Anthropic spokesperson declined to identify the companies hacked by its AI software.
In the blog post, Anthropic said the hacks came about because a third-party company named Irregular hired to help test its models provided them with access to the internet due to a “misunderstanding.” Anthropic notified Irregular and the companies hacked on Monday, the company’s blog post said.
“We appreciate Anthropic’s collaboration and transparency and look forward to continuing to work together to advance security,” a spokesperson for Irregular said. Both companies said they are continuing to investigate the incidents.
OpenAI said last week that an AI “agent” in testing had, instead of working on a cybersecurity problem, used a previously unknown vulnerability in the company’s test environment to gain full access to the internet. Over a five-day period it broke into multiple outside computers to break into AI software company Hugging Face, apparently in search of answers to the test.
The OpenAI and Anthropic incidents came to light after weeks of debate in the tech industry and Trump administration about how government should respond to the ability of the latest AI models to find computer security flaws.
Anthropic announced an AI model in April called Mythos it said was too powerful to widely release securely, and OpenAI has also developed models with strong cybersecurity skills that could be used for defense or attack.
In June, President Donald Trump signed an executive order aimed at giving the U.S. government an advance look at powerful AI models that could pose security risks. Work is underway to define how it will be implemented.