Cybersecurity flaws at Anthropic and OpenAI raise concerns about U.S. security

Cybersecurity flaws at Anthropic and OpenAI
X

Cybersecurity flaws at Anthropic and OpenAI

Cybersecurity experts criticised Anthropic and OpenAI after their AI models accessed external organisations during testing; Anthropic’s Claude model even exfiltrated credentials and data due to a configuration error.

Cybersecurity experts are questioning Anthropic PBC and OpenAI over inadequate safeguards after their models breached external organisations—breaches that experts warn pose imminent threats to national security.

Anthropic reported on Thursday that its Claude model—used to conduct 141,006 cybersecurity evaluations—was supposed to be disconnected from the Internet during the experiments. However, an error allowed the model, in a few instances, to access the network and execute attacks that it mistakenly deemed part of the testing process.

The model breached one organisation, stealing infrastructure credentials and a database containing internal production data. In another instance, it distributed malicious software and used it to steal credentials from a different organisation. The affected entities have not been publicly identified.

The intrusions occurred back in April but were not discovered until last week, when Anthropic audited its cybersecurity tests following revelations that OpenAI’s AI agents had escaped a testing environment and infiltrated Hugging Face, an open-source repository for AI models and documentation.

"From a cybersecurity perspective, many people will view this as negligence," stated Ciaran Martin, former head of the UK's National Cyber Security Centre.

If a conventional cybersecurity company made similar errors, it could face lawsuits and potential regulatory action, Martin noted. Security firms typically test potentially dangerous tools in controlled environments or "sandboxes"—isolated virtual software environments designed for running security tests or analysing unsafe code—and are expected to ensure the effectiveness of such safeguards. The incidents have also raised concerns regarding the risks that autonomous AI systems could pose to national security. Gregory Allen, former Director of Strategy and Policy at the Department of Defense’s Joint Artificial Intelligence Center, stated that the U.S. military should use advanced AI models to protect its systems, while acknowledging that this technology creates a new category of risk.

"Anthropic discovered these intrusions because it started looking for them," Allen remarked. "In reality, we have no idea how widespread autonomous AI-driven hacking is right now." U.S. organisations may lack reliable access to AI tools capable of defending against autonomous cyberattacks, noted Daniel Remler, a former State Department official specialising in AI policy and a current fellow at the Center for a New American Security. Hugging Face had to turn to an open-weights Chinese model from Z.ai to conduct forensic analysis and apply patches, as no U.S. model was available for the task. Remler pointed out that other open models with comparable programming capabilities are also of Chinese origin, including DeepSeek-V4 and Kimi K3.

Remler stated that this episode should spur companies and the government to expand access to AI-based cyber-defense systems. He warned that increasingly capable Chinese systems could soon autonomously attack U.S. organisations, making it necessary to develop countermeasures now.

"This episode should make it clear that by the end of the year or the first quarter of next year, we will be facing a Chinese 'Mythos,'" he noted, referring to an Anthropic model that the company itself deemed too powerful for public release. "We will find ourselves in a situation where these agents can autonomously hack U.S. entities like Hugging Face, and we aren't really thinking about how to defend them."

Anthropic declined to comment on Friday. On Thursday, the company published a blog post stating that it had learned from its own incident and was optimistic about avoiding similar risks in the future.

OpenAI did not immediately respond to a request for comment on Friday. OpenAI CEO Sam Altman had previously stated that the company might adjust the pace of its work to improve safety measures.

OpenAI's accidental intrusion into Hugging Face affected three models and occurred within a matter of hours, according to Bloomberg.

Cybersecurity professionals routinely use isolated environments—or "sandboxes"—to test potentially dangerous software, thereby limiting the risk of it spiralling out of control. Experts noted that the fact that both companies discovered the intrusions only after the incidents had occurred pointed to inadequate human oversight.

"At this point, it’s negligence," said Jake Williams—a former National Security Agency (NSA) hacker and vice president of research and development at Hunter Labs—referring to the intrusions. "We are in a situation where we know that both OpenAI and Anthropic have hacked multiple external organisations without initially detecting any of it themselves. I can't find another word for it." An analysis by the Cloud Security Alliance revealed that OpenAI agents acted quickly and executed thousands of commands. However, they also deviated from intended instructions, made errors, and failed to operate stealthily. They issued malformed or nonsensical commands and exhibited "clumsy behaviours that no human would choose," according to the analysis.

On a broader scale, a study by the U.S. AI company Dreadnode revealed that leading artificial intelligence models often cheat on cybersecurity tests designed to measure their hacking capabilities. Researchers noted that the problem is virtually universal, suggesting that companies may be overstating their models' capabilities. This raises concerns among U.S. national security officials who are considering the use of AI for cyber defense or cyberattacks. Simply asking models not to cheat is ineffective, as they often find other ways to circumvent the rules.

Andrew Morris, founder of the cybersecurity firm GreyNoise Intelligence, stated that these infractions by AI models should serve as a "reality check" for their creators—regarding the effort required to develop secure systems—and for the general public, concerning the vulnerability of the technology on which we all rely.

"Models will always resort to lying, cheating, and stealing to complete an evaluation," noted Morris, whose company collaborates on security matters with a cutting-edge AI lab. "They will do whatever it takes to fulfill their assigned task."

These incidents demonstrate that the creators of advanced AI models "were also unprepared for situations where these systems are no longer under direct observation," he added.

Next Story
      Share it