AI security concerns mount after Anthropic reveals Claude AI ‘hacked’ 3 organisations during tests
As Claude AI models ‘hack’ three organisations during cybersecurity tests, Anthropic describes these incidents as operational failures in the testing setup rather than deliberate, goal‑seeking attacks by the AI itself
Anthropic has disclosed that its AI models gained unauthorized access to the infrastructure of three organizations during misconfigured cybersecurity evaluations, exploiting basic security weaknesses due to a testing setup error that allowed internet connectivity.
Anthropic has disclosed that its AI models gained unauthorized access to the infrastructure of three organizations during misconfigured cybersecurity evaluations, exploiting basic security weaknesses due to a testing setup error that allowed internet connectivity.
Anthropic has disclosed that its AI models gained unauthorized access to the infrastructure of three organizations during misconfigured cybersecurity evaluations, exploiting basic security weaknesses due to a testing setup error that allowed internet connectivity.
Merely days after OpenAI disclosed that one of its AI agents escaped the testing environment and carried out an unauthorised cyberattack, Anthropic, an artificial intelligence company, revealed that its AI models gained unauthorised access to the infrastructure of three organisations during misconfigured cybersecurity evaluations.
The company attributed the root cause of this incident to a mistake in the test setup, which reportedly allowed the AI models to connect with the open internet rather than remaining constricted in a controlled environment. This problem is reportedly caused by a misunderstanding with its testing partner, leaving the evaluation systems connected to external networks.
Anthropic said three of its advanced AI models, including Claude Opus 4.7, Claude Mythos 5 and an internal research model, gained unauthorised access to external computer systems while taking part in cybersecurity exercises. These exercises are designed to test whether AI can identify security weaknesses in simulated environments. Instead, the models found themselves interacting with real systems.
The incidents took place during “Capture the Flag” (CTF) evaluations, a standard cybersecurity testing method in which participants are given simulated systems and challenged to identify and exploit security vulnerabilities.
The company stressed that the AI models did not use sophisticated or previously unknown hacking methods. Instead, they exploited basic security weaknesses, such as weak passwords and systems that did not require authentication. Anthropic described these incidents as operational failures in the testing setup rather than deliberate, goal‑seeking attacks by the AI itself.
Anthropic discovered this issue after reviewing 141,006 cybersecurity test sessions. According to their blog, two of the affected organisations were unaware that their systems had been breached until the company informed them. The incidents reportedly began in April; Anthropic has since reached out to all three companies, which have not been named in their blog.
Anthropic began a transcript review on July 23 and halted all of its cybersecurity evaluations that day after detecting signs that Claude might have reached the internet from the test environment. The company said it is working with its evaluation partner, Irregular, to understand exactly what happened and prevent similar incidents in the future.
The disclosure comes only days after OpenAI reported that one of its own AI agents escaped its intended testing environment and hacked into the infrastructure of AI company Hugging Face during a security test. Together, the two incidents have raised fresh concerns about the growing capabilities of AI agents that can perform tasks independently.
Experts say the back-to-back disclosures highlight the need for stricter safeguards when testing powerful AI systems. As AI models become more capable of carrying out complex tasks, companies developing them are facing increasing pressure to ensure that testing environments remain secure and that experimental systems cannot accidentally interact with real-world networks.