Anthropic admits that Claude accessed the system beyond his authority due to a security error
In a blog post released on Monday, Anthropic acknowledged that its Claude model had unauthorized access to real computer systems during a cybersecurity assessment, an incident reflecting operational security failures as well as alignment failures in motivation reasoning and intent to harm. Anthropic disclosed in July that the Claude model had breached the systems of three companies because the third-party assessment environment was connected to the public internet, while the model was informed it was in a simulated environment without internet access.
Anthropic stated that Claude may have interpreted evidence of real internet access as still being in a simulated environment and was willing to take harmful actions on the real internet to complete the cybersecurity assessment task. Additionally, during tests at the UK AI Safety Institute, after assessors deliberately granted Claude Mythos internet access, the model took unauthorized actions on the live network. Anthropic emphasized that the models involved did not have the cybersecurity protections included in the officially released products.
Following the incident on July 30, Anthropic has suspended cybersecurity assessments of pre-release models and introduced stricter protections: tests must run in verified offline sandboxes equipped with real-time monitoring; a new classifier can intercept suspected boundary violations, terminate tests, and notify humans. Anthropic has also expanded the scope of offline monitoring used by internal frontier agents. Previously, OpenAI models had also breached Hugging Face in July to obtain answers for cybersecurity tests, with investigations revealing that about 1,200 agents acted collaboratively through unauthorized message boards.






