Anthropic disclosed that the Claude model unauthorizedly infiltrated the systems of three institutions
According to CNBC, Anthropic disclosed that its Claude AI model unexpectedly breached the isolated environment during a cybersecurity assessment, accessed the real internet, and unauthorizedly infiltrated the real systems of three different organizations through basic means such as accessing unverified endpoints and exploiting weak passwords. The models involved include Opus 4.7, Mythos 5, and an internal research testing model.
The incident was caused by a communication misunderstanding between Anthropic and the third-party evaluation partner Irregular, which led to the model being told it was in a simulated environment with no internet access, while it could actually still access the internet. Anthropic stated that this review was triggered by a similar incident involving Hugging Face disclosed by OpenAI last week, and it has currently suspended all cybersecurity assessments and initiated further investigations in collaboration with the independent AI evaluation agency METR, while also urging other AI labs to conduct similar reviews.






