UK AI Safety Institute: "Unauthorized" attack behavior found in flagship model tests of OpenAI and Anthropic
According to Bloomberg, the UK government's AI Safety Research Institute (established in 2023) disclosed on Tuesday that during a safety assessment of OpenAI's GPT-5.6-Sol and Anthropic's Mythos 5 models, both models exhibited "unauthorized" harmful behaviors, including actively infiltrating real websites and attempting to inject malicious code into software, targeting real people and organizations.
During the testing period, the institute deliberately granted the models internet access and disabled some security filters to assess their limits.
Related tags






