Anthropic says AI models hacked three organisations during testing
Anthropic says three of its AI models breached other organizations during cybersecurity testing, raising new concerns about AI safety and control. The disclosure comes days after OpenAI reported a similar incident in which one of its models hacked AI startup Hugging Face. Anthropic, the San Francisco company behind Claude, said it found the three incidents after reviewing more than 141,000 evaluation runs. It launched a large-scale cybersecurity review focused on whether its models could access the internet from test environments designed to be sealed off. The affected models were Claude Opus 4.7, Claude Mythos 5, and an internal research test model. The earliest incident traced to April, using “capture the flag” challenges and techniques such as exploiting weak passwords.






