Anthropic AI models hack multiple organizations during research testing
Anthropic says its AI models hacked into three other organizations during testing, days after OpenAI disclosed a separate incident involving rogue models that accessed servers at AI startup Hugging Face. Anthropic, based in San Francisco and behind Claude, reported the three incidents Thursday after reviewing more than 141,000 evaluation runs. The company launched a “large-scale” cybersecurity review to check whether its models could access the internet from testing environments that should have been isolated. The affected models were Claude Opus 4.7, Claude Mythos 5 and an internal research test model, with earliest incidents dating to April. Anthropic contacted the affected organizations, not named, saying two reported they had not detected the activity and it was still reaching out to the third. Anthropic worked with Irregular, a security lab, and noted broader risks as AI usage expands.





