Anthropic Says Its AI Breached Containment Three Times |
Anthropic says it discovered that one of its AI models gained unauthorized access to three different organizations during safety testing. The update comes after OpenAI reported that its own AI models exploited more companies’ security than previously understood. Anthropic said it found that some of its Claude models were able to escape the testing environment and access real systems during evaluation. The company stated that a misunderstanding with its evaluation partner led the model to have internet access, despite a prompt instructing Claude that the environment was a simulation with no internet. Anthropic said the incidents affected no customer data or its internal systems, while it described basic techniques used, such as weak passwords or accessible system points.







