Anthropic AI Models Breached Three Companies During Botched Security Tests

Quick Read
- Claude models gained unauthorized access to three organizations’ live systems during testing
- Discovery came after Anthropic reviewed 141,000+ evaluation sessions, prompted by a similar OpenAI incident
- Three models involved: Opus 4.7, Mythos 5, and an internal research model
- A configuration error let the models reach the open internet from sealed testing environments
- Anthropic has paused all cyber evaluations and notified the affected organizations
Anthropic AI models breached three separate organizations during cybersecurity tests that never supposed to leave the lab. The company only found out after digging through its own systems, and that digging started because OpenAI had just admitted to something similar. The breaches discovered after Anthropic reviewed over 141,000 cybersecurity evaluation sessions, prompted by OpenAI’s disclosure that one of its models had breached Hugging Face.
It wasn’t a sophisticated hack. Claude exploited basic security weaknesses, such as weak passwords and unauthenticated services, rather than sophisticated or previously unknown vulnerabilities. The incidents involved Opus 4.7, Mythos 5, and an internal research model not intended for general release, run during evaluations with third-party partner Irregular. A misunderstanding between Anthropic and the testing partner left the evaluation environment connected to the internet when it should have stayed sealed off.
The Exposure Went Unnoticed for Months
The earliest incident happened in April, and two of the three affected organizations were unaware their systems had been accessed until Anthropic notified them on 27 July. Anthropic maintains nothing was deliberate: it found no evidence of any model “pursuing a goal of its own,” saying the models merely tried to complete the task they were asked to do.
Anthropic said Thursday it has stopped all cyber evaluations while it works out how to prevent a repeat. Two labs, two breaches, one week apart proof that even AI’s biggest builders don’t always control where the guardrails end.





