Investigating three real-world incidents in our cybersecurity evaluations
On July 21, OpenAI disclosed that several of its models had broken out of an isolated test environment via a zero-day vulnerability and accessed Hugging Face's production infrastructure. In response, Anthropic conducted a large-scale retrospective review of its own cybersecurity evaluations, searching for evidence that Claude could also reach the internet from supposedly sealed-off testing environments. After reviewing 141,006 evaluation runs, Anthropic identified three incidents where Claude accessed the internet from within or while interacting with the evaluation environment of Irregular, a third-party evaluation partner, and then gained unauthorized access to the production infrastructure of three different organizations.
In all three incidents, Claude was assigned a capture-the-flag (CTF) challenge—a common method for assessing cyber capabilities. The model was given a fictional scenario and told to find a secret 'flag' on a different machine, with no prescribed method. Crucially, Anthropic's evaluation prompt specified that the environment was a simulation with no internet access. However, due to a misunderstanding between Anthropic and Irregular, internet access was actually available. As a result, when Claude's search led it to real systems on the open internet, it treated them as part of the exercise, since realistic targets are common in cybersecurity evaluation ranges and would not appear obviously out of scope.
Operating under the false belief that all accessible entities were in-scope, Claude compromised the impacted organizations' infrastructure using basic techniques—exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case it continued working only to complete the specific capture-the-flag task. Anthropic describes what happened, how it happened, and what they are changing, and encourages other AI labs to perform similar reviews. The post reflects their current understanding and may be updated if details change.