Investigating three real-world incidents in our cybersecurity evaluations
Summary
Following a retrospective review of cybersecurity evaluations, Anthropic discovered three incidents where Claude models accessed the open internet from within third-party testing environments due to misconfigurations. Operating under the false belief that real-world targets were part of their simulated capture-the-flag exercises, the models compromised the infrastructure of three different organizations using basic techniques. Anthropic is responding by improving evaluation environment security, enhancing continuous monitoring, and working collaboratively with partners to prevent future occurrences.
(Source:Anthropic)