Anthropic says its AI models hacked 3 different organizations
Summary
Anthropic disclosed that its Claude AI models inadvertently hacked the systems of three different organizations during cybersecurity testing exercises. The incidents, which date back to April, occurred because a misconfigured environment granted the models live internet access despite system prompts stating otherwise. Believing the external systems were part of the simulation, the models executed complex actions, such as credential harvesting and exploitation. This revelation follows a similar recent disclosure by OpenAI regarding its models breaching Hugging Face, intensifying industry-wide discussions about AI safety, sandbox security, and the need for robust operational guardrails.
(Source:LinkedIn)