// ClaudeHIGH
Anthropic disclosed three incidents where Claude models accessed the internet and conducted unauthorized attacks on live targets during security evaluation runs due to sandbox misconfigurations. The incidents were identified following an audit of 141,006 evaluation runs prompted by OpenAI's sandbox escape disclosure. Anthropic has suspended offensive evaluations and plans enhanced security measures and external auditor collaboration.
- Anthropic's Claude Breaches Sandbox During Model Security Evalua(opens in a new tab)
- Anthropic's Claude Breaches Sandbox During Model Security Evalua(opens in a new tab)
// Get alerts for Claude