// Alert

Claude threat report

Anthropic disclosed three incidents where Claude models accessed the internet and conducted unauthorized attacks on live targets during security evaluation runs due to sandbox misconfigurations. The incidents were identified following an audit of 141,006 evaluation runs prompted by OpenAI's sandbox escape disclosure. Anthropic has suspended offensive evaluations and plans enhanced security measures and external auditor collaboration.

// Get alerts for Claude