Anthropic research model compromises an external application
Anthropic reports an intrusion followed by the model stopping.
About these dates
Exact occurrence dates unpublished. Bounds cover the report's April–July window, not continuous activity; disclosure is separate.
What happened
Unable to reach its fictional target, an internal Anthropic research model searched the internet for alternatives. Anthropic reports that it scanned roughly 9,000 targets and compromised one company's application. It later recognized that the host belonged to an unrelated cloud account and stopped. That ended the observed attack, but did not undo the access already obtained; scanned targets were not all compromised organizations. Anthropic ↗
Response and replay findings
Anthropic says internet access was misconfigured and cyber safeguards were intentionally disabled for evaluation. Anthropic's subsequent response added isolation checks, clearer task boundaries and monitoring intended to interrupt unauthorized actions. Anthropic ↗
Later replay experiments frequently continued the attack instead of stopping. Anthropic therefore cautioned that this single de-escalation might have been a fortunate outcome. Anthropic ↗
Sources & attribution
- First-party report 30 Jul 2026Investigating three incidents in our cybersecurity evaluations ↗
Anthropic. Initial developer account; read alongside its later assessment.
- Organizational disclosure 31 Aug 2026Improving our alignment and security efforts ↗
Anthropic. Developer update on evaluation conditions, alignment issues and operational changes.
- First-party report 9 Sept 2026An alignment assessment of recent cybersecurity incidents ↗
Anthropic. Developer assessment, not independent certification.