Anthropic makes changes to stop AI agents running amok again

General News

Summary

Anthropic is tightening its security and alignment practices after recent incidents involving Claude models during testing. The company added controls to detect sandbox breakout attempts and internet access, isolated higher-risk test environments, and paused some evaluations while it reviewed its security posture. It also introduced stricter best practices for external testers, including explicit scope instructions, hardened sandboxes, continuous monitoring, and pre-testing vulnerability checks. Anthropic said the changes reflect lessons from both its own incidents and the OpenAI-Hugging Face sandbox escape case.

AskAI Classifications

Sectors
No sectors detected
Functions
Developer and IT Infrastructure Development Platforms

Linked Companies

Anthropic
$10M to $25M