VYPR
advisoryPublished Aug 10, 2026· 2 sources

Anthropic Defaults Claude Code to AI-Powered Security Review

Anthropic is making AI-driven auto mode the default for reviewing Claude Code actions on its Pro, Max, and Team plans, citing significant improvements in detecting dangerous commands over human review.

Anthropic is implementing a significant shift in its AI security posture by making an AI-powered auto mode the default for reviewing Claude Code actions across its Pro, Max, and Team subscription plans, effective August 14th. This move aims to bolster security by leveraging artificial intelligence to identify and block potentially harmful commands more effectively than human reviewers.

In a controlled experiment involving 1,053 paid professional testers, Anthropic found that human reviewers only managed to catch 13.6% of dangerous commands. In stark contrast, the AI-driven auto mode demonstrated a remarkable success rate, identifying and blocking 89% of such commands. This substantial difference underscores the perceived advantage of AI in this critical security function.

While auto mode is becoming the default for many users, it remains optional on the Claude Enterprise platform, the Claude API, Anthropic Platform, Amazon Bedrock, Google Cloud’s Agent Platform, and Microsoft Foundry. This phased rollout allows administrators on these enterprise-level services time to evaluate the feature before full integration. Anthropic plans to extend the default auto mode across these platforms within the next month and will absorb the costs associated with the additional compute power required for its safety classifier.

The auto mode functions by routing every tool call through a specialized classifier designed to identify and block irreversible, destructive, or unauthorized actions targeting resources outside the user's defined environment. When the classifier flags an action, Claude either seeks a safer alternative or prompts the user for explicit approval. The system is designed to revert to manual approvals if it encounters persistent blocking issues, either after three consecutive blocks or 20 blocks within a single session.

Despite the enhanced security, Anthropic cautions that auto mode does not entirely eliminate risk, as it still relies on AI to make judgments about command safety. The company strongly advises users to continue reviewing Claude's actions, particularly when dealing with high-risk operations such as modifications to production systems or critical infrastructure.

Anthropic's internal testing and external evaluations, including red-teaming exercises and prompt-injection assessments, have consistently shown auto mode to be as safe or safer than manual review. This includes a benchmark study where auto mode outperformed manual approvals in preventing harmful actions that users had not explicitly requested. Furthermore, an independent evaluation by Trajectory Labs found that Claude models running in auto mode successfully defended against 720 simulated prompt injection attacks, a feat not matched by other leading AI models in the same benchmark.

The introduction of auto mode also addresses concerns raised by developers regarding the security and privacy of AI coding assistants. By automating the review of potentially dangerous commands, Anthropic aims to reduce user overhead, increase productivity, and provide a more robust safety net against both accidental missteps and malicious exploitation, such as prompt injection attacks.

Anthropic's internal use of auto mode has already yielded positive results, preventing incidents like off-network data leaks and destructive mass operations. The company has further fortified auto mode with features such as hard denies, granular data access controls, pre-action Git status checks, and enhanced prompt injection screening, reinforcing its commitment to balancing AI's capabilities with stringent security measures.

Anthropic is making auto mode the default for Claude Code on August 14th for its Pro, Max, and Team plans, with plans to extend this to other services within the month. The company has also ceased charging for the extra tokens consumed by the classifier in auto mode for these plans, with similar changes planned for other platforms. This move follows extensive testing where auto mode reportedly blocked 89% of dangerous commands, significantly outperforming human testers.

Synthesized by Vypr AI