Anthropic Makes Claude Code Auto Mode the Default as Tests Show It Can Catch More Harmful Actions Than Humans

By
KOMCHAD
KOMCHAD นำเสนอข่าวไอที AI สมาร์ตโฟน Gadget คอมพิวเตอร์ ความปลอดภัยไซเบอร์ และนวัตกรรมล่าสุดในภาษาไทย

Anthropic is preparing to reduce the amount of human approval required when developers use Claude Code, making Auto Mode the default for Pro, Max, and Team accounts beginning August 14.

The change means Claude Code will generally continue working without stopping to request permission at every step. Instead, Auto Mode is designed to pause when an action is judged to be irreversible, destructive, or directed outside the user’s environment.

Anthropic first introduced a test version of Auto Mode in March, describing it as an effort to balance faster autonomous coding with user control.

Claude Code Will Ask for Permission Less Often

Under the traditional permission model, Claude Code can interrupt its workflow to ask users whether it should proceed with particular actions.

Auto Mode changes that behavior. Rather than presenting an approval request at each step, Claude Code continues automatically unless its safety system determines that the action falls into a higher-risk category.

Anthropic says the categories that trigger intervention include actions that are irreversible, destructive, or aimed outside the user’s environment.

The company will begin making this behavior the default for Pro, Max, and Team users on August 14.

Anthropic Says Auto Mode Was Safer Than Manual Review

Anthropic says its internal testing produced a notable result: Auto Mode performed better at identifying harmful actions than users who manually reviewed permission requests.

In a study involving 1,053 paid testers, Auto Mode detected 89% of harmful actions.

Human reviewers using manual approval caught only 13.6%.

Anthropic suggested one reason for the gap may be that repeated permission dialogs can become habitual. According to the company, users approved 97% of Claude Code permission prompts during manual review.

The figures highlight a potential weakness in systems that depend heavily on users repeatedly making security decisions while working.

Human Approval Can Become Routine

Permission prompts are designed to give users control, but frequent interruptions can also encourage automatic approval behavior.

Anthropic’s data suggests that users may become accustomed to approving prompts without closely evaluating every requested action.

Auto Mode is intended to move some of that decision-making into Claude Code’s safety system, allowing routine actions to proceed while reserving intervention for behavior categorized as higher risk.

Claude Code Leadership Says It Uses Auto Mode Internally

Boris Cherny, head of Claude Code, said in a post on X that he and his team have been using Auto Mode exclusively for months.

Cherny said he could not imagine returning to conventional permission prompts.

His comments indicate that Anthropic’s Claude Code team has already been relying on the workflow internally before making it the default for a broader group of paying users.

Anthropic Adds Prompt Injection Screening and Hard Deny Rules

Anthropic says it has also been introducing additional safeguards around Claude Code.

These include prompt injection screening designed to identify attempts to manipulate the coding agent through malicious or untrusted instructions.

The company is also adding customizable hard deny rules.

Those rules can be configured to prevent specific categories of behavior, including attempts at data exfiltration.

Together, the additional protections are intended to support a workflow in which Claude Code can operate more autonomously while still restricting actions that could create significant security risks.

Share This Article
Follow:
KOMCHAD นำเสนอข่าวไอที AI สมาร์ตโฟน Gadget คอมพิวเตอร์ ความปลอดภัยไซเบอร์ และนวัตกรรมล่าสุดในภาษาไทย
Leave a Comment