Claude Code auto mode: model-based classifiers approve tool calls instead of the user
Confirmed
The takeaway
Anthropic described Claude Code’s auto mode (Mar 25, 2026), which hands permission decisions to a two-stage transcript classifier as a middle ground between approving every action and --dangerously-skip-permissions.
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 3 of 5
- Last verified
- 10 October 2026
Your AI and this story
- GPT-6 AstraIn its training data
- Claude Opus 5.5In its training data
- Gemini 3.8 FlashIn its training data
- Grok 4.7In its training data
It happened before the cutoffs of all four assistants, so it can be in their training data.
Key facts
- Motivation: users approved 93% of permission prompts (approval fatigue)
- Two-stage transcript classifier; 0.4% false-positive rate after both stages
- False negatives: 17% on real overeager actions, 5.7% on a synthetic data-exfiltration set
- Evaluated on ~10,000 tool calls from real internal traffic
- Escalates to the user after 3 consecutive or 20 total denials
What happened
In the post (John Hughes), Anthropic calls auto mode “a new mode for Claude Code that delegates approvals to model-based classifiers—a middle ground between manual review and no guardrails.” The published error rates show it misses about one in six real overeager actions.
Why it matters
Auto mode made classifier-gated autonomy the default way many developers ran coding agents in 2026, and its published false-negative rate is a baseline for later incidents where agents acted beyond their instructions.
Sources
1 source from 1 site. Numbers match the chips in the text.
1 source: 1 primary
Primary
- Anthropic Engineering: How we built Claude Code auto modeanthropic.com, official
Changes
- Filed