Bible Network Crypto DeFi Onchain RWA AI Agent Stablecoin Chain SAFU CryptoTax DeFAI AGI Claude Me Claude Skill Claude Design Claude Cowork
Independent Media
Not affiliated with any project
Learn Claude Skills. Do Everything Better.
claudeskill-me.com
LATEST
Your Hook Says "Blocking Error" But the File Still Changed? PostToolUse and PreToolUse Don't Actually Block the Same Way  ·  Enabled Auto Mode and Thought You Were Safe? Permission Modes and Sandbox Boundaries Are Two Different Layers — Conflating Them Is How Things Break  ·  Messages API Adds On-Demand Compaction in Beta: Developers Decide When to Compact, Not the System  ·  Claude Cowork and Chat Officially Merge, Launching Claude Docs and Claude Slides Alongside It  ·  Why a Single "hi" Can Burn 20,000+ Tokens in Claude Code: Breaking Down the Fixed Startup Overhead  ·  How to Write Your First Custom Slash Command in Claude Code: A Working Example From Scratch
advanced

Enabled Auto Mode and Thought You Were Safe? Permission Modes and Sandbox Boundaries Are Two Different Layers — Conflating Them Is How Things Break

30-Second Version · For the impatient
The sandbox doesn't care what the model chose to run — only what the process actually touches. That's the fundamental difference from permission modes.

Full Explanation +
01 · Why did this happen?

If I already have Auto Mode enabled, do I still need to configure the sandbox separately?

Yes, and the two cannot substitute for each other. Auto Mode controls whether a given tool call should execute, based on a classifier's judgment of the action's content; the sandbox controls what filesystem paths and network domains a command can reach once it's executing, enforced at the OS level. The two layers cover different risks: an action Auto Mode fails to catch can still have its actual damage contained if the sandbox boundary is tight enough. Conversely, no matter how strict the sandbox boundary is, it doesn't govern tool types that don't spawn child processes — Read, Edit, WebFetch, MCP calls — which remain entirely dependent on permission rules for gating. Enabling only one layer leaves the other layer's entire attack surface exposed.

02 · What is the mechanism?

The sandbox's auto-allow mode and the permission layer's Auto Mode sound so similar — are they just two names for the same thing?

No, and the official documentation specifically clarifies this. The sandbox's auto-allow mode lets Bash commands run without asking each time because the command is already contained within the sandbox boundary — even if it tries something harmful, what it can actually reach is limited. The permission layer's Auto Mode lets tool calls run without asking each time because a classifier reviews the action's content and judges whether it's safe.

The two can run simultaneously and independently, but what replaces "asking" is completely different in each case — one is a hard boundary that blocks out-of-scope access, the other is a soft content judgment. Assuming they're the same layer just because both names contain "auto" is an easy way to misjudge your actual protection scope.

03 · How does it affect me?

If I use --dangerously-skip-permissions, does that mean the sandbox gets skipped too?

No, and this is exactly where the flag is easy to misread. --dangerously-skip-permissions governs the permission layer — it removes the approval requirement for every tool call, even skipping protected-path checks, and what replaces the prompt is nothing at all. But the sandbox is an entirely separate mechanism from the permission layer, and whether it's enabled depends on your own sandbox configuration (like sandbox.enabled), independent of whether you've added this flag.

This also points to a risk that's often overlooked: if you skip the permission layer and haven't separately enabled or tightened your sandbox settings, both defenses fail at once — no tool call gets filtered, and there's no boundary at all on what filesystem or network access is actually reachable. The flag's name literally says "dangerously" for a reason — confirm the sandbox layer is independently holding its own before using it.

04 · What should I do?

I'm not a security expert — I just want to check whether my automation pipeline has fallen into this trap. What's the simplest way to check?

You don't need to memorize every setting key upfront — just do one thing first: find the list of tools you currently let Claude execute automatically, and for any item on that list that invokes Bash, PowerShell, or a similar command-line tool, confirm whether the sandbox has restricted it to a clear, safe directory scope, rather than defaulting to open access across the whole system. If the answer is "I haven't specifically adjusted the sandbox settings," that likely means you're currently relying solely on the permission layer for protection, and your actual sandbox scope may be far wider than you assume.

Once you've checked the settings, set up a safe test environment and deliberately simulate a scenario where "the tool call is approved, but attempts to touch a path or domain it shouldn't" — and actually observe whether the sandbox stops it. That test will tell you where your real defense line sits far more reliably than reading through any configuration document.

Full Content +

"I turned on Auto Mode, so nothing too bad should happen." Behind that sentence sits a common misunderstanding: treating "whether Claude is allowed to act" and "what Claude can reach once it acts" as the same thing to manage. These are actually two entirely independent layers of defense — one decides whether a given tool call needs your approval first, the other decides what files and network domains that command can actually touch once it's approved to run. Conflating these two layers is one of the most common, and most damaging, mistakes teams make when setting up Claude Code automation.

Permission modes decide "whether to ask"; the sandbox decides "what happens once it runs"

The official documentation defines this division of labor clearly: permission rules are evaluated before any tool actually runs, apply to nearly every tool type — Bash, Read, Edit, WebFetch, MCP, and others — and decide whether a given tool call should be allowed to execute at all. The sandbox, by contrast, is an OS-level enforced boundary that applies only to Bash, PowerShell, and Monitor commands and the child processes they spawn, and decides which filesystem paths and network domains that command can actually reach once it's running.

The two are also enforced differently. Permission decisions happen before a command runs, based on the command string itself, and, in auto mode, a separate classifier's judgment about whether that command is safe. The sandbox boundary, on the other hand, is enforced directly by the operating system on the running process — meaning it holds regardless of what the model chose to run, and even if an approved command does more than its name suggests. It doesn't care what the model "decided" to execute; it only cares what the process actually touches.

A specific combination that goes wrong: allowing the tool call without tightening the sandbox scope

The community writeup "Claude Code Permission Modes in 2026" (by jsmanifest) gives a direct failure case: a configuration permits the run_command tool to be invoked, but forgets to restrict the Bash sandbox to a safe directory tree — the result is that the agent can execute something like rm -rf / without friction. The key point here is that "allowing this tool to be invoked" and "what corner of the system that tool can actually reach" are two entirely separate settings — approving the former doesn't mean the latter has been tightened too.

The same article documents another production case: a team configures Prompt mode (asking approval each time) for a code review workflow, approves an npm audit fix request, but the sandbox mounts /node_modules as read-write with root privileges. Permission review passed cleanly at that layer, but sandbox isolation failed to contain whatever malicious behavior a package installation might trigger within the scope it should have been confined to. This case makes a point worth internalizing: however carefully you've configured the permission layer, if the sandbox boundary isn't aligned with it, the defense gets bypassed just the same.

How the official documentation defines the division of labor

Claude Code's official documentation lays it out directly in a comparison table: /sandbox controls what a Bash command can access once it runs, and what replaces the per-action prompt is the sandbox boundary itself (in auto-allow mode); Auto mode (a permission mode) controls whether each tool call runs at all, and what replaces the per-action prompt is a classifier that reviews the action's content; and --dangerously-skip-permissions also controls whether each tool call runs, but what replaces the prompt is nothing at all — even protected-path checks get skipped, leaving only the small set of actions no mode ever auto-approves still restricted.

The documentation also stresses that the sandbox's auto-allow mode and the permission layer's Auto mode are two independent mechanisms that don't substitute for each other. Auto-allow lets Bash commands pass without a prompt because the sandbox boundary already contains them; Auto mode lets tool calls pass without a prompt because a classifier judged the action's content as safe. The two can run simultaneously and independently — but that also means if one layer is loosely configured, the other layer won't automatically compensate for it.

What specifically the two layers each correspond to in practice

If you're actually adjusting settings, the official documentation's mapping is quite concrete. On the filesystem side, sandbox.filesystem.allowWrite grants subprocess write access to paths outside the working directory, while sandbox.filesystem.denyWrite and denyRead Block access to specific paths — these are configured separately from the permission layer's Edit allow rules and Read/Edit deny rules, but ultimately get merged into the sandbox's final configuration. On the network side, the sandbox's allowedDomains/deniedDomains control which domains Bash commands can reach, which is again a separate set of settings from the permission layer's WebFetch(domain:...) allow/deny rules. In other words, the same underlying question — "can this path be written to" — is governed by both permission rules and sandbox settings simultaneously, and you need to check both sides to know where the actual boundary sits.

So is the sandbox simply more secure than permission modes?

Not quite the right way to frame it. The sandbox boundary's enforcement genuinely is harder — it's an OS-level restriction that doesn't care what the model chose to run, only what the process actually touches, meaning even an approved command that does more than expected still gets stopped at the boundary. But the sandbox only covers Bash, PowerShell, and Monitor — the tool types that spawn child processes. Read, Edit, WebFetch, MCP calls, and other tool types remain entirely dependent on the permission-rule layer for gating. The two layers cover different attack surfaces, and neither is dispensable — there's no valid shortcut version of "as long as the sandbox is strict enough, permission rules don't matter."

How to actually audit your own setup

Rather than trying to memorize every setting key by name, it's more useful to ask two concrete questions: if Claude is approved to call a given tool, is the filesystem scope that tool can actually reach narrow enough; and even with a tightly configured sandbox boundary, is there some tool category (MCP calls, for instance) that bypasses the sandbox entirely and relies solely on permission rules — and if so, are those permission rules themselves tight enough. Before putting any automation pipeline into production, it's worth running an actual test against both questions — deliberately letting an approved action attempt to touch something it shouldn't, and personally confirming whether it was genuinely blocked, rather than assuming a setting works simply because it was written down.

Sources: Configure the sandboxed Bash tool - Claude Code Docs, Permission modes - Claude Code Docs, Claude Code Permission Modes in 2026: What allowedTools, Whitelists, and Sandbox Boundaries Actually Restrict, Configure permissions - Claude Code Docs
Diagram
權限模式與沙盒邊界的分工對照兩層各管不同範圍,任何一層設寬都不會被另一層自動補起來Permission Mode vs Sandbox — Two Independent LayersPermission ModesControls: whether a tool call runsApplies to: Bash, Read, Edit,WebFetch, MCP — nearly all toolsAuto mode: classifier reviews action--dangerously-skip-permissions: nothingSandbox BoundaryControls: what a running commandcan access (files, domains)Applies to: Bash, PowerShell,Monitor + child processes onlyOS-level enforcement, not model choiceKnown failure combinationrun_command allowed + sandbox not scoped to safe dir= agent can run rm -rf / without frictionNeither layer compensates for the other being looseClaude Skill Me · claudeskill-me.com
Feel free to share. Please credit the source.
Ask a Question
Please enter at least 10 characters
Related Articles
Subagents Aren't Smarter Mini-Claudes — They Solve Isolation, Not Capability
advanced · Aug 31
Your Hook Says "Blocking Error" But the File Still Changed? PostToolUse and PreToolUse Don't Actually Block the Same Way
practice · Sep 28
Why a Single "hi" Can Burn 20,000+ Tokens in Claude Code: Breaking Down the Fixed Startup Overhead
advanced · Sep 05
Why Your Claude Code Skill Silently Stops Triggering: The 15,000-Character Description Budget No One Warns You About
practice · Sep 02
Related News
More Related Topics