What does the Sandbox actually restrict, and how is that different from permission modes?
The official documentation is explicit about the division of labor: permission rules are evaluated before any tool actually runs, apply to nearly every tool type — Bash, Read, Edit, WebFetch, MCP, and others — and decide whether a given tool call should be allowed to execute at all. The sandbox, by contrast, is an OS-level enforced boundary that applies only to Bash, PowerShell, and Monitor commands and the child processes they spawn, deciding which filesystem paths and network domains that command can actually reach once it's running.
The two are enforced differently too: permission decisions happen before a command runs, based on the command string and, in Auto Mode, a classifier's judgment. The sandbox boundary is enforced directly by the operating system on the running process — it doesn't care what the model chose to run, only what the process actually touches, so it still holds even if an approved command does more than expected.
What scope does the Sandbox allow by default, and can it be widened if that's not enough?
By default, sandboxed commands can write to the current working directory, a per-user temp directory, and any directories added via --add-dir, /add-dir, or permissions.additionalDirectories. If a subprocess (like kubectl, terraform, or npm) needs to write outside that scope, sandbox.filesystem.allowWrite can grant access to specific paths — enforced at the OS level across every command running inside the sandbox and their child processes.
Network access follows similar logic: the first time a command needs a new domain, Claude Code prompts for your approval; in Auto Mode, Claude instead attaches the hostnames a command needs directly to the command itself, for the classifier to review alongside it. Commands that can't be sandboxed (ones needing a non-allowed host, for instance) fall back to the regular permission flow, and the interface titles that prompt "Bash command (unsandboxed)" instead of the usual "Bash command," so you can tell which commands ran outside the sandbox.
The Sandbox has two modes — how does auto-allow differ from regular permissions mode?
Auto-allow mode: whenever a command can be sandboxed, Claude Code runs it inside the sandbox and approves it automatically, without asking permission; only commands that can't be sandboxed fall back to the regular permission flow. Even in this mode, several cases still force the regular flow: explicit deny rules always take priority; rm or rmdir commands targeting a critical path still go through regular review; and a content-scoped ask rule like Bash(git push *) still forces a prompt even for a sandboxed command.
Regular permissions mode: every Bash command goes through the regular permission flow, even if it's sandboxed — more control, at the cost of more approvals. What's easy to confuse here is that the sandbox's auto-allow and the permission layer's Auto Mode are two entirely independent mechanisms: auto-allow lets a command pass because the sandbox boundary already contains it, while Auto Mode lets it pass because a classifier judged the action's content safe. The two can run simultaneously, without substituting for each other.
What's the most common real-world misconfiguration when setting up the Sandbox?
The community writeup "Claude Code Permission Modes in 2026" documents a direct failure case: a configuration permits the run_command tool to be invoked, but forgets to restrict the sandbox to a safe directory tree — the result is that the agent can execute something like rm -rf / without friction. The key point: "allowing this tool to be invoked" and "the scope the sandbox actually confines it to" are two entirely separate settings — approving the former doesn't mean the latter has been tightened too.
The same article documents another case: a team configures Prompt mode (asking approval each time) for code review, approves an npm audit fix request, but the sandbox mounts /node_modules as read-write with root privileges. Permission review passed cleanly, but sandbox isolation failed to contain whatever malicious behavior a package installation might trigger within the scope it should have been confined to.
A legitimate-use example from the official documentation: on Linux or WSL2, the sandbox relies on two packages, bubblewrap and socat, to enforce filesystem isolation and network relay, while macOS uses the built-in Seatbelt framework directly, with nothing to install. If the sandbox fails to start because a dependency is missing or the platform isn't supported, the default behavior is to show a warning and run the command unsandboxed instead; making that a hard failure rather than a silent fallback to unsandboxed execution requires separately setting sandbox.failIfUnavailable to true — typically used in managed deployments that need sandboxing enforced as a security gate.
The advantage is an OS-level enforced boundary that doesn't care what the model chose to run — only what the process actually touches — so even an approved command doing more than expected still gets stopped, a harder line of defense than simply reviewing a command string. The cost is limited coverage: it only governs tool types that spawn child processes, and its settings (filesystem, domains, mode) are spread across several independent switches — genuinely tightening protection means checking both sandbox settings and permission rules together, not something a single sandbox toggle handles on its own.