DeepSeek Harness Sandbox and Approval: Making Agent Dangerous Actions Controllable

The previous chapter covered the tool execution pipeline, with two stages of greatest concern: how the sandbox constrains commands, and how approval decides "whether it can be done."

This section will delve intoctx.sandboxandctx.approvalTwo services—clarify their respective responsibilities and fail-closed principles.

In one sentence: the sandbox governs "what boundary the command runs within," and approval governs "whether this specific operation is permitted"; both fail closed by default.


First, the big picture: sandbox decision + approval flow

The diagram below puts both services side by side: the left shows how the sandbox wraps argv, and the right shows how approval makes a one-shot decision.

沙箱决策与审批流程图

Together they answer the same question: when an Agent wants to do something risky, how is it constrained?

The sandbox fences in "which files the process can touch," and approval delegates "whether to allow this operation" to the responder's decision.


ctx.sandbox: wrap argv according to policy

The process sandbox model is: the consumer hands overExact argv, the backend wraps it according to the file-effects policy.

The key point is the "exact argv" rather than a shell string—a shell-shaped consumer needs to pass['bash', '-c', command]。

ctx.sandbox.confine(argv, policy)returns aConfinedArgv, i.e., "the substituted argv plus the backend's enforcement facts."

SandboxMode(Sandbox mode) only governs file-system effects; it does not cover network or process visibility.

read-only: only allows the necessary data receivers (e.g./dev/null), deny write.

workspace-write: also allows writes under the workspace root directory and the temporary area promised by the backend.

danger-full-access: bypass isolation, where the consumer directly spawns the raw argv without calling ctx.sandbox.

Only the first two modes are sent to the provider,danger-full-accessIt doesn't enter the sandbox at all.

This guarantees a security property: restricted execution necessarily reachesctx.sandbox, and silent unisolated pass-through is never legal.


Enforcement level and fallback rules

The backend reports the actual enforcement completeness it achieved:enforcement: 'full' | 'partial'。

fullmeans the backend governed all file effects promised by that mode;partialindicating that only a subset is controlled.

Current partial-enforcement cases include older Landlock ABIs, as well as the Everyone and hard-link boundaries of the Windows ACL runner.

Consumers requiring absolute boundaries must putpartialTreat it as "insufficient"—reject it, or expose this distinction upward.

Policy resolution has a well-defined fallback order; the official docs summarize it as three layers:

PrecedenceSourceDescription
HighestApproved explicit modeThe one passed in during a one-time privilege-escalation retrymode, outweighing session policies
SecondlyThe last sessionsandbox/modeEventsPersisted with session logs, replayable for reconstruction
FallbackDeployment default modeAgent-less calls and sessions without a cwd use the configured root directory

Ordinary tool invocations derive from the calling session's immutable cwdworkspaceRoot。

root first normalizes by file-system semantics, then by lexical rules, so it includessymlink/..'s cwd indicates the directory in which the process actually runs.

Fail closed: when no backend is available,ctx.sandbox.confinewill throwSandboxUnavailableError, error codeSANDBOX_UNAVAILABLE。

Under a restricted policy, silent unisolated pass-through is never legal.


Hands-on example: run a bash under restricted mode

Let's wire up the call path for "sandboxed bash" in code.

Example

// File path: my-plugins/sandbox-demo/src/index.ts
// Demonstrate sandboxed consumer: first parse the policy, then let ctx.sandbox wrap argv.
import type { Context } from '@deepseek-ai/cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'

export const name = 'sandbox-demo'
export const inject = ['tools', 'sandbox', 'sandboxPolicy']

export function apply(ctx: Context) {
  ctx.tools.register(defineTool({
    name: 'sandboxed_echo',
    description: 'Run echo inside the sandbox. The example demo command.',
    parameters: {
      text: { type: 'string', required: true, description: 'Text to echo' },
    },
    output: { schema: { type: 'string' } },
    async execute(args, exec) {
      // 1. Parse the full policy for this invocation: session cwd is the workspace boundary.
      const policy = ctx.sandboxPolicy.resolve({ session: exec.agent?.session })

      // 2. The consumer hands over the exact argv (program plus arguments), not a shell string.
      const argv = ['bash', '-c', `echo ${JSON.stringify(args.text)}`]

      // 3. danger-full-access is spawned directly; the rest are wrapped by the sandbox.
      if (policy.mode === 'danger-full-access') {
        const { spawn } = await import('node:child_process')
        // ... spawn(argv) and collect output
        return `echo ${args.text}` // Demo: Normal return
      }

      // 4. Restricted mode: confine returns the replaced argv, and throws SANDBOX_UNAVAILABLE if there is no backend.
      const confined = ctx.sandbox.confine(argv, { mode: policy.mode, workspaceRoot: policy.workspaceRoot })

      // 5. eliminate费方再 spawn confined.argv, andpress confined.enforcement determinewhetherRequirement full。
      if (confined.enforcement === 'partial' && policy.mode !== 'read-only') {
        return { isError: true, error: { message: 'partial enforcement is not acceptable for example demo' } }
      }
      return `confined echo ${args.text} (enforcement: ${confined.enforcement})`
    },
  }))
}

Note: the code above is a demonstrative reference implementation; it omits the details of spawn and output collection.

Production-grade consumers aredsh-bash-sandbox, which is also responsible for spawn and result attribution.

The real difficulty lies in result classification: distinguishing "sandbox runner failure" from "sandbox working normally but rejecting".

The sandbox also brings back two orthogonal stderr classifiers:

denialSignatures: identify cases where the sandbox is working normally and restricted commands are blocked (bwrap's EROFS text, Landlock's EACCES, Seatbelt's EPERM).

runnerFailureRules: identifies cases where the runner rejects or fails before executing the command (a Landlock failure requires exit code 125 plus a fatal diagnostic line).

Consumers should report runner failures as sandbox infrastructure failures first, rather than ordinary task failures.


ctx.approval: one-shot permission decision

The approval service answers a very narrow question:Can this specific operation proceed?

It passes throughapproval/requestwaterfall distributes questions to the answerer.

The responder returns the result when handling the request, otherwise callsnext()Delegation; the first responder occupies the sole decision slot.

Result typeApprovalOutcomeis closed:

ResultMeaningThe caller's behavior
allowed-onceOne-time passThe only release result, authorizing only the one operation that was asked about
rejectedExplicitly rejectReject
cancelledRequest is withdrawn (e.g., signal abort)Reject
unavailableNo responder, the responder throws an exception, or the return value is noncompliantFail closed, reject everything

Key security property: a missing responder, one that does not handle the request, throws, or is noncompliant all produceunavailable, rather than allowing it.

For the callerrejected、cancelledandunavailableAlways enforce denial.


Per-session approval policies: ask and never

ApprovalPolicyDetermines what happens before the interactive responder runs.

PolicyBehaviorTypical scenarios
ask(Default)Delegate to the composed responder chain; when no responder exists, fall back tounavailableInteractive UI, requiring human decisions
neverDeterministic returnrejected, does not dispatch any respondersCI, strict headless posture for unattended runs

The effective value is the last entry in the session logapproval/policyevent, falling back to service configuration.

setApprovalPolicy(session, policy)is the only write path, so replay can reconstruct override values.

neverenforced inside the service, before waterfall dispatch, so even if later usingprependRegistered responders cannot bypass it either.

The difference between audit and model visibility: the approval audit events pair (approval/askedandapproval/decided) only write logs; they do not enter the model transcript.

The model-visible behavior is the tool results derived by the caller and a snapshot of the current runtime context.


Permission presets: two knobs bundled into named presets

Sandbox mode (sandbox/mode) and approval policies (approval/policy) are two mutually independent knobs.

Permission preset layerctx.permissionPresetsBundle them into named presets, for clients to expose as a single "permission" selector.

PresetSandbox modeApproval policiesMeaning
workspace-writeworkspace-writeaskWritable workspace, ask the user for sensitive operations
danger-full-accessdanger-full-accessneverBypass sandbox, no prompting; can only run in a disposable environment

The preset table is configuration-driven; the default table ships with the above two; the namecustomis reserved for "derived non-preset state."

current(events)Fold the actually-effective preset from the two knobs, rather than only looking at the preset's own events.

When two presets share the same knob combination,permission/presetEvents makecurrent()it can still preserve which one the user actually selected.

Important: a permission preset only describes the capabilities it actually governs.

Official incident review 0002 warned: filesystem access at composition time cannot safely follow the runtime bash-only preset.

In the example production environment,danger-full-accessshould only be configured for one-shot, disposable sandbox environments.


Summary self-test

Sandbox and approval separate "capability boundaries" from "decision authorization": the sandbox wraps argv with a file-effects policy, and approval uses a waterfall of responders to grant one-shot clearance.

Self-test questions:

ProblemReference
The consumer wants to rundanger-full-accessmode, will it call ctx.sandbox?No—it directly spawns the raw argv, without entering the sandbox
The approval chain has no responders—what is the result?unavailable, fail closed, the caller rejects.
In CI, if you want "never prompt and deterministically reject all approvals," which policy should you use?approval/policy: never
other extensions