5.1. Tool execution harness

Six modules of harness/ are the security and command-execution substrate: tool_primitives.py (modes, capabilities, ToolResult), command_policy.py, workspace.py, sandbox.py, resource_control.py and tool_runtime.py (CommandRunner, AuditLogger). The provider-neutral lifecycle above it is split between ToolRegistry, ToolRouter, AgentRuntime and TaskController; none imports rag_chat. The substrate remains independently testable.

5.1.1. Two tool surfaces

The registry serves two vocabularies, and a turn sees exactly one of them:

The coding core’s six tools — read_file, search_files, patch, write_file, delete_file, terminal — on every request with no standard engaged, and in the implementation pass of a MIXED request (it needs an OpenAI-compatible endpoint). The core reaches the harness through SpearHost (control_plane.py, Implementation mode), which calls the same router authorization, workspace resolution, target policy, command policy and sandbox described below. External capabilities and workspace knowledge add no tool: they are reached with the spear-capability and spear-knowledge host commands, which SpearHost answers itself and never hands to a shell (The control plane in front of the coding core).

The normative runtime’s tools — the standard.* tools, the file and command tools of that runtime, search_corpus, and the web pair — in a session where a standard is engaged, filtered further by what the turn asks about: a purely normative question is offered no tool that reaches a working tree.

Neither surface sees the other’s tools, and the coding core never sees a normative one.

5.1.2. The web pair

search_internet finds pages; fetch_url reads one, or saves it with save_as. Both run in the chat process rather than the sandboxed shell, so reading works in SAFE; saving is a mutation and takes the ordinary write authorization. A long PDF comes back in slices and each names the pages range that continues it — told only that the answer was truncated, a model re-fetches the same url.

Their exposure follows web_enabled alone. It used to depend on matching the objective against \b(web|internet|online|latest|current release)\b, which withheld both tools whenever the phrasing missed — “please get the complete Code-G pdf” matched nothing, and so did a bare fetch https://…. The two schemas that gate saved measure 155 tokens of a 65536-token window; a tool the model cannot see is one it narrates instead of using.

Tool execution harness architecture

Fig. 5.1 The three layers of the substrate: what may run, where it runs, how it is confined.

5.1.3. Three layers

The substrate is organised as three layers with a one-way dependency: the decision layer never knows how confinement works, and the execution layer never knows why a command was allowed.

5.1.3.1. Decision layer — what may run at all

Type

Role

ExecutionMode

SAFE · ASK · AUTO.

CapabilityPolicy

Immutable per-mode capability set. for_mode() is the only way to ask what a mode grants.

CommandPolicy.classify()

Classifies a command into a CommandAssessment: READ_ONLY, WORKSPACE_MUTATING, SHELL_COMPLEX or DANGEROUS.

CommandPolicy.authorize() → AuthorizationResult

The outcome: the granted capabilities, or a terminal denied / cancelled result; in ASK the confirmation is asked here.

ExecutionProfile

The pure execution contract derived from granted capabilities: workspace_read, workspace_write, shell_complex, network, the vetted host_read_paths, and the not-yet-implemented gpu / ssh / container_runtime / secrets_allowed.

5.1.3.2. Boundary layer — where it runs

Workspace

A canonical set of roots that every filesystem tool path must resolve inside. resolve() rejects traversal, host absolute paths into the launch directory (unless explicitly allowed) and — importantly — symlink escapes, by resolving the whole path including the parent of a not-yet-existing file before checking containment.

CommandRunner

Owns the resource contracts (resource_limits, cgroup_limits) and the sandbox instance. This is the ownership boundary: rag_chat knows nothing about MemoryMax, TasksMax, CPUQuota, systemd-run or unit names.

AuditLogger

Append-only, metadata only. KEY=value assignments that look like secrets are redacted, and leading environment assignments are stripped before the executable summary is derived.

5.1.3.3. Execution layer — how it is confined

BubblewrapSandbox

build_argv() composes the structured bwrap argv; run() executes the closed-network path; _run_with_slirp() executes the network path.

ResourceLimits / CgroupLimits

Two orthogonal resource contracts. See Resource control for why there are two and why neither replaces the other.

SystemdScopeRunner

Knows systemd and nothing else. Given an argv and a CgroupLimits it produces a wrapped argv; it has no idea what a capability is.

ToolResult

The single result type: status, stdout, stderr, exit_code, changed_paths. status is one of ok, denied, cancelled, invalid_path, not_found, timeout, failed.

5.1.4. Life of a tool call

  1. ToolExposurePolicy supplies the role’s model-visible registry view.

  2. The model proposes a tool call and ToolRouter validates its schema.

  3. The router’s hard gates run: the request scope and the target policy (sibling targets, generated and snapshot targets, what a shell command writes, searches outside the project), then the execution modes the tool’s spec declares (execution_modes).

  4. The router invokes the registered handler and observer hooks.

  5. CommandPolicy.classify() classifies the command.

  6. The classification is intersected with the capabilities the current mode grants. Anything not granted ends here.

  7. In ASK mode a confirmation is requested for mutating or network work.

  8. CommandRunner.ensure_sandbox() preflights Bubblewrap — once, cached — and for network work also preflights the full slirp path.

  9. BubblewrapSandbox.run() checks, before spawning anything: the bwrap binary, prlimit if resource limits are active, the cgroup mechanism if cgroup limits are active, and for network work the slirp helper plus its pinned-namespace support.

  10. The command runs, wrapped as described in Sandbox.

  11. The substrate returns a ToolResult; the router normalizes it as a ToolResultEnvelope.

  12. Large safe output is kept in ResultStore while only a bounded preview is rendered into model context.

  13. Grounded mutations update WorkingState and mutating attempts are recorded by AuditLogger.

Every one of the step-9 checks returns a failed ToolResult rather than proceeding in a degraded mode. There is no code path from “mechanism unavailable” to “run it anyway”.

A refusal attaches to the file, not to the tool. Generated output (a generated/ or build/tmp/ path, or a header saying the file is generated) and snapshot or third-party copies are refused to write_file and patch, to delete_file – a deleted file could be written afresh – and to every write a terminal command makes: a redirection, tee, cp, install, truncate, ln, and any path an inline program mentions. The decision is taken on the path as written and on the file it reaches, so a link or a .. does not change it.

A refusal is deterministic, so asking again cannot change it. When a turn asks for the same refused operation five times – the same tool, target and refusal once spacing, quoting, a leading ./ and trailing slashes are set aside – it is stopped with an answer that names the refusal, and the stop is audited (repeated_refusal_stopped). SPEAR_REFUSAL_REPEATS sets the limit. One model response is bounded too: SPEAR_RESPONSE_MAX_TOKENS (16384) caps what a single response may generate, and a response cut there is reported as truncated.

The core’s terminal is bounded the same way: a call runs for 180 s unless it asks for another timeout, and one above 600 s is refused; its output is cut to 50 000 characters, keeping the head and the tail (agent/tools.py). The timeout replaces the sandbox’s own default (45 s) for that call only.

5.1.5. Once the sandbox is known to be down

The checks above decide one command at a time. One conclusion outlives the command that reached it: when a command’s output reports the sandbox missing, _registered_command (the normative runtime’s command handler) records it on the turn’s cache, and every later file write or deletion in that turn is refused explicitly instead of performed — the coding core’s write and delete ports read the same mark.

The reason is not sandbox purity but verifiability. Without the sandbox nothing the model writes can be read back, compiled or run, so an edit made from retrieved context alone is a change nobody can check — and the answer the user needs is that the sandbox is down, not a plausible patch. DENIED and not FAILED: the tool did not break, it declined.

The guard predates the router and was carried across it deliberately. Keeping “our side” of that merge would have dropped it silently, since the monolithic execute_tool it lived in no longer exists.

5.1.6. Preflight caching

Preflight runs a real, minimal sandbox rather than probing version strings — it executes a small /bin/sh that asserts $PWD, workspace writability, $HOME and $TMPDIR. Probing what the kernel actually permits is the only honest test; a version number does not tell you whether unprivileged user namespaces are enabled.

Results are cached on the sandbox instance, and terminal failures (ABSENT, INEXECUTABLE, REFUSED) are remembered so a broken environment is not re-probed on every call. The network preflight is cached per (sandbox, workspace) pair and invalidated when either changes.

5.1.7. Containment poisoning

The network path has one failure mode that must never be retried blindly: the harness asked the sandbox child to die, and could not confirm that it did.

If _wait_pidfd_exit() cannot observe the child’s termination, the sandbox sets network_containment_failed and every subsequent network request is refused for the lifetime of that sandbox object. A namespace whose death is unverified may still hold a network namespace that a helper is attached to; continuing would mean starting a second helper against unknown state.

5.1.8. Why the pidfd

Every place the harness signals or observes the sandbox child, it does so through a pidfd, never a numeric PID:

signal.pidfd_send_signal(pidfd, signal.SIGKILL)   # not os.kill(pid, ...)

A numeric PID can be reused between the moment it is read and the moment it is signalled. On a machine that spawns processes as fast as a build does, that is not a theoretical concern. _stop_namespace_child() retries once through the same pidfd and has deliberately no os.kill fallback.

The pidfd and the namespace handles introduced in Network backend have distinct roles that are worth keeping straight:

Handle

Identity it pins

pidfd

the process — liveness and safe signalling

ns/net fd, owner userns fd

the namespaces — stable targets for the network helper

Neither substitutes for the other, and the code does not conflate them.