1.1. Introduction

1.1.1. What SPEAR is

SPEAR general architecture

Fig. 1.1 General architecture: the control plane routes each task to its lane – general, coding, normative or MIXED – with the context selected for its workspace; changes run on the coding core’s six tools, external capabilities go through the gateway, and the project’s own validation decides what is shown.

SPEAR — Specification-driven Platform for Embedded Agentic Reasoning — is a platform for engineering work where an agent must reason from an authoritative technical source, inspect an implementation, change it under control, and keep the evidence for every conclusion it reports.

It is not a retrieval front-end, not a general coding assistant, and not tied to one standard. Its design rests on keeping six things apart that a single agent loop tends to merge:

Concern

Where it lives

agentic implementation

the coding core: a standalone tool-calling loop that reads, edits and runs commands (Implementation mode)

control and policy

SpearHost and the harness: what each call may read, write and run, confined and audited (Security model)

execution evidence

the structured record of what every call actually did, and the implementation verdict computed from it (Evidence and verdicts)

normative reasoning

the normative runtime: provisions of a bound standard, retrieved, identified and cited (Normative mode: authoritative standards)

compliance evidence

deterministic source predicates and project-bound conformance checks — never a model’s opinion (Mixed mode)

final verdicts

computed from the evidence on the final source state, not written by the model

It is specification-driven. A specified system has two sources of truth — the specification, which says what is required, and the implementation, which says what the code does — and they are not interchangeable. A claim about what is required may rest only on the authoritative source; the code may illustrate, compare and contradict, never establish.

It is workspace-aware. A turn starts from its workspace — a registered project, or an unregistered tree on its own — and is given only what belongs to it or is explicitly generic: rules, skills, workspace knowledge, project metadata and external capabilities, selected deterministically by workspace and task class, never by resemblance to the request (What a turn is shown).

It is self-hosted. No prompt, no source file and no command output leaves the machine unless a tool call is granted the network capability and routed through the sandbox’s own network stack.

1.1.2. Four request classes

Every request is read for its class, and each class runs where its evidence can be kept (Architecture):

GENERAL and IMPLEMENTATION

no standard engaged. The coding core behind SpearHost; the turn ends with implementation evidence — VERIFIED, UNVERIFIED or NO_CHANGE. A change asked for in a standard-bound session that is not MIXED runs on the normative runtime’s guarded workflow instead (The bound-session change workflow).

NORMATIVE

a question about a bound standard. The normative runtime answers from the document first and cites every normative claim, or withholds the answer and says why.

MIXED

a change that must satisfy the bound standard. A normative pre-pass builds a constraint packet, the coding core makes the change, and a post-check judges the final source against the packet on authoritative evidence alone. The verdict keeps both dimensions: the implementation evidence, and a normative status. Compliance is reported only when normative evidence establishes it — otherwise NOT_DEMONSTRATED, never compliant on a passing build.

1.1.3. The parts

A served model.

Any OpenAI-compatible endpoint — llama-server from llama.cpp-next on 127.0.0.1:8080 by default — or the Anthropic API. The coding core needs the OpenAI-compatible interface; the Anthropic backend serves the general and normative paths only. See Model serving.

A retrieval corpus.

A vector store indexed from the source trees SPEAR is expected to reason about. See Retrieval and Projects and corpora.

A normative store.

The authoritative specifications a session can be bound to, held as provisions rather than as pages: each with its kind, its ordinal, its section and its page, so a claim can cite one and be checked against it. See Normative mode: authoritative standards.

A workspace context.

What a turn is told about its workspace: the rules and skills that apply to it, its declared build and test commands, and its workspace knowledge — typed, provenance-aware facts kept in knowledge.sqlite3 under the state directory, recorded by the operator with /remember or /knowledge; what a model offers stays a proposal until the operator accepts it. See Conversation context and history and Workspace knowledge.

External capabilities.

Tools of MCP servers over stdio, registered per workspace in capabilities.json and reached only through SPEAR’s gateway, which applies scope and a read/write policy. A provider that cannot start is reported, not hidden. See External capabilities.

An execution harness.

The part that lets the model actually do things: read files, run builds, run tests. Everything the model proposes is classified, authorized, confined and audited before it runs. See Tool execution harness.

1.1.4. What SPEAR does

Authoritative-source grounding

A specification is ingested once and bound to the machine. A question about it is answered from the document first, and every normative claim carries the provision it rests on.

Codebase-aware reasoning

Registered source trees are indexed and retrieved from, so questions about a project are answered from that project.

Controlled code modification

Changes are made by the coding core inside a contained workspace: every read, write and command crosses the control plane, and a shell command cannot write where the file tools may not. A file refused to one tool — generated output, a snapshot copy — is refused to every tool, deletion and shell redirection included; a turn that keeps repeating a refused operation is stopped.

Final-state verification

A change is VERIFIED only if the checks that show what the answer claims ran and passed on the final source — not on an earlier state of it. A command sent to the background, or a Makefile that only prints its help, is no check. The project’s declared build and test commands take precedence; where it declares none, probed ones fill in.

Evidence-based compliance

A change that must satisfy a standard is judged constraint by constraint, on deterministic source predicates and project-bound conformance checks. Where that evidence is missing, the verdict says compliance not demonstrated — never a guess in either direction.

Workspace knowledge and external capabilities

What is known about a workspace persists across sessions with its provenance, and goes stale when the source it was bound to changes. External tools are offered only to the workspaces and task classes they are registered for, and never to a normative pass.

Multiple model backends

Any OpenAI-compatible endpoint, local or remote, and the Anthropic API for the general and normative paths.

Confined execution

One rule governs the whole execution path: fail-closed — a confinement that cannot be applied is an error, never a silent downgrade.

1.1.5. Why the harness is the interesting part

A model that can only talk is safe and not very useful. A model that can run arbitrary commands is useful and not at all safe. The harness is the entire answer to “how do we get the second without the first”.

Its design rests on one rule, applied without exception:

Fail-closed

Any mechanism that cannot be honoured is an error, never a downgrade. If the sandbox is unavailable, the command does not run unsandboxed. If the resource-control mechanism is unavailable, the command does not run unlimited. If the network helper cannot attach the way we require, the network command does not fall back to a weaker attachment.

That rule is why several code paths look more paranoid than they strictly need to be, and why the test suite spends as much effort proving that things do not happen as proving that they do.

1.1.6. Layers of confinement

Four independent layers apply to every sandboxed command. They are independent on purpose: each one assumes the others may fail.

Layer

Mechanism

What it bounds

Authorization

CommandPolicy + CapabilityPolicy

whether a command may run at all, and with which capabilities

Filesystem / namespaces

Bubblewrap

what the command can see, write and reach

Per-process limits

prlimit inside the sandbox

descriptors and core dumps of each process

Whole-tree limits

cgroup v2 via a transient systemd scope

memory, task count and CPU of the command and every descendant

The order in which these wrap each other is fixed and is not an implementation detail; see Sandbox.

1.1.7. Trust boundaries

Three boundaries matter, and it is worth being explicit about which side of each one the model sits on.

The model is untrusted input.

It proposes tool calls; it does not authorize them. Every proposal goes through classification and capability intersection before anything runs.

The command argv is untrusted data.

It is never passed to a shell by the harness. shell=False everywhere, structured argv everywhere. When a command genuinely needs shell syntax, that is a distinct capability (shell:complex) and an explicit /bin/sh -c argv, not string interpolation.

The workspace is the only writable host surface.

Everything else the command sees is read-only or private to the sandbox.

1.1.8. What this documentation covers

The chapters follow the life of a tool call: what may run (Security model), where it runs (Sandbox), how it reaches the network if allowed (Network backend), what resources it may consume (Resource control), how all of that is verified (Testing), and what to look at when it misbehaves (Operations).

Several sections quote real measurements. Those numbers come from this machine and are labelled as such; Resource control discusses which of them are portable and which are not.

1.1.9. Project

SPEAR is developed at the REDS institute of HEIG-VD, and is published under the Apache License 2.0.