1.2. Architecture
1.2.1. Infrastructure overview
Before the internals, the whole platform on one page. Read it from the left:
an operator’s question or change request enters through spear-chat and
reaches the reasoning and control plane, which draws on two sources of
evidence kept apart on purpose — the authoritative, normative one (standards,
in warm) and the implementation one (repositories, corpora, local tools and
build outputs, in blue). The model backends only reason; everything that
touches a tree goes through the execution environment and stays inside its
workspace and read/write boundaries. What comes out on the right is either
grounded — an answer, a plan, a change, a validation result — or an explicit
refusal to answer when the evidence is insufficient.
Fig. 1.2 The SPEAR infrastructure. The two sources of truth — specification and implementation — are kept separate and reconciled only in the runtime.
What a turn is told about its workspace — rules, skills, workspace knowledge, project metadata — and the external capabilities it may reach are selected by the control plane per workspace and task class; they are described below (What a turn is shown, Workspace knowledge, External capabilities).
1.2.2. Request classes and execution paths
Every request is classified before anything runs, from the request’s own words and from whether a standard is engaged for the session:
Class |
What it is |
|---|---|
|
names no subject of its own, and no standard is engaged |
|
about the working tree: read it, change it, build it, test it |
|
about what a bound standard requires |
|
a change to the tree that must satisfy the bound standard |
Fig. 1.3 One request, one class, one path.
The class decides the path:
The coding core, behind SpearHost — every request with no standard engaged.
A standalone tool-calling loop with six tools (read_file,
search_files, patch, write_file, delete_file, terminal)
whose every call crosses SpearHost, the control plane. The turn ends with
implementation evidence computed from the record of what the calls did:
VERIFIED only when a build or test ran and passed after the last change to
the source, in the final source epoch. A command sent to the background, or a
Makefile that only prints its help, is no check. The project’s declared
commands (build_commands, test_commands in projects.json) and the
ones SPEAR probes from the tree may both verify; a declared kind wins, and a
probe only fills a kind left undeclared. The coding core needs an
OpenAI-compatible endpoint; the Anthropic backend serves the general and
normative paths. See Implementation mode.
The normative runtime — requests in a session where a standard is engaged.
The provider-neutral AgentRuntime loop with the standard’s tools, provision
records and the answer guards. It answers normative questions, and holds a
change asked for in such a session that is not MIXED to a five-stage workflow.
See Normative mode: authoritative standards and The bound-session change workflow.
The MIXED orchestration — a change that must satisfy the bound standard.
The normative runtime runs first, read-only; its cited provisions become a
constraint packet; the coding core makes the change; a post-check judges the
final source against the packet by deterministic evidence providers. The
verdict composes the implementation evidence with a normative status:
compliance only from normative evidence, otherwise NOT_DEMONSTRATED. See
Mixed mode.
The three paths share the harness underneath — workspace, command policy, sandbox, audit — and none of them lets a model’s statement stand as evidence.
Each path is also given only its own context. A turn starts from its workspace
— the registered project, or an unregistered tree on its own — and a
deterministic selector (context_selection.py) keeps the context its class
and pass call for, of this workspace or explicitly generic: rules, skills,
workspace knowledge, project metadata and external capabilities. Every
eligible skill is offered — there is no similarity gate, only a bound of eight
per turn — and knowledge and capabilities reach the passes that run on the
coding core, never a normative one. The selector also names the tool family;
every decision is audited and none is shown to the model
(What a turn is shown).
1.2.3. Runtime components
CLI / application (rag_chat.py)
|
v
TaskController --- request class, standard binding
|
context selection --- rules, skills, knowledge, metadata, capabilities
| | |
v v v
coding core AgentRuntime MixedOrchestrator
(agent/) (normative runtime) normative pre-pass
| ModelBackend NormativeConstraintSet
v WorkingState coding core
SpearHost ContextEngine post-check
(control_plane) VerificationPolicy composite verdict
target policy CheckpointManager
refusal breaker SessionStore
capability gateway --- MCP providers (stdio)
| |
v v
ToolRouter -- ToolRegistry -- ResultStore
|
v
CommandRunner -- Bubblewrap
The coding core (client/agent/) imports nothing of SPEAR: it calls a
Host interface for every read, write and command, and SpearHost
implements that interface with SPEAR’s policy. Its evidence is turned into the
implementation verdict by completion.py. In front of every mutation,
target_policy.py refuses generated output and snapshot copies to every tool
— delete_file and shell redirection included — and refusal_breaker.py
stops a turn that repeats the same refused operation five times. The capability
gateway (capability_gateway.py) answers the spear-capability host
command itself, never through a shell. The MIXED orchestration
(mixed_orchestration.py) composes the two runtimes and the normative
evidence modules (normative_constraints, normative_coverage,
normative_predicates, normative_evidence); it changes neither.
The AgentRuntime can also run Explorer and Reviewer roles with isolated
contexts and read-only tools; they are experimental and off by default.
1.2.4. Task ownership
On the normative runtime, WorkingState is authoritative task truth. It changes only through typed,
grounded events. Conversation, retrieved context, durable memory and compacted
summaries are context sources, not alternative task-state stores.
TaskController owns the bounded synchronous lifecycle: optional planning,
Main execution, verification, and checkpoint finalization by default. Explorer,
Reviewer and their bounded repair cycle are explicit experimental opt-ins. It
receives an already constructed
AgentContext and injected callbacks for project bench execution, diff
evidence, child sessions, and the remaining code-block compatibility path. It
does not read terminal input, render output, choose a provider, or construct a
filesystem persistence implementation.
1.2.5. Context and memory
On the normative runtime, ContextEngine composes the context. It accounts
for system rules, project rules, WorkingState projection, conversation
summaries, recent conversation, retrieval and tool evidence. The coding core
builds its own request from the selected rules, project metadata, skills,
workspace knowledge and admitted capabilities, and reads the tree through its
tools; it does not compact, it stops at half the window and asks for a
summary. Each of its model responses is bounded by
SPEAR_RESPONSE_MAX_TOKENS (16384 by default).
Workspace knowledge (workspace_knowledge.py) is what a workspace keeps
across sessions: typed records with a provenance, a verification and a
lifecycle (proposed, active, stale, revoked), stored in knowledge.sqlite3
under the state directory. /remember and /knowledge record
operator-confirmed knowledge; what a model offers — the remember tool, or
spear-knowledge propose on the coding core — is only a proposal, inert
until the operator accepts it. A record bound to a
source file goes stale when the file changes, and two active records that
disagree are reported as a conflict. Knowledge describes; it never overrides
the request, the rules, the configuration or WorkingState. The legacy
memories-*.md notes are no longer shown to a turn;
/knowledge migrate-remember migrates them.
External capabilities (capabilities.py, capability_gateway.py,
mcp_provider.py) are tools of MCP servers over stdio, registered in
capabilities.json (SPEAR_CAPABILITIES_FILE) with the workspaces and
task classes they apply to and a read/write policy. A small family is shown
whole; a large one as an index that spear-capability list,
describe and invoke read on demand. A provider that cannot start is
reported to the operator and to the turn.
1.2.6. Persistence ownership and retention
Persistent artifacts have distinct owners:
audit/runtime-trace.jsonlis optional observability and may be rotated or deleted without affecting resume.audit/sessions/contains resumable state and references tool results and checkpoints. Do not remove an active session’s referenced evidence.audit/tool-results/holds full bounded-out-of-context tool evidence.audit/checkpoints/holds exact pre-mutation bytes for safe rollback.trajectories.jsonland history files are training and user-facing projections, not resumable runtime state. Recording is deliberately wider than judging: a turn the project’s bench judged is recordedpassorfail, and a turn nothing judged is recordedunratedrather than dropped. Gating the recording on a verdict is what once left the dataset holding a single trajectory — one project declares a bench — so the width is the point, not an oversight./goodpromotes what was right.knowledge.sqlite3is durable workspace knowledge; legacymemories-*.mdfiles are kept only until they are migrated.
No automatic garbage collector currently runs. Closed session bundles may be archived or deleted manually as a unit after their checkpoint and result references are no longer needed. Deleting a trace or trajectory never repairs or invalidates a session; deleting referenced results/checkpoints makes the associated evidence unavailable.
1.2.7. Training data as a by-product
Fifteen of the modules belong to a subsystem the rest of the harness feeds rather than calls: canonical episode capture, SFT and preference curation, governance, readiness, the frozen bundle, and the operator control plane that launches a job. It sits outside the turn — nothing in it is model-visible, no tool reaches it, and freezing a bundle executes no training. It has its own chapter, Training and fine-tuning.
The coupling that does exist runs one way and is deliberate: AgentRuntime
and TaskController emit trajectories, rag_chat owns the /finetune
operator command, and training_handoff is allowed to stop the inference
service because on a single-GPU host the card that trains is the card that
serves.
1.2.8. Dependency direction
Lower runtime modules do not import rag_chat or the cli/ modules it is
built from. The client owns backend construction, project selection, terminal
commands, human confirmation and presentation. TaskController depends on provider-neutral policies and
protocols; AgentRuntime does not depend on TaskController or on
Explorer/Reviewer orchestration.
1.2.9. Model capability boundary
ModelBackend.complete remains the intentionally small required protocol.
Canonical ModelTurn carries stop reason and optional usage, while context
and output limits are explicit AgentContext configuration. Cancellation
is checked immediately before and after calls for every backend. No mandatory
capability object was added during consolidation: the current adapters cannot
truthfully promise transport-level cancellation or exact token counting, and
making those flags mandatory would couple the runtime to provider behavior.
Future adapters may expose optional native capabilities without changing the
canonical turn contract.
1.2.10. Default profile
Requests with no standard engaged run on the coding core. --fresh starts a
session without its stored conversation; knowledge, rules and configuration
still apply. Requests in a
standard-bound session run on the normative runtime, or through the MIXED
orchestration when they ask for a change in the standard’s terms. On the
normative runtime, planning remains deterministic and conservative; Explorer,
Reviewer and reviewer repair are disabled unless the caller explicitly enables
them.
1.2.11. Provenance
The coding core is SPEAR’s own module, built by porting the coding loop and the
coding tools of Hermes Agent (Nous Research, MIT license) function by function,
and verifying the port against results captured from Hermes itself. What sits
around it — SpearHost, the command policy and sandbox, the evidence plane, the
normative runtime and the MIXED orchestration — is SPEAR’s. Every file copied
or adapted, with its upstream revision, its license and its destination, is
listed in THIRD_PARTY_NOTICES.md at the repository root.