2.4. Container

The container exists to hand the assistant to someone else: the harness, its dependencies, the embedder and a prebuilt retrieval index, in one image. They mount their own checkouts and point it at a model endpoint.

This chapter is for whoever builds the images. To run a published one, read Running the public image instead.

The image is defined in docker/; the scripts that build, run and publish it are in scripts/docker/, and scripts/spear-image drives them (Managing images).

File

Purpose

docker/Dockerfile

The image. Do not build it by hand — see Two build contexts.

docker/entrypoint.sh

Refuses to start on the two failures that are otherwise silent.

docker/projects.docker.json

The corpus registry, in relative paths. Generated by gen-registry.py and not in the repository.

docker/README.md

The short version, for whoever receives the image.

scripts/spear-image

The operator’s entry point: build, list, inspect, run, save, load, push, erase.

scripts/docker/build.sh

Builds, wrapping both contexts.

scripts/docker/spear-docker.sh

Runs, deriving the mounts and carrying the three --security-opt flags.

scripts/docker/push.sh

Publishes, refusing a private image without --allow-push.

scripts/docker/stage-*.py, gen-registry.py

What build.sh stages into its contexts.

2.4.1. Managing images

Put the tree’s commands on PATH once per shell, from the repository root:

$ cd ~/spear && . ./env.sh

Only scripts/ goes on PATH. scripts/docker/ holds a build.sh, and a shell that has also sourced another tree’s env.sh with a build.sh of its own would then have two.

$ spear-image build private               # spear:<version>-private
$ spear-image build private --bake so3,so3-doc
$ spear-image build private --knowledge so3      # common state, see below
$ spear-image build public                # spear:<version>-public
$ spear-image list                        # profile and label of each
$ spear-image inspect private
$ spear-image run private -- --auto       # spear-docker.sh options, then -- harness args
$ spear-image save private                # -> spear-private.tar.zst
$ spear-image load spear-private.tar.zst
$ spear-image push ghcr.io/<org>/spear:<version>-public
$ spear-image erase private [--cache]     # or public, all, a tag

A profile names the tag build.sh gives by default, spear:<version>-<profile>, where the version is the release the tree is (scripts/spearversion.sh: the latest v* tag, else the fallback the release sets, such as 0.3.0-rc1); anything containing a : is taken as a tag. spear-image decides nothing the scripts under it would not: build.sh still decides what a profile may carry and push.sh still refuses a private image without --allow-push.

erase --cache also prunes the build cache. Removing a private image does not remove the layers it was built from — the licensed documents and customer trees it copied stay in the cache until that is pruned.

2.4.2. Daily use

spear-docker --reds --auto        # what `spear-chat --reds --auto` does

~/.local/bin/spear-docker is the sibling of spear-chat and takes the same arguments; everything after them is passed to the harness unchanged. Three things the runner does that the native launcher does not have to:

The SSH tunnel stays on the host. --reds opens it exactly as spear-chat.sh does, reading the same reds.conf, then hands the container SPEAR_API_BASE=http://127.0.0.1:8082/v1. Putting the tunnel inside would mean shipping the keys and ~/.ssh/config into an image meant to be handed around; --network host makes 127.0.0.1 the same thing on both sides anyway.

Every tree is bound at its own absolute path. The harness runs its tools in the cwd, whatever the corpus (Retrieval), and build systems record absolute paths — CMake caches, BitBake stamps, toolchain locations — so a tree mounted elsewhere builds against compilers that are not there. An identity mount keeps both working, and the host cwd needs no translation: it is passed as the container’s working directory unchanged. This is the workstation mode; a colleague without the repository mounts under /corpora instead (Running the public image).

Session state is written outside the image, as the host user. History, workspace knowledge, trajectories and the audit trail accumulate; docker run --rm would throw them away, and a container running as root would leave them owned by root and unreadable to the harness running natively. The default is ~/.spear/state, overridable with --state DIR or SPEAR_STATE_DIR.

That separation is also a property of the harness itself: STATE_DIR covers every path that accumulates and defaults to the application directory, so a workstation launch is unaffected. The paths that matter most are the ones that make a session reconstructible rather than merely readable:

Path under STATE_DIR

What is lost with it

audit/sessions/

the resumable session — an interrupted task cannot be picked up again

audit/tool-results/

the full tool output kept out of model context; only the previews the model saw would survive

audit/checkpoints/

the exact pre-mutation bytes, so a rollback has nothing to restore from

audit/runtime-trace.jsonl

the spans: which tool ran, how long, with what outcome

audit/tool-actions.jsonl

the metadata-only record of every mutating attempt

knowledge.sqlite3

the workspace knowledge: every recorded fact, its provenance and its state

history*.json, history-archive.jsonl, trajectories.jsonl, .input_history

conversation, and the trajectories a future fine-tune would train on

Evidence inside the image is evidence lost with the container that produced it, which is precisely the case it exists for.

External capabilities are not baked: capabilities.json is a machine’s own file, and the MCP servers it starts must exist where the harness runs. A container that needs them is given the file through a mount and SPEAR_CAPABILITIES_FILE (External capabilities).

2.4.3. What is baked, and what is not

Content

Where

Why

harness + venv

image

pinned, reproducible

embedder (bge-m3)

image, 4.5 G

fetching it on first run is a surprise on a machine that may have no Hugging Face access at all

ChromaDB index

image if the building host has one, 4.5 G

re-indexing takes hours and needs every corpus tree present — the one thing a newcomer does not have

rules, skills, benches, notes corpus

image if present

a deployment’s own content; see below

corpora/

image, 38 M

cross-cutting references, attached to every session without duplicating them into each project index

source trees

mounted

working copies that change daily; an image would be stale the next morning

model weights

neither

the harness talks to an endpoint, it does not host a model

2.4.4. Two build contexts

build.sh passes two, and the reason is size:

. (default) → client/

Harness code and the prebuilt index. Narrow on purpose: a wider context would be re-transferred on every build and would invalidate the 4.5 GB embedder layer.

repo (named) → the repository root

Only corpora/, docker/entrypoint.sh and the generated docker/projects.docker.json are taken from it. A named context is fetched lazily — BuildKit transfers only the paths actually COPY-ed — so pointing it at a tree holding 122 GB of weights costs nothing.

A bare docker build therefore fails on the missing --from=repo.

2.4.5. The optional inputs

Five of the things the image would like to carry are not in the repository and cannot be: the retrieval index is built on the host and gitignored, and the rules, the skills, the benches and the shared notes corpus are a deployment’s own content. Requiring them meant the documented build failed on a clean clone — on the first missing one, with a message about a directory the reader had no way to produce.

Each is now a named context of its own, resolved by build.sh in three steps: the environment variable, else the in-tree directory, else an empty directory.

Context

Override

In-tree default

index

SPEAR_INDEX_DIR

client/chromadb

rules

SPEAR_RULES_DIR

client/rules.d

skills

SPEAR_SKILLS_DIR

client/skills

benches

SPEAR_BENCH_DIR

client/benches

notes

SPEAR_NOTES_DIR

claude/

build.sh prints which of the five it found and which it did not, so an image that carries less says so at build time rather than at the first question. An empty context still creates the directory, so the harness finds an empty directory rather than no directory — which is the difference between “no rules” and a stack trace.

The first four variables are the same ones the harness itself reads at runtime (Content the deployment owns), so a deployment that keeps its rules outside the checkout points one variable at them and both the native run and the image follow.

Two further contexts are staged rather than pointed at, because neither is a directory that happens to be in the right shape already: the normative store is filtered per document, and the corpus trees are a named subset of a registry that may run to hundreds of gigabytes.

2.4.6. Two profiles

An image that carries a normative store is one somebody can be handed. What it may carry is decided per build, and the default is the one that is safe to give to anyone.

$ scripts/docker/build.sh --profile public
$ scripts/docker/build.sh --profile private --bake so3,acme-firmware

Profile

What it carries

public

only normative documents that declare themselves PUBLIC; of rules, skills, benches and notes only the files the repository tracks; no retrieval index (it holds chunks of every corpus on the host) and no baked tree or workspace knowledge (--bake and --knowledge are refused). Anything withheld is printed as WITHHELD. Labelled redistributable=true.

private

everything the building host has, licensed documents and the original PDFs included where the store retained them. Labelled redistributable=false.

Which document is which is never a list kept in this repository. Every ingested document already records source_origin and raw_pdf_retained in its manifest, and stage-standards.py reads them. A list here would have to name a customer’s standard in order to exclude it, would go stale on the next ingestion, and would leave the whole decision one forgotten edit away from shipping a licensed document. A manifest that declares nothing is treated as licensed.

Important

The active binding travels only if the document it names travelled. A binding pointing at an absent store is worse than none: the session opens looking bound and answers from nothing. Where the image carries exactly one document, the entrypoint binds it at startup — through the harness, so the fingerprints are computed rather than fabricated.

2.4.7. Baking the trees

--bake copies named registered corpora into the image, at the paths the generated registry already resolves them to. Nothing is baked by default, and a bind mount on /corpora still shadows whatever was — so a workstation keeps working from its own checkouts, and only the handed-over container relies on what is inside.

A baked corpus must be registered and indexed on the host first. The image ships the index; a tree whose collection is absent answers with no retrieval at all, which is the whole reason the tool exists.

$ spear-corpus add acme-firmware ~/src/acme-firmware
$ spear-index ~/src/acme-firmware
$ scripts/docker/build.sh --profile private --bake so3,acme-firmware

When --bake is used the image’s registry is restricted to what the image actually carries. Otherwise the recipient opens the container to a list of corpora they do not have and cannot get.

2.4.8. Common and local state

An image carries a common state, /opt/spear/common, that everyone who runs it reads and nobody writes: the normative store and, when the build is asked for it, the workspace knowledge of named projects. Each user keeps their own state in the directory mounted on /state, where everything a session writes goes.

$ spear-image build private --knowledge so3,so3-doc
$ spear-image build private --knowledge all

Only active records travel, each at its current version: proposals nobody accepted, revoked records and the earlier wording of an amended one stay with the user who built the image. Knowledge of an unregistered tree never travels, since it is named after a path on the building host. A recipient’s record follows the project by its registered name, so the project must be registered under the same name in the image’s registry.

A user’s change to a common record (revoking it, amending it, or the record going stale because their tree differs from the builder’s) is copied into their own state and shadows the common one there; the image is not changed. Shipping new knowledge is a new build.

2.4.8.1. Consolidating

The team keeps its common store on the building host, in the directory SPEAR_COMMON_STATE_DIR names (machine.env is the place to set it), and --knowledge ships from it; without one, the builder’s own store is shipped. spear-consolidate merges users’ stores into it, so that what each of them learns reaches everyone at the next build:

$ spear-consolidate ~/.local/state/spear/knowledge.sqlite3 alice.sqlite3
$ spear-consolidate ~/.local/state/spear/knowledge.sqlite3 alice.sqlite3 --apply
$ spear-image build private --knowledge so3

A user’s store is knowledge.sqlite3 in their state directory: the one mounted on /state for a container, ~/.local/state/spear on a workstation. Without --apply the command only reports.

Outcome

When

added

an active record the common store does not have

updated · revoked

a common record the user amended, accepted or revoked; the common store takes it with its history

duplicate

the same fact, already common under another id

conflict

not merged: a common record changed since the user copied it, or a new fact that contradicts an active common one. Settle it in either store and run again; the command exits 1 while any remains

skipped

a proposal (the user accepts it first), or a record stale in the user’s tree only

Knowledge describes the building host’s trees, customer code included, so --knowledge needs --profile private.

2.4.9. A deployment’s own content, and how it gets in

A deployment keeps what is specific to it outside the checkout: its rules, its skills, its benches, the trees it works on and the normative documents it answers from. client/machine.env is the one untracked file that says where those live, and it is read by the launcher, by spear-corpus and — since it decides what an image carries — by scripts/docker/build.sh.

Everything therefore reaches an image by one of three routes, and none of them is a second repository:

What

How it gets in

rules, skills, benches

machine.env sets SPEAR_RULES_DIR and the other two; build.sh resolves them as named contexts (The optional inputs)

the normative store

staged per document by profile (Two profiles)

the corpus trees

--bake, once each is registered and indexed on the host

workspace knowledge

--knowledge, as the image’s common state (Common and local state)

Important

build.sh reads machine.env for exactly this reason. Without it the three variables are unset, the table falls back to the in-tree directories — a README in each — and the build reports them as found, because they are directories and they exist. The image then ships a harness with no rules and no skills and announces neither: the failure machine.env exists to prevent on a workstation, reproduced in the artefact handed to someone else, where it is harder to notice and impossible to fix from inside.

2.4.9.1. There is no second image

A deployment’s private material is content, not an application: there is no harness in it, nothing to execute, and so nothing to build an image around. It is not packaged separately — it is what makes a private image a private image.

Start to finish, on the machine that has the content:

$ spear-corpus add acme-firmware ~/src/acme-firmware   # register…
$ spear-index ~/src/acme-firmware                      # …and index
$ scripts/docker/build.sh --profile private --bake so3,acme-firmware

What the build reports is what the image carries:

settings  …/client/machine.env
rules     …/rules.d
skills    …/skills
benches   …/benches
standards <licensed document>  (LICENSED_STANDARD)
standards bound on open: <licensed document>
corpora   acme-firmware -> /corpora/…

The recipient opens a container already bound, with the rules, the skills, the index and the trees, and nothing to mount.

Note

The variables that build a corpus rather than use one — the path to a source PDF, the working tree a pilot drives — are deliberately not baked. They name host paths that do not exist in a container, and the corpus they produce is already inside it. They belong on the machine that builds the image, not in what is handed over.

2.4.10. Publishing, and not publishing

$ docker tag spear:<version>-public ghcr.io/smartobjectoriented/spear:<version>-public
$ scripts/docker/push.sh ghcr.io/smartobjectoriented/spear:<version>-public
$ scripts/docker/push.sh <private-registry>/spear:<version>-private            # refused
$ scripts/docker/push.sh --allow-push <private-registry>/spear:<version>-private

A public image goes to the GitHub container registry of the project, where Running the public image tells colleagues to pull it from. A private image goes only to the private registry of the organisation that owns its content; where that is, and how its users run it, is documented with that content, not here.

The guard reads the label, not the tag. A tag gets retyped, shortened and reused; a label travels with the bytes through docker save, a registry and back. An image not built by build.sh carries no label at all and is refused rather than guessed at.

Warning

--allow-push on a private image is a decision about a licence and a contract, not about a registry: it puts a licensed corpus and a customer’s source on that registry’s infrastructure, under whatever access policy it has today. It costs a deliberate word for that reason.

2.4.11. Relative corpus paths

projects.docker.json is build output: gen-registry.py derives it from this machine’s projects.json on every build, and it is gitignored. A committed copy would be a published map of one developer’s disk — corpus names, federations and the collection hash of each — and stale the moment a corpus was added. A clone with no projects.json gets an empty registry and an image that registers nothing, which is the honest answer: the operator registers their own trees. The one registry the repository carries by hand is client/projects.example.json, and it is a template.

It registers every corpus relative to SPEAR_CORPUS_ROOT (/corpora in the image), where a workstation registry uses absolute paths under the operator’s home. That is what makes one image work for someone whose checkouts live elsewhere; resolve_corpus_path() leaves absolute paths untouched, so the workstation keeps behaving exactly as before (Retrieval).

A tree the host does not have is not an error. The entrypoint prints what resolved and what did not, at startup, rather than letting a missing corpus surface three questions later as an empty retrieval.

2.4.12. Mounts are derived, not listed

Without --corpora DIR, spear-docker.sh computes one bind per registered corpus from the registry itself, dropping any path already inside another.

This is not tidiness. Binding the whole SPEAR checkout — the obvious single mount — would put models/ (75 G) and the served GGUFs (47 G) inside the container read-write, in the one tree the sandbox grants write access to, for no retrieval value. Weights are not corpus. Deriving the list also means a corpus added to the registry is mounted without touching this script.

2.4.13. The three flags that are not optional

--security-opt seccomp=unconfined --security-opt apparmor=unconfined
--security-opt systempaths=unconfined

The harness runs every command inside bubblewrap and refuses to run any without it (Sandbox). Docker’s default seccomp profile blocks clone(CLONE_NEWUSER), so bwrap cannot start, and the failure surfaces as sandbox unavailable — which reads like a broken harness rather than a missing run flag. bwrap also mounts a fresh /proc for each command, which Docker’s masked system paths prevent unless systempaths=unconfined is given. spear-docker.sh passes all three, and entrypoint.sh tests bwrap first and prints exactly this if it cannot.

None of the flags grants the container new privileges on the host: all three are about letting an unprivileged namespace be created inside it. The sandbox is still what confines the model’s commands, and it is still doing its job.

2.4.14. Refreshing the index

The baked index is a snapshot. Re-index on the workstation, then rebuild: chromadb/ is copied late in the Dockerfile, so only that layer and the ones after it are rebuilt.

spear-index /path/to/tree     # updates chromadb/ on the workstation
scripts/docker/build.sh --profile private

2.4.15. Why retrieval is not served from the GPU host

The obvious economy is to move the embedder and the index onto the machine that already serves the model, and keep a thin client here. It was costed and declined; the numbers are worth keeping, because the idea comes back every time someone looks at the image size.

What it would save, measured on a development workstation:

Removed from the local install

Size

nvidia/ + torch/ + transformers + sentence-transformers

4.0 GB

the bge-m3 weights

4.3 GB

the ChromaDB index

4.3 GB

total, leaving ~0.2 GB of client

12.6 GB

The image would fall from 17.1 GB to roughly 5 GB, and the ten-second embedder load on the first query of a session would disappear, since the model would stay resident server-side on a GPU. Latency is not the objection: a round trip through the SSH tunnel that is already open measures 17 ms, against ~100 ms for a local CPU embedding. (The existing deploy/embed_worker.py could not be used as it stands: it opens one SSH connection per batch, 350 ms, which is right for indexing 45 000 chunks once and wrong for three embeddings a turn. It would need a persistent service behind the tunnel.)

It was declined for three reasons, in increasing order of weight:

  • It ends offline operation. Retrieval is what turns 18 % into 90 % on the build-system questions. Making it require a VPN and a reachable GPU host means a session on a train is not a degraded session, it is a different assistant.

  • It empties the container of its purpose. The image exists so that retrieval works the moment it starts, on a machine that may have no Hugging Face access at all. A 5 GB image that needs a tunnel to answer anything is a different product, not a smaller one.

  • A GPU host is often a shared login. Collections are named from the corpus’s absolute path, so two people indexing the same path under one account would write into the same collection. Fixing that means per-user prefixes and a shared ChromaDB on a filesystem other users already fill.

The work itself is small – about a day: a retrieve endpoint doing embedding and query in one round trip, a second port forward in spear-chat.sh, ten PersistentClient call sites and five embedding ones. Cheapness was never the question.