2.4. Container
The container exists to hand the assistant to someone else: the harness, its dependencies, the embedder and a prebuilt retrieval index, in one image. They mount their own checkouts and point it at a model endpoint.
This chapter is for whoever builds the images. To run a published one, read Running the public image instead.
The image is defined in docker/; the scripts that build, run and
publish it are in scripts/docker/, and scripts/spear-image drives them
(Managing images).
File |
Purpose |
|---|---|
|
The image. Do not build it by hand — see Two build contexts. |
|
Refuses to start on the two failures that are otherwise silent. |
|
The corpus registry, in relative paths. Generated by
|
|
The short version, for whoever receives the image. |
|
The operator’s entry point: build, list, inspect, run, save, load, push, erase. |
|
Builds, wrapping both contexts. |
|
Runs, deriving the mounts and carrying the three |
|
Publishes, refusing a |
|
What |
2.4.1. Managing images
Put the tree’s commands on PATH once per shell, from the repository root:
$ cd ~/spear && . ./env.sh
Only scripts/ goes on PATH. scripts/docker/ holds a build.sh,
and a shell that has also sourced another tree’s env.sh with a
build.sh of its own would then have two.
$ spear-image build private # spear:<version>-private
$ spear-image build private --bake so3,so3-doc
$ spear-image build private --knowledge so3 # common state, see below
$ spear-image build public # spear:<version>-public
$ spear-image list # profile and label of each
$ spear-image inspect private
$ spear-image run private -- --auto # spear-docker.sh options, then -- harness args
$ spear-image save private # -> spear-private.tar.zst
$ spear-image load spear-private.tar.zst
$ spear-image push ghcr.io/<org>/spear:<version>-public
$ spear-image erase private [--cache] # or public, all, a tag
A profile names the tag build.sh gives by default,
spear:<version>-<profile>, where the version is the release the tree is
(scripts/spearversion.sh: the latest v* tag, else the fallback the
release sets, such as 0.3.0-rc1);
anything containing a : is taken as a tag. spear-image decides nothing the scripts under it
would not: build.sh still decides what a profile may carry and push.sh
still refuses a private image without --allow-push.
erase --cache also prunes the build cache. Removing a private image
does not remove the layers it was built from — the licensed documents and
customer trees it copied stay in the cache until that is pruned.
2.4.2. Daily use
spear-docker --reds --auto # what `spear-chat --reds --auto` does
~/.local/bin/spear-docker is the sibling of spear-chat and takes the
same arguments; everything after them is passed to the harness unchanged.
Three things the runner does that the native launcher does not have to:
The SSH tunnel stays on the host. --reds opens it exactly as
spear-chat.sh does, reading the same reds.conf, then hands the
container SPEAR_API_BASE=http://127.0.0.1:8082/v1. Putting the tunnel
inside would mean shipping the keys and ~/.ssh/config into an image meant
to be handed around; --network host makes 127.0.0.1 the same thing on
both sides anyway.
Every tree is bound at its own absolute path. The harness runs its tools
in the cwd, whatever the corpus (Retrieval), and build systems
record absolute paths — CMake caches, BitBake stamps, toolchain locations — so
a tree mounted elsewhere builds against compilers that are not there. An
identity mount keeps both working, and the host cwd needs no translation: it is
passed as the container’s working directory unchanged. This is the
workstation mode; a colleague without the repository mounts under /corpora
instead (Running the public image).
Session state is written outside the image, as the host user. History,
workspace knowledge, trajectories and the audit trail accumulate; docker run --rm
would throw them away, and a container running as root would leave them
owned by root and unreadable to the harness running natively. The default is
~/.spear/state, overridable with --state DIR or
SPEAR_STATE_DIR.
That separation is also a property of the harness itself: STATE_DIR covers
every path that accumulates and defaults to the application directory, so a
workstation launch is unaffected. The paths that matter most are the ones that
make a session reconstructible rather than merely readable:
Path under |
What is lost with it |
|---|---|
|
the resumable session — an interrupted task cannot be picked up again |
|
the full tool output kept out of model context; only the previews the model saw would survive |
|
the exact pre-mutation bytes, so a rollback has nothing to restore from |
|
the spans: which tool ran, how long, with what outcome |
|
the metadata-only record of every mutating attempt |
|
the workspace knowledge: every recorded fact, its provenance and its state |
|
conversation, and the trajectories a future fine-tune would train on |
Evidence inside the image is evidence lost with the container that produced it, which is precisely the case it exists for.
External capabilities are not baked: capabilities.json is a machine’s own
file, and the MCP servers it starts must exist where the harness runs. A
container that needs them is given the file through a mount and
SPEAR_CAPABILITIES_FILE (External capabilities).
2.4.3. What is baked, and what is not
Content |
Where |
Why |
|---|---|---|
harness + venv |
image |
pinned, reproducible |
embedder (bge-m3) |
image, 4.5 G |
fetching it on first run is a surprise on a machine that may have no Hugging Face access at all |
ChromaDB index |
image if the building host has one, 4.5 G |
re-indexing takes hours and needs every corpus tree present — the one thing a newcomer does not have |
rules, skills, benches, notes corpus |
image if present |
a deployment’s own content; see below |
|
image, 38 M |
cross-cutting references, attached to every session without duplicating them into each project index |
source trees |
mounted |
working copies that change daily; an image would be stale the next morning |
model weights |
neither |
the harness talks to an endpoint, it does not host a model |
2.4.4. Two build contexts
build.sh passes two, and the reason is size:
.(default) →client/Harness code and the prebuilt index. Narrow on purpose: a wider context would be re-transferred on every build and would invalidate the 4.5 GB embedder layer.
repo(named) → the repository rootOnly
corpora/,docker/entrypoint.shand the generateddocker/projects.docker.jsonare taken from it. A named context is fetched lazily — BuildKit transfers only the paths actuallyCOPY-ed — so pointing it at a tree holding 122 GB of weights costs nothing.
A bare docker build therefore fails on the missing --from=repo.
2.4.5. The optional inputs
Five of the things the image would like to carry are not in the repository and cannot be: the retrieval index is built on the host and gitignored, and the rules, the skills, the benches and the shared notes corpus are a deployment’s own content. Requiring them meant the documented build failed on a clean clone — on the first missing one, with a message about a directory the reader had no way to produce.
Each is now a named context of its own, resolved by build.sh in three
steps: the environment variable, else the in-tree directory, else an empty
directory.
Context |
Override |
In-tree default |
|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
build.sh prints which of the five it found and which it did not, so an
image that carries less says so at build time rather than at the first
question. An empty context still creates the directory, so the harness finds
an empty directory rather than no directory — which is the difference between
“no rules” and a stack trace.
The first four variables are the same ones the harness itself reads at runtime (Content the deployment owns), so a deployment that keeps its rules outside the checkout points one variable at them and both the native run and the image follow.
Two further contexts are staged rather than pointed at, because neither is a directory that happens to be in the right shape already: the normative store is filtered per document, and the corpus trees are a named subset of a registry that may run to hundreds of gigabytes.
2.4.6. Two profiles
An image that carries a normative store is one somebody can be handed. What it may carry is decided per build, and the default is the one that is safe to give to anyone.
$ scripts/docker/build.sh --profile public
$ scripts/docker/build.sh --profile private --bake so3,acme-firmware
Profile |
What it carries |
|---|---|
|
only normative documents that declare themselves |
|
everything the building host has, licensed documents and the original
PDFs included where the store retained them. Labelled
|
Which document is which is never a list kept in this repository. Every
ingested document already records source_origin and raw_pdf_retained
in its manifest, and stage-standards.py reads them. A list here would have
to name a customer’s standard in order to exclude it, would go stale on the
next ingestion, and would leave the whole decision one forgotten edit away from
shipping a licensed document. A manifest that declares nothing is treated as
licensed.
Important
The active binding travels only if the document it names travelled. A binding pointing at an absent store is worse than none: the session opens looking bound and answers from nothing. Where the image carries exactly one document, the entrypoint binds it at startup — through the harness, so the fingerprints are computed rather than fabricated.
2.4.7. Baking the trees
--bake copies named registered corpora into the image, at the paths the
generated registry already resolves them to. Nothing is baked by default, and a
bind mount on /corpora still shadows whatever was — so a workstation keeps
working from its own checkouts, and only the handed-over container relies on
what is inside.
A baked corpus must be registered and indexed on the host first. The image ships the index; a tree whose collection is absent answers with no retrieval at all, which is the whole reason the tool exists.
$ spear-corpus add acme-firmware ~/src/acme-firmware
$ spear-index ~/src/acme-firmware
$ scripts/docker/build.sh --profile private --bake so3,acme-firmware
When --bake is used the image’s registry is restricted to what the image
actually carries. Otherwise the recipient opens the container to a list of
corpora they do not have and cannot get.
2.4.8. Common and local state
An image carries a common state, /opt/spear/common, that everyone who
runs it reads and nobody writes: the normative store and, when the build is
asked for it, the workspace knowledge of named projects. Each user keeps their
own state in the directory mounted on /state, where everything a session
writes goes.
$ spear-image build private --knowledge so3,so3-doc
$ spear-image build private --knowledge all
Only active records travel, each at its current version: proposals nobody accepted, revoked records and the earlier wording of an amended one stay with the user who built the image. Knowledge of an unregistered tree never travels, since it is named after a path on the building host. A recipient’s record follows the project by its registered name, so the project must be registered under the same name in the image’s registry.
A user’s change to a common record (revoking it, amending it, or the record going stale because their tree differs from the builder’s) is copied into their own state and shadows the common one there; the image is not changed. Shipping new knowledge is a new build.
2.4.8.1. Consolidating
The team keeps its common store on the building host, in the directory
SPEAR_COMMON_STATE_DIR names (machine.env is the place to set it), and
--knowledge ships from it; without one, the builder’s own store is
shipped. spear-consolidate merges users’ stores into it, so that what each
of them learns reaches everyone at the next build:
$ spear-consolidate ~/.local/state/spear/knowledge.sqlite3 alice.sqlite3
$ spear-consolidate ~/.local/state/spear/knowledge.sqlite3 alice.sqlite3 --apply
$ spear-image build private --knowledge so3
A user’s store is knowledge.sqlite3 in their state directory: the one
mounted on /state for a container, ~/.local/state/spear on a
workstation. Without --apply the command only reports.
Outcome |
When |
|---|---|
|
an active record the common store does not have |
|
a common record the user amended, accepted or revoked; the common store takes it with its history |
|
the same fact, already common under another id |
|
not merged: a common record changed since the user copied it, or a new fact that contradicts an active common one. Settle it in either store and run again; the command exits 1 while any remains |
|
a proposal (the user accepts it first), or a record stale in the user’s tree only |
Knowledge describes the building host’s trees, customer code included, so
--knowledge needs --profile private.
2.4.9. A deployment’s own content, and how it gets in
A deployment keeps what is specific to it outside the checkout: its rules,
its skills, its benches, the trees it works on and the normative documents it
answers from. client/machine.env is the one untracked file that says where
those live, and it is read by the launcher, by spear-corpus and — since it
decides what an image carries — by scripts/docker/build.sh.
Everything therefore reaches an image by one of three routes, and none of them is a second repository:
What |
How it gets in |
|---|---|
rules, skills, benches |
|
the normative store |
staged per document by profile (Two profiles) |
the corpus trees |
|
workspace knowledge |
|
Important
build.sh reads machine.env for exactly this reason. Without it the
three variables are unset, the table falls back to the in-tree directories
— a README in each — and the build reports them as found, because
they are directories and they exist. The image then ships a harness with no
rules and no skills and announces neither: the failure machine.env
exists to prevent on a workstation, reproduced in the artefact handed to
someone else, where it is harder to notice and impossible to fix from
inside.
2.4.9.1. There is no second image
A deployment’s private material is content, not an application: there is no
harness in it, nothing to execute, and so nothing to build an image around. It
is not packaged separately — it is what makes a private image a
private image.
Start to finish, on the machine that has the content:
$ spear-corpus add acme-firmware ~/src/acme-firmware # register…
$ spear-index ~/src/acme-firmware # …and index
$ scripts/docker/build.sh --profile private --bake so3,acme-firmware
What the build reports is what the image carries:
settings …/client/machine.env
rules …/rules.d
skills …/skills
benches …/benches
standards <licensed document> (LICENSED_STANDARD)
standards bound on open: <licensed document>
corpora acme-firmware -> /corpora/…
The recipient opens a container already bound, with the rules, the skills, the index and the trees, and nothing to mount.
Note
The variables that build a corpus rather than use one — the path to a source PDF, the working tree a pilot drives — are deliberately not baked. They name host paths that do not exist in a container, and the corpus they produce is already inside it. They belong on the machine that builds the image, not in what is handed over.
2.4.10. Publishing, and not publishing
$ docker tag spear:<version>-public ghcr.io/smartobjectoriented/spear:<version>-public
$ scripts/docker/push.sh ghcr.io/smartobjectoriented/spear:<version>-public
$ scripts/docker/push.sh <private-registry>/spear:<version>-private # refused
$ scripts/docker/push.sh --allow-push <private-registry>/spear:<version>-private
A public image goes to the GitHub container registry of the project, where
Running the public image tells colleagues to pull it from. A private image
goes only to the private registry of the organisation that owns its content;
where that is, and how its users run it, is documented with that content, not
here.
The guard reads the label, not the tag. A tag gets retyped, shortened and
reused; a label travels with the bytes through docker save, a registry and
back. An image not built by build.sh carries no label at all and is refused
rather than guessed at.
Warning
--allow-push on a private image is a decision about a licence and a
contract, not about a registry: it puts a licensed corpus and a customer’s
source on that registry’s infrastructure, under whatever access policy it
has today. It costs a deliberate word for that reason.
2.4.11. Relative corpus paths
projects.docker.json is build output: gen-registry.py derives it
from this machine’s projects.json on every build, and it is gitignored. A
committed copy would be a published map of one developer’s disk — corpus names,
federations and the collection hash of each — and stale the moment a corpus was
added. A clone with no projects.json gets an empty registry and an image
that registers nothing, which is the honest answer: the operator registers
their own trees. The one registry the repository carries by hand is
client/projects.example.json, and it is a template.
It registers every corpus relative to SPEAR_CORPUS_ROOT (/corpora
in the image), where a workstation registry uses absolute paths under the
operator’s home. That is what makes one
image work for someone whose checkouts live elsewhere;
resolve_corpus_path() leaves absolute paths untouched, so the workstation
keeps behaving exactly as before (Retrieval).
A tree the host does not have is not an error. The entrypoint prints what resolved and what did not, at startup, rather than letting a missing corpus surface three questions later as an empty retrieval.
2.4.12. Mounts are derived, not listed
Without --corpora DIR, spear-docker.sh computes one bind per registered
corpus from the registry itself, dropping any path already inside another.
This is not tidiness. Binding the whole SPEAR checkout — the obvious
single mount — would put models/ (75 G) and the served GGUFs (47 G) inside
the container read-write, in the one tree the sandbox grants write access
to, for no retrieval value. Weights are not corpus. Deriving the list also
means a corpus added to the registry is mounted without touching this script.
2.4.13. The three flags that are not optional
--security-opt seccomp=unconfined --security-opt apparmor=unconfined
--security-opt systempaths=unconfined
The harness runs every command inside bubblewrap and refuses to run any
without it (Sandbox). Docker’s default seccomp profile blocks
clone(CLONE_NEWUSER), so bwrap cannot start, and the failure surfaces
as sandbox unavailable — which reads like a broken harness rather than a
missing run flag. bwrap also mounts a fresh /proc for each command,
which Docker’s masked system paths prevent unless systempaths=unconfined
is given. spear-docker.sh passes all three, and entrypoint.sh tests
bwrap first and prints exactly this if it cannot.
None of the flags grants the container new privileges on the host: all three are about letting an unprivileged namespace be created inside it. The sandbox is still what confines the model’s commands, and it is still doing its job.
2.4.14. Refreshing the index
The baked index is a snapshot. Re-index on the workstation, then rebuild:
chromadb/ is copied late in the Dockerfile, so only that layer and the
ones after it are rebuilt.
spear-index /path/to/tree # updates chromadb/ on the workstation
scripts/docker/build.sh --profile private
2.4.15. Why retrieval is not served from the GPU host
The obvious economy is to move the embedder and the index onto the machine that already serves the model, and keep a thin client here. It was costed and declined; the numbers are worth keeping, because the idea comes back every time someone looks at the image size.
What it would save, measured on a development workstation:
Removed from the local install |
Size |
|---|---|
|
4.0 GB |
the bge-m3 weights |
4.3 GB |
the ChromaDB index |
4.3 GB |
total, leaving ~0.2 GB of client |
12.6 GB |
The image would fall from 17.1 GB to roughly 5 GB, and the ten-second embedder
load on the first query of a session would disappear, since the model would
stay resident server-side on a GPU. Latency is not the objection: a round trip
through the SSH tunnel that is already open measures 17 ms, against ~100 ms
for a local CPU embedding. (The existing deploy/embed_worker.py could not
be used as it stands: it opens one SSH connection per batch, 350 ms, which
is right for indexing 45 000 chunks once and wrong for three embeddings a
turn. It would need a persistent service behind the tunnel.)
It was declined for three reasons, in increasing order of weight:
It ends offline operation. Retrieval is what turns 18 % into 90 % on the build-system questions. Making it require a VPN and a reachable GPU host means a session on a train is not a degraded session, it is a different assistant.
It empties the container of its purpose. The image exists so that retrieval works the moment it starts, on a machine that may have no Hugging Face access at all. A 5 GB image that needs a tunnel to answer anything is a different product, not a smaller one.
A GPU host is often a shared login. Collections are named from the corpus’s absolute path, so two people indexing the same path under one account would write into the same collection. Fixing that means per-user prefixes and a shared ChromaDB on a filesystem other users already fill.
The work itself is small – about a day: a retrieve endpoint doing embedding
and query in one round trip, a second port forward in spear-chat.sh, ten
PersistentClient call sites and five embedding ones. Cheapness was never
the question.