7.1. Operations
What to look at when something misbehaves, and where each state is reported.
7.1.1. Health check
host $ systemctl status spear-inference
host $ curl -s http://127.0.0.1:8010/health
{"status":"ok"}
host $ ~/spear/scripts/bootstrap-runtime.sh --root ~/spear-runtime --verify
client $ curl -s http://127.0.0.1:8082/v1/models | head -c 200 # through the tunnel
client $ cd ~/spear/client && ./bin/python -m unittest tests.test_tool_runtime tests.test_command_policy tests.test_sandbox
The server’s side is Setting up an inference host; its step 9 is the full check.
7.1.2. Reading a failure
Every harness failure surfaces as a ToolResult whose summary names the
mechanism that refused. The vocabulary is deliberately narrow.
Summary contains |
Meaning and first thing to check |
|---|---|
|
|
|
Preflight ran and failed. Usually unprivileged user namespaces are
disabled: |
|
|
|
Cgroup limits active, |
|
No |
|
|
|
|
|
The helper is too old. There is no fallback; see Network backend. |
|
The backend is poisoned: a previous run could not confirm the sandbox child died. Restart the chat session. |
|
The pin or the type check failed. The command never ran. |
|
The helper did not signal within |
|
Wall-clock timeout. The scope was killed as a tree. |
7.1.2.1. No user bus
Failed to connect to bus: No medium found
The harness fails closed, by design. If SPEAR must run without an
interactive session, the deployment decision is
loginctl enable-linger <user> — taken by an administrator. The runtime
never does it (Resource control).
7.1.3. Inspecting a live scope
Scopes are short-lived and --collect removes them promptly, so catching one
alive means sampling while the command runs:
$ watch -n0.1 "systemctl --user list-units --all 'spear-tool-*' --no-legend"
# from a known bwrap pid
$ cat /proc/<pid>/cgroup
$ base=/sys/fs/cgroup$(cut -d: -f3 /proc/<pid>/cgroup)
$ cat $base/memory.peak $base/pids.peak $base/cpu.stat
memory.peak and pids.peak are monotonic high-water marks, so the last
read before the cgroup disappears is the meaningful one.
7.1.4. Cleaning up after an interrupted session
$ systemctl --user list-units --all 'spear-tool-*' --no-legend
$ systemctl --user kill --kill-whom=all --signal=KILL <unit>
$ pgrep -a bwrap
$ pgrep -a slirp4netns
An orphan bwrap blocked on its block-fd is the signature of a
supervisor that died between spawning the sandbox and releasing the command.
It is harmless — the command never started — but it holds a namespace, so kill
it.
7.1.5. Switching model
host $ cp -p ~/spear-runtime/config/server.conf{,.bak-$(date +%F)}
host $ $EDITOR ~/spear-runtime/config/server.conf # SPEAR_SERVER_MODEL / _LORA
host $ sudo systemctl restart spear-inference
Or, for one foreground run without persisting:
host $ SPEAR_SERVER_MODEL=/path/to/other.gguf SPEAR_SERVER_ROOT=~/spear-runtime \
~/spear/server/inference/serve.sh
7.1.6. Logs and audit
The paths below are relative to the state directory: SPEAR_STATE_DIR, or
client/ when it is not set.
- Server log
Server-side: model load, context, slot activity. In the journal (
sudo journalctl -u spear-inference), or in the file the unit was rendered with--log.audit/tool-actions.jsonlOne record per mutating tool attempt. Metadata only — no file contents, no command output, no environment values. This is the file to read to answer “what did the assistant change”, and it is intentionally useless for answering “what was in it”.
history*.jsonConversation transcripts, per project.
7.1.6.1. Runtime tracing
Tracing is disabled by default. Enable provider-neutral JSONL traces for a benchmark run with:
$ spear-chat --trace # or SPEAR_TRACE=1
Events are appended to audit/runtime-trace.jsonl; --trace-file
(SPEAR_TRACE_FILE) chooses another location. Traces record timing, counts, provider and model
identifiers, normalized outcomes and safe tool metadata. They do not record
raw prompts, model responses, command strings, tool content, query or note
values, environment variables, or authorization data. Workspace-relative file
paths and tool argument names are kept, because without them a trace cannot
be used to debug or to analyse a benchmark run.
7.1.7. Known environment quirks
Symptom |
Explanation |
|---|---|
|
Expected. |
A tool needs a file from the host |
It will not find it. Only |
|
a non-snap build exporting headless can drop the labels. Export with
the drawio snap, staged under |
Network tests flaky under heavy load |
the slirp readiness wait is a wall-clock timeout; the attachment itself is timing-independent (Network backend). |
Suspend breaks a running CUDA job |
unrelated to the harness: do not let the machine suspend during a fine-tuning run. |
7.1.7.1. Training-data capture
Normal harness tasks capture a provider-neutral, redacted training episode in
$SPEAR_STATE_DIR/audit/training-data. Capture is enabled by default and
can be disabled with SPEAR_TRAINING_CAPTURE=0. Drafts are updated only at
model/tool boundaries; finalized episodes are immutable, checksummed JSON with
an append-only metadata manifest. These records are source evidence for a
future selective exporter, not ready-to-train SFT or preference data.