6.1. Setting up an inference host
This chapter builds, from an empty machine, the host that serves the model to
every SPEAR client: llama-server with the profile’s model, supervised by
systemd, reachable through SSH, with the embedding worker beside it. Follow
it in order; each step ends with a check, and the next one assumes it passed.
The figures quoted are those of the reference deployment. Rebuilding the runtime explains why the runtime is built the way it is; this page is the procedure.
6.1.1. What you are building
GPU host each client
───────────────────────────────────────── ───────────────────────
~/spear/ checkout (tens of MB) spear-chat --reds
~/spear-runtime/ │ SSH tunnel
llama.cpp/ pinned build │ 127.0.0.1:8082
models/gguf/ the weights, 79 GiB ▼
embed/ embedding worker venv ◄──── ssh … run-worker.sh
config/ server.conf, gpu.conf
log/
spear-inference.service ─► serve.sh ─► llama-server 127.0.0.1:<port>
The server listens on 127.0.0.1 only. Clients never reach it over the
network: they open an SSH tunnel, and the embedder is started per request
over SSH. Nothing on the host is exposed but sshd.
6.1.2. 1. Host prerequisites
Item |
Requirement |
Reference deployment |
|---|---|---|
GPU |
enough VRAM for the weights and the KV cache (below) |
RTX PRO 6000 Blackwell, 96 GB (97.9 GB usable) |
NVIDIA driver |
CUDA 13 capable (the embedder’s torch is a |
595.84 |
CUDA toolkit |
|
13.0, at |
OS |
a systemd Linux |
Ubuntu 24.04 |
Packages |
|
|
Disk, under the account’s home |
~100 GB: model 85 GB (79 GiB), embedder venv 5 GB, embedder weights 4.3 GB, llama.cpp build 1 GB |
|
RAM |
not critical once the weights are on the GPU |
46 GB |
Account |
one account that owns the checkout and the runtime |
see below |
Sizing the GPU. The profile serves Qwen3-Coder-Next Q8_0: 79 GiB of
weights. Its KV cache is small because only 12 of its 48 layers use full
attention, with 2 KV heads of 256: at q8_0 that is about 13 KB per token,
6.8 GB for the 524 288-token window. Measured: 89.5 of 97.9 GB in use at
512K. On a smaller card, lower SPEAR_SERVER_CTX (step 4); on a card that
cannot hold the weights at all, SPEAR_SERVER_NCPUMOE keeps the
mixture-of-experts weights in system RAM, at a large cost in speed.
The account. Everything below lives in one account’s home and needs no
root, except installing the system unit (step 6). On a machine shared with
other people under the same account, the isolation is a matter of
discipline rather than permissions: pin the card by UUID (step 3), keep every
cache inside ~/spear-runtime (step 7), and never stop a process you did
not start by pattern — pkill -f matches other people’s command lines, and
its own.
Check:
$ nvidia-smi --query-gpu=name,uuid,driver_version,memory.total --format=csv
$ /usr/local/cuda/bin/nvcc --version | tail -1
$ df -h ~
6.1.3. 2. The checkout
$ git clone https://github.com/smartobjectoriented/spear.git ~/spear
$ cd ~/spear && git checkout <release>
Deploy a release, not whatever main holds (Release process).
The host runs its launcher from this checkout, so it must stay a clean
checkout of a known commit: never edit a file in it on the host. A launcher
that matches no commit is a deployment nobody can reproduce.
6.1.4. 3. The runtime
Find the UUID of the card this deployment may use, then let the bootstrap build everything the manifest describes:
$ nvidia-smi -L
GPU 0: NVIDIA RTX PRO 6000 Blackwell … (UUID: GPU-xxxxxxxx-…)
$ cd ~/spear
$ scripts/bootstrap-runtime.sh --root ~/spear-runtime --gpu GPU-xxxxxxxx-… --dry-run
$ scripts/bootstrap-runtime.sh --root ~/spear-runtime --gpu GPU-xxxxxxxx-…
The dry run says what it would do and writes nothing. The real run clones
and builds the pinned llama.cpp, downloads the four model shards (budget
about an hour per 60 GB), builds the embedder’s virtualenv, and writes
config/server.conf and config/gpu.conf. It is idempotent: interrupt
it, run it again, and it keeps what is already correct.
Pin the card, even on a single-GPU host. Unpinned, llama-server
spreads its layers over every visible card, which on a shared host means
somebody else’s.
Check, and prove the weights byte for byte once:
$ scripts/bootstrap-runtime.sh --root ~/spear-runtime --verify
$ scripts/bootstrap-runtime.sh --root ~/spear-runtime --verify --checksum
--verify exits 0 only when the runtime satisfies the manifest, and ends
with the exact command line the launcher would run.
6.1.5. 4. The configuration
~/spear-runtime/config/server.conf is generated once from the profile and
is the deployment’s from then on; no later bootstrap overwrites it. Every key
is described in server/config/server.conf.example. The ones to review:
SPEAR_SERVER_CTX=524288The context window. The model is trained at 262 144 tokens, so the profile serves twice that with YaRN, enabled by the next two keys. Lower it if the card is smaller: at or below 262 144 no scaling is used.
SPEAR_SERVER_NATIVE_CTX=262144andSPEAR_SERVER_ARCH=qwen3nextThe trained length and the model architecture. When
SPEAR_SERVER_CTXexceeds the first,serve.shadds--rope-scaling yarn --rope-scale 2 --yarn-orig-ctx 262144and overridesqwen3next.context_length. The override is not optional:llama-servercaps each slot at the trained length it reads from the GGUF, so without it a 512K request serves 256K and says so only in its log. YaRN is static — it applies to short prompts too — which is why it is switched on only past the trained length.SPEAR_SERVER_PORT=8010,SPEAR_SERVER_HOST=127.0.0.1Keep the host on loopback. Clients use the port in their tunnel (step 8).
SPEAR_SERVER_NCPUMOE=0Everything on the GPU. Unset, it is guessed from the model’s file name.
6.1.6. 5. First run, in the foreground
Get it serving by hand before supervising it — a unit that restarts a misconfigured server hides the error behind a restart loop:
$ SPEAR_SERVER_ROOT=~/spear-runtime ~/spear/server/inference/serve.sh
It prints the card it pinned and, past the trained length, the YaRN factor. Loading takes a minute or two. From a second shell:
$ until curl -sf http://127.0.0.1:8010/health >/dev/null; do sleep 2; done
$ curl -s http://127.0.0.1:8010/props | python3 -c \
'import sys,json;print(json.load(sys.stdin)["default_generation_settings"]["n_ctx"])'
524288
If that prints 262144, the architecture override is missing (step 4). Stop the server with Ctrl-C.
6.1.7. 6. Supervision
A system unit is the one that survives reboots and logouts without
changing how the machine treats the account. render-unit.sh adapts the
template for it — account, home, boot target — and prints the result:
$ ~/spear/server/inference/render-unit.sh --system --account "$USER" \
| sudo tee /etc/systemd/system/spear-inference.service >/dev/null
$ sudo systemctl daemon-reload
$ sudo systemctl enable --now spear-inference
Add --log ~/spear-runtime/log/serve.log to send the output to a file
instead of the journal. The unit holds no serving decision: model, port,
context and card all stay in config/, so changing one is an edit there and
a restart, never an edit of the unit.
A user unit (render-unit.sh --user into ~/.config/systemd/user/) needs
no root, but survives a logout only with loginctl enable-linger — a
decision for the host’s administrator. server/inference/README.md weighs
the two.
At boot the NVIDIA driver can come up after the service. serve.sh
therefore waits for the configured card (SPEAR_SERVER_GPU_WAIT, 300 s by
default) and exits rather than start llama-server without it: llama.cpp would
otherwise serve the whole model from the CPU, two orders of magnitude slower,
while /health answers ok. SPEAR_SERVER_NGL=0 serves from the CPU on
purpose.
active means started, not serving; wait on /health as in step 5.
Day to day:
$ sudo systemctl restart spear-inference
$ systemctl status spear-inference
$ sudo journalctl -u spear-inference -f # or tail the --log file
6.1.8. 7. The embedding worker
Clients offload corpus indexing to this host’s GPU. The bootstrap already
built the worker’s virtualenv and wrote its launcher,
~/spear-runtime/embed/run-worker.sh, which pins the same card and keeps
the embedder’s weights in ~/spear-runtime/embed/hf rather than in the
account-wide Hugging Face cache. It is not a daemon: each client request
starts it over SSH and it exits when done.
Prove it end to end — the real launcher, the real protocol, one text:
$ scripts/bootstrap-runtime.sh --root ~/spear-runtime --verify --probe
The first probe downloads BAAI/bge-m3 (4.3 GB). The client side of
the worker is in step 8; the protocol is in server/embed/README.md.
6.1.9. 8. Client access
Each colleague needs an SSH key accepted by the host account
(~/.ssh/authorized_keys) and a host entry on their machine:
# ~/.ssh/config
Host spear-host
HostName <host address>
User <account>
IdentityFile ~/.ssh/<key>
IdentitiesOnly yes
IdentitiesOnly matters on hosts with a low MaxAuthTries: without it
ssh offers every key it has and is disconnected before the right one.
The model. In the client checkout, client/reds.conf (from
reds.conf.example):
REDS_HOST=spear-host
REDS_PORT=8010
REDS_MODEL=qwen3
spear-chat --reds then opens the tunnel 127.0.0.1:8082 → host:8010
itself and checks that something answers. A container client uses the same
tunnel with SPEAR_API_BASE=http://127.0.0.1:8082/v1
(Running the public image). By hand:
$ ssh -o ExitOnForwardFailure=yes -L 8082:localhost:8010 -N -f spear-host
$ curl -s http://127.0.0.1:8082/v1/models
The embedder. Two files in the client/ directory:
active-embed-remote.conf line 1: spear-host
line 2: extra ssh options, if any
active-embed-remote-cmd.conf ~/spear-runtime/embed/run-worker.sh
A leading ~/ is expanded on the host. Indexing (/reindex,
spear-index) then encodes on the host’s GPU.
6.1.10. 9. Verifying the whole
host $ scripts/bootstrap-runtime.sh --root ~/spear-runtime --verify --probe
host $ curl -s http://127.0.0.1:8010/props | grep -o '"n_ctx":[0-9]*'
client $ spear-chat --reds # the banner shows the model and the window
What the reference deployment measured after its move to 512K:
VRAM in use |
89.5 of 97.9 GB |
Generation |
156 tokens/s (157 at 98K: YaRN costs nothing measurable) |
Prompt processing |
~2 100 tokens/s — a 512K prompt takes about four minutes |
Retrieval past the trained length |
a fact placed 40 % into a 311 744-token prompt, recalled exactly |
6.1.11. 10. Updating and rolling back
An update is a new release of the checkout, then whatever the manifest now says, then a restart:
$ cd ~/spear && git fetch && git checkout <new-release>
$ scripts/bootstrap-runtime.sh --root ~/spear-runtime --dry-run
$ scripts/bootstrap-runtime.sh --root ~/spear-runtime
$ sudo systemctl restart spear-inference
The dry run shows whether the release moved the llama.cpp pin or the model;
if it did not, the bootstrap only verifies. config/ is never rewritten,
so a key a new release introduces (as SPEAR_SERVER_ARCH was) has to be
added by hand — its release notes say so, and --verify reports the command
line it would now produce.
Before editing server.conf, copy it: cp -p server.conf
server.conf.bak-<what>-<date>.
To roll back, check out the previous release, run the bootstrap (a moved pin is rebuilt, a kept one is left alone), restore the configuration backup if the change touched it, and restart. Weights of a previous model are not kept: a rollback across a model change downloads them again.
6.1.12. If the harness itself also runs on this host
Everything above is the server side. A host that also runs spear-chat
needs what any client needs to sandbox commands — unprivileged user namespaces
for bubblewrap (Ubuntu 24.04 restricts them through AppArmor and needs a
bwrap profile) and delegated cgroup controllers, which a plain SSH session
does not have. Operations lists the symptoms, and
Sandbox and Resource control the mechanisms.