6.3. Model backends
SPEAR talks to a model through a provider-neutral interface. Everything below the application layer — the runtime, the tool lifecycle, the guards — imports no provider and behaves identically whichever backend is selected.
Two providers are supported:
|
Talks to |
|---|---|
|
any OpenAI-compatible chat-completions endpoint (the default) |
|
the Anthropic API |
The OpenAI-compatible path covers a locally served model, a model served on another machine, and hosted endpoints that speak the same protocol. What changes between them is a URL.
Note
The coding core needs an OpenAI-compatible endpoint: it keeps its own OpenAI-format history and calls the endpoint directly. The Anthropic backend serves normative and general questions only.
6.3.1. Selecting a backend
Launched without a backend flag, an interactive session lists the configured
choices and preselects the one used last — remembered in
active-backend.conf in the state directory (SPEAR_STATE_DIR, else
client/). A flag skips the picker; off a terminal, the
last choice is reused silently.
$ spear-chat --local # the locally served model
$ spear-chat --remote # a remote endpoint over an SSH tunnel
$ spear-chat --reds # a configured shared GPU server
$ spear-chat --provider anthropic --model claude-sonnet-5
The remote flags open an SSH tunnel and point the client at the local end of
it. --reds does not start anything on the far side: the server there is
yours to start. --remote (a rented pod) pushes the pod’s serve script if it
is missing and starts it.
6.3.2. Pointing at an endpoint directly
Any endpoint can be addressed without touching a configuration file:
$ spear-chat --api-base https://inference.example.org/v1 \
--model qwen3 \
--ctx 32768
or through the environment, which is what client/machine.env is for:
# client/machine.env — untracked, this machine only
export SPEAR_API_BASE="http://localhost:8080/v1"
export SPEAR_MODEL_NAME="qwen3"
export SPEAR_CTX=32768
Note
SPEAR_MODEL_NAME matters for some servers and not others. A
llama.cpp server ignores the model name in the request; a vLLM server
matches it against its own --served-model-name and rejects a
mismatch.
6.3.3. The context window
--ctx (SPEAR_CTX) tells the client how much room it has. It is not a
request to the server — it is what the client budgets against, and setting it
above what the server actually serves produces requests the server refuses
whole.
When SPEAR_CTX is not set, the client asks the endpoint at startup —
llama.cpp’s /props (n_ctx, per slot) or vLLM’s /v1/models
(max_model_len) — and uses what it reports. If neither answers, it falls
back to 32768 and says so. The startup line names the source:
context window: 524288 (server /props)
context window: 32768 (default: server did not report a window — /props: URLError; ...)
Two settings interact with it:
Setting |
Effect |
|---|---|
|
cap on one reply (default 8192); reserved out of the window so a long edit completes instead of being truncated mid-call |
|
the fraction of the window the prompt may occupy before compaction (default 0.95) |
These two, and the sampling settings below, govern the normative and general
runtime. The coding core (Implementation mode) reserves its own
output budget, sends no sampling parameters of its own, and does not compact:
it stops at half the window and asks for a summary. One coding-core response
is capped at SPEAR_RESPONSE_MAX_TOKENS (default 16384), so a reply that
degenerates into a repeating tool call ends as truncated instead of running
to the full reservation.
6.3.4. Sampling
--temp (SPEAR_TEMP) defaults to 0.25. A low temperature gives more
conservative, more consistent edits; a model that falls into repetition at a
low temperature wants it raised rather than lowered.
Alongside the temperature, the normative and general runtime sends
top_p 0.8, top_k 20, a repetition penalty of 1.05 and disables the
model’s thinking mode. SPEAR_SAMPLING=server sends none of them, so the
server’s own configuration applies — which is what the coding core always
does.
6.3.5. Recording and replaying
Two flags make a backend optional:
$ spear-chat --record turns.jsonl # write down every model turn
$ spear-chat --replay turns.jsonl # answer from the recording
On replay, the tools, the files and the gates all run for real; only the model is a recording. A change to the platform can therefore be judged in minutes and at no inference cost. A turn that needs more rounds than were recorded simply ends.
Setting |
Meaning |
|---|---|
|
OpenAI-compatible endpoint URL |
|
|
|
model id sent with the request |
|
context window in tokens (default: asked of the server, else 32768) |
|
sampling temperature (default 0.25) |
|
cap on one reply (default 8192) |
|
cap on one coding-core response (default 16384) |
|
|
|
credential for the Anthropic provider |
Warning
Credentials belong in the environment or in an untracked
client/machine.env, never in a tracked file.
See also