6. The model and its weights
Where the answers come from: which model is served, how it is brought up, how it is trained on the harness’s own usage, and — the page that exists because the others would otherwise repeat its lessons — what the earlier attempts measured before this one was chosen.
- 6.1. Setting up an inference host
- 6.1.1. What you are building
- 6.1.2. 1. Host prerequisites
- 6.1.3. 2. The checkout
- 6.1.4. 3. The runtime
- 6.1.5. 4. The configuration
- 6.1.6. 5. First run, in the foreground
- 6.1.7. 6. Supervision
- 6.1.8. 7. The embedding worker
- 6.1.9. 8. Client access
- 6.1.10. 9. Verifying the whole
- 6.1.11. 10. Updating and rolling back
- 6.1.12. If the harness itself also runs on this host
- 6.2. Model serving
- 6.2.1. Current profile
- 6.2.2. The flags that matter
- 6.2.3. The launcher holds no flag values
- 6.2.4. Sampling
- 6.2.5. Changing the model or the adapter
- 6.2.6. Service placement
- 6.2.7. Backends
- 6.2.8. Prompt caching on the Anthropic path
- 6.2.9. Authenticating against Anthropic
- 6.2.10. The
stop_reasoncontract - 6.2.11. Smoke test
- 6.3. Model backends
- 6.4. Rebuilding the runtime
- 6.5. Training and fine-tuning
- 6.5.1. The five stages
- 6.5.2. Capture is wider than judging
- 6.5.3. Governance: where a sample may come from
- 6.5.4. Readiness is evidence, not a feeling
- 6.5.5. The frozen bundle
- 6.5.6. The MoE lesson
- 6.5.7. Will it fit? Measure it
- 6.5.8. The operator surface
- 6.5.9. Where it runs today
- 6.5.10. Related material
- 6.6. Model history and what it cost
Setting up an inference host is the procedure: an empty GPU machine to a supervised server with clients attached. Model serving covers the serving side: the quantised weights, the GPU/CPU split and the flags that decide throughput. Rebuilding the runtime is what has to be true before the first token, and what the harness does when it is not. Model backends covers the endpoints the client can talk to and how it is pointed at one.
Training and fine-tuning is the fine-tuning side, and it stands on its own: how the harness turns its own usage into a governed dataset, and what it takes to run a job on the card that is also serving the model.
Model history and what it cost is the record of what was tried and what it cost — including the experiment that rented a B200 to learn something a smaller run would have shown. It is kept because a choice whose alternatives are forgotten gets re-litigated every year.