Separate the two decisions

Model residency is an operating-policy choice; response telemetry is evidence about one completed response. Do not promote a model merely because a response appears inexpensive, and do not treat a residency setting as proof of a response-time outcome.

Ollama normally retains a model in memory for five minutes before unloading it. The keep_alive API parameter accepts a positive duration or seconds count for a bounded interval, a negative value for continued loading, and zero for unloading after generation.

For an application policy, first choose the retention mode that fits the workload you define, then retain the response evidence needed to review that decision. Workload labels, thresholds, evidence-retention periods, and remediation limits below are proposed application conventions, not provider-prescribed rules.

Interpret response telemetry at the right boundary

Usage responses can report overall duration, loading duration, prompt token counts, cached prompt token counts, prompt evaluation duration, generated-token count, and generation duration. The duration fields use nanoseconds.

With a streaming response, collect usage from the terminal item identified by done being true. An earlier item is not a sound basis for converting absent metrics into a numeric zero.

The ollama ps display provides a separate check of whether a model is presently loaded. Its processor output can distinguish full GPU placement, full system-memory placement, and split CPU/GPU placement; keep this state check separate from per-response telemetry.

Four policy choices in one review table

Illustrative application decision record: four residency cases
Workload assumptionResidency choiceEvidence to retainMetric-collection pointollama ps verificationRejected inferenceBounded application-owned action
Requests are not classified as needing custom retention.Accept the default residency window.Terminal response metrics when available, plus the policy decision.For a stream, read the item where done is true.Check whether the model is currently listed in memory and inspect its processor placement.Do not infer that the five-minute default proves a particular request was fast.Keep the default while the application gathers review evidence.
The application expects a defined reuse interval.Choose a bounded explicit duration.Chosen interval, terminal metrics, and review rationale.Record telemetry only after response completion.Confirm current loading separately when needed.Do not infer that a configured duration guarantees lower latency.Revisit the selected duration at the application's scheduled review.
The application deliberately requires continued availability.Keep a model loaded with a negative value.Approval rationale and terminal response observations.Use the final streaming item, not an intermediate item.Observe current residency and CPU/GPU placement with ollama ps.Do not infer that continued loading demonstrates a small load duration for every response.Apply a locally defined review boundary before continuing the exception.
The application treats the work as one-off.Unload immediately with zero.Policy rationale and available terminal metrics.Capture the final item when it declares done true.Use a later ollama ps observation to inspect current state.Do not treat configured intent as an ollama ps observation.Use the zero setting unless a documented local review changes the policy.

The rows are deliberately decision records, not performance predictions. The listed actions bound what the application does while keeping a distinction between configured intent, a completed response's measurements, and the model state observed at a particular moment.

Hypothetical failure walkthrough

Consider an illustrative collector that reads an early streaming item, maps absent usage fields to zero load duration, and changes its local policy to indefinite residency. Later, an operator checks ollama ps and sees that the model remains loaded. This is a hypothetical process failure, not a reported Ollama incident or a measurement.

The correction is procedural: wait for the item whose done value is true, record unavailable telemetry as unavailable rather than zero, and temporarily return to a bounded local residency rule pending review. This response is an application-owned safeguard; it is not a claim that Ollama selects that fallback.

Limits and review triggers

This guidance is limited to the supplied FAQ and usage excerpts. It makes no release-specific compatibility assertion, no benchmark claim, and no latency guarantee. Reassess the design when the official source material changes, a relevant version is retired, or a target release behaves incompatibly.

Decide locally how to classify demand, how long to retain telemetry, what privacy controls apply, what to do after an interrupted stream, and which bounded fallback duration is appropriate. Treat those matters as application-policy choices requiring application context.