This guide is limited to the supplied Ollama excerpts. It separates response-level usage evidence from model-placement evidence so that a review does not turn a duration into a hardware claim.

Keep the two evidence streams separate

Usage responses describe loading, prompt handling, cache reads, and generation for a response. All duration values in those responses use nanoseconds. By contrast, ollama ps lists loaded models and provides a Processor value that describes their memory placement.

A sound review therefore retains both records: response telemetry answers what work the request reported, while the Processor observation addresses CPU, GPU, or split residency. A duration alone is not sufficient evidence for full GPU placement.

Four-case evidence decision table

Proposed application review matrix; the ownership and follow-up columns are local policy, not Ollama requirements.
CaseRelevant usage fieldsStreaming collection pointConclusion supported by those fieldsRequired placement evidence from ollama psInference to rejectEvidence ownerBounded application-owned action
model loadingload_durationStore it from the closing streamed fragment when done is true.Records the model-load portion reported for that response.Capture the model’s Processor entry separately.Do not treat load duration as proof of GPU residency.Request telemetry collector; placement observer.Mark the review incomplete if either record is absent.
uncached prompt evaluationprompt_eval_count and prompt_eval_durationStore it from the closing streamed fragment when done is true.Describes input-token evaluation reported for the response.Capture the model’s Processor entry separately.Do not infer a processor location from prompt evaluation timing.Request telemetry collector; placement observer.Keep phase classification separate from placement review.
cached prompt reuseprompt_eval_cached_count, alongside prompt_eval_countStore it from the closing streamed fragment when done is true.Identifies prompt tokens reported as read from cache.Capture the model’s Processor entry separately.Do not treat cache reuse as evidence of GPU residency.Request telemetry collector; placement observer.Record both counters when cache classification is needed.
output generationeval_count and eval_durationStore it from the closing streamed fragment when done is true.Describes reported output-token generation work.Capture the model’s Processor entry separately.Do not use eval_duration as proof of full GPU placement.Request telemetry collector; placement observer.Escalate conflicting records for human review.

Hypothetical review failure

Consider a deliberately hypothetical incident. A team files a request record and labels its eval_duration as evidence that the model had full GPU placement. Later, its separately captured ollama ps row shows 100% CPU. Under the documented mapping, 100% CPU denotes exclusive use of system memory for the loaded model.

  1. Keep eval_duration as output-generation telemetry for the request; it is still useful phase evidence.
  2. Record the Processor observation as independent placement evidence.
  3. Withdraw the full-GPU conclusion, because the placement record contradicts it.
  4. Flag the correlation between the request and the placement observation for reviewer assessment rather than silently reconciling the conflict.

This is a proposed handling rule for an illustrative scenario, not a provider-prescribed workflow.

Choose an operating approach

A telemetry-first approach is useful when the immediate question concerns request work and token activity. A placement-first approach is better when the immediate question is where loaded-model memory resides. In either approach, collect the other evidence stream before making a conclusion that crosses the boundary between phases and placement.

Proposed evidence policy and limits

The following controls are application choices. Retain the response identifier, relevant usage fields, stream-completion status, a placement observation, collection time, and the reviewer decision. Define a local rule for missing final streamed usage data, uncertain request-to-observation matching, and mixed CPU/GPU Processor values. Recheck this design when the supplied source material changes or when a release makes it incompatible.

The supplied excerpts do not establish a correlation mechanism between a particular response and a particular ollama ps observation. They also do not provide a performance benchmark, a hardware promise, or behavior for configurations outside the stated excerpt scope. Treat a mixed Processor result as placement evidence only; do not convert it into a performance conclusion.