Choose evidence by the decision at hand, rather than treating every evaluation artifact as interchangeable. This article proposes an application policy for evidence owners, gates, and escalation; it does not present those local choices as provider requirements. It uses the supplied OpenAI documentation excerpts as shared guidance, not independent corroboration, and does not recommend implementation through the deprecated Evals platform.

Start with the decision, not the artifact

The supplied agent guidance describes a trace as a record of one workflow run that can include model activity, tool activity, guardrails, and transfers between agents. That makes a trace useful when a failure is not yet understood. When the team can state the property to check, structured grading can score selected traces against that property. For change comparison that must be repeatable, move to a dataset and evaluation runs rather than relying on a single observed run.

This is a practical distinction in evidence strength, not a claim that one surface replaces another. Preserve diagnostic traces because they explain a concrete path; use a representative collection when the decision concerns behavior across cases.

Four-case evidence decision table

Proposed application policy for selecting workflow evidence
Decision contextEvaluation surfaceEvidence unitSuitable criterionConclusion supportedInference to rejectEvidence ownerBounded application-owned next step
An unexpected workflow failure is being investigated.Trace inspectionOne end-to-end workflow recordLocate the tool call, guardrail, or agent transfer associated with the symptom.A specific run provides a diagnosis lead.One inspected run proves the change is ready for release.Workflow engineerKeep the trace, state the suspected rule, and nominate cases for later evaluation.
A known workflow rule must be checked across selected runs.Trace gradingA set of traces scored with a defined ruleDid routing select the appropriate tool or make the expected transfer?The selected traces received scores for the declared rule.A score on selected traces establishes behavior for unrepresented inputs.Evaluation ownerReview the scoring rule and add failures to the candidate dataset.
A prompt, route, or workflow change requires a repeatable comparison.Dataset and evaluation runsA collection of test cases and their evaluation resultsCompare the changed workflow against explicit task metrics on representative cases.The comparison supports a bounded decision for the chosen dataset.A favorable result removes the need to inspect surprising failures.Release reviewerSet a local gate, record exceptions, and expand cases when new risks appear.
Added agent routing and handoffs are being considered.Handoff-focused traces plus dataset evaluationTraces and cases that exercise routing and transfersCheck whether the request reaches the suitable specialist and returns correctly when the topic changes.The team has targeted evidence for or against the proposed routing design.Additional agents are justified merely because specialization sounds useful.Architecture ownerWithhold adoption until local handoff criteria, escalation ownership, and representative cases are agreed.

The owner and next-step columns are deliberately local policy. The documentation supports the evaluation surfaces and example workflow questions, while the table's assignments, thresholds, and release handling are proposed operating choices.

Why handoffs need dedicated evidence

The supplied guidance identifies tool selection and extraction of tool arguments as evaluation dimensions for agents. It also notes that adding specialized agents introduces routing and transfer behavior that can vary between runs. Therefore, a team considering extra routing should create cases that test both the initial destination and a change of topic that requires a return or new transfer.

Use comparison-oriented judgments where possible: classification, a paired choice, or scoring against named conditions are better suited to a stated decision than an unconstrained request for a judge to generate commentary. If a model grades outputs, validate its agreement with human labels before making it a local release input; ordering and response length can bias model judges.

Hypothetical routing-failure walkthrough

This scenario is illustrative and reports no executed evaluation. A team changes its router after investigating a support conversation. One post-change trace reaches the intended specialist, so the team initially treats that trace as proof that the regression is resolved. That conclusion is too broad for the evidence unit.

The team then runs a repeatable, representative dataset evaluation that includes ordinary requests, topic switches, ambiguous wording, and adversarial routing prompts. In this hypothetical exercise, edge cases reveal failed handoffs. The team retains the successful trace as diagnostic context, pauses its release decision, adds those edge cases to the dataset, defines handoff criteria, and reruns the applicable evaluation.

The scenario is a proposed response pattern, not a documented OpenAI incident or a measured result. Its rationale is that a trace represents one run, whereas datasets and evaluation runs are the supplied guidance's surface for repeated comparisons over time.

Limits and review triggers

The supplied excerpts are sufficient to distinguish traces, graders, datasets, and evaluation runs, but they do not establish an application's case count, dataset mix, acceptance threshold, escalation path, or release authority. Those choices need local risk review. They also do not independently verify this article's recommendations; the evidence base is one provider's documentation family.

Revisit the policy when the source guidance changes or when a version retirement or incompatible release affects the chosen evaluation surface. For every workflow change, record which decision is being made, who owns the evidence, which cases are missing, and what conclusion the available evidence cannot support.