Choose evidence by the decision at hand, rather than treating every evaluation artifact as interchangeable. This article proposes an application policy for evidence owners, gates, and escalation; it does not present those local choices as provider requirements. It uses the supplied OpenAI documentation excerpts as shared guidance, not independent corroboration, and does not recommend implementation through the deprecated Evals platform.
Start with the decision, not the artifact
The supplied agent guidance describes a trace as a record of one workflow run that can include model activity, tool activity, guardrails, and transfers between agents. That makes a trace useful when a failure is not yet understood. When the team can state the property to check, structured grading can score selected traces against that property. For change comparison that must be repeatable, move to a dataset and evaluation runs rather than relying on a single observed run.
This is a practical distinction in evidence strength, not a claim that one surface replaces another. Preserve diagnostic traces because they explain a concrete path; use a representative collection when the decision concerns behavior across cases.
Four-case evidence decision table
| Decision context | Evaluation surface | Evidence unit | Suitable criterion | Conclusion supported | Inference to reject | Evidence owner | Bounded application-owned next step |
|---|---|---|---|---|---|---|---|
| An unexpected workflow failure is being investigated. | Trace inspection | One end-to-end workflow record | Locate the tool call, guardrail, or agent transfer associated with the symptom. | A specific run provides a diagnosis lead. | One inspected run proves the change is ready for release. | Workflow engineer | Keep the trace, state the suspected rule, and nominate cases for later evaluation. |
| A known workflow rule must be checked across selected runs. | Trace grading | A set of traces scored with a defined rule | Did routing select the appropriate tool or make the expected transfer? | The selected traces received scores for the declared rule. | A score on selected traces establishes behavior for unrepresented inputs. | Evaluation owner | Review the scoring rule and add failures to the candidate dataset. |
| A prompt, route, or workflow change requires a repeatable comparison. | Dataset and evaluation runs | A collection of test cases and their evaluation results | Compare the changed workflow against explicit task metrics on representative cases. | The comparison supports a bounded decision for the chosen dataset. | A favorable result removes the need to inspect surprising failures. | Release reviewer | Set a local gate, record exceptions, and expand cases when new risks appear. |
| Added agent routing and handoffs are being considered. | Handoff-focused traces plus dataset evaluation | Traces and cases that exercise routing and transfers | Check whether the request reaches the suitable specialist and returns correctly when the topic changes. | The team has targeted evidence for or against the proposed routing design. | Additional agents are justified merely because specialization sounds useful. | Architecture owner | Withhold adoption until local handoff criteria, escalation ownership, and representative cases are agreed. |
The owner and next-step columns are deliberately local policy. The documentation supports the evaluation surfaces and example workflow questions, while the table's assignments, thresholds, and release handling are proposed operating choices.
Why handoffs need dedicated evidence
The supplied guidance identifies tool selection and extraction of tool arguments as evaluation dimensions for agents. It also notes that adding specialized agents introduces routing and transfer behavior that can vary between runs. Therefore, a team considering extra routing should create cases that test both the initial destination and a change of topic that requires a return or new transfer.
Use comparison-oriented judgments where possible: classification, a paired choice, or scoring against named conditions are better suited to a stated decision than an unconstrained request for a judge to generate commentary. If a model grades outputs, validate its agreement with human labels before making it a local release input; ordering and response length can bias model judges.
Hypothetical routing-failure walkthrough
This scenario is illustrative and reports no executed evaluation. A team changes its router after investigating a support conversation. One post-change trace reaches the intended specialist, so the team initially treats that trace as proof that the regression is resolved. That conclusion is too broad for the evidence unit.
The team then runs a repeatable, representative dataset evaluation that includes ordinary requests, topic switches, ambiguous wording, and adversarial routing prompts. In this hypothetical exercise, edge cases reveal failed handoffs. The team retains the successful trace as diagnostic context, pauses its release decision, adds those edge cases to the dataset, defines handoff criteria, and reruns the applicable evaluation.
The scenario is a proposed response pattern, not a documented OpenAI incident or a measured result. Its rationale is that a trace represents one run, whereas datasets and evaluation runs are the supplied guidance's surface for repeated comparisons over time.
Limits and review triggers
The supplied excerpts are sufficient to distinguish traces, graders, datasets, and evaluation runs, but they do not establish an application's case count, dataset mix, acceptance threshold, escalation path, or release authority. Those choices need local risk review. They also do not independently verify this article's recommendations; the evidence base is one provider's documentation family.
Revisit the policy when the source guidance changes or when a version retirement or incompatible release affects the chosen evaluation surface. For every workflow change, record which decision is being made, who owns the evidence, which cases are missing, and what conclusion the available evidence cannot support.
The single-trace limit is right, but should the release gate also reject a dataset result that covers only successful routing? The supplied guidance defines a trace as one end-to-end run and directs teams to datasets and eval runs for repeatable comparisons; neither surface alone demonstrates recovery from an ambiguous side effect. The evaluation guide
Consider a routed tool call where the downstream system commits an operation, but the workflow records a timeout before receiving the result. A successful post-change trace diagnoses the intended route; a representative routing dataset may show high selection accuracy. Neither result establishes that retrying will avoid a duplicate operation or that an unknown outcome will be detected and reconciled.
For release readiness involving consequential tools, add a recovery case alongside routing cases: inject completion loss after the external commit, require a stable operation identity, verify reconciliation against authoritative downstream state, and assert either one intended side effect or an explicit
unknownoutcome with dispatch halted. This makes the boundary clearer: traces identify the suspected path, datasets compare behavior across cases, and an injected-failure test supplies the missing containment-and-recovery evidence.