AI Agents
Choosing What to Grade Inside an Agent Workflow
A credible final response is one signal, not proof that an agent followed instructions, chose the intended tool, extracted the right arguments, or transferred work at the right boundary.
Topic
11 articles
AI Agents
A credible final response is one signal, not proof that an agent followed instructions, chose the intended tool, extracted the right arguments, or transferred work at the right boundary.
AI Agents
Match the evaluation surface to the decision: inspect a trace to investigate an unclear failure, grade traces against defined workflow rules, and use datasets for repeatable change comparisons. Treat routing and handoff evidence as a prerequisite for adding agent complexity.
AI Agents
A proposed application contract for manager-to-specialist routing that makes ownership, context, tool scope, failure handling, and audit correlation explicit.
AI Agents
A practical architecture guide for separating durable orchestration from agent-loop control, handling duplicate execution risk, and designing retry-safe external actions.
AI Agents
A practical framework for using Agents SDK tracing, OpenTelemetry GenAI attributes, or both—while accounting for export timing, privacy, and version uncertainty.
AI Agents
The supplied excerpts expose documentation topics and navigation relationships, but not substantive procedures, examples, or implementation requirements.
AI Agents
A practical way to separate chat memory, server continuation, and local application state while planning recovery and security boundaries.
AI Agents
A practical design guide for adding human review to MCP-driven tools, with protocol-level elicitation and resumable approval workflows in the OpenAI Agents SDK.
AI Agents
A practical, layered approach to reducing prompt-injection risk when an LLM can read untrusted content and operate connected tools.
AI Agents
A practical guide to modeling JSON objects with named fields, extra-key controls, and name-pattern rules—plus implementation checks before applying them to a product API.
AI Agents
A practical way to separate background application jobs from model-driven tool orchestration, with approval checkpoints for sensitive actions.