Start with ownership, not routing
The first design decision is who is accountable for the next user-visible reply. In an SDK handoff, a specialist assumes conversational control for the selected branch. By contrast, an agent used as a tool remains a bounded helper while its manager retains responsibility for the reply.
For support triage, use a handoff when the runbook specialist should answer the customer directly. Use a manager-plus-helper arrangement when triage must combine the specialist’s research with other information before replying. This separation is a proposed application design, not a provider-prescribed implementation.
Do not make tool names or prompt text the authorization system. Treat them as declared intent, then independently enforce scope and permissions at each tool boundary where an external effect could occur.
Proposed handoff contract
The following is an illustrative, application-owned record. Its purpose is to make the delegation decision reviewable before a specialist receives work; it is not an SDK schema or an authority grant.
| Field | Illustrative type | Example value | Application validation |
|---|---|---|---|
| task_identity | string enum | support_runbook_lookup | Accept only registered task classes. |
| goal | string | Find a read-only recovery procedure for the reported symptom. | Require a bounded, customer-safe objective. |
| context_references | array of opaque IDs | ticket_4821, runbook_ref_17 | Resolve references under tenant and access checks; do not trust pasted identifiers. |
| allowed_tools | array of string enum | read_runbook | Compare every proposed call with this allowlist at the tool boundary. |
| delegated_scope | object | Read published support procedures; no infrastructure mutation. | Reject targets, arguments, or requested effects outside the declared scope. |
| response_owner | string enum | runbook_specialist | Require exactly one owner for the next customer response. |
| failure_destination | string enum | triage_router | Route malformed output, unavailable references, and denials to an explicit application handler. |
| correlation_id | opaque string | case_4821_route_03 | Create it at intake and attach it to application audit events and trace links. |
Failure routing and correlation format are deliberately local conventions. The contract should be validated before dispatch, retained with the routing decision, and checked again by any tool that could create a side effect.
Worked path: successful read-only delegation
Illustrative path: triage classifies ticket_4821 as a runbook lookup, creates the contract above, and transfers the branch to the runbook specialist. The specialist may use only the declared read-only capability, retrieves the approved procedure, and becomes the response owner for that branch. Its answer and the routing record share case_4821_route_03 so the application can join its own audit events to runtime traces.
Tracing can capture a structured account of model activity, tool activity, handoffs, guardrails, and custom spans. A correlation identifier complements that record; the identifier’s format and retention rules are application choices.
Worked path: denied production change
Illustrative path: the customer asks the specialist to apply a production configuration change. Scope validation finds that the requested effect exceeds read-only delegation. The application returns a denied outcome to triage_router, records the reason against case_4821_route_03, and does not invoke a production-change capability. Triage can provide a safe explanation or route the request into a separately authorized review workflow.
This denial is an application routing result, not evidence that a run occurred. For sensitive capabilities, place validation and any approval decision next to the tool that would cause the effect, rather than relying solely on an agent-level check.
Keep pauses separate from failures
A review pause is neither a fresh customer turn nor necessarily an error. When a tool needs approval, the run can return interruptions together with resumable state; after a decision, the application continues that same run from its saved state. Delayed review can therefore use stored serialized state, subject to the application’s access and expiry controls.
Failures need a separate route. This illustrative policy assumes a hypothetical category named uncertain_completion, defined as the application not having established whether its requested side effect completed. The category is application-owned and is not presented as an observed SDK outcome. Under this assumption, the design forbids automatic retry: preserve evidence, ask for reconciliation, and create a new authorized action only after the application can establish an appropriate next step.
Agent-level checks have limited placement: input checks apply at the initial agent, output checks apply at the agent producing the final output, and tool checks apply to the function tools carrying them. Attach critical controls to each relevant tool, including tools reached through nested specialist work.
Operational limits and alternatives
Begin with one agent when a separate contract does not improve capability separation, policy separation, prompt clarity, or trace readability. Add a specialist only when its branch has a materially different responsibility. If the manager must always synthesize the answer, agents as tools are the simpler ownership model; if a specialist must own the conversation branch, use a handoff.
This article uses conceptual SDK guidance within the stated observation scope and does not claim language-specific API stability, deployment validation, measured reliability, or a complete authorization model. The contract types, retry taxonomy, correlation format, and storage controls described here are application-owned choices requiring application-specific design and testing.
Boundary enforcement is necessary, but “sufficient” is stronger than the evidence supports. An
allowed_toolscheck alone does not establish that the declared scope constrains the actual operation: the boundary also needs a canonical proposed-action record and validation of target, action, arguments, calling identity, and applicable time window against that scope. The approved guidance explicitly calls for those per-call checks, not merely a tool-name decision.Make sufficiency testable with a repeatable negative-case matrix: identical delegated contracts should be rejected when only the target, argument-derived effect, caller identity, or engagement window becomes out of scope; the same cases should also be exercised through nested work. Record whether the tool was invoked and whether any side effect occurred. Without that evidence, the contract is an intent declaration with a boundary check, rather than a demonstrated authority constraint. Approved guidance