AI Cert Prep
Type to search documentation.

API Developer Path

D4 · Designing and Building Agentic Systems

The three agent runtimes, sessions and environments, hosted versus self-hosted sandboxes, tools and MCP, multi-agent and subagents, guardrails and approvals, tracing, and the Agents API constraints.

This domain shares top weight with the Responses API – roughly 11 of 60 items (18%) – and mirrors the Academy Design and Build Agentic Systems course (110 min, the longest API course). It tests whether you can choose among the three agent runtimes, give an agent tools and boundaries, orchestrate multiple agents or subagents, add guardrails and approvals, observe it with tracing, and respect the constraints of the managed Agents API. The single most-tested distinction is which runtime for which problem.

What you need to know

An agent is a model with instructions, tools and a loop that lets it act, observe and continue toward a goal. OpenAI gives you three ways to run one: the Responses API + tools where you own the loop; the open-source Agents SDK you run yourself for code-first orchestration; and the Agents API where OpenAI runs a managed Codex harness – sessions, orchestration, context compaction and recovery – while you supply tools and pick the environment. Agents work in sessions inside an environment (hosted or self-hosted sandbox), use tools including MCP connections, can delegate to subagents, and need guardrails, approvals and tracing. The Agents API is US-data-residency-only with no ZDR, and a self-hosted sandbox does not change that.

Learning objectives

By the end of this page you should be able to:

  1. Choose the right runtime – Responses + tools, Agents SDK, or Agents API – for a given problem.
  2. Explain sessions, environments and hosted vs self-hosted sandboxes.
  3. Give an agent tools, including MCP connections, safely.
  4. Design multi-agent systems and subagent delegation.
  5. Add guardrails, approvals, and human-in-the-loop controls.
  6. Use tracing and observability, and respect the Agents API constraints.

4.1 The three runtimes

This is the heart of the domain. Learn the table cold.

RuntimeWhat it isChoose it when
Responses API + toolsYou own the loop, call by callSimple tool-using assistants, full control, no session/sandbox needs
Agents SDKOpen-source Python/TS framework you run: agent definitions, orchestration, guardrails, sandboxing, tracing, agent evalsYou want code-first control, custom orchestration, self-hosted execution
Agents API (/v1/agents/sessions, header OpenAI-Beta: agents=v1, beta)OpenAI-managed Codex harness: OpenAI runs sessions, orchestration, context compaction, recovery; you supply tools and pick the environmentDurable cloud agents, hosted sandboxes, long-running work, artifacts
text
Do you need OpenAI to run and recover a durable, long session for you?
│
┌────┴─────────────────────────┐
│ Yes │ No
▼ ▼
Agents API (managed harness) Do you want a framework for
hosted/self-hosted sandbox, code-first orchestration you run?
subagents, compaction │
┌───────┴────────┐
│ Yes │ No
▼ ▼
Agents SDK Responses API + tools
(you run it) (you own the loop)

Assessment signal

“OpenAI runs the session / recovers it / manages compaction” → Agents API. “Open-source framework we run / custom orchestration in our code” → Agents SDK. “Simple assistant, we control each call” → Responses + tools.

4.2 Building an agent in the Agents SDK

python
from agents import Agent, Runner, function_tool
@function_tool
def lookup_order(order_id: str) -> dict:
return orders_db.get(order_id)
support_agent = Agent(
name="Support",
instructions="Help with orders. Use tools; never invent an order.",
model="gpt-5.6-terra",
tools=[lookup_order],
)
result = Runner.run_sync(support_agent, "Where is order 8842?")
print(result.final_output)

4.3 Starting an Agents API session

The Agents API is a beta surface; you opt in with a header and OpenAI runs the harness.

json
POST /v1/agents/sessions
OpenAI-Beta: agents=v1
{
"agent": {
"model": "gpt-5.6-terra",
"instructions": "Migrate the test suite to the new framework.",
"tools": [{ "type": "shell" }],
"mcp_servers": [{ "url": "https://mcp.internal/repo" }]
},
"environment": { "type": "hosted" },
"multi_agent": { "enabled": true, "max_concurrent_subagents": 3 }
}

You then follow progress via streaming or webhooks, and steer or continue the session. OpenAI handles orchestration, context summarisation, subagent delegation, and session resumption.

4.4 Sessions and environments

ConceptMeaning
AgentModel, instructions, tools, MCP servers
EnvironmentOpenAI-hosted or self-hosted sandbox: files/artifacts, lifecycle, security
SessionA durable instance: create → task → follow via streaming/webhooks → continue or steer
Events and itemsThe stream of what the agent did (tool calls, messages, results)

Hosted vs self-hosted sandbox: a hosted sandbox is OpenAI-run (billed at container rates) and simplest; a self-hosted sandbox runs in your infrastructure for control over the execution environment. Crucially, self-hosting the sandbox does not change the Agents API residency/ZDR constraints – see 4.8.

4.5 Tools and MCP

Agents act through tools: function calling, web search, file search, code interpreter, shell, computer use, image generation, apply patch, and MCP connections that expose external systems as tools. MCP lets an agent reach a database, a repo, or a SaaS system through a standard protocol instead of bespoke glue.

Least privilege for tools

Give an agent only the tools and MCP scopes it needs for its task. A support agent that only reads orders should not hold write access to billing. Over-provisioned tools are the largest agentic blast-radius risk.

4.6 Multi-agent and subagents

Two patterns solve different problems:

PatternShapeUse for
Multi-agent / handoffSpecialist agents pass control (triage → billing → refunds)Distinct skills, clear routing
SubagentsA lead agent delegates parallel sub-tasks (multi_agent: { enabled, max_concurrent_subagents })Parallelisable work under one goal

Subagents in the managed harness let the lead delegate concurrent work and gather results; keep max_concurrent_subagents bounded so cost and rate limits stay predictable.

4.7 Guardrails, approvals and tracing

ControlWhat it doesExample
Input/output guardrailsValidate or block unsafe inputs/outputsReject prompt-injection, PII leakage
Approvals (human-in-the-loop)Pause before a high-stakes actionConfirm before issuing a refund or deleting data
Tracing / observabilityRecord every step for debugging and auditInspect which tool call went wrong

An agent that can take irreversible actions needs an approval gate on those actions and tracing so you can see what it did and why.

4.8 Agents API constraints

Know these – they are frequent, high-value items.

  • US data residency only. The Agents API processes in the US; there is no non-US residency option today.
  • No ZDR (Zero Data Retention). ZDR is not available for the Agents API.
  • Self-hosting the sandbox does not lift these. Running your own sandbox controls the execution environment, not the API’s residency or retention terms.
  • Beta. It ships behind OpenAI-Beta: agents=v1; treat it as beta in production planning.
  • Billing. Model API rates + standard tool rates + container rates for hosted sandboxes.
text
If the requirement is EU-only residency or ZDR:
Agents API is NOT eligible (even self-hosted sandbox).
→ Use Agents SDK (you run it, your residency/retention) or
Responses + tools with your own controls.

Decision framework

Use RUNTIME-FIT to choose and secure an agent design.

StepQuestionIf yes
Managed durabilityDo you need OpenAI to run/recover long sessions?Agents API (check residency/ZDR first)
Residency/ZDRDo you need EU residency or ZDR?Not Agents API → SDK or Responses
Custom orchestrationDo you want code-first control of the loop?Agents SDK
SimplicitySimple tool assistant, you own each call?Responses + tools
Tool scopeWhat is the least privilege that works?Grant only needed tools/MCP scopes
High-stakes actionsAny irreversible action?Add an approval gate + tracing

The decisive early check is residency/ZDR: it can eliminate the Agents API before any other consideration.

Common mistakes

MistakeWhy it happensWhat to do instead
Using the Agents API where EU residency or ZDR is requiredManaged harness is attractiveIt is US-only, no ZDR; use the SDK or Responses instead
Assuming a self-hosted sandbox grants ZDR“Self-hosted = my rules”Self-hosting controls execution, not API residency/retention
Reaching for a runtime before scopingRuntimes are excitingChoose by durability, orchestration and simplicity needs
Over-provisioning agent toolsConvenienceGrant least-privilege tools and MCP scopes
No approval gate on irreversible actionsAutonomy feels like the pointGate refunds, deletes, sends behind human approval
No tracingIt works in the demoTrace every step for debugging and audit
Unbounded subagents“More parallelism is better”Bound max_concurrent_subagents for cost and rate limits
Building a framework by handNot knowing the SDK existsUse the Agents SDK for orchestration you would otherwise reinvent

Scenario challenge

Scenario. A European bank wants a durable agent that migrates a large legacy codebase over many hours, running shell commands and reading an internal repo via MCP, with progress dashboards. Data must remain in the EU and the bank’s policy requires zero data retention. An architect proposes the Agents API with a self-hosted sandbox “so the data stays in our EU infrastructure and ZDR is satisfied”.

Expert reasoning trace.

  1. Check residency and ZDR first. The bank requires EU residency and ZDR. The Agents API is US-only with no ZDR – that is a hard constraint, checked before anything else.
  2. Reject the self-hosted-sandbox reasoning. A self-hosted sandbox controls where execution happens, but the Agents API’s residency and retention terms are properties of the API, not the sandbox. So self-hosting does not make the Agents API EU-resident or ZDR-compliant. The architect’s proposal fails on a factual constraint.
  3. The durability need is real, but it must be met another way. The Agents SDK, which the bank runs in its own EU infrastructure, gives durable orchestration under the bank’s own residency and retention controls – the model calls still go to the API, so the bank must also confirm the model endpoint meets its data terms, but the harness is theirs.
  4. Keep least privilege. Shell and repo MCP access are powerful; scope the MCP connection to read-only where possible and gate any write/commit behind approval.
  5. Add tracing for the multi-hour run so the migration is auditable and recoverable.
  6. Consider subagents carefully: parallelising per-module migration is attractive, but in the SDK the bank owns that orchestration and its cost/rate implications.

Exam-correct decision: the Agents API is ineligible because of EU residency and ZDR (self-hosting the sandbox does not change that); build on the Agents SDK in the bank’s own infrastructure with least-privilege MCP, approval gates on writes, and full tracing. Not “Agents API with a self-hosted sandbox satisfies ZDR”.

Assessment traps

TrapWhy it is temptingThe discriminator
“Self-hosted sandbox gives you ZDR on the Agents API”Self-hosting sounds like controlSelf-hosting controls execution, not the API’s residency/retention
“Use the Agents API for the EU project”Managed harness is convenientAgents API is US-only, no ZDR
“Give the agent every tool for flexibility”Fewer blockersLeast privilege limits blast radius
“Autonomy means no human approvals”Full automation feels advancedIrreversible actions need an approval gate
“Build our own agent loop from scratch”Feels like controlThe Agents SDK is the code-first framework for this
“Responses + tools can’t do agents”Agents seem to need a special APIResponses + tools is a valid runtime when you own the loop
“More subagents is always faster”Parallelism intuitionUnbounded subagents blow cost and rate limits; bound them

Practice questions

Each item states how many responses to select. Commit before revealing.

Q1 · You want OpenAI to run and recover a durable, hours-long agent session with managed context compaction. Which runtime fits BEST? (Select one)

A. Responses API + tools. B. The Agents API with its managed Codex harness. C. Chat Completions. D. Fine-tuning.

Answer: B. The Agents API runs sessions, orchestration, compaction and recovery for you – exactly the managed-durability case. Responses + tools (A) means you own the loop, Chat Completions (C) is legacy, and fine-tuning (D) is not a runtime.

Q2 · A team wants code-first control of a custom multi-step orchestration they run on their own servers. Which runtime? (Select one)

A. The Agents API. B. The open-source Agents SDK. C. The Assistants API. D. Batch API.

Answer: B. The Agents SDK is the open-source framework for code-first orchestration you run yourself. The Agents API (A) is OpenAI-managed, the Assistants API (C) is legacy, and Batch (D) is for bulk offline jobs.

Q3 · An EU bank requires data residency in the EU and zero data retention. Is the Agents API eligible? (Select one)

A. Yes, if you enable EU mode. B. No; the Agents API is US-data-residency-only with no ZDR, and a self-hosted sandbox does not change that. C. Yes, as long as you use a self-hosted sandbox. D. Yes, ZDR is on by default.

Answer: B. The Agents API is US-only with no ZDR, and self-hosting the sandbox controls execution, not the API’s residency/retention. There is no EU mode (A), the self-hosted sandbox does not lift the constraint (C), and ZDR is not available (D).

Q4 · What is the correct way to let an agent read an internal repository through a standard protocol? (Select one)

A. Paste the whole repo into the prompt. B. Connect it via an MCP server exposed to the agent as a tool. C. Fine-tune on the repo. D. Email the files to OpenAI.

Answer: B. MCP connects external systems to an agent as tools through a standard protocol. Pasting the repo (A) is impractical and stale, fine-tuning (C) bakes in a snapshot, and emailing files (D) is not a mechanism.

Q5 · An agent can issue refunds. What control is REQUIRED before it acts? (Select one)

A. Higher reasoning effort. B. A human approval gate on the refund action, plus tracing of the decision. C. A larger context window. D. Nothing; the agent is trusted.

Answer: B. Irreversible, high-stakes actions need a human approval gate and tracing. Effort (A) and context (C) do not make an irreversible action safe, and blanket trust (D) removes the necessary control.

Q6 · Which TWO statements about the Agents API managed harness are TRUE? (Select two)

A. OpenAI runs sessions, orchestration, context compaction and recovery. B. It supports subagent delegation via multi_agent with max_concurrent_subagents. C. It guarantees EU data residency. D. It offers ZDR by default. E. It requires no beta header.

Answer: A and B. The managed harness runs the session lifecycle and supports bounded subagent delegation. It is US-only (contradicting C), has no ZDR (D), and requires OpenAI-Beta: agents=v1 (E).

Q7 · A simple assistant needs to call one internal function and return an answer, with full control over each step and no session or sandbox. Which runtime is simplest and correct? (Select one)

A. Agents API with a hosted sandbox. B. Responses API + tools, owning the loop yourself. C. Agents SDK with subagents. D. Fine-tuning.

Answer: B. For a simple tool-using assistant where you control each call and need no session/sandbox, Responses + tools is the right, lightest runtime. The Agents API (A) and SDK subagents (C) add machinery you do not need, and fine-tuning (D) is unrelated.

Q8 · What is the difference between multi-agent handoffs and subagents? (Select one)

A. They are identical. B. Handoffs pass control between specialist agents for routing; subagents let a lead agent delegate parallel sub-tasks under one goal. C. Handoffs are cheaper because they use no models. D. Subagents only exist in Chat Completions.

Answer: B. Handoffs route between specialists; subagents parallelise sub-tasks under a lead. They are not identical (A), handoffs still use models (C), and subagents are a harness/SDK feature, not a Chat Completions one (D).

Q9 · Why bound `max_concurrent_subagents` in a managed agent session? (Select one)

A. It is required syntactically or the call fails. B. To keep cost and rate-limit consumption predictable while still parallelising work. C. Because subagents cannot run in parallel at all. D. To disable tracing.

Answer: B. Bounding concurrency controls cost and rate-limit pressure while retaining parallelism. It is a tuning choice, not a syntactic requirement (A); subagents can run in parallel (C); and it has nothing to do with tracing (D).

Q10 · A support agent needs to read orders but a developer grants it write access to billing 'just in case'. What principle is violated and what is the fix? (Select one)

A. None; more access is safer. B. Least privilege; grant only the read-orders tool/scope the task needs. C. Statelessness; store nothing. D. Streaming; enable it.

Answer: B. Over-provisioning tools violates least privilege and enlarges blast radius; scope to what the task needs. More access is not safer (A), and statelessness (C) and streaming (D) are unrelated to the access issue.

Q11 · Which TWO capabilities make the Agents SDK a good fit for a team that must keep all processing under its own governance? (Select two)

A. It is open-source and runs in the team’s own infrastructure. B. It gives code-first control of orchestration, guardrails and sandboxing. C. It forces all data through OpenAI-managed US sessions. D. It removes the need to call any model. E. It disables observability.

Answer: A and B. The SDK is self-run and code-first, so the team controls orchestration and execution under its own governance. It does not force managed US sessions (C), it still calls models (D), and it includes observability rather than disabling it (E).

Q12 · An agent occasionally takes a wrong tool action and nobody can tell why afterwards. What is missing? (Select one)

A. A bigger model. B. Tracing/observability that records every step, tool call and result for inspection. C. More subagents. D. Lower reasoning effort.

Answer: B. Tracing records the agent’s steps so failures can be diagnosed and audited. A bigger model (A) does not explain past actions, subagents (C) add complexity, and lowering effort (D) does not add visibility.

Q13 · A hosted sandbox in the Agents API is billed how, relative to a self-hosted one? (Select one)

A. Both are entirely free. B. Model API rates plus standard tool rates plus container rates for the hosted sandbox; self-hosting shifts execution to your infrastructure. C. Only a flat monthly fee. D. Self-hosted sandboxes cost more than hosted ones in every case.

Answer: B. Agents API billing is model rates + tool rates + container rates for hosted sandboxes; self-hosting moves execution (and its cost) to you. Neither is free (A), there is no flat-fee-only model (C), and self-hosted is not universally more expensive (D).

Q14 · A prompt-injection payload in a retrieved document tries to make an agent exfiltrate data via a tool. Which control MOST directly addresses this? (Select one)

A. Increase the context window. B. Input/output guardrails that detect and block injected instructions and unsafe tool use. C. Switch to the Agents API. D. Add more subagents.

Answer: B. Guardrails validate inputs and outputs and can block injected instructions and unsafe tool actions. Context size (A), runtime choice (C) and subagents (D) do not by themselves stop prompt-injection-driven exfiltration.

Q15 · A startup wants the fastest path to a durable cloud agent with hosted execution and does not need EU residency or ZDR. Which runtime is the BEST fit? (Select one)

A. Responses + tools, hand-building session recovery. B. The Agents API with a hosted sandbox. C. Batch API. D. Fine-tuning a model.

Answer: B. With no residency/ZDR constraint, the Agents API’s managed durability and hosted sandbox are the fastest path. Hand-building recovery on Responses (A) reinvents the harness, Batch (C) is for offline bulk jobs, and fine-tuning (D) is not a runtime.

Q16 · You need specialist routing: a triage agent hands billing questions to a billing agent and technical ones to a tech agent. Which pattern is this? (Select one)

A. Subagent parallel delegation. B. Multi-agent handoff between specialists. C. Prompt caching. D. Structured outputs.

Answer: B. Passing control between specialist agents by topic is the multi-agent handoff pattern. Subagents (A) parallelise sub-tasks under one goal, and caching (C) and structured outputs (D) are unrelated features.

Key takeaways

  • Three runtimes: Responses + tools (you own the loop), Agents SDK (code-first, you run it), Agents API (OpenAI-managed Codex harness).
  • Choose by managed durability, custom orchestration and simplicity – and check residency/ZDR first.
  • The Agents API is US-only with no ZDR; a self-hosted sandbox controls execution, not those terms.
  • Agents act through tools including MCP; grant least privilege to limit blast radius.
  • Use multi-agent handoffs for routing and bounded subagents for parallel sub-tasks.
  • Gate irreversible actions behind human approvals and add guardrails against injection.
  • Always add tracing/observability so agent behaviour is auditable and debuggable.

Last updated Sep 18, 2026