AI Cert Prep
Type to search documentation.

API Developer Path

D2 · The Responses API and Model Selection

The Responses API request and response shape, conversation state, streaming, background mode, structured outputs, function calling, reasoning effort, file inputs, compaction, token counting and choosing the right model.

This is the heaviest domain on the OAI-API mock alongside agentic systems – roughly 11 of 60 items (18%). It tests whether you can drive the Responses API fluently: build a request, read the output items, keep conversation state, stream, run work in the background, get structured JSON, call functions, set reasoning effort, pass files, manage long context with compaction, count tokens, and pick the cheapest model that meets the bar. The Responses API is the primary interface for new work; Chat Completions is legacy.

What you need to know

The Responses API takes an input (a string or a list of typed items) and returns a response whose output is an ordered list of items – messages, function_calls, reasoning, and tool results. State is either stateless (you resend history) or server-held (store: true plus previous_response_id). You choose a reasoning effort per request and follow the rule use the lowest effort that gets the result. You get machine-readable output with structured outputs (JSON Schema), extend the model with function calling and hosted tools, run slow work in background mode, and keep long conversations inside the context window with compaction. Model selection is a four-way choice – Astra, Sol, Terra, Luna – driven by task difficulty, volume and cost.

Learning objectives

By the end of this page you should be able to:

  1. Construct a Responses API request and parse the output item list.
  2. Manage conversation state statelessly and with previous_response_id.
  3. Use streaming and background mode appropriately.
  4. Produce structured outputs and handle function calling end to end.
  5. Set reasoning effort by the lowest-effort-that-works rule.
  6. Select a model from Astra / Sol / Terra / Luna using price, context and difficulty.

2.1 The request and response shape

A minimal call names a model and an input and reads output_text.

python
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-5.6-terra",
input="Summarise this ticket in one sentence: ...",
)
print(resp.output_text) # convenience accessor for the text

The full response carries an ordered output array. Do not assume item zero is your text – reasoning items, tool calls and messages can interleave.

text
resp.output = [
{ type: "reasoning", ... }, # may appear for reasoning models
{ type: "function_call", name, args }, # if the model called a tool
{ type: "message", role: "assistant",
content: [ { type: "output_text", text: "..." } ] }
]

Assessment signal

Stems mentioning output_text, output items, “the primary interface”, or “which API for new development” point here. The Responses API is correct for new work; Chat Completions is the legacy distractor.

2.2 Conversation state

You have two ways to keep a multi-turn conversation, and the choice has cost and privacy consequences.

ApproachHowTrade-off
StatelessResend the whole history in input each turnFull control, no server retention; you pay input tokens for the whole history every turn
Server-heldstore: true, then previous_response_id on the next callLess data resent, simpler; response is retained server-side
python
first = client.responses.create(
model="gpt-5.6-terra", input="My name is Sam.", store=True,
)
second = client.responses.create(
model="gpt-5.6-terra",
input="What did I say my name was?",
previous_response_id=first.id,
)

2.3 Streaming and background mode

text
latency need
│
┌────────────────┼─────────────────┐
▼ ▼ ▼
User is waiting Long task, user Fire-and-forget,
for tokens can wait a bit hours-long
│ │ │
stream=True stream, or background=True
(token-by- background + + poll or webhook
token UI) webhook for completion
  • Streaming (stream=True) emits server-sent events so a UI shows tokens as they arrive – use it whenever a human waits.
  • Background mode (background=True) starts the response and returns immediately; you poll the response ID or receive a webhook when it finishes. Use it for long reasoning or agentic runs that outlive an HTTP request.

2.4 Structured outputs

When another system consumes the output, force a schema rather than parsing prose.

python
schema = {
"type": "object",
"properties": {
"queue": {"type": "string", "enum": ["billing", "tech", "sales"]},
"priority": {"type": "integer", "minimum": 1, "maximum": 4},
},
"required": ["queue", "priority"],
"additionalProperties": False,
}
resp = client.responses.create(
model="gpt-5.6-luna",
input="Route this ticket: 'I was charged twice this month.'",
text={"format": {"type": "json_schema", "name": "routing",
"schema": schema, "strict": True}},
)

strict: True guarantees the output validates against the schema – you no longer defensively parse. This is the right pattern for classification, extraction and any machine-consumed result.

2.5 Function calling

Function calling lets the model request that your code run, then continue with the result. The loop is: model emits a function_call → you execute it → you return a function_call_output referencing the call_id.

python
tools = [{
"type": "function",
"name": "get_weather",
"description": "Get current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"], "additionalProperties": False,
},
}]
resp = client.responses.create(
model="gpt-5.6-terra", input="What's the weather in Oslo?", tools=tools,
)
call = next(o for o in resp.output if o.type == "function_call")
result = get_weather(**json.loads(call.arguments)) # your code runs
final = client.responses.create(
model="gpt-5.6-terra",
previous_response_id=resp.id,
input=[{"type": "function_call_output",
"call_id": call.call_id, "output": json.dumps(result)}],
)

Assessment signal

“The model needs live/private data”, “call an internal system”, or “take an action” point to function calling. Options that say “fine-tune the data in” or “paste the data every request” are the distractors.

2.6 Reasoning effort and the lowest-effort rule

The GPT-5.6 family exposes reasoning efforts from none to max; Astra adds xhigh. Higher effort means more internal reasoning tokens – better on hard problems, but slower and more expensive (you pay for reasoning tokens as output).

EffortUse forCost/latency
none / lowExtraction, classification, formatting, simple Q&ACheapest, fastest
mediumGeneral reasoning, most agentic stepsBalanced
high / xhigh / maxHard multi-step reasoning, planning, tricky codeSlowest, priciest

The rule from the docs: use the lowest reasoning effort that gets the result. There is no exact mapping from old GPT-5.5 efforts to GPT-5.6 – re-tune per task with your eval, do not assume.

2.7 File inputs, compaction and token counting

  • File inputs: pass PDFs and images as input items (file uploads or base64/URL for images) so the model reads them directly; you do not paste raw text you already have as a file.
  • Compaction: long conversations eventually approach the 1.05M-token window. Compaction summarises earlier turns so the thread continues without losing the plot; the Agents API does this automatically, and you can do it yourself in a Responses loop by replacing old turns with a summary.
  • Token counting: every response reports usage (input_tokens, output_tokens, total_tokens). Use it to track cost and to know how close you are to the context limit.
text
Context budget (1.05M window)
┌───────────────────────────────────────────────────────────┐
│ system + tools │ retrieved context │ history │ headroom │
└───────────────────────────────────────────────────────────┘
as history grows → compact old turns → keep headroom for output

2.8 The model line-up (September 2026)

All four latest models: text + image input, text output, multilingual, vision, 1.05M context, 128K max output, served through the Responses API.

ModelIDReasoning effortsInput $/MTokOutput $/MTokContextCutoff
GPT-6 Astragpt-6-astralow…max, plus xhigh$10$501.05M30 Apr 2026
GPT-5.6 Solgpt-5.6-sol (alias gpt-5.6)none…max$4$201.05M16 Feb 2026
GPT-5.6 Terragpt-5.6-terranone…max$2$121.05M16 Feb 2026
GPT-5.6 Lunagpt-5.6-lunanone…max$0.20$1.201.05M16 Feb 2026

Previous generation still available: gpt-5.5 ($5 / $30) and gpt-5.5-pro ($30 / $180).

Selection guidance, from the docs: Astra for the hardest end-to-end work (sustained reasoning, judgment, multi-tool); Sol for complex, open-ended, high-value work; Terra the pragmatic all-rounder and natural replacement for GPT-5.5 workloads; Luna for clear, repeatable, high-volume tasks (extraction, classification, transformation, structured summaries).

Decision framework

Use the LADDER for model-and-effort selection: start low, climb only when an eval proves you must.

RungChoiceMove up when
Luna, low effortHigh-volume, well-defined tasksThe eval shows quality below target
A step to TerraGeneral work, GPT-5.5 replacementsTerra misses on genuinely complex reasoning
Dial effort upSame model, medium then highFailures are reasoning depth, not model class
Deploy SolComplex, open-ended, high-valueSol still misses the hardest end-to-end tasks
Escalate to AstraHardest sustained reasoning + multi-tool—
Re-measureAfter every change, re-run the evalNever assume; there is no fixed effort mapping

The habit the course prizes: change one thing, then re-measure. Jumping straight to Astra at max is the anti-pattern this framework exists to prevent.

Common mistakes

MistakeWhy it happensWhat to do instead
Defaulting to Astra at max“Best model, no risk”Start on the LADDER’s lowest rung and climb with eval evidence
Using Chat Completions for new workFamiliarity, old tutorialsUse the Responses API; Chat Completions is legacy
Assuming output[0] is the textIt usually is in demosIterate output; reasoning and tool-call items interleave
Resending full history every turn on huge threadsSimplest to codeUse store + previous_response_id, or compact old turns
Parsing JSON out of proseHabit from older modelsUse structured outputs with strict: true
Fine-tuning to inject live dataConfusing knowledge with data accessUse function calling / retrieval for live and private data
Copying GPT-5.5 effort levels to GPT-5.6Assumes a mapping existsRe-tune effort per task; there is no exact mapping
Blocking a UI on a long responseNot knowing streaming/background existStream for waiting users; background mode for long jobs

Scenario challenge

Scenario. Your team ships a customer-support assistant. It must (a) answer FAQs instantly in a chat widget, (b) look up the customer’s current plan from an internal billing service, (c) occasionally handle a gnarly multi-account billing dispute that needs careful reasoning, and (d) return a structured {intent, plan, next_action} object to the surrounding app. A junior engineer proposes: “Use gpt-6-astra at max for everything, resend the full transcript each turn, and parse the JSON out of the reply text.”

Expert reasoning trace.

  1. Split the workload by difficulty. Most turns are FAQs and lookups – well-defined, high-volume work that gpt-5.6-luna or gpt-5.6-terra handles at low/medium effort. Only the rare multi-account dispute needs escalation. One-model-fits-all at max burns roughly 50× the input price for the common path.
  2. Live plan data is a function call, not model knowledge. The billing lookup must be a function_call to the internal service; recalling it or fine-tuning it in would be stale and unverifiable.
  3. Structured output, not prose parsing. The app consumes {intent, plan, next_action}, so use text.format with a JSON Schema and strict: true. Parsing prose is fragile and the junior’s plan would break the first time the model adds a friendly sentence.
  4. State: don’t resend everything. For a chat widget, store: true + previous_response_id keeps input tokens down; for long threads, compact older turns.
  5. Latency: stream the FAQ answers so the widget feels instant; use high effort only on the dispute path, ideally in background mode if the user can wait.
  6. Escalation is eval-driven. Route the dispute path to Sol (or Astra at high) only after an eval shows Terra fails it – LADDER, not default.

Exam-correct decision: route common turns to Luna/Terra at low effort with streaming, use function calling for the billing lookup, force structured output with a schema, keep state server-side, and escalate the rare hard case up the LADDER on eval evidence. Not Astra-at-max-for-everything, not prose JSON parsing, not full-transcript resends.

Assessment traps

TrapWhy it is temptingThe discriminator
“Use gpt-6-astra at max to be safe”Best model feels lowest-riskLowest effort that works; climb the LADDER with eval evidence
“Chat Completions is fine for the new service”It still worksResponses API is the primary interface for new development
“Fine-tune so it knows the customer’s plan”Sounds like teaching the modelLive/private data is a function call, not training
“Parse the JSON from the text reply”It often works in testingStructured outputs with strict: true guarantee the shape
“Map GPT-5.5 high to GPT-5.6 high”Symmetry looks safeNo exact mapping; re-tune with your eval
“Resend the whole transcript each turn”Simple and explicitprevious_response_id and compaction control cost on long threads
“Read output[0] for the answer”Works in the happy pathReasoning and tool-call items interleave; iterate output

Practice questions

Each item states how many responses to select. Commit before revealing.

Q1 · You are starting a new API integration. Which interface should you build on? (Select one)

A. The Chat Completions API, because it is the most established. B. The Responses API, the primary interface for new development. C. The Assistants API. D. The legacy Completions endpoint.

Answer: B. The Responses API is the primary interface for new work. Chat Completions (A) and the older Completions endpoint (D) are legacy for new development, and the Assistants API (C) is now grouped under Legacy APIs.

Q2 · A classifier must return `{queue, priority}` that the surrounding app parses reliably. What is the BEST way to guarantee the shape? (Select one)

A. Ask the model in the prompt to ‘only return JSON’ and parse the text. B. Use structured outputs with a JSON Schema and strict: true. C. Post-process the prose with a regular expression. D. Fine-tune the model to emit JSON.

Answer: B. Structured outputs with a strict JSON Schema guarantee the response validates against the shape. Prompt instructions (A) and regex (C) are not guaranteed, and fine-tuning (D) is unnecessary and does not guarantee validity.

Q3 · An assistant needs the customer's current subscription plan from an internal billing service to answer. What is the correct mechanism? (Select one)

A. Fine-tune the model on all customer plans nightly. B. Define a function tool the model calls, execute it, and return a function_call_output. C. Paste the entire customer database into every prompt. D. Rely on the model’s training knowledge.

Answer: B. Live, private data is accessed via function calling: the model requests the call, your code runs it, and you return the result. Nightly fine-tuning (A) is stale and costly, pasting the database (C) is expensive and non-authoritative, and training knowledge (D) does not contain live account data.

Q4 · A high-volume extraction task runs 2M times per day and `gpt-5.6-luna` passes your eval at 97%. Which model should you deploy? (Select one)

A. gpt-6-astra at max effort for maximum quality. B. gpt-5.6-luna, because it meets the metric at the lowest cost for repeatable high-volume work. C. gpt-5.6-sol to be safe. D. gpt-5.5-pro for reliability.

Answer: B. Luna is built for clear, repeatable, high-volume tasks and it passes the eval, so it is the lowest-cost model that works. Astra-at-max (A), Sol (C) and 5.5-pro (D) all cost far more per token with no measured quality gain.

Q5 · A user is waiting in a chat UI for an answer that takes several seconds to generate. What improves the experience MOST? (Select one)

A. Enable background mode and poll every 30 seconds. B. Stream the response so tokens appear as they are generated. C. Switch to gpt-6-astra at max effort. D. Increase the max output tokens.

Answer: B. Streaming shows tokens as they arrive, which is the right pattern when a human waits. Background mode (A) suits fire-and-forget long jobs, a bigger model (C) is slower not faster, and raising max output (D) does not change perceived latency.

Q6 · What is TRUE about reasoning effort on the GPT-5.6 family? (Select one)

A. GPT-5.5 effort levels map exactly onto GPT-5.6. B. You should use the lowest effort that produces the required result, re-tuning per task with an eval. C. Higher effort is always better and should be the default. D. Effort has no effect on cost or latency.

Answer: B. The rule is lowest effort that works, verified per task; there is no exact mapping from 5.5 (A). Higher effort costs more time and tokens (contradicting C and D).

Q7 · You are building a long-running research agent whose run outlives a single HTTP request. Which feature fits? (Select one)

A. stream=True on a synchronous request. B. Background mode: start the response, then poll the response ID or receive a webhook on completion. C. Increase the HTTP timeout to one hour. D. Split the work into 60 separate synchronous calls.

Answer: B. Background mode is designed for work that outlives a request; you poll or receive a webhook. Streaming (A) still holds a connection, extending timeouts (C) is fragile, and manual splitting (D) loses the model’s continuity.

Q8 · Which TWO practices control input-token cost on a long multi-turn conversation? (Select two)

A. Store the response server-side and continue with previous_response_id. B. Compact older turns into a summary as the thread grows. C. Resend the full transcript verbatim every turn. D. Raise reasoning effort to max. E. Switch every turn to gpt-6-astra.

Answer: A and B. Server-held state plus previous_response_id avoids resending, and compaction shrinks old history – both cut input tokens. Resending everything (C) increases cost, and raising effort (D) or model class (E) raises it further.

Q9 · When reading a Responses API result that used a reasoning model and a tool, why is `resp.output[0]` unreliable for the answer text? (Select one)

A. The output array is always empty. B. Reasoning items and function-call items can appear in the output list before the assistant message. C. Text is only in resp.error. D. Reasoning models never return text.

Answer: B. The output array is an ordered list that can interleave reasoning, function calls and messages, so you must find the message item rather than assume index zero. The array is not empty (A), text is not in an error field (C), and reasoning models do return text (D).

Q10 · A task is a complex, open-ended, high-value analysis that Terra fails on your eval, and it is low volume. Which model is the BEST next choice? (Select one)

A. gpt-5.6-luna to save cost. B. gpt-5.6-sol, positioned for complex, open-ended, high-value work. C. gpt-5.5 for stability. D. Stay on Terra and lower the effort.

Answer: B. Sol is the documented choice for complex, open-ended, high-value work, and low volume means its higher price is affordable. Luna (A) is for simple high-volume tasks, 5.5 (C) is previous-generation, and lowering Terra’s effort (D) would make the failing task worse.

Q11 · You must keep a two-turn conversation without resending the first turn's text. Which call structure is correct? (Select one)

A. Set stream=True on both calls. B. Set store=True on the first call and pass its id as previous_response_id on the second. C. Set background=True on the second call. D. Fine-tune the model on the first turn.

Answer: B. Server-held state is created with store=True and continued via previous_response_id. Streaming (A) and background mode (C) address latency, not state, and fine-tuning (D) is unrelated to conversation continuity.

Q12 · You have a 900-page PDF the model must read to answer questions. What is the correct input handling? (Select one)

A. Copy the PDF text into the prompt as one giant string manually. B. Provide the PDF as a file input item so the model reads it directly. C. Screenshot each page and describe it in words. D. Fine-tune on the PDF first.

Answer: B. File inputs let the model read documents directly rather than you hand-pasting text. Manual copy (A) is error-prone, screenshots-as-prose (C) loses content, and fine-tuning (D) is the wrong tool for a one-document question.

Q13 · Your conversation is approaching the 1.05M-token context window. What is the appropriate action? (Select one)

A. Switch to a model with a smaller context to force brevity. B. Compact older turns into a summary to free room while preserving the thread’s meaning. C. Delete the system instructions. D. Truncate the newest turns.

Answer: B. Compaction summarises earlier turns so the conversation continues within the window. A smaller-context model (A) makes it worse, dropping system instructions (C) breaks behaviour, and truncating the newest turns (D) discards the current task.

Q14 · Which TWO signals in a request scenario point toward function calling rather than model recall? (Select two)

A. The answer depends on live inventory in an internal system. B. The model must trigger an action such as creating a ticket. C. The user wants a poem about autumn. D. The task is to translate a paragraph the user provided. E. The user asks for a definition of a common word.

Answer: A and B. Live private data (A) and taking an action (B) both require the model to call your code. A poem (C), translating supplied text (D) and defining a common word (E) are pure generation the model does without tools.

Q15 · A team reports GPT-5.6 Terra 'feels worse than GPT-5.5 at high effort' after copying their effort settings across. What is the MOST likely explanation and fix? (Select one)

A. Terra is a worse model; switch back permanently. B. Effort levels do not map exactly between generations; re-tune Terra’s effort against the eval instead of copying the old setting. C. The context window shrank; reduce inputs. D. Pricing changed; there is nothing to do.

Answer: B. There is no exact effort mapping between GPT-5.5 and GPT-5.6, so copied settings mislead; re-tune with the eval. Terra is not inherently worse (A), the window is unchanged at 1.05M (C), and pricing (D) does not explain quality.

Q16 · An app parses `output_text` and occasionally crashes when the model calls a tool mid-response. What is the ROOT fix? (Select one)

A. Wrap the parse in a try/except and ignore failures. B. Handle the full output item list: detect and execute function_call items, return function_call_output, then read the final message. C. Disable all tools so only text is returned. D. Retry the request until it returns only text.

Answer: B. The crash comes from ignoring tool-call items in output; the correct design executes them and returns their results before reading the final message. Swallowing errors (A) hides the bug, disabling tools (C) removes needed capability, and retrying (D) is nondeterministic and wasteful.

Key takeaways

  • The Responses API is the primary interface for new work; iterate the output item list rather than assuming index zero.
  • Keep state statelessly (resend) or server-side (store + previous_response_id); compact long threads.
  • Stream for waiting users; use background mode for work that outlives a request.
  • Force machine-consumed results with structured outputs and strict: true.
  • Access live and private data with function calling, not fine-tuning or pasted dumps.
  • Use the lowest reasoning effort that works; there is no exact GPT-5.5 → GPT-5.6 mapping.
  • Climb the LADDER for model choice – Luna → Terra → dial effort → Sol → Astra – re-measuring with an eval after each change.

Last updated Sep 18, 2026