API Developer Path
D2 · The Responses API and Model Selection
The Responses API request and response shape, conversation state, streaming, background mode, structured outputs, function calling, reasoning effort, file inputs, compaction, token counting and choosing the right model.
This is the heaviest domain on the OAI-API mock alongside agentic systems – roughly 11 of 60 items (18%). It tests whether you can drive the Responses API fluently: build a request, read the output items, keep conversation state, stream, run work in the background, get structured JSON, call functions, set reasoning effort, pass files, manage long context with compaction, count tokens, and pick the cheapest model that meets the bar. The Responses API is the primary interface for new work; Chat Completions is legacy.
What you need to know
The Responses API takes an input (a string or a list of typed items) and returns a response whose output is an ordered list of items – messages, function_calls, reasoning, and tool results. State is either stateless (you resend history) or server-held (store: true plus previous_response_id). You choose a reasoning effort per request and follow the rule use the lowest effort that gets the result. You get machine-readable output with structured outputs (JSON Schema), extend the model with function calling and hosted tools, run slow work in background mode, and keep long conversations inside the context window with compaction. Model selection is a four-way choice – Astra, Sol, Terra, Luna – driven by task difficulty, volume and cost.
Learning objectives
By the end of this page you should be able to:
- Construct a Responses API request and parse the
outputitem list. - Manage conversation state statelessly and with
previous_response_id. - Use streaming and background mode appropriately.
- Produce structured outputs and handle function calling end to end.
- Set reasoning effort by the lowest-effort-that-works rule.
- Select a model from Astra / Sol / Terra / Luna using price, context and difficulty.
2.1 The request and response shape
A minimal call names a model and an input and reads output_text.
from openai import OpenAIclient = OpenAI()
resp = client.responses.create( model="gpt-5.6-terra", input="Summarise this ticket in one sentence: ...",)print(resp.output_text) # convenience accessor for the textimport OpenAI from 'openai';const client = new OpenAI();
const resp = await client.responses.create({ model: 'gpt-5.6-terra', input: 'Summarise this ticket in one sentence: ...',});console.log(resp.output_text);The full response carries an ordered output array. Do not assume item zero is your text – reasoning items, tool calls and messages can interleave.
resp.output = [ { type: "reasoning", ... }, # may appear for reasoning models { type: "function_call", name, args }, # if the model called a tool { type: "message", role: "assistant", content: [ { type: "output_text", text: "..." } ] }]Assessment signal
Stems mentioning output_text, output items, “the primary interface”, or “which API for new development” point here. The Responses API is correct for new work; Chat Completions is the legacy distractor.
2.2 Conversation state
You have two ways to keep a multi-turn conversation, and the choice has cost and privacy consequences.
| Approach | How | Trade-off |
|---|---|---|
| Stateless | Resend the whole history in input each turn | Full control, no server retention; you pay input tokens for the whole history every turn |
| Server-held | store: true, then previous_response_id on the next call | Less data resent, simpler; response is retained server-side |
first = client.responses.create( model="gpt-5.6-terra", input="My name is Sam.", store=True,)second = client.responses.create( model="gpt-5.6-terra", input="What did I say my name was?", previous_response_id=first.id,)2.3 Streaming and background mode
latency need │ ┌────────────────┼─────────────────┐ ▼ ▼ ▼ User is waiting Long task, user Fire-and-forget, for tokens can wait a bit hours-long │ │ │ stream=True stream, or background=True (token-by- background + + poll or webhook token UI) webhook for completion- Streaming (
stream=True) emits server-sent events so a UI shows tokens as they arrive – use it whenever a human waits. - Background mode (
background=True) starts the response and returns immediately; you poll the response ID or receive a webhook when it finishes. Use it for long reasoning or agentic runs that outlive an HTTP request.
2.4 Structured outputs
When another system consumes the output, force a schema rather than parsing prose.
schema = { "type": "object", "properties": { "queue": {"type": "string", "enum": ["billing", "tech", "sales"]}, "priority": {"type": "integer", "minimum": 1, "maximum": 4}, }, "required": ["queue", "priority"], "additionalProperties": False,}
resp = client.responses.create( model="gpt-5.6-luna", input="Route this ticket: 'I was charged twice this month.'", text={"format": {"type": "json_schema", "name": "routing", "schema": schema, "strict": True}},)strict: True guarantees the output validates against the schema – you no longer defensively parse. This is the right pattern for classification, extraction and any machine-consumed result.
2.5 Function calling
Function calling lets the model request that your code run, then continue with the result. The loop is: model emits a function_call → you execute it → you return a function_call_output referencing the call_id.
tools = [{ "type": "function", "name": "get_weather", "description": "Get current weather for a city.", "parameters": { "type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"], "additionalProperties": False, },}]
resp = client.responses.create( model="gpt-5.6-terra", input="What's the weather in Oslo?", tools=tools,)
call = next(o for o in resp.output if o.type == "function_call")result = get_weather(**json.loads(call.arguments)) # your code runs
final = client.responses.create( model="gpt-5.6-terra", previous_response_id=resp.id, input=[{"type": "function_call_output", "call_id": call.call_id, "output": json.dumps(result)}],)Assessment signal
“The model needs live/private data”, “call an internal system”, or “take an action” point to function calling. Options that say “fine-tune the data in” or “paste the data every request” are the distractors.
2.6 Reasoning effort and the lowest-effort rule
The GPT-5.6 family exposes reasoning efforts from none to max; Astra adds xhigh. Higher effort means more internal reasoning tokens – better on hard problems, but slower and more expensive (you pay for reasoning tokens as output).
| Effort | Use for | Cost/latency |
|---|---|---|
none / low | Extraction, classification, formatting, simple Q&A | Cheapest, fastest |
medium | General reasoning, most agentic steps | Balanced |
high / xhigh / max | Hard multi-step reasoning, planning, tricky code | Slowest, priciest |
The rule from the docs: use the lowest reasoning effort that gets the result. There is no exact mapping from old GPT-5.5 efforts to GPT-5.6 – re-tune per task with your eval, do not assume.
2.7 File inputs, compaction and token counting
- File inputs: pass PDFs and images as input items (file uploads or base64/URL for images) so the model reads them directly; you do not paste raw text you already have as a file.
- Compaction: long conversations eventually approach the 1.05M-token window. Compaction summarises earlier turns so the thread continues without losing the plot; the Agents API does this automatically, and you can do it yourself in a Responses loop by replacing old turns with a summary.
- Token counting: every response reports
usage(input_tokens,output_tokens,total_tokens). Use it to track cost and to know how close you are to the context limit.
Context budget (1.05M window)┌───────────────────────────────────────────────────────────┐│ system + tools │ retrieved context │ history │ headroom │└───────────────────────────────────────────────────────────┘ as history grows → compact old turns → keep headroom for output2.8 The model line-up (September 2026)
All four latest models: text + image input, text output, multilingual, vision, 1.05M context, 128K max output, served through the Responses API.
| Model | ID | Reasoning efforts | Input $/MTok | Output $/MTok | Context | Cutoff |
|---|---|---|---|---|---|---|
| GPT-6 Astra | gpt-6-astra | low…max, plus xhigh | $10 | $50 | 1.05M | 30 Apr 2026 |
| GPT-5.6 Sol | gpt-5.6-sol (alias gpt-5.6) | none…max | $4 | $20 | 1.05M | 16 Feb 2026 |
| GPT-5.6 Terra | gpt-5.6-terra | none…max | $2 | $12 | 1.05M | 16 Feb 2026 |
| GPT-5.6 Luna | gpt-5.6-luna | none…max | $0.20 | $1.20 | 1.05M | 16 Feb 2026 |
Previous generation still available: gpt-5.5 ($5 / $30) and gpt-5.5-pro ($30 / $180).
Selection guidance, from the docs: Astra for the hardest end-to-end work (sustained reasoning, judgment, multi-tool); Sol for complex, open-ended, high-value work; Terra the pragmatic all-rounder and natural replacement for GPT-5.5 workloads; Luna for clear, repeatable, high-volume tasks (extraction, classification, transformation, structured summaries).
Decision framework
Use the LADDER for model-and-effort selection: start low, climb only when an eval proves you must.
| Rung | Choice | Move up when |
|---|---|---|
| Luna, low effort | High-volume, well-defined tasks | The eval shows quality below target |
| A step to Terra | General work, GPT-5.5 replacements | Terra misses on genuinely complex reasoning |
| Dial effort up | Same model, medium then high | Failures are reasoning depth, not model class |
| Deploy Sol | Complex, open-ended, high-value | Sol still misses the hardest end-to-end tasks |
| Escalate to Astra | Hardest sustained reasoning + multi-tool | — |
| Re-measure | After every change, re-run the eval | Never assume; there is no fixed effort mapping |
The habit the course prizes: change one thing, then re-measure. Jumping straight to Astra at max is the anti-pattern this framework exists to prevent.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
Defaulting to Astra at max | “Best model, no risk” | Start on the LADDER’s lowest rung and climb with eval evidence |
| Using Chat Completions for new work | Familiarity, old tutorials | Use the Responses API; Chat Completions is legacy |
Assuming output[0] is the text | It usually is in demos | Iterate output; reasoning and tool-call items interleave |
| Resending full history every turn on huge threads | Simplest to code | Use store + previous_response_id, or compact old turns |
| Parsing JSON out of prose | Habit from older models | Use structured outputs with strict: true |
| Fine-tuning to inject live data | Confusing knowledge with data access | Use function calling / retrieval for live and private data |
| Copying GPT-5.5 effort levels to GPT-5.6 | Assumes a mapping exists | Re-tune effort per task; there is no exact mapping |
| Blocking a UI on a long response | Not knowing streaming/background exist | Stream for waiting users; background mode for long jobs |
Scenario challenge
Scenario. Your team ships a customer-support assistant. It must (a) answer FAQs instantly in a chat widget, (b) look up the customer’s current plan from an internal billing service, (c) occasionally handle a gnarly multi-account billing dispute that needs careful reasoning, and (d) return a structured {intent, plan, next_action} object to the surrounding app. A junior engineer proposes: “Use gpt-6-astra at max for everything, resend the full transcript each turn, and parse the JSON out of the reply text.”
Expert reasoning trace.
- Split the workload by difficulty. Most turns are FAQs and lookups – well-defined, high-volume work that
gpt-5.6-lunaorgpt-5.6-terrahandles atlow/mediumeffort. Only the rare multi-account dispute needs escalation. One-model-fits-all atmaxburns roughly 50× the input price for the common path. - Live plan data is a function call, not model knowledge. The billing lookup must be a
function_callto the internal service; recalling it or fine-tuning it in would be stale and unverifiable. - Structured output, not prose parsing. The app consumes
{intent, plan, next_action}, so usetext.formatwith a JSON Schema andstrict: true. Parsing prose is fragile and the junior’s plan would break the first time the model adds a friendly sentence. - State: don’t resend everything. For a chat widget,
store: true+previous_response_idkeeps input tokens down; for long threads, compact older turns. - Latency: stream the FAQ answers so the widget feels instant; use
higheffort only on the dispute path, ideally in background mode if the user can wait. - Escalation is eval-driven. Route the dispute path to Sol (or Astra at
high) only after an eval shows Terra fails it – LADDER, not default.
Exam-correct decision: route common turns to Luna/Terra at low effort with streaming, use function calling for the billing lookup, force structured output with a schema, keep state server-side, and escalate the rare hard case up the LADDER on eval evidence. Not Astra-at-max-for-everything, not prose JSON parsing, not full-transcript resends.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
“Use gpt-6-astra at max to be safe” | Best model feels lowest-risk | Lowest effort that works; climb the LADDER with eval evidence |
| “Chat Completions is fine for the new service” | It still works | Responses API is the primary interface for new development |
| “Fine-tune so it knows the customer’s plan” | Sounds like teaching the model | Live/private data is a function call, not training |
| “Parse the JSON from the text reply” | It often works in testing | Structured outputs with strict: true guarantee the shape |
“Map GPT-5.5 high to GPT-5.6 high” | Symmetry looks safe | No exact mapping; re-tune with your eval |
| “Resend the whole transcript each turn” | Simple and explicit | previous_response_id and compaction control cost on long threads |
“Read output[0] for the answer” | Works in the happy path | Reasoning and tool-call items interleave; iterate output |
Practice questions
Each item states how many responses to select. Commit before revealing.
Q1 · You are starting a new API integration. Which interface should you build on? (Select one)
A. The Chat Completions API, because it is the most established. B. The Responses API, the primary interface for new development. C. The Assistants API. D. The legacy Completions endpoint.
Answer: B. The Responses API is the primary interface for new work. Chat Completions (A) and the older Completions endpoint (D) are legacy for new development, and the Assistants API (C) is now grouped under Legacy APIs.
Q2 · A classifier must return `{queue, priority}` that the surrounding app parses reliably. What is the BEST way to guarantee the shape? (Select one)
A. Ask the model in the prompt to ‘only return JSON’ and parse the text.
B. Use structured outputs with a JSON Schema and strict: true.
C. Post-process the prose with a regular expression.
D. Fine-tune the model to emit JSON.
Answer: B. Structured outputs with a strict JSON Schema guarantee the response validates against the shape. Prompt instructions (A) and regex (C) are not guaranteed, and fine-tuning (D) is unnecessary and does not guarantee validity.
Q3 · An assistant needs the customer's current subscription plan from an internal billing service to answer. What is the correct mechanism? (Select one)
A. Fine-tune the model on all customer plans nightly.
B. Define a function tool the model calls, execute it, and return a function_call_output.
C. Paste the entire customer database into every prompt.
D. Rely on the model’s training knowledge.
Answer: B. Live, private data is accessed via function calling: the model requests the call, your code runs it, and you return the result. Nightly fine-tuning (A) is stale and costly, pasting the database (C) is expensive and non-authoritative, and training knowledge (D) does not contain live account data.
Q4 · A high-volume extraction task runs 2M times per day and `gpt-5.6-luna` passes your eval at 97%. Which model should you deploy? (Select one)
A. gpt-6-astra at max effort for maximum quality.
B. gpt-5.6-luna, because it meets the metric at the lowest cost for repeatable high-volume work.
C. gpt-5.6-sol to be safe.
D. gpt-5.5-pro for reliability.
Answer: B. Luna is built for clear, repeatable, high-volume tasks and it passes the eval, so it is the lowest-cost model that works. Astra-at-max (A), Sol (C) and 5.5-pro (D) all cost far more per token with no measured quality gain.
Q5 · A user is waiting in a chat UI for an answer that takes several seconds to generate. What improves the experience MOST? (Select one)
A. Enable background mode and poll every 30 seconds.
B. Stream the response so tokens appear as they are generated.
C. Switch to gpt-6-astra at max effort.
D. Increase the max output tokens.
Answer: B. Streaming shows tokens as they arrive, which is the right pattern when a human waits. Background mode (A) suits fire-and-forget long jobs, a bigger model (C) is slower not faster, and raising max output (D) does not change perceived latency.
Q6 · What is TRUE about reasoning effort on the GPT-5.6 family? (Select one)
A. GPT-5.5 effort levels map exactly onto GPT-5.6. B. You should use the lowest effort that produces the required result, re-tuning per task with an eval. C. Higher effort is always better and should be the default. D. Effort has no effect on cost or latency.
Answer: B. The rule is lowest effort that works, verified per task; there is no exact mapping from 5.5 (A). Higher effort costs more time and tokens (contradicting C and D).
Q7 · You are building a long-running research agent whose run outlives a single HTTP request. Which feature fits? (Select one)
A. stream=True on a synchronous request.
B. Background mode: start the response, then poll the response ID or receive a webhook on completion.
C. Increase the HTTP timeout to one hour.
D. Split the work into 60 separate synchronous calls.
Answer: B. Background mode is designed for work that outlives a request; you poll or receive a webhook. Streaming (A) still holds a connection, extending timeouts (C) is fragile, and manual splitting (D) loses the model’s continuity.
Q8 · Which TWO practices control input-token cost on a long multi-turn conversation? (Select two)
A. Store the response server-side and continue with previous_response_id.
B. Compact older turns into a summary as the thread grows.
C. Resend the full transcript verbatim every turn.
D. Raise reasoning effort to max.
E. Switch every turn to gpt-6-astra.
Answer: A and B. Server-held state plus previous_response_id avoids resending, and compaction shrinks old history – both cut input tokens. Resending everything (C) increases cost, and raising effort (D) or model class (E) raises it further.
Q9 · When reading a Responses API result that used a reasoning model and a tool, why is `resp.output[0]` unreliable for the answer text? (Select one)
A. The output array is always empty.
B. Reasoning items and function-call items can appear in the output list before the assistant message.
C. Text is only in resp.error.
D. Reasoning models never return text.
Answer: B. The output array is an ordered list that can interleave reasoning, function calls and messages, so you must find the message item rather than assume index zero. The array is not empty (A), text is not in an error field (C), and reasoning models do return text (D).
Q10 · A task is a complex, open-ended, high-value analysis that Terra fails on your eval, and it is low volume. Which model is the BEST next choice? (Select one)
A. gpt-5.6-luna to save cost.
B. gpt-5.6-sol, positioned for complex, open-ended, high-value work.
C. gpt-5.5 for stability.
D. Stay on Terra and lower the effort.
Answer: B. Sol is the documented choice for complex, open-ended, high-value work, and low volume means its higher price is affordable. Luna (A) is for simple high-volume tasks, 5.5 (C) is previous-generation, and lowering Terra’s effort (D) would make the failing task worse.
Q11 · You must keep a two-turn conversation without resending the first turn's text. Which call structure is correct? (Select one)
A. Set stream=True on both calls.
B. Set store=True on the first call and pass its id as previous_response_id on the second.
C. Set background=True on the second call.
D. Fine-tune the model on the first turn.
Answer: B. Server-held state is created with store=True and continued via previous_response_id. Streaming (A) and background mode (C) address latency, not state, and fine-tuning (D) is unrelated to conversation continuity.
Q12 · You have a 900-page PDF the model must read to answer questions. What is the correct input handling? (Select one)
A. Copy the PDF text into the prompt as one giant string manually. B. Provide the PDF as a file input item so the model reads it directly. C. Screenshot each page and describe it in words. D. Fine-tune on the PDF first.
Answer: B. File inputs let the model read documents directly rather than you hand-pasting text. Manual copy (A) is error-prone, screenshots-as-prose (C) loses content, and fine-tuning (D) is the wrong tool for a one-document question.
Q13 · Your conversation is approaching the 1.05M-token context window. What is the appropriate action? (Select one)
A. Switch to a model with a smaller context to force brevity. B. Compact older turns into a summary to free room while preserving the thread’s meaning. C. Delete the system instructions. D. Truncate the newest turns.
Answer: B. Compaction summarises earlier turns so the conversation continues within the window. A smaller-context model (A) makes it worse, dropping system instructions (C) breaks behaviour, and truncating the newest turns (D) discards the current task.
Q14 · Which TWO signals in a request scenario point toward function calling rather than model recall? (Select two)
A. The answer depends on live inventory in an internal system. B. The model must trigger an action such as creating a ticket. C. The user wants a poem about autumn. D. The task is to translate a paragraph the user provided. E. The user asks for a definition of a common word.
Answer: A and B. Live private data (A) and taking an action (B) both require the model to call your code. A poem (C), translating supplied text (D) and defining a common word (E) are pure generation the model does without tools.
Q15 · A team reports GPT-5.6 Terra 'feels worse than GPT-5.5 at high effort' after copying their effort settings across. What is the MOST likely explanation and fix? (Select one)
A. Terra is a worse model; switch back permanently. B. Effort levels do not map exactly between generations; re-tune Terra’s effort against the eval instead of copying the old setting. C. The context window shrank; reduce inputs. D. Pricing changed; there is nothing to do.
Answer: B. There is no exact effort mapping between GPT-5.5 and GPT-5.6, so copied settings mislead; re-tune with the eval. Terra is not inherently worse (A), the window is unchanged at 1.05M (C), and pricing (D) does not explain quality.
Q16 · An app parses `output_text` and occasionally crashes when the model calls a tool mid-response. What is the ROOT fix? (Select one)
A. Wrap the parse in a try/except and ignore failures.
B. Handle the full output item list: detect and execute function_call items, return function_call_output, then read the final message.
C. Disable all tools so only text is returned.
D. Retry the request until it returns only text.
Answer: B. The crash comes from ignoring tool-call items in output; the correct design executes them and returns their results before reading the final message. Swallowing errors (A) hides the bug, disabling tools (C) removes needed capability, and retrying (D) is nondeterministic and wasteful.
Key takeaways
- The Responses API is the primary interface for new work; iterate the
outputitem list rather than assuming index zero. - Keep state statelessly (resend) or server-side (
store+previous_response_id); compact long threads. - Stream for waiting users; use background mode for work that outlives a request.
- Force machine-consumed results with structured outputs and
strict: true. - Access live and private data with function calling, not fine-tuning or pasted dumps.
- Use the lowest reasoning effort that works; there is no exact GPT-5.5 → GPT-5.6 mapping.
- Climb the LADDER for model choice – Luna → Terra → dial effort → Sol → Astra – re-measuring with an eval after each change.
Last updated Sep 18, 2026