API Developer Path
API Developer Path – Mock Exam 1
A 60-item independent mock exam for the OpenAI API developer path, blueprint-weighted across seven domains, with full explanations and a readiness readout.
This is a full-length, blueprint-weighted independent mock exam for the API developer path built from publicly available OpenAI learning objectives and product documentation. It is not an official OpenAI assessment and not the Academy assessment: Academy badges and pathway certificates are not certifications. All 60 items are new and do not repeat the domain-page questions. Use this as your diagnostic sitting before the harder Mock Exam 2.
Instructions
- Time: 90 minutes for 60 items — the pace our mock uses; OpenAI does not publish a public exam length.
- Items: 60, single-answer and multiple-response. Each item states how many answers to select.
- Selection: for a Select two item you must pick both correct options and no incorrect one to earn the mark. There is no guessing penalty, so answer every question.
- Target: aim for at least 80% raw (about 48 of 60) before you sit the real Academy assessment, which passes at the same 80% line.
- Work each question before expanding the answer.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| D1 | Scoping AI Solutions | 7 |
| D2 | The Responses API and Model Selection | 11 |
| D3 | Evaluating AI Applications | 10 |
| D4 | Designing and Building Agentic Systems | 11 |
| D5 | Retrieval-Augmented Generation | 9 |
| D6 | Performance, Latency and Cost | 7 |
| D7 | Production Safety and Operations | 5 |
Total: 7 + 11 + 10 + 11 + 9 + 7 + 5 = 60 items.
Readiness interpretation
This is an independent readiness indicator, not an official score.
| Raw score (of 60) | Band | Interpretation |
|---|---|---|
| 54–60 (90%+) | Strong readiness | Confident across all domains |
| 48–53 (80–89%) | Assessment ready | At or above the 80% Academy line; tidy up weak domains |
| 42–47 (70–79%) | Building confidence | Close; target the domains that cost you marks |
| under 42 (less than 70%) | Keep learning | Revisit the domain pages before re-attempting |
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
60 questions · one at a time · 90-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A stakeholder describes a goal as 'make the chatbot smarter'. Before writing any code, what turns this into a scopable problem?
Show answer
Answer: A.
Scoping converts a vague wish into a task with defined input, output, user and a measurable success metric, which is what every later decision depends on. B jumps to a model before the problem is known and wastes budget. C presumes fine-tuning is needed before any task definition exists. D skips the definition that a prototype would be built to test, so feedback would be unfocused.
Which of these is a deterministic requirement that should NOT rely on a language model as the source of truth?
Show answer
Answer: A.
An exact balance must come from arithmetic over authoritative data, not a model that can approximate; the model can present or explain it but not compute it as truth. B, C and D are judgment or language tasks where an approximate, reviewable output is acceptable, which is exactly where models fit.
A team must decide between building a custom transcription pipeline and buying a hosted product for support-call transcription. Which factor MOST favours buying?
Show answer
Answer: A.
Buy when the capability is a commodity that provides no differentiation and mature products exist, so effort goes to the product instead. B is a preference, not a business reason. C is a constraint that would argue against buying, not for it. D suggests build capacity but capacity is not a reason to build something with no differentiation value.
You are scoping a summariser for 40,000 documents per day at ~1,200 input and ~200 output tokens each;
gpt-5.6-lunameets the quality bar. Luna costs $0.20 per million input tokens and $1.20 per million output tokens. Roughly what is the daily cost?Show answer
Answer: A.
Input: 40,000 x 1,200 = 48M tokens x $0.20/M = $9.60. Output: 40,000 x 200 = 8M tokens x $1.20/M = $9.60. Total ~$19.20/day, so A. B undercounts by an order of magnitude, C uses Terra-like rates, and D is ten times too high.
A stakeholder wants an AI feature that auto-approves expense reports and posts them to the ledger with no human review. Which TWO points should shape your scope?
Show answer
Answer: A and B.
A recognises an irreversible action needs a gate or reversible staging, and B keeps arithmetic deterministic - both reshape the design. C is false: a bigger model does not make generated arithmetic authoritative. D is wrong because automation raises, not removes, the need for a metric. E dismisses evaluation, which is exactly what a money-moving feature requires.
A director wants
gpt-6-astramandated for every feature 'to be safe'. A high-volume extraction task passes your eval at 96% ongpt-5.6-luna. What is the BEST response?Show answer
Answer: A.
Model selection is per task and driven by eval results against the quality bar and cost; presenting evidence lets the team choose Luna where it suffices. B wastes large multiples of cost with no quality benefit. C is disproportionate. D doubles cost and latency on every request for no production value.
During scoping you learn a feature must respond while a user waits AND another must clear a nightly backlog of 3M items. What does correct scoping conclude?
Show answer
Answer: A.
Interactive and bulk workloads have opposite latency profiles, so scoping treats them separately and can pick Batch for the backlog and low-latency processing for the interactive path. B forces a single tier onto incompatible needs. C would make the interactive feature unusably slow. D would waste money and offer no benefit on the backlog.
For a new API integration in September 2026, which interface should you build on for the long term?
Show answer
Answer: A.
The Responses API is the primary interface for new work; Chat Completions is legacy for new builds. B misstates Chat Completions as new. C and D are wrong: Assistants and Agent Builder are grouped under Legacy APIs, not the recommended default.
You need
gpt-5.6-terrato run with the least thinking that still passes your eval. Which principle applies?Show answer
Answer: A.
The documented guidance is to use the lowest reasoning effort that gets the result, because higher effort adds latency and cost. B over-spends by default. C is wrong: there is no exact mapping from GPT-5.5 efforts to GPT-5.6. D is false since more reasoning generally means more tokens, latency and cost.
A classification service must emit
{"queue": "...", "priority": 1}that downstream code parses without defensive cleanup. What is the BEST mechanism?Show answer
Answer: A.
Structured outputs bound to a schema guarantee the parseable shape. B is a request the model can still violate. C is brittle and fails on edge cases. D makes output less predictable, the opposite of what is needed.
An assistant must fetch a customer's live subscription tier from an internal billing service to answer. What is the correct mechanism?
Show answer
Answer: A.
Live, per-customer data must be fetched at request time through a tool the model calls, i.e. function calling. B is impossible: the model has no knowledge of live account state. C is infeasible and a data-exposure risk. D confuses output length with data access.
A high-volume extraction task runs 3M times per day and
gpt-5.6-lunapasses your eval at 97%. Which model should you deploy?Show answer
Answer: A.
Luna is designed for clear, repeatable, high-volume tasks and it already meets the quality bar, so it is the cheapest correct choice. B, C and D all cost multiples more for a task Luna passes, with no quality justification, and D also picks a costlier previous-generation model.
You must maintain a two-turn conversation without resending the first turn's text on the second call. Which Responses API feature supports this?
Show answer
Answer: A.
The Responses API supports server-side conversation state so a follow-up references the prior response without resending its text. B is unrelated to state. C is the thing conversation state exists to avoid. D is a training operation, not a way to carry a single conversation.
A run must continue after a single HTTP request ends and report progress via streaming or webhooks. Which Responses API feature fits?
Show answer
Answer: A.
Background mode lets a long-running task continue beyond one HTTP request and be followed via streaming or webhooks. B reduces cost on repeated prefixes but does not detach the run. C speeds edits, unrelated to lifetime. D only affects response length.
Which TWO practices reduce input-token cost on a very long multi-turn conversation without dropping needed context?
Show answer
Answer: A and B.
Compaction summarises stale history and conversation state avoids resending prior text, both cutting input tokens. C increases spend and latency. D changes randomness, not tokens. E raises the per-token price, worsening cost.
When you read a Responses API result from a reasoning model that also called a tool, why can indexing the first output item be unreliable for the answer text?
Show answer
Answer: A.
A response is a list of typed items; reasoning and tool-call items can precede the message, so code should use the aggregated output text helper rather than a fixed index. B, C and D are false generalisations that would break in the common case where text and tool items coexist.
A complex, open-ended, high-value analysis fails on
gpt-5.6-terrain your eval and volume is low. Which is the BEST next model to try?Show answer
Answer: A.
Sol is positioned for complex, open-ended, high-value work and is the natural step up from Terra, especially at low volume where its price matters less. B moves down the capability ladder for a task Terra already failed. C steps back a generation. D changes randomness, not capability.
What is TRUE about reasoning effort levels on the GPT-5.6 family?
Show answer
Answer: A.
Per the model table, the GPT-5.6 models support none through max and Astra additionally offers xhigh. B is false since Sol, Terra and Luna also take effort. C is explicitly denied by the docs. D is invented;
noneandloware available.A developer edits a prompt, checks two examples that look better, and wants to ship. What should they do FIRST?
Show answer
Answer: A.
Two examples cannot show whether a change helps overall; a representative eval compared to a baseline can. B ships blind. C is anecdotal. D changes cost without evidence the prompt change is even good.
You must grade whether extracted invoice fields exactly match ground truth. Which grader is BEST?
Show answer
Answer: A.
Exact field matching against ground truth is a deterministic check, so a string/exact-match grader is precise and cheap. B adds noise and cost for a task with a definite answer. C does not scale. D would pass near-misses that must be exact.
What is the difference between offline and online evaluation?
Show answer
Answer: A.
Offline evaluation uses a curated dataset before release; online evaluation observes live traffic to catch inputs the dataset never contained. B mislabels their purposes. C denies a real distinction. D is wrong: online complements, not replaces, offline datasets.
Before trusting an LLM-as-judge grader that scores summaries for faithfulness, what must you do?
Show answer
Answer: A.
An automated judge is only trustworthy once it has been calibrated against human labels on a sample. B is unearned trust. C risks the judge favouring its own style with no verification. D removes the controlled comparison that calibration needs.
A team runs only offline evals and keeps getting production failures on phrasings never in their dataset. What is missing?
Show answer
Answer: A.
Failures on unseen phrasings are exactly what online evaluation and traffic sampling catch, feeding new cases back into the dataset. B, C and D change model behaviour but do nothing to surface the inputs the offline set never covered.
Which TWO are legitimate reasons to include human review in an evaluation process?
Show answer
Answer: A and B.
Human review calibrates judges and resolves high-stakes or ambiguous cases. C does not scale and is not the goal of review. D discards the reusable dataset humans help build. E is not a reason at all.
Which TWO purposes does keeping a baseline eval score serve?
Show answer
Answer: A and B.
A baseline is the reference that lets you see both improvements and regressions from a change. C invents a platform requirement. D is cosmetic. E is wrong: baselines and online evaluation address different things and one does not replace the other.
How should you describe the Evals API given the current OpenAI docs structure?
Show answer
Answer: A.
The docs group Evals under Legacy APIs while evals remain conceptually important, so A is the accurate framing. B overstates its status. C is wrong: legacy is not the same as removed. D misplaces it in the Responses core.
Your eval dataset has 10 items and scores swing widely between runs. What is the MOST likely problem?
Show answer
Answer: A.
With only 10 items, a single flip moves the percentage sharply, so the instability is a sample-size problem; enlarge the set. B, C and D are possible in general but the described symptom - large swings on a tiny set - points squarely at insufficient data.
A grounded-answer eval checks whether answers are supported by retrieved context. When should it be re-run?
Show answer
Answer: A.
Any change to retrieval, chunking, prompt or model can alter grounding, so the eval is re-run on each such change. B and C freeze evaluation while the system evolves. D is reactive and lets regressions ship unnoticed.
You want OpenAI to run and recover a durable, hours-long agent session with managed context compaction. Which runtime fits BEST?
Show answer
Answer: A.
The Agents API is the OpenAI-managed harness that runs sessions, orchestration, context compaction and recovery, which matches durable hours-long work. B and C put the durability and recovery burden on you. D has no session, recovery or compaction machinery.
A team wants code-first control of a custom multi-step orchestration they run on their own servers. Which runtime is the BEST fit?
Show answer
Answer: A.
The Agents SDK gives code-first, self-hosted orchestration and guardrails, matching the requirement to run on their own servers. B is OpenAI-managed, not self-hosted. C is for bulk offline jobs. D lacks the orchestration and guardrail primitives the SDK provides.
An EU bank requires EU data residency and zero data retention. Is the Agents API eligible?
Show answer
Answer: A.
The Agents API is US data residency only with no ZDR, and choosing a self-hosted sandbox does not lift those constraints. B invents a region header. C is false; ZDR is not offered. D confuses network privacy with residency and retention.
What is the correct way to let an agent read an internal repository through a standard protocol?
Show answer
Answer: A.
MCP is the standard protocol for connecting tools and data sources to an agent, so an MCP server is the correct mechanism. B does not scale and leaks context budget. C is costly and stale between runs. D is an insecure credential-handling anti-pattern.
An agent can issue refunds. What control is REQUIRED before it acts on a high-value refund?
Show answer
Answer: A.
Irreversible, high-value actions require a human approval gate before execution. B, C and D may affect quality or capacity but none of them stops the agent from committing an unwanted irreversible action.
Which TWO statements about the Agents API managed harness are TRUE?
Show answer
Answer: A and B.
The managed harness runs the session lifecycle and supports subagent delegation via
multi_agentwithmax_concurrent_subagents. C is false: it is US-only with no ZDR. D contradicts the managed recovery it provides. E describes the Agents SDK, not the Agents API.A simple assistant needs to call one internal function and return an answer, with full control over each step and no session or sandbox. Which runtime is simplest and correct?
Show answer
Answer: A.
For a single tool call with full step-by-step control and no session or sandbox needs, the Responses API with function calling is the simplest fit. B and C add managed sessions or framework machinery the task does not need. D is a training operation, not a runtime for tool use.
What is the difference between multi-agent handoffs and subagents?
Show answer
Answer: A.
A handoff passes control to a different agent, while a subagent is delegated a bounded piece of work and returns its result to the parent. B denies a real distinction. C misdescribes delegation. D places handoffs in an unrelated API.
A support agent needs to read orders, but a developer grants it write access to billing 'just in case'. What principle is violated?
Show answer
Answer: A.
Granting more access than the task needs violates least privilege; the fix is read-only order access. B, C and D are real engineering concepts but none of them describes over-broad permissions.
A hosted sandbox in the Agents API is billed how, relative to a self-hosted one?
Show answer
Answer: A.
Agents API billing is model rates plus standard tool rates plus container rates for hosted sandboxes. B ignores the container charge. C invents a flat fee. D is wrong because the container rate is exactly what a hosted sandbox adds over self-hosting.
An agent occasionally takes a wrong tool action and afterwards nobody can tell why. What is missing?
Show answer
Answer: A.
Being unable to explain a past action is an observability gap; tracing steps and tool calls provides the audit trail. B, C and D add capability, parallelism or budget but none of them records what happened for later diagnosis.
A knowledge base of 50,000 internal documents changes weekly and answers must cite sources. What is the BEST approach?
Show answer
Answer: A.
RAG retrieves current documents at query time and can attach citations, handling weekly change cheaply. B is expensive, slow to update and hard to cite. C is infeasible and wasteful. D cannot know private, changing internal content.
Users search by exact error codes like
ERR-4021and pure semantic retrieval keeps missing them. What is the fix?Show answer
Answer: A.
Exact tokens like error codes are matched by lexical search, so hybrid retrieval combining keyword and semantic recall fixes the misses. B tweaks embeddings but still favours meaning over exact strings. C and D change generation, not retrieval.
Clinicians must only retrieve protocols for their own department. Where must this restriction be enforced?
Show answer
Answer: A.
Authorisation must be enforced in retrieval so out-of-scope documents are never fetched; a prompt request is bypassable. B relies on the model behaving, which is not a security control. C and D affect quality or cost, not access.
A retrieved chunk lacks the answer, yet the model responds with a confident number from training knowledge. How do you prevent this?
Show answer
Answer: A.
Grounding requires instructing the model to answer only from context and to abstain otherwise, verified by a grounding eval. B may still hallucinate. C affects length, not grounding. D removes the very context that anchors the answer.
Chunks are set to 4,000 tokens each and answers now include lots of irrelevant text. What is the MOST likely cause?
Show answer
Answer: A.
Oversized chunks bundle unrelated passages, so retrieval brings in noise; smaller, focused chunks help. B is unlikely to manifest as bloated-but-relevant retrieval. C and D do not explain why retrieved text is off-topic.
Why is overlap added between adjacent chunks when indexing?
Show answer
Answer: A.
Overlap preserves context that straddles a chunk boundary so a boundary-spanning answer is still retrievable. B is false; overlap increases storage. C and D are unrelated to what overlap does.
Which TWO metrics should a grounded-answer eval measure?
Show answer
Answer: A and B.
A grounded-answer eval measures whether retrieval fetched the right context and whether the answer is faithful to it. C is a throughput stat, not answer quality. D is unrelated. E is a storage metric, not a quality measure.
A developer proposes putting a whole 900-page manual in the prompt every request because the window is 1.05M tokens. Why is RAG usually better here?
Show answer
Answer: A.
Even when the manual fits, stuffing it every request pays for and processes irrelevant tokens on each call; RAG retrieves only what a query needs. B overclaims universal accuracy. C is false since 900 pages fit in 1.05M tokens. D is untrue.
What does file search over a managed vector store give you out of the box?
Show answer
Answer: A.
Managed file search provides hosted chunking, embedding, indexing and retrieval exposed as a tool. B is a different, training operation. C is not something any retrieval layer can guarantee. D invents a pricing claim.
Many requests share a 5,000-token system prompt followed by a short user message. What reduces cost and latency MOST directly?
Show answer
Answer: A.
A large, repeated prefix is exactly what prompt caching targets, cutting the cost and time of reprocessing it. B and C increase spend and latency. D affects output length, not the repeated input.
A nightly job classifies 5M documents and latency does not matter. Which processing choice is MOST cost-effective?
Show answer
Answer: A.
Latency-insensitive bulk work is the Batch API's purpose and it is the cheapest tier for it. B pays a premium model synchronously. C and D optimise for interactivity the job does not need, at higher cost.
A summariser runs 200,000 times per day at 1,000 input and 200 output tokens on
gpt-5.6-luna($0.20/M input, $1.20/M output). Roughly what is the daily cost?Show answer
Answer: A.
Input: 200,000 x 1,000 = 200M x $0.20/M = $40. Output: 200,000 x 200 = 40M x $1.20/M = $48. Total ~$88/day, so A. B is a tenth of the real figure, C uses Terra-scale rates, and D is far too high.
An interactive feature 'feels slow'. Before changing anything, what should you do FIRST?
Show answer
Answer: A.
You cannot optimise what you have not measured, so profiling latency and time-to-first-token comes first. B, C and D are blind changes that may not touch the real bottleneck and could make it worse.
Time-to-first-token dominates latency for an interactive feature running
gpt-5.6-solonhigheffort. Which TWO levers most directly help?Show answer
Answer: A and B.
Lower effort cuts the thinking that delays the first token, and streaming shows output as it arrives, improving perceived latency. C lengthens generation. D is for offline bulk work, not interactive latency. E adds input processing, increasing latency.
You are editing a large document where about 95% of the output equals the input. Which feature cuts latency MOST?
Show answer
Answer: A.
Predicted outputs accelerate generation when most of the output is already known, exactly the edit-with-small-changes case. B is for offline bulk jobs. C adds latency. D affects capacity, not generation speed for near-identical output.
A latency-tolerant workload wants lower cost than standard processing but cannot wait the hours the Batch API takes. Which tier fits?
Show answer
Answer: A.
Flex processing trades some latency for lower cost without the multi-hour turnaround of Batch, fitting a latency-tolerant but not offline workload. B optimises for speed at higher cost. C targets real-time interaction. D is the standard tier it is trying to beat on cost.
An API call returns a
429. What is the correct handling?Show answer
Answer: A.
A 429 is a rate-limit signal handled by exponential backoff with jitter. B worsens the overload. C overreacts to a transient condition. D does not address rate limiting and may just move the problem.
Which error should you NOT automatically retry?
Show answer
Answer: A.
A 400 means the request itself is wrong, so retrying the same payload will fail again; fix the request instead. B, C and D are transient or throttling conditions where backoff and retry are appropriate.
A team fears a runaway agent loop could generate a huge bill. Which control MOST directly caps the dollar cost?
Show answer
Answer: A.
Spend limits directly cap dollar exposure regardless of loop behaviour. B, C and D have no bearing on total spend and D would increase it. Rate limits help throttle but spend limits cap the money most directly.
A Kubernetes workload should call the API without holding a long-lived API key. Which TWO practices are correct?
Show answer
Answer: A and B.
Workload identity federation avoids long-lived keys by issuing short-lived credentials, and least-privilege scoping limits what those credentials can do. C embeds a secret in an image that can leak. D widens blast radius and complicates rotation. E removes observability and does nothing about the key itself.
An agent processes untrusted customer messages and can take actions. Which TWO controls address prompt-injection risk BEST?
Show answer
Answer: A and B.
Least privilege limits the blast radius of a successful injection and approval gates stop irreversible actions from firing automatically. C changes randomness, D does not prevent injection, and E only affects output length.
Last updated Sep 18, 2026