API Developer Path
API Developer Path – Mock Exam 2
A harder 60-item independent mock exam for the OpenAI API developer path, with multi-constraint and judgment items, full explanations and a readiness readout.
This is the second full-length independent mock exam for the API developer path, built from publicly available OpenAI learning objectives and product documentation. It is not an official OpenAI assessment: Academy badges and pathway certificates are not certifications. Mock 2 is deliberately harder than Mock Exam 1 — more multi-constraint stems, more FIRST / BEST / MOST cost-effective / TWO qualifiers, and more scenario framing. Treat it as your readiness gate: all 60 items are new and distinct from Mock 1 and the domain pages.
Instructions
- Time: 90 minutes for 60 items — the pace our mock uses; OpenAI does not publish a public exam length.
- Items: 60, single-answer and multiple-response. Each item states how many answers to select.
- Selection: for a Select two item you must pick both correct options and no incorrect one to earn the mark. There is no guessing penalty, so answer every question.
- Target: aim for at least 80% raw (about 48 of 60) here before you sit the real Academy assessment, which passes at the same 80% line. Clearing this harder mock at 80%+ is a strong sign of readiness.
- Work each question before expanding the answer.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| D1 | Scoping AI Solutions | 7 |
| D2 | The Responses API and Model Selection | 11 |
| D3 | Evaluating AI Applications | 10 |
| D4 | Designing and Building Agentic Systems | 11 |
| D5 | Retrieval-Augmented Generation | 9 |
| D6 | Performance, Latency and Cost | 7 |
| D7 | Production Safety and Operations | 5 |
Total: 7 + 11 + 10 + 11 + 9 + 7 + 5 = 60 items.
Readiness interpretation
This is an independent readiness indicator, not an official score.
| Raw score (of 60) | Band | Interpretation |
|---|---|---|
| 54–60 (90%+) | Strong readiness | Handling the hardest multi-constraint items well |
| 48–53 (80–89%) | Assessment ready | At or above the 80% Academy line on the harder set |
| 42–47 (70–79%) | Building confidence | Close; drill the judgment and cost-arithmetic items |
| under 42 (less than 70%) | Keep learning | Revisit the domain pages and retake Mock 1 first |
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
60 questions · one at a time · 90-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A VP wants a fully autonomous incident-handling system live in one week, with no labelled incidents and a requirement that no customer-facing action fire without sign-off. What is the MOST appropriate FIRST step?
Show answer
Answer: A.
With no labelled data and a hard human-sign-off requirement, the responsible first step is a narrow assisted slice with a metric, earning autonomy through evals. B contradicts the sign-off requirement and ships risk. C manufactures data that may not reflect reality and still skips evaluation. D leans on model size to paper over an absence of data and metrics, which capability cannot fix.
You are scoping a feature and want to elicit only requirements that would CHANGE the architecture. Which TWO qualify?
Show answer
Answer: A and B.
Irreversibility drives approval gates and staging, and interactive-versus-bulk drives model, tier and streaming choices, so both change the architecture. C, D and E are cosmetic or organisational details that do not affect how the system is built.
Two candidate metrics are proposed for a triage classifier: (i) 'users are happier' and (ii) 'macro-F1 ≥ 0.85 on a held-out labelled set with misroute rate below 3%'. Which is the BETTER success metric and why?
Show answer
Answer: A.
A good metric is specific and measurable against ground truth with a business-relevant threshold, which (ii) provides. (i) is unmeasurable as stated and cannot gate a release. C is false. D ignores that only (ii) can actually be computed and acted on.
A workload does 90,000 classifications per day at 800 input and 120 output tokens. On
gpt-5.6-luna($0.20/$1.20 per M) it passes eval; a director prefersgpt-5.6-terra($2/$12 per M). Roughly how much MORE per day does Terra cost for the same workload?Show answer
Answer: A.
Daily input 90,000 x 800 = 72M tokens; output 90,000 x 120 = 10.8M tokens. Luna: 72 x $0.20 + 10.8 x $1.20 = $14.40 + $12.96 = ~$27. Terra: 72 x $2 + 10.8 x $12 = $144 + $129.60 = ~$274. The gap is about $247/day, so A. B is roughly the Luna total, not the gap; C is false since Terra is ten times pricier per token; D is ten times too high.
During scoping you find the task's output auto-cancels shipments, an irreversible action. How should the solution plan BEST reflect this?
Show answer
Answer: A.
Irreversibility is a hard constraint that must shape the design with an approval gate or reversible staging before commit. B defers a safety-critical requirement. C cannot make an irreversible action safe by capability alone. D only records damage after it happens, which does not prevent it.
A commodity capability with a mature hosted product exists, your team is small, and speed-to-value matters, but the data is highly sensitive and must never leave your network. What does this MOST likely change about the buy decision?
Show answer
Answer: A.
A hard data-residency constraint can override the usual buy-the-commodity heuristic, favouring build or an in-network deployment. B ignores the constraint that changes the calculus. C leaps to fine-tuning, which residency does not require. D is a non sequitur.
Which item belongs specifically in the 'Success metric' line of a solution plan, rather than the requirements or risks lines?
Show answer
Answer: A.
A success metric states the measurable quality bar and how it is measured, which A does. B is a requirement/constraint, C is a risk, and D is a latency requirement, none of which is the quality success metric.
A pipeline needs strict schema output AND a live tool call AND to persist conversation state across turns, all on the primary interface. Which single API supports all three directly?
Show answer
Answer: A.
The Responses API is the primary interface and supports structured outputs, function calling and conversation state together. B is legacy for new work. C is grouped under Legacy APIs. D is an offline processing tier, not an interactive interface with these features.
A scenario states: 'the model must return today's exchange rate and the customer's current balance'. Which TWO signals point toward tool calls rather than model recall?
Show answer
Answer: A and B.
Constantly changing external data and live private state are both outside model knowledge and must come from tools. C, D and E are surface features that say nothing about whether the needed information exists inside the model.
A team is on
gpt-5.6-solathigheffort. Latency is unacceptable and their eval passes atmedium. What is the MOST cost-effective correct change that preserves quality?Show answer
Answer: A.
The lowest effort that passes the eval is the target, and medium passes while cutting latency and cost. B increases both cost and latency. C adds parallel work unrelated to per-call effort latency. D does not reduce reasoning latency and harms determinism.
An application parses
output_textand intermittently crashes when the model emits a tool call before any message. What is the ROOT fix, not a workaround?Show answer
Answer: A.
The root cause is treating the output as text-first when it is a sequence of typed items; handling each item type fixes it structurally. B hides errors and loses tool results. C removes a needed capability. D is a flaky workaround that wastes calls.
A conversation is approaching the 1.05M-token context window mid-task and must keep going. What is the appropriate action?
Show answer
Answer: A.
Compaction summarises stale context so the task continues within the window while preserving needed information. B risks discarding essential context. C makes the problem worse. D concerns output length, not the input context pressure.
A task is complex, open-ended and the single highest-value workflow in the product, requiring sustained multi-tool reasoning and judgment, at low volume. Which model is the BEST fit?
Show answer
Answer: A.
Astra is positioned for the hardest sustained-reasoning, judgment and multi-tool work, and low volume means its price is tolerable. B is for clear high-volume tasks, C is a mid all-rounder likely to fall short on the hardest work, and D steps back a generation.
A long-running research agent must survive well beyond one HTTP request AND deliver incremental progress to a UI. Which combination is correct?
Show answer
Answer: A.
Background mode detaches the run and streaming or webhooks carry incremental progress, matching both needs. B optimises cost and edit speed, not run lifetime. C and D address capacity and length, not detachment or progress reporting.
Which statement about the GPT-5.6 family is accurate?
Show answer
Answer: A.
The lineup shows all four with 1.05M context and 128K max output, with Astra the most expensive at $10/$50. B is wrong: they share the same context. C is false: Terra ($2/$12) is cheaper than Sol ($4/$20). D is false: Astra is the most expensive, not cheaper than Terra.
A team keeps a two-turn conversation by re-sending the entire first turn's text on every follow-up 'to be safe', and their input costs are climbing. What is the BEST correction?
Show answer
Answer: A.
Conversation state removes the need to resend prior text, directly cutting the growing input cost. B treats a symptom while keeping the wasteful pattern. C does not address input tokens. D wrongly assumes the cost cannot be avoided.
You must choose the pragmatic all-rounder that is the natural replacement for existing GPT-5.5 workloads at moderate cost. Which model is it?
Show answer
Answer: A.
Terra is described as the pragmatic all-rounder and the natural replacement for GPT-5.5 workloads. B is the top-tier flagship for the hardest work. C is for clear high-volume tasks. D is a costlier previous-generation model, not the replacement.
A guarantee is needed that a response is machine-parseable AND that a specific enum field only ever contains one of three allowed values. What MOST reliably enforces both?
Show answer
Answer: A.
A schema with an enum constraint guarantees both parseability and the restricted value set. B is guidance the model can violate. C catches errors after the fact and wastes calls. D reduces variance but does not enforce an enum.
After a model upgrade, the eval score drops on a cluster of long-input questions but is flat elsewhere. What is the MOST appropriate response?
Show answer
Answer: A.
A localised regression should be diagnosed on the affected cluster to understand cause before acting. B discards a potential improvement without evidence. C ignores a real regression. D hides the problem by removing the very cases that reveal it.
A prompt change fixes the single complaint you received, but you have no dataset. Before shipping, which TWO steps are the SAFEST?
Show answer
Answer: A and B.
Creating a small representative set with the failing case and comparing both prompts gives evidence the fix generalises without regressions. C ships on a single anecdote. D changes cost without evidence. E tests in production, exactly what an eval avoids.
A grader for a support-triage classifier is being chosen. The task has known correct categories. Which grading approach is MOST appropriate and cheapest?
Show answer
Answer: A.
With known categories, comparing predicted to true labels is deterministic, precise and cheap. B introduces judge noise for a task with definite answers. C does not scale. D is meant for open text, not fixed categories.
A team wants to trust an LLM-as-judge for faithfulness at scale. What is the correct ORDER of operations?
Show answer
Answer: A.
Calibration precedes reliance: label a sample, verify agreement, then scale only if it is adequate. B scales an unvalidated judge first. C skips the calibration entirely. D risks self-preference bias with no independent check.
Which is the STRONGEST reason online evaluation is needed even when offline evals pass?
Show answer
Answer: A.
Real traffic surfaces distributions the offline set could not foresee, which is precisely what online evaluation catches. B is not the reason and is not generally true. C is false. D is wrong: baselines still matter online.
The prompt optimizer is proposed to improve a prompt. What does it REQUIRE to be genuinely useful?
Show answer
Answer: A.
Optimising a prompt meaningfully needs an eval with examples and graders to measure whether changes help. B is too little signal to optimise against. C is unrelated to optimisation feedback. D is not a prerequisite for optimisation.
An eval dataset of 12 items gives scores that jump between 58% and 83% across identical runs. What is the MOST likely problem and fix?
Show answer
Answer: A.
On 12 items each flip moves the score by roughly eight points, so the swings are a small-sample artefact fixed by enlarging the set. B misattributes normal variance to a broken model. C removes the measurement. D does not address sample size.
Given the current OpenAI docs, which statement about evals and fine-tuning is accurate?
Show answer
Answer: A.
The docs group Evals and fine-tuning under Legacy APIs while evals still matter conceptually. B confuses legacy grouping with removal. C misstates fine-tuning's status. D misplaces evals inside the Agents API core.
A grounded-answer eval passes at launch. Which change should trigger re-running it BEFORE shipping the change?
Show answer
Answer: A.
Grounding depends on chunking, retrieval, prompt and model, so any of those changes warrants re-running the eval. B, C and D do not touch the pipeline that determines whether answers stay supported by context.
A team optimises cost by switching from Sol to Luna but does not re-run the eval. What is the BEST description of the risk?
Show answer
Answer: A.
Changing the model without re-evaluating means any quality regression goes undetected until users feel it. B is not necessarily true and misses the point. C is false; the family shares 1.05M context. D ignores that cost cuts can trade away quality.
A startup needs the FASTEST path to a durable cloud agent with hosted sandboxed execution and long-running artifacts, and does not need EU residency or ZDR. Which runtime is the BEST fit?
Show answer
Answer: A.
The Agents API gives OpenAI-managed durable sessions, hosted sandboxes and artifacts with no self-hosting, and the US-only/no-ZDR constraint is acceptable here. B requires building and hosting the runtime. C and D lack managed durability, sandboxes and recovery.
Which TWO capabilities make the Agents SDK the right choice for a firm that must keep ALL processing under its own governance and infrastructure?
Show answer
Answer: A and B.
Self-hosted execution and code-first orchestration/guardrails keep processing under the firm's control. C, D and E all describe the OpenAI-managed Agents API, which is the opposite of keeping everything under the firm's own governance.
An EU healthcare provider wants durable cloud agents but has a firm EU-residency and ZDR mandate. What is the correct conclusion about the Agents API, even with a self-hosted sandbox?
Show answer
Answer: A.
The Agents API is US-residency only with no ZDR, and a self-hosted sandbox does not lift that; the mandate rules it out. B misreads the sandbox choice as changing residency. C and D address network security, not residency or retention.
In a managed agent session, why should you bound
max_concurrent_subagentsrather than leave it unbounded?Show answer
Answer: A.
Bounding concurrency caps parallel resource use, cost and the blast radius of delegated work. B is false since subagents can run concurrently. C confuses concurrency with context size. D is unrelated and undesirable.
A prompt-injection payload hidden in a retrieved document tries to make an agent call a tool that exfiltrates data. Which control MOST directly limits the damage if the injection succeeds?
Show answer
Answer: A.
If tools are scoped to least privilege, a hijacked agent simply cannot reach the exfiltration capability, containing the damage. B and C may reduce but cannot guarantee resistance to injection. D is guidance an injection can override.
You are building a triage-and-specialist agent system. Which TWO design elements correctly match their purpose?
Show answer
Answer: A and B.
A handoff transfers control to a specialist, and a subagent handles a bounded sub-task and returns to the parent. C and D misdescribe the mechanics, and E places handoffs in an unrelated offline API.
A developer chooses the Agents API for a simple assistant that makes one internal function call, needs full per-step control, and has no session or sandbox requirement. Why is this the WRONG runtime?
Show answer
Answer: A.
For a single tool call with full control and no session/sandbox needs, the managed harness is overkill; the Responses API with function calling is the right, simpler tool. B and C are false. D wrongly ties function calling to training.
An agent takes an occasional wrong action and the team cannot reconstruct the decision path afterward. Which capability BEST closes this gap?
Show answer
Answer: A.
Reconstructing a past decision requires traces of steps, tool calls and their inputs/outputs. B, C and D add capacity, parallelism or budget but none of them produces the audit trail needed to explain a past action.
A team wants OpenAI to summarise context and resume sessions after interruptions for a multi-hour agent. Which runtime provides these as managed features?
Show answer
Answer: A.
The Agents API's managed harness provides context summarisation and session resumption out of the box. B leaves these for you to implement. C requires you to hand-build durability. D is an offline bulk tier with none of these.
For an agent that can take irreversible actions, which combination is REQUIRED before it acts autonomously in production?
Show answer
Answer: A.
Irreversible actions need least-privilege scoping to limit reach and a human gate to prevent unwanted commits. B, C and D adjust capability, parallelism or verbosity but none of them prevents an unwanted irreversible action.
How is a hosted-sandbox Agents API session billed compared with using a self-hosted sandbox, holding model and tools constant?
Show answer
Answer: A.
Hosted sandboxes add container rates on top of the model and tool rates, which self-hosting avoids. B ignores the container charge. C and D misstate the billing model, which keeps model and tool rates and only adds the container cost.
Protocols are revised monthly and the previous version must NEVER be quoted. Re-indexing the new version is not enough. What else is REQUIRED?
Show answer
Answer: A.
If the old version remains retrievable, it can still be quoted, so it must be removed or superseded in the index. B relies on the model obeying and is bypassable. C and D affect retrieval granularity and generation, not whether stale content is present.
Retrieval recall is high (the right chunks are fetched) but faithfulness is low (answers add unsupported claims). What does this MOST likely indicate?
Show answer
Answer: A.
High recall with low faithfulness isolates the problem to generation not adhering to context, so grounding must be tightened and evaluated. B contradicts the high recall. C and D target retrieval, which is already working.
Which TWO are the STRONGEST reasons to prefer RAG over fine-tuning for a frequently changing internal knowledge task that must cite sources?
Show answer
Answer: A and B.
RAG handles frequent change by re-indexing and can cite retrieved sources, both decisive here. C is false as a blanket claim. D is wrong: RAG still needs grounding evals. E overpromises; grounding reduces but cannot guarantee zero hallucination.
A grounded assistant cites the wrong section for a correct-sounding answer. What is the BEST diagnostic step?
Show answer
Answer: A.
Diagnosing a mis-citation means examining the retrieved chunks and checking whether the cited one supports the claim, which localises retrieval versus generation fault. B and C change behaviour without diagnosis. D hides the very signal that exposes the problem.
Semantic retrieval works for prose but fails on queries containing exact SKUs like
SKU-88231. What is the MOST direct remedy?Show answer
Answer: A.
Exact identifiers are matched by lexical search, so a hybrid retriever fixes the SKU misses directly. B is heavy and still meaning-biased. C fetches more chunks but not the right exact match. D affects generation, not retrieval.
A team validates RAG quality only by reading a handful of answers that 'sound right'. What is the BEST description of what is missing?
Show answer
Answer: A.
Sounding right is not measured grounding; a grounded-answer eval on a representative set is what is missing. B, C and D change infrastructure but none of them measures whether answers are actually supported by retrieved context.
Answers are bloated with unrelated text after chunk size was raised to 5,000 tokens. What is the BEST first adjustment?
Show answer
Answer: A.
Oversized chunks carry unrelated content, so shrinking them to focused passages reduces the noise. B worsens it. C changes generation, not the noisy retrieval input. D risks splitting boundary-spanning answers without fixing bloat.
A clinician-facing RAG system must restrict each user to their own department's protocols. A developer proposes enforcing this in the system prompt. Why is that insufficient?
Show answer
Answer: A.
Access control is a security boundary that must live in retrieval; a prompt rule can be circumvented and out-of-scope documents would still be retrievable. B and C are false. D is not the reason; correctness and security are.
A developer wants to stuff a whole 900-page manual into every prompt because it fits in 1.05M tokens. Which is the STRONGEST argument for RAG instead?
Show answer
Answer: A.
Even when it fits, paying to process the entire manual on every call is wasteful and can distract the model; RAG sends only what the query needs. B is false since it fits. C is untrue. D overpromises perfection.
A feature is both too slow and too expensive on
gpt-5.6-solathigheffort. What is the correct ORDER of operations to fix it?Show answer
Answer: A.
Measure first, then reduce effort/model to the cheapest that still passes the eval, then apply caching or batching where appropriate. B ignores interactivity needs and skips measurement. C changes model with no quality check. D adds latency and cost.
An eval shows
gpt-5.6-lunameets the same targetgpt-5.6-solcurrently hits on a high-volume task. Which TWO actions are the MOST cost-effective and safe?Show answer
Answer: A and B.
Luna passes, so switching cuts cost, and continued monitoring guards against regression, which is both cost-effective and safe. C keeps paying more for no quality gain. D increases cost further. E removes the very check that makes the switch safe.
Which workload is the WRONG fit for the Batch API?
Show answer
Answer: A.
Batch trades latency for cost and is unsuitable for a synchronous feature where users wait in real time. B, C and D are latency-tolerant offline jobs that are exactly what Batch is for.
A workload runs 500,000 requests per day at 1,500 input and 500 output tokens on
gpt-5.6-terra($2/M input, $12/M output). Roughly what is the daily cost?Show answer
Answer: A.
Input: 500,000 x 1,500 = 750M x $2/M = $1,500. Output: 500,000 x 500 = 250M x $12/M = $3,000. Total ~$4,500/day, so A. B is a tenth too small, C omits the output cost, and D is ten times too high.
A latency-tolerant service wants cheaper-than-standard processing but cannot accept Batch's hours-long turnaround. Which tier is the BEST fit and why?
Show answer
Answer: A.
Flex processing is the middle tier that reduces cost for latency-tolerant work without the multi-hour Batch wait. B optimises speed and costs more. C is for real-time interaction. D is the baseline the service is trying to beat on price.
Time-to-first-token is the dominant latency term for an interactive feature on a reasoning model at
higheffort. Which TWO levers help MOST directly?Show answer
Answer: A and B.
Lower effort shortens the thinking before the first token, and caching the shared prefix removes reprocessing time, both cutting time-to-first-token. C lengthens output. D adds input to process. E is offline and irrelevant to interactive latency.
A document-editing feature regenerates a 20,000-token document where roughly 97% is unchanged. Which feature cuts latency MOST, and why?
Show answer
Answer: A.
Predicted outputs speed generation when most of the output equals a known reference, exactly the near-identical edit case. B helps repeated inputs, not near-identical outputs. C is offline. D adds capacity but not generation speed for known output.
Under sustained load your service receives frequent
429responses. Beyond exponential backoff, which control BEST prevents overrunning your allocation?Show answer
Answer: A.
Shaping outbound traffic to your rate limits prevents repeatedly hitting 429s in the first place, complementing backoff. B intensifies the overload. C and D affect quality and length, not request rate.
A sensitive workload must reach the API WITHOUT traversing the public internet. Which control fits BEST?
Show answer
Answer: A.
Private Link provides a private network path that avoids the public internet. B restricts which public IPs may connect but still uses the public internet. C caps cost. D changes throughput, neither of which provides a private path.
You are writing a pre-launch checklist for an agent that can take irreversible actions on customer accounts. Which TWO items MUST be on it?
Show answer
Answer: A and B.
An approval gate stops unwanted irreversible actions and least-privilege scopes plus spend limits contain reach and cost. C and D raise capacity or cost without addressing safety. E removes the observability needed to investigate incidents.
What is the key difference between red teaming and misalignment monitoring?
Show answer
Answer: A.
Red teaming is a proactive pre-launch probe for weaknesses, while misalignment monitoring is an ongoing watch on deployed behaviour. B denies a real distinction. C reverses their timing. D miscategorises safety practices as billing.
A read-only reporting job uses an API key that also carries permission to delete resources. Which principle is violated and what is the correct fix?
Show answer
Answer: A.
A reporting job holding delete rights violates least privilege; the fix is a read-only scoped key. B, C and D are legitimate concepts but none describes granting more permissions than the task needs.
Last updated Sep 18, 2026