AI Cert Prep
Type to search documentation.

Applied AI Foundations

Applied AI Foundations · Mock Exam 2

A second, harder 50-item independent mock exam for the Applied AI Foundations track, with new items, multi-constraint stems and a readiness indicator.

This is Mock Exam 2 for the Applied AI Foundations track: 50 new items across the six domains, none repeating Mock Exam 1 or the domain-page questions. It is an independent mock exam built from publicly available OpenAI learning objectives — not an official OpenAI assessment and not endorsed by OpenAI. Mock Exam 2 is deliberately harder: more multi-constraint stems and more FIRST / BEST / MOST cost-effective / TWO qualifiers, so you must weigh competing considerations rather than spot a single keyword. Use it as your readiness gate: reach 80%+ here, timed, before you take the real OpenAI Academy assessment.

Instructions

  • Time: 60 minutes; sit it timed, under mock conditions, once you have cleared Mock Exam 1 and revised your weak domains.
  • Items: 50, single-response and multiple-response. Each item states how many answers to select.
  • Selection: for a select-one item choose exactly one option; for a select-two item you must choose both correct options and no incorrect one to score the item.
  • No guessing penalty: answer every item — an unanswered item scores the same as a wrong one.
  • Target: aim for ≥ 80% raw (≈ 40/50) before you take the real Academy assessment, which itself passes at 80%. Because this mock runs harder, a comfortable pass here is a stronger readiness signal than the same score on Mock Exam 1.

Domain distribution

#DomainWeightItems here
1Finding and Scoping Opportunities16%8
2Decomposing Work into Steps18%9
3Inputs, Outputs and Contracts16%8
4Choosing the Right Capability20%10
5Review Points and Human Oversight16%8
6Repeatability and Improvement14%7

Total: 8 + 9 + 8 + 10 + 8 + 7 = 50 items.

Readiness interpretation

This is an independent readiness indicator, never an official score.

Raw score (of 50)BandInterpretation
40+ (80%+)Assessment readyOn the harder mock this is a strong signal you are ready for the Academy assessment
35–39 (70–79%)Building confidenceNearly there; target the domains where you lost the most marks
under 70%Keep learningReturn to the domain pages, especially D2 and D4, before re-attempting
45+ (90%+)Strong readinessExcellent margin on the harder items

Take the mock exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it. Compare your per-domain results with Mock Exam 1 — a domain strong on Mock 1 but weak here signals shallow understanding worth revisiting.

Interactive mode

Take the practice exam

50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Finding and Scoping OpportunitiesSelect one

    Four candidates pass the value screen, but you can build only one this quarter. A saves 6 h/week with a light monthly sample; B saves 14 h/week but every output is an irreversible external filing needing per-item approval; C saves 7 h/week but touches regulated data your policy forbids leaving a sanctioned tool you do not yet have; D saves 5 h/week, fully correctable, low-sensitivity. Which should you build FIRST?

    • A. B, because 14 hours is the largest raw saving
    • B. A, because its net payback after light review is high and its risk is low, and it is buildable now
    • C. C, because regulated work is the most valuable
    • D. D, because the smallest task is always safest
    Show answer

    Answer: B.

    A has strong risk-adjusted payback, low risk, and no blocking dependency, so it is the best first build. B's per-item approval of irreversible filings erodes most of its 14 hours (A ignores that). C is policy-blocked until you have the sanctioned tool (C). D is buildable but its payback is the smallest of the buildable options (D).

  2. Q2D1 · Finding and Scoping OpportunitiesSelect one

    A recurring task happens 200 times a week and saves real time, but roughly one in fifty outputs is an irreversible customer-facing action while the rest are internal and correctable. What is the MOST appropriate scoping decision?

    • A. Do not automate at all, because some outputs are irreversible
    • B. Automate the whole thing end-to-end because most outputs are correctable
    • C. Automate the correctable majority with sampling, and route the irreversible customer-facing slice to a mandatory human gate
    • D. Automate everything and add a review only after the actions have happened
    Show answer

    Answer: C.

    The right move splits the stream: sample the correctable majority and gate the irreversible slice, capturing value while containing risk. Abandoning it (A) throws away the correctable value. End-to-end automation (B) puts irreversible actions in the model's hands, and a review after the fact (D) is not a gate at all.

  3. Q3D1 · Finding and Scoping OpportunitiesSelect one

    A stakeholder insists the FIRST automation should be the CEO's quarterly investor deck because 'it is the most important thing we produce'. Which counter-argument is strongest?

    • A. Investor decks cannot be produced with AI at all
    • B. At four instances a year the workflow never amortises and its low error tolerance argues for manual assistance, not a productised workflow
    • C. The deck is too short to be worth automating
    • D. The CEO would object to using AI
    Show answer

    Answer: B.

    Importance is not frequency: a four-times-a-year, low-tolerance showpiece is the classic impressive-but-rare trap and should be AI-assisted manually. Decks can be AI-assisted (A). Length (C) and a presumed objection (D) are not the analytical reason it fails the screen.

  4. Q4D1 · Finding and Scoping OpportunitiesSelect one

    You have scoped a workflow to 'draft first-reply text for known ticket types'. Over three weeks it quietly grows to also tag priority, send the reply, and issue small credits. What scoping element failed, and what is the fix?

    • A. The trigger failed; add a second trigger
    • B. The out-of-scope boundary failed; restate what the workflow must not do (no sending, no credits) and hold the line
    • C. The model was too small; upgrade it
    • D. Nothing failed; scope creep is healthy growth
    Show answer

    Answer: B.

    Silent expansion into sending and issuing credits is scope creep that a written out-of-scope boundary is meant to prevent. The trigger (A) governs when it runs, not what it must avoid. Model size (C) is irrelevant, and unmanaged creep into irreversible actions (D) is a risk, not healthy growth.

  5. Q5D1 · Finding and Scoping OpportunitiesSelect two

    Which TWO characteristics most strongly signal that a task should keep a human fully in the loop rather than be automated end-to-end, even when it is frequent and time-consuming?

    • A. The action is irreversible once taken
    • B. The output is internal and easily corrected before use
    • C. The task involves a regulated decision with legal exposure
    • D. The data is public marketing copy
    • E. The task is repetitive and tedious
    Show answer

    Answer: A and C.

    Irreversibility and regulated legal exposure are the risk signals that keep a human in the loop regardless of volume. Internal correctable output (B) and public data (D) lower risk, and tedium (E) is a reason to automate, not to keep it manual.

  6. Q6D1 · Finding and Scoping OpportunitiesSelect one

    Two tasks each save about 10 hours a week. Task P is fully repeatable end to end; task Q has a large, variable human-judgment core that must be exercised on every instance. Which is the BETTER automation candidate and why?

    • A. Q, because judgment tasks are more valuable
    • B. P, because a bigger stable repeatable core means more of the 10 hours is actually removed, whereas Q's variable judgment stays manual
    • C. They are identical since both save 10 hours
    • D. Neither, because 10 hours is too small
    Show answer

    Answer: B.

    The realizable saving depends on how much of the task is a stable repeatable core; Q's large variable judgment part cannot be automated, so less of its 10 hours is actually removed. Judgment tasks are not automatically more valuable (A), the two are not identical once you separate core from judgment (C), and 10 hours is a substantial saving (D).

  7. Q7D1 · Finding and Scoping OpportunitiesSelect one

    A one-time migration will take three days of tedious manual reformatting. A colleague proposes building a reusable workflow for it. What is the MOST cost-effective decision?

    • A. Build the reusable workflow because three days is a large saving
    • B. Do it ad hoc with AI assistance, because at a frequency of one the workflow overhead never pays back
    • C. Skip AI entirely and do it by hand
    • D. Build the workflow and schedule it to run monthly just in case
    Show answer

    Answer: B.

    Frequency of one means a productised workflow never runs again, so ad hoc AI assistance captures the saving without the build overhead. Building a reusable workflow (A) or scheduling it monthly (D) wastes effort on a task that will not recur, and doing it by hand (C) forgoes the available AI help.

  8. Q8D2 · Decomposing Work into StepsSelect one

    A teammate proposes fixing an unreliable mega-prompt that produces a launch pack by (1) upgrading to the most expensive model AND (2) adding three more paragraphs of instructions. Why is decomposition the BETTER first move?

    • A. The expensive model is banned by policy
    • B. Decomposition makes each failure visible and testable and lets cheaper models handle each step, whereas a bigger model plus more instructions hides drops at higher cost
    • C. Cheaper models are always more accurate than expensive ones
    • D. Decomposition removes the need for any review
    Show answer

    Answer: B.

    The problem is structural, not capability; decomposing exposes and isolates failures and often lets lighter models run each step. The expensive model is not banned (A), cheaper models are not always more accurate (C), and decomposition does not remove review (D).

  9. Q9D2 · Decomposing Work into StepsSelect one

    You are designing a national brief from twelve regional reports that must reconcile terminology and totals, and each report buries the figures in prose. Which composed structure is BEST?

    • A. Sequential only, one report after another
    • B. Fan-out with extract-then-transform on each report, then a fan-in draft-then-critique that synthesises and reconciles
    • C. Draft-then-critique on all twelve raw reports at once
    • D. One mega-prompt with all twelve pasted in
    Show answer

    Answer: B.

    Independent reports fan out (each extracted to a mini-structure) and the fan-in synthesises with a critique that reconciles terms and totals — patterns composed to fit. Pure sequential (A) ignores independence, critiquing all twelve raw (C) skips per-report structure, and one mega-prompt (D) is the anti-pattern.

  10. Q10D2 · Decomposing Work into StepsSelect one

    A workflow was split into a dozen tiny steps, each doing a fraction of a job, and it is now slow and loses context between handoffs while adding no new checks. What is the BEST corrective action?

    • A. Split it into even more steps for robustness
    • B. Merge steps that share a single job and need no intermediate check, keeping only splits that separate distinct jobs or add an inspection point
    • C. Collapse everything back into one mega-prompt
    • D. Upgrade the model to remove the handoffs
    Show answer

    Answer: B.

    Over-splitting adds latency and context loss without new inspection, so the fix is to merge steps that share one job while retaining meaningful splits. More steps (A) worsens it, one mega-prompt (C) is the opposite anti-pattern, and a bigger model (D) does not remove handoffs.

  11. Q11D2 · Decomposing Work into StepsSelect one

    A contract-review workflow drafts a memo then checks it in the same reply, and reviewers notice the self-check almost never flags anything. What single change MOST improves the critique's objectivity?

    • A. Ask the model to try harder in the same reply
    • B. Make the critique a separate step that scores the draft against an explicit named checklist
    • C. Increase the max output tokens
    • D. Lower the temperature of the drafting step
    Show answer

    Answer: B.

    A same-pass self-review defends its own draft; a separate checklist-driven critique step is the fix that makes review objective. 'Try harder' (A) is not a structural change, token limits (C) do not affect objectivity, and drafting temperature (D) does not fix the missing separate pass.

  12. Q12D2 · Decomposing Work into StepsSelect one

    Which ordering is correct for a robust client proposal built from a discovery transcript, where numbers must be checked and the summary must match the final body?

    • A. Executive summary, then body, then pricing, then a final check
    • B. Extract scope and dates, produce timeline and pricing with a totals check, draft the body from those, write the summary last, then run a terms check
    • C. One prompt producing everything at once
    • D. Pricing, then summary, then scope, then body
    Show answer

    Answer: B.

    Extract first, produce the numbers with a check where they are created, draft from those, summarise last, then check against terms — each step inspectable and the summary matching the final body. Summary-first (A) drifts, one prompt (C) is the mega-prompt, and pricing before scope (D) has no inputs to price.

  13. Q13D2 · Decomposing Work into StepsSelect one

    A task involves reading a dense supplier email and then deciding, from the extracted terms, whether it breaches a policy threshold. A single prompt sometimes flags a breach the numbers do not support. What is the ROOT-CAUSE decomposition?

    • A. Draft-then-critique the email
    • B. Extract the terms into a structure first, then apply the threshold rule to that structure so the extraction and the rule are separately checkable
    • C. Ask for a longer, more careful single answer
    • D. Fan out across the email's paragraphs
    Show answer

    Answer: B.

    Blending extraction and the threshold decision hides whether an error is a misread term or a misapplied rule; extract-then-transform separates them for checking. Draft-then-critique (A) reviews prose, a longer single answer (C) keeps the errors blended, and fanning over paragraphs (D) does not separate extraction from the rule.

  14. Q14D2 · Decomposing Work into StepsSelect two

    A single mega-prompt produces a five-part deliverable that drops parts and contradicts itself. Which TWO structural fixes address these two distinct symptoms at the root?

    • A. Generate each part as its own named fan-out step so a dropped part cannot hide
    • B. Add a sentence telling the model not to drop parts
    • C. Extract a shared fact sheet first so every part draws on one source of truth
    • D. Switch to the most expensive available model
    • E. Generate the whole thing twice and keep the longer output
    Show answer

    Answer: A and C.

    Named fan-out steps make dropped parts visible, and a shared extracted fact sheet removes the contradictions at the root. An instruction not to drop parts (B) is prompt padding, a bigger model (D) hides rather than exposes failures, and generating twice (E) doubles cost without fixing either cause.

  15. Q15D3 · Inputs, Outputs and ContractsSelect one

    A downstream billing step requires a numeric total, but the producing step's contract only promises a formatted string like '$1,204.00'. What is the BEST fix?

    • A. Parse the string downstream and hope the format never changes
    • B. Change the producer's output contract to return the total as a number, matching the consumer's required input type
    • C. Add a bigger model to the consumer
    • D. Ask the producer to be more consistent
    Show answer

    Answer: B.

    The clean fix aligns the producer's output type to the consumer's required numeric input rather than relying on fragile downstream parsing. Parsing and hoping (A) is brittle to format changes, a bigger model (C) does not fix a type mismatch, and 'be more consistent' (D) is not a contract change.

  16. Q16D3 · Inputs, Outputs and ContractsSelect one

    An extraction step outputs sentiment as free text like 'mostly positive' instead of the expected enum, and a colleague running it gets yet other phrasings. Which change MOST directly fixes both problems?

    • A. Ask the model to be more careful
    • B. Constrain the output to an enumerated set {positive, neutral, negative} in the schema and return only that structure
    • C. Increase max tokens
    • D. Add more example emails to the prompt
    Show answer

    Answer: B.

    An enumerated value set in the schema forces one of the allowed values for every user and input; free text means the value set was never constrained. 'Be careful' (A) is not a contract, max tokens (C) is unrelated, and more examples (D) do not constrain the output enum.

  17. Q17D3 · Inputs, Outputs and ContractsSelect one

    A step that classifies feedback outputs a product area even when the email names no product, polluting the CRM with guessed values. What is the BEST contract change?

    • A. Let the model keep guessing but log the guesses
    • B. Add a missing-input rule: if no product is identifiable, output 'unknown' and set needs_review, and constrain product_area to the allowed list plus 'unknown'
    • C. Remove the product_area field entirely
    • D. Raise the temperature for more variety
    Show answer

    Answer: B.

    An explicit missing-input rule (unknown plus a review flag) combined with a constrained value set stops the fabrication. Guess-and-log (A) still pollutes with fabricated values, removing the field (C) drops a required output, and higher temperature (D) worsens variance.

  18. Q18D3 · Inputs, Outputs and ContractsSelect one

    You are choosing the output shape for a step whose result a human reads and must check for completeness. Which shape is MOST appropriate?

    • A. A strict JSON schema with only the JSON returned
    • B. A fixed template with named sections in a fixed order, so a missing section is visible
    • C. Free-form prose in whatever style the model prefers
    • D. A single unstructured paragraph
    Show answer

    Answer: B.

    For a human-read deliverable checked for completeness, a fixed template makes a missing section visible. A strict JSON schema (A) suits a parser, not a reader, and free-form prose (C) or a single paragraph (D) makes completeness hard to verify.

  19. Q19D3 · Inputs, Outputs and ContractsSelect one

    A colleague says 'the output looked fine in my three tests, so the shape is defined'. Why is this reasoning unsafe?

    • A. Three tests is a policy violation
    • B. An unpinned output shape can still vary across other inputs and other users; a few good samples are not a specification
    • C. The model must be retrained after three tests
    • D. Testing always proves the contract is complete
    Show answer

    Answer: B.

    A handful of passing samples does not pin the shape; the format can drift on unseen inputs or with another user unless the contract specifies it. There is no such policy (A), no retraining requirement (C), and testing does not prove a contract is complete (D).

  20. Q20D3 · Inputs, Outputs and ContractsSelect one

    A required 'known_customers' reference is absent when a lookup step runs, and correctness fully depends on it. What is the correct missing-input behaviour?

    • A. Default to the most common customer
    • B. Fail loudly with a clear reason, since no safe default exists and correctness depends on the reference
    • C. Guess a plausible customer to keep moving
    • D. Return the first customer alphabetically
    Show answer

    Answer: B.

    When correctness depends on a value and no safe default exists, the step must fail loudly rather than fabricate. A most-common default (A), a guess (C), and an alphabetical pick (D) all substitute invented data for a required input.

  21. Q21D3 · Inputs, Outputs and ContractsSelect one

    Why does writing the step contract matter MOST for handover and repeatability?

    • A. It makes the prompt shorter
    • B. The contract is the interface: it tells a colleague exactly what inputs to give, what output to expect, and when the step is done
    • C. It lets you skip acceptance criteria
    • D. It removes the need for any review
    Show answer

    Answer: B.

    A written contract is the runnable interface a colleague follows to get the same result; an undocumented step lives only in the author's head. It does not shorten prompts (A), skip criteria (C), or remove review (D).

  22. Q22D3 · Inputs, Outputs and ContractsSelect two

    You are hardening a two-step chain where step 1 feeds step 2. Which TWO checks MOST reduce flakiness at the boundary?

    • A. Confirm step 1's output shape exactly matches step 2's required input, including field names and types
    • B. Confirm both steps use the same temperature
    • C. Define what step 2 does when a required field from step 1 is missing or malformed
    • D. Confirm both prompts are the same length
    • E. Confirm the two steps run on the same day
    Show answer

    Answer: A and C.

    Matching the producer's output to the consumer's required input, and defining step 2's missing/malformed-input behaviour, are the two boundary checks that remove most flakiness. Matching temperature (B), prompt length (D) and run day (E) are not contract-boundary concerns.

  23. Q23D4 · Choosing the Right CapabilitySelect one

    A manager says 'just build one API application' for a bundle that is (a) your own recurring digest, (b) a version reps run themselves in ChatGPT, and (c) a future scheduled Slack post. What is the BEST right-sizing?

    • A. One API application for all three
    • B. A Project for (a), a shared custom GPT for (b), and an API application only for (c)
    • C. A custom GPT for all three
    • D. Keep everything as weekly prompts
    Show answer

    Answer: B.

    Each need maps to a different rung: recurring solo context is a Project, others self-serving is a custom GPT, and scheduled integrated automation is an API application. One API app for all (A) over-builds (a) and (b), a custom GPT for all (C) cannot do scheduled Slack posting, and weekly prompts (D) neither scale nor self-serve.

  24. Q24D4 · Choosing the Right CapabilitySelect one

    A stem says the output 'must be a sourced, cited market report that is THEN refined into a polished brief over several edits'. Which single choice best matches the FULL requirement?

    • A. Search only, for a quick current fact
    • B. Deep research for the sourced report, then Canvas for the iterative refinement
    • C. Data analysis over an uploaded spreadsheet
    • D. A workspace agent to answer one question
    Show answer

    Answer: B.

    A sourced, cited report calls for deep research, and refining it across several edits calls for Canvas — both parts of the requirement. Search (A) under-delivers the report, data analysis (C) fits no file-computation need here, and a one-question agent (D) does not match a multi-edit deliverable.

  25. Q25D4 · Choosing the Right CapabilitySelect one

    You keep re-pasting a 4-page background into fresh chats for a monthly task and only you run it. Which is the MOST cost-effective fix?

    • A. Build a full API application immediately
    • B. Move the background into a Project's knowledge files with standing instructions
    • C. Build a custom GPT even though no one else runs it
    • D. Keep re-pasting but use a bigger model
    Show answer

    Answer: B.

    Recurring standing context for a solo task is exactly a Project, the lightest rung that removes the re-pasting. An API app (A) and a custom GPT (C) over-reach when only you run it, and a bigger model (D) does not fix the re-pasting or the missing standing context.

  26. Q26D4 · Choosing the Right CapabilitySelect one

    A workflow currently done as weekly prompts must eventually run automatically every Monday and post to a Slack channel. Which rung does THAT specific requirement justify, and why?

    • A. A saved instruction, because it repeats
    • B. A Project, because it needs standing context
    • C. An API application, because scheduled, automated, Slack-integrated execution must run in code
    • D. A custom GPT, because others might run it
    Show answer

    Answer: C.

    Scheduled, automated, integrated-with-Slack execution is embedding at scale, the justification for an API application. A saved instruction (A) and a Project (B) do not run on a schedule wired into Slack, and a custom GPT (D) is for others to self-serve, not scheduled posting.

  27. Q27D4 · Choosing the Right CapabilitySelect one

    A colleague reaches for deep research to answer 'what is our competitor's stock price right now?'. What is the BEST correction?

    • A. Deep research is correct because it is thorough
    • B. Use search: a single current fact is a quick lookup, and deep research is reserved for sourced multi-source reports
    • C. Use data analysis on a spreadsheet
    • D. Use Canvas to iterate on the answer
    Show answer

    Answer: B.

    A single current fact is a search job; deep research is heavier and reserved for cited, multi-source synthesis. Deep research here (A) wastes time, data analysis (C) needs a file to compute over, and Canvas (D) is an editing surface, not a lookup tool.

  28. Q28D4 · Choosing the Right CapabilitySelect one

    Two tasks are proposed: (1) a repeatable assistant that colleagues chat with turn by turn, and (2) a multi-step task delegated to run semi-autonomously with checkpoints. Which mapping is correct?

    • A. (1) workspace agent, (2) custom GPT
    • B. (1) custom GPT, (2) workspace agent
    • C. Both are API applications
    • D. Both are one-off prompts
    Show answer

    Answer: B.

    A turn-by-turn assistant others chat with is a custom GPT, while delegated multi-step execution under oversight is a workspace agent. Swapping them (A) inverts answer-vs-execute, and neither is inherently an API app (C) or a one-off prompt (D).

  29. Q29D4 · Choosing the Right CapabilitySelect one

    A Project that produces a weekly digest has started giving stale, subtly wrong outputs even though the prompt is unchanged. What is the MOST likely cause?

    • A. The model silently downgraded itself
    • B. The Project's knowledge files have gone stale and are misleading the output
    • C. The temperature drifted upward
    • D. Projects cannot produce reliable output
    Show answer

    Answer: B.

    Heavier rungs carry a maintenance cost; stale Project knowledge files silently mislead the output even when the prompt is untouched. Models do not silently downgrade (A), temperature does not drift on its own (C), and Projects are perfectly capable of reliable output when maintained (D).

  30. Q30D4 · Choosing the Right CapabilitySelect two

    Which TWO tasks are best served by an in-conversation capability rather than by moving to a heavier packaging rung?

    • A. Finding outliers and plotting a trend in an uploaded spreadsheet (data analysis)
    • B. Letting fifty colleagues self-serve the same configured assistant
    • C. Getting this week's current news on three named accounts (search)
    • D. Running the workflow automatically on a nightly schedule
    • E. Embedding the workflow inside your own product's UI
    Show answer

    Answer: A and C.

    Computation over a file (data analysis) and a current-facts lookup (search) are in-conversation capabilities. Fifty colleagues self-serving (B) is a custom GPT, and nightly scheduling (D) or embedding in a product UI (E) are API-application needs.

  31. Q31D4 · Choosing the Right CapabilitySelect one

    A team is about to build a custom GPT for a task that, on inspection, only its author ever runs and simply needs the same three files each time. What is the BEST advice?

    • A. Proceed with the custom GPT; it feels more capable
    • B. Use a Project instead — solo recurring context is the lighter, sufficient rung, and a custom GPT over-reaches
    • C. Build an API application for future-proofing
    • D. Keep re-pasting the files each run
    Show answer

    Answer: B.

    When only the author runs it and it needs standing context, a Project is the correct lighter rung; a custom GPT is justified only when others run it. Proceeding with the GPT (A) over-reaches, an API app (C) over-reaches further, and re-pasting (D) is the waste a Project removes.

  32. Q32D5 · Review Points and Human OversightSelect one

    A support workflow auto-sends replies for 'simple' categories. One category quotes a delivery date pulled from a field that can be stale, and a wrong date reached a customer. What is the BEST targeted fix?

    • A. Gate every reply of every category
    • B. Route replies that quote dynamic data to a gate or verify the data before sending, while keeping self-contained replies automated
    • C. Stop using AI for support entirely
    • D. Lower the temperature
    Show answer

    Answer: B.

    The risk is specific to replies depending on stale-able data, so gate or verify those while leaving genuinely self-contained replies automated. Gating everything (A) is the throttling over-correction, abandoning AI (C) is disproportionate, and temperature (D) does not fix stale data.

  33. Q33D5 · Review Points and Human OversightSelect one

    A workflow shows 97% aggregate accuracy, so a manager proposes removing the gate before an irreversible legal filing. What is the BEST response?

    • A. Agree — 97% is high enough
    • B. Keep the gate: aggregate accuracy hides the irreversible, regulated tail where a single error is catastrophic
    • C. Remove the gate but add one after filing
    • D. Raise the accuracy target to 99% and then remove the gate
    Show answer

    Answer: B.

    A high average does not cover the low-frequency, high-consequence errors on an irreversible regulated step, so the mandatory gate stays. 97% 'good enough' (A) ignores the tail, a gate after filing (C) is no gate, and a higher target (D) still cannot make an irreversible regulated action safe to leave ungated.

  34. Q34D5 · Review Points and Human OversightSelect two

    A reviewer is approving 500 items an hour and only glancing at formatting. Which TWO changes MOST improve oversight without simply doing more work?

    • A. Switch most of the stream to sampling so each remaining gate is meaningful
    • B. Give the reviewer the acceptance criteria so they check substance, not just format
    • C. Increase the review target to 1,000 items an hour
    • D. Remove all review
    • E. Use a bigger model so review is unnecessary
    Show answer

    Answer: A and B.

    Fewer, meaningful gates via sampling, plus criteria-driven substantive review, fix rubber-stamping. Doubling the load (C) worsens fatigue, removing review (D) is reckless, and a bigger model (E) does not eliminate the need for oversight on risky steps.

  35. Q35D5 · Review Points and Human OversightSelect one

    A gate has been placed on the draft produced early in a workflow, but the draft is substantially rewritten by two later steps before the irreversible send. What is wrong with this placement?

    • A. Nothing — earlier review is always more proactive
    • B. The reviewer is approving an artefact that changes before the risky action, so the gate should move to just before the send
    • C. The gate should be removed entirely
    • D. The draft step should be deleted
    Show answer

    Answer: B.

    A gate must review the final artefact at the last reversible moment; reviewing an early draft that later changes wastes the review. Earlier is not always better (A), the gate is needed (C), and the draft step is legitimate (D) — only the gate's placement is wrong.

  36. Q36D5 · Review Points and Human OversightSelect one

    For a high-volume categorisation step, sampled error rate has been climbing for two weeks on high-value items specifically. What is the BEST response?

    • A. Do nothing; sampling already covers it
    • B. Tighten sampling and add a gate on the high-value slice where the drift is concentrated
    • C. Stop sampling to save reviewer time
    • D. Gate every item across all slices
    Show answer

    Answer: B.

    Drift concentrated in a risky slice calls for tighter sampling and a targeted gate on that slice, not blanket action. Doing nothing (A) ignores the drift, stopping sampling (C) removes the signal, and gating every item (D) is the throttling over-correction.

  37. Q37D5 · Review Points and Human OversightSelect one

    Which mode-to-step mapping is correct?

    • A. Gate a stable low-stakes internal note; sample an irreversible external publication
    • B. Gate an irreversible external publication; sample a high-volume correctable categorisation; spot-check a stable low-stakes note
    • C. Spot-check an irreversible payment; gate a low-volume brainstorm
    • D. Sample a regulatory filing; gate internal brainstorming
    Show answer

    Answer: B.

    Grave steps are gated, high-volume correctable work is sampled, and stable low-stakes work is spot-checked. The other options invert the mapping, e.g. sampling an irreversible publication (A), spot-checking an irreversible payment (C), or sampling a regulatory filing (D).

  38. Q38D5 · Review Points and Human OversightSelect two

    A team over-reacted to one bad auto-sent reply by gating all 20,000 daily replies. Which TWO consequences make this over-correction self-defeating?

    • A. It throttles throughput to human speed, erasing the value the automation created
    • B. It makes the underlying model less accurate
    • C. It breeds reviewer fatigue and rubber-stamping, so real errors slip through anyway
    • D. It violates OpenAI usage policy
    • E. It automatically documents the workflow
    Show answer

    Answer: A and C.

    A blanket gate on high-volume work throttles throughput and induces rubber-stamping, defeating the control. It does not change model accuracy (B), violate policy (D), or self-document the workflow (E).

  39. Q39D6 · Repeatability and ImprovementSelect one

    A false-positive regression appeared after many untracked prompt edits. You now adopt versioning. What is the BEST recovery?

    • A. Delete the workflow and rebuild from memory
    • B. Restore the last version with a low false-positive rate, then re-apply the individual changes one at a time, measuring after each
    • C. Keep all edits and lower the temperature
    • D. Add more rules to overwhelm the false positives
    Show answer

    Answer: B.

    Rolling back to a known-good version and re-applying changes one at a time with measurement isolates the offending edit and restores quality. Rebuilding from memory (A) loses the good history, keeping the edits and changing temperature (C) does not isolate the cause, and adding rules (D) compounds the noise.

  40. Q40D6 · Repeatability and ImprovementSelect one

    A support-reply prompt was edited to 'be more concise' and customer satisfaction dropped because replies now omit a needed next step. With versioning in place, what is the BEST response?

    • A. Rewrite the whole workflow from scratch
    • B. Roll back to the prior version, then change only the length guidance while keeping the next-step instruction
    • C. Keep the concise version; satisfaction will recover on its own
    • D. Add more knowledge files
    Show answer

    Answer: B.

    Versioning lets you revert the regression and re-apply a single isolated change. Starting over (A) discards working history, keeping a version that measurably hurt satisfaction (C) is not evidence-based, and knowledge files (D) do not address the omitted next step.

  41. Q41D6 · Repeatability and ImprovementSelect one

    Before optimising a workflow, you measure that 12 of its 18 minutes per run are spent in a manual review step. What is the MOST effective next move?

    • A. Spend the effort shortening the prompt
    • B. Redesign the review step, since measurement shows it is the real bottleneck
    • C. Switch to a cheaper model to save cost
    • D. Add more steps to the automated portion
    Show answer

    Answer: B.

    Cycle-time measurement locates the bottleneck; with review consuming most of the run, redesigning review is the biggest lever. Shortening the prompt (A) or adding automated steps (D) does not touch the dominant review time, and model price (C) is not the bottleneck here.

  42. Q42D6 · Repeatability and ImprovementSelect two

    Which TWO properties distinguish evidence-based iteration from vibes-based 'improvement'?

    • A. It changes one variable at a time so effects are attributable
    • B. It decides to keep or revert on a measured metric, with a rollback path
    • C. It changes several settings at once to improve faster
    • D. It judges success by whether the output feels better
    • E. It never keeps any change regardless of the metric
    Show answer

    Answer: A and B.

    Single-variable changes plus metric-based keep-or-revert with rollback are what make iteration attributable and reversible. Changing several settings at once (C) and judging by feel (D) are the vibes anti-pattern, and evidence-based iteration keeps changes that improve the metric (E).

  43. Q43D6 · Repeatability and ImprovementSelect one

    A workflow cut a task from 120 to 20 minutes per run. What does this measurement PRIMARILY support?

    • A. Nothing — time is irrelevant to workflow value
    • B. It quantifies the payback that justified the opportunity and gives a baseline for further improvement
    • C. It proves the output quality is high
    • D. It sets the model temperature
    Show answer

    Answer: B.

    Cycle-time savings quantify the payback screened at the opportunity stage and give a baseline to improve against. Time is highly relevant (A), it does not by itself prove quality (C) — that needs the acceptance-criteria metric — and it has nothing to do with temperature (D).

  44. Q44D6 · Repeatability and ImprovementSelect one

    A colleague offers 'record a video of me doing it once' as the documentation for handover. Why is a runbook the BETTER choice?

    • A. Videos are against company policy
    • B. A runbook captures trigger, inputs, ordered steps, review points, definition of done and failure handling in a form a colleague can execute unaided; a single video rarely covers those reliably
    • C. Videos cannot be stored
    • D. A runbook removes the need for review points
    Show answer

    Answer: B.

    A runbook is the structured, executable record of the workflow, whereas a one-off video usually misses inputs, failure handling and the definition of done. There is no such policy (A), videos can be stored (C), and a runbook records review points rather than removing them (D).

  45. Q45D6 · Repeatability and ImprovementSelect one

    Which of these is the correct order of the DRIVE maintenance loop for a live workflow?

    • A. Vary many things, then document, then evaluate by feel
    • B. Document, Record versions, Instrument, Vary one thing, Evaluate on the metric
    • C. Evaluate, then guess, then ship
    • D. Instrument, then change everything at once, then hope
    Show answer

    Answer: B.

    DRIVE is Document, Record versions, Instrument, Vary one thing, Evaluate — a disciplined single-change, measured cycle. The other options either change many variables at once, judge by feel, or skip measurement, which is the vibes anti-pattern.

  46. Q46D2 · Decomposing Work into StepsSelect one

    A weekly workflow summarises ten separate customer-call transcripts and must produce one consistent themes report. A colleague simply concatenates the ten summaries. What is the MOST accurate critique?

    • A. Concatenation is exactly what fan-in means, so it is correct
    • B. The fan-in must reconcile, dedupe and total across the ten summaries; stapled summaries are not a synthesis
    • C. There should have been eleven transcripts
    • D. The fan-out should have been sequential instead
    Show answer

    Answer: B.

    Fan-in is a synthesis step that resolves contradictions, dedupes themes and reconciles across branches; concatenation skips the actual combining work. Concatenation is not fan-in (A), the transcript count (C) is irrelevant, and running the independent transcripts sequentially (D) would misfit the parallel structure.

  47. Q47D2 · Decomposing Work into StepsSelect two

    A single prompt is asked to (i) extract figures from an invoice and (ii) apply reimbursement rules, and wrong totals cannot be traced to either cause. Which TWO changes best restore traceability?

    • A. Extract the invoice fields into a structure as a first, independently checkable step
    • B. Add a sentence asking the model to double-check its arithmetic
    • C. Apply the reimbursement rules to the structured fields as a second, separately testable step
    • D. Switch to the most expensive model
    • E. Increase the maximum output tokens
    Show answer

    Answer: A and C.

    Splitting into an extraction step and a rule-application step (extract-then-transform) makes each independently checkable, so a wrong total traces to reading or logic. A self-check sentence (B) keeps the errors blended, and a bigger model (D) or more tokens (E) does not separate the two failure sources.

  48. Q48D4 · Choosing the Right CapabilitySelect one

    A task needs the same reference files each run AND must compute totals over an uploaded spreadsheet each time. What is the BEST combination?

    • A. A one-off prompt with data analysis
    • B. A Project holding the reference files, with data analysis switched on for the computation
    • C. A custom GPT even though only you run it
    • D. An API application because a spreadsheet is involved
    Show answer

    Answer: B.

    The packaging need (standing files, solo) is a Project, and the compute need is the in-conversation data analysis capability — the two axes combine. A one-off prompt (A) loses the standing files, a custom GPT (C) over-reaches when only you run it, and an API app (D) is unjustified merely because a spreadsheet is involved.

  49. Q49D5 · Review Points and Human OversightSelect one

    A workflow drafts and then a human sends replies, but leadership wants to add a second gate 'just in case' immediately after the send. Why is this the WRONG design?

    • A. Two gates are always better than one
    • B. A gate after the irreversible send cannot prevent the mistake; oversight must precede the point of no return
    • C. The second gate should replace the first
    • D. Post-send review lowers model accuracy
    Show answer

    Answer: B.

    Review after an irreversible action is not a gate, because the harm has already occurred; the effective gate is the one before the send. More gates are not always better (A), the pre-send gate should stay (C), and post-send review does not affect model accuracy (D).

  50. Q50D1 · Finding and Scoping OpportunitiesSelect one

    A candidate task scores high on frequency and time cost, but its output must be published externally with almost no error tolerance. On the FTED screen, what is the MOST appropriate build decision?

    • A. Skip it, because low error tolerance means it has no value
    • B. Automate the drafting up to the external step and keep a mandatory human gate before anything is published
    • C. Automate it fully because the value axes score high
    • D. Choose a cheaper model to offset the risk
    Show answer

    Answer: B.

    High value with low error tolerance means you capture the value by automating the drafting but keep a mandatory gate before the external, hard-to-retract action. Skipping (A) throws away real value, full automation (C) puts an unforgiving external step in the model's hands, and model price (D) does not address the error-tolerance risk.

Last updated Sep 18, 2026