AI Cert Prep
Type to search documentation.

Agents and Workflows

Agents – Mock Exam 2

A harder 50-item, domain-weighted independent mock exam for the Agents and Workflows track, designed as a timed readiness gate.

This is the second full-length independent mock exam for the Agents and Workflows track, built from publicly available OpenAI learning objectives and not an official OpenAI assessment. It is deliberately harder than mock exam 1 — more multi-constraint stems and more FIRST / BEST / MOST cost-effective / TWO qualifiers. All 50 questions are new and distinct from both mock exam 1 and the domain-page items. Use this as your readiness gate under timed conditions once mock exam 1 and your revision are behind you.

Instructions

  • Time: 60 minutes. Sit this one timed, in a single uninterrupted block.
  • Items: 50 multiple-choice and multiple-response questions. Each item states how many answers to select.
  • Selection: for a multiple-response item you must select all correct options and no incorrect ones to earn the mark; there is no partial credit.
  • No guessing penalty: answer every question — a wrong answer costs nothing beyond the mark.
  • Target: aim for at least 80% raw (40 of 50) under timed conditions before you take the real Academy Agents and Workflows assessment. The 80% line matches the Academy badge threshold.
  • Read multi-constraint stems carefully: the qualifier (FIRST, BEST, TWO) usually decides between two defensible options.

Domain distribution

#DomainWeightItems here
1What an Agent Is and When to Use One16%8
2Defining Objectives and Tasks18%9
3Context, Tools and Permissions18%9
4Boundaries and Guardrails16%8
5Reviewing and Verifying Agent Work16%8
6Reliability and Iteration16%8

Total: 8 + 9 + 9 + 8 + 8 + 8 = 50 items.

Readiness interpretation

This is an independent readiness indicator, not an official score.

Raw score (of 50)BandInterpretation
45–5090%+Strong readiness; you are ready for the real assessment
40–4480–89%Assessment ready; review any weak domain
35–3970–79%Building confidence; another revision pass advised
Below 35under 70%Keep learning; this harder set has found real gaps

Because this mock is harder than mock exam 1, a score in the 80–89% band here is a strong signal of readiness. Treat 40 of 50 as your minimum gate.

Take the mock exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your readiness indicator, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · What an Agent Is and When to Use OneSelect one

    A team lead has a task whose steps are fully known and identical each run, but which touches an irreversible external send. Considering path, access, duration and reversibility together, which build is MOST appropriate?

    • A. An agent with autonomous send, because the send is the tedious part
    • B. A workflow with a human approval gate before the send
    • C. A single prompt that does everything at once
    • D. An agent with broad connectors so it never gets stuck
    Show answer

    Answer: B.

    A fixed, known path points to a workflow rather than an agent, and the irreversible send must sit behind a human gate. Autonomous send (A) removes the only control that matters, a single prompt (C) cannot carry a sequenced process, and a broad-access agent (D) adds autonomy the fixed path does not need plus unnecessary blast radius.

  2. Q2D1 · What an Agent Is and When to Use OneSelect one

    Two tasks look similar in difficulty. Task X has steps that depend on what earlier steps find and runs unattended; task Y has a fixed sequence you review between. Which is the BEST classification?

    • A. Both are agents, because both are hard
    • B. X is an agent and Y is a workflow, because the deciding axis is whether the path can be scripted, not difficulty
    • C. Both are prompts, because each could start with one request
    • D. X is a workflow and Y is an agent
    Show answer

    Answer: B.

    Fit is decided by whether the path is knowable in advance: X must plan (agent), Y is fixed with reviews (workflow). Treating difficulty as decisive (A) is the trap, both being prompts (C) ignores their multi-step nature, and D reverses the correct mapping.

  3. Q3D1 · What an Agent Is and When to Use OneSelect one

    A manager wants to automate a recurring task end to end, including a public publish step, arguing it is the most tedious part. What is the FIRST thing you should establish?

    • A. Which model is cheapest for the task
    • B. Whether the irreversible publish can be placed behind an approval gate; if it cannot, the dangerous step should not be automated
    • C. How many words the brief should contain
    • D. Whether a competitor uses agents for this
    Show answer

    Answer: B.

    Reversibility governs whether automation is safe: an un-gateable irreversible step is a signal not to automate that step. Model cost (A) and brief length (C) are secondary, and what a competitor does (D) is irrelevant to whether this action can be safely delegated.

  4. Q4D1 · What an Agent Is and When to Use OneSelect one

    A stakeholder insists that because a task will take an agent twenty unattended minutes, it must be 'more powerful' than a two-minute prompt. What is the MOST accurate correction?

    • A. They are right; longer runs are always more capable
    • B. Duration changes oversight, not power: a longer unwatched run means you must design boundaries and verification because you cannot correct it live
    • C. Duration is purely a cost concern with no oversight implications
    • D. A two-minute prompt cannot produce useful work
    Show answer

    Answer: B.

    Duration is an oversight axis: longer unwatched runs shift control from live observation to designed boundaries and after-the-fact verification. It does not equal power (A), it is more than a cost concern (C), and short prompts routinely produce useful work (D).

  5. Q5D1 · What an Agent Is and When to Use OneSelect one

    A task is genuinely multi-step and benefits from tools, but you cannot articulate what a finished, correct result looks like. What does this MOST strongly imply?

    • A. Delegate it to an agent anyway; it will infer success
    • B. It is not yet delegable — the missing definition of done must be resolved before an agent runs, regardless of how tool-friendly it seems
    • C. Use the largest model to compensate for the vagueness
    • D. Grant every tool so it has more ways to succeed
    Show answer

    Answer: B.

    Without a definition of done you can neither brief nor verify the work, so multi-step and tool-friendly do not make it delegable yet. Delegating regardless (A) yields unverifiable output, and a bigger model (C) or more tools (D) cannot substitute for a missing acceptance test.

  6. Q6D1 · What an Agent Is and When to Use OneSelect two

    A knowledge worker is weighing an agent for a quarterly board-pack assembly. Which TWO factors most strongly argue AGAINST an agent and FOR a workflow with gates?

    • A. The assembly steps are the same every quarter and fully documented
    • B. The final distribution to the board is irreversible and external
    • C. The task is long and involves several documents
    • D. The requester finds the task boring
    • E. The output is only ever a private draft
    Show answer

    Answer: A and B.

    A fixed, documented path (A) makes a workflow cheaper and more verifiable than an agent, and the irreversible external distribution (B) demands a gate rather than autonomy. Length and multiple documents (C) do not by themselves require an agent, boredom (D) is irrelevant, and a private draft (E) would lower risk rather than argue against an agent.

  7. Q7D1 · What an Agent Is and When to Use OneSelect one

    Which reasoning BEST distinguishes an agent from a very long, highly detailed prompt?

    • A. The prompt is longer, so it has more autonomy
    • B. In the prompt you still decide each next step; the agent chooses its own steps toward a goal, so autonomy — not length — is the divider
    • C. The agent always uses more expensive models
    • D. A long prompt cannot include any constraints
    Show answer

    Answer: B.

    Autonomy is who chooses the next step: a long prompt is still low-autonomy if you drive it, whereas an agent plans the path itself. Length does not confer autonomy (A), model cost (C) is unrelated, and prompts can certainly include constraints (D).

  8. Q8D1 · What an Agent Is and When to Use OneSelect one

    A five-minute manual task would need an hour of setup and verification each time it ran as an agent. Applying the reversibility-and-cost view, what is the BEST decision?

    • A. Automate it; automation is always worth it eventually
    • B. Do not delegate: when verifying the unwatched output costs more than the work, delegation is a net loss
    • C. Delegate but skip verification to save the hour
    • D. Use two agents in parallel to be safe
    Show answer

    Answer: B.

    When checking the output costs more than simply doing the task, an agent is a net loss — a core 'wrong tool' signal. Automating regardless (A) ignores the cost, skipping verification (C) trades a small saving for unbounded risk, and a second agent (D) increases both cost and verification burden.

  9. Q9D2 · Defining Objectives and TasksSelect one

    An agent returned a report that did something adjacent to what was wanted, silently dropped a section, and cited a stale figure. Applying the brief-first discipline, which single change should you make FIRST?

    • A. Upgrade to the most capable model
    • B. Interrogate and repair the brief — restate the outcome, add a coverage-complete definition of done and name the source of truth — before touching the model
    • C. Re-run the identical brief several times and average the outputs
    • D. Grant broad connector access to reduce future stalls
    Show answer

    Answer: B.

    All three symptoms map to brief defects — a task-shaped goal, a missing definition of done and an unnamed source — so the disciplined first move is to repair the brief. A bigger model (A) reproduces the same shaped errors, re-running the same brief (C) reproduces the defect, and broad access (D) addresses none of these causes.

  10. Q10D2 · Defining Objectives and TasksSelect one

    Several supplied sources give conflicting revenue figures. What should a well-formed brief provide so the agent stays grounded?

    • A. Nothing — let the agent average the numbers
    • B. A ranked conflict-resolution order stating which source wins and when to flag remaining uncertainty
    • C. An instruction to always trust the public web
    • D. A larger context window
    Show answer

    Answer: B.

    Ranking the sources and telling the agent when to flag uncertainty keeps the output grounded when data disagrees. Averaging (A) blends good and bad data, always trusting the web (C) is often the least authoritative choice, and a bigger window (D) does not resolve which source to believe.

  11. Q11D2 · Defining Objectives and TasksSelect one

    A brief for a bulk-processing job over 200 items has a clear outcome and definition of done but no checkpoints. Which SINGLE checkpoint adds the MOST safety for the least review cost?

    • A. Pause after every single item for individual approval
    • B. Pause after the first item so the pattern can be approved before the remaining items run
    • C. Pause every thirty seconds regardless of progress
    • D. Never pause, to finish faster
    Show answer

    Answer: B.

    The 'one before many' checkpoint prevents a 200-way mistake at the cost of a single review. Pausing on every item (A) defeats the point of delegating, time-based pauses (C) do not align with risk, and no pauses (D) removes all steering on a bulk action.

  12. Q12D2 · Defining Objectives and TasksSelect one

    A manager wrote 'prefer a concise style', 'never share draft financials externally' and 'aim to finish today' into a brief. Which is the MOST important to promote from wording into an enforced control, and why?

    • A. 'Prefer a concise style', because tone matters most to readers
    • B. 'Never share draft financials externally', because it is a hard constraint whose violation is irreversible and sensitive
    • C. 'Aim to finish today', because deadlines are the priority
    • D. None; all three are equal preferences
    Show answer

    Answer: B.

    The confidentiality rule is a hard constraint guarding an irreversible, sensitive disclosure, so it warrants system enforcement rather than mere wording. Style (A) and the timing wish (C) are preferences, and treating all three as equal (D) misses that only one protects against real, hard-to-reverse harm.

  13. Q13D2 · Defining Objectives and TasksSelect one

    Which objective is BEST written as a delegable outcome rather than a task list?

    • A. Open each folder, read the notes, then think about what to write
    • B. Produce a one-page onboarding checklist covering the five required steps, sourced to the current HR policy, with any gap flagged
    • C. Go through everything and improve it
    • D. Do the usual onboarding thing you always do
    Show answer

    Answer: B.

    Option B names a concrete artefact, a bounded scope, an authoritative source and a failure behaviour, so it is both executable and checkable. A describes keystrokes rather than an outcome, and C and D are vague wishes with no acceptance test.

  14. Q14D2 · Defining Objectives and TasksSelect two

    A project lead delegates migrating formatting across 150 documents with a clear outcome and definition of done. Which TWO checkpoints add the MOST safety?

    • A. Pause after the first document so the pattern can be approved before the rest run
    • B. Pause before any step that would irreversibly overwrite an original file
    • C. Pause after every single document for individual approval
    • D. Pause every minute regardless of progress
    • E. Never pause, to finish faster
    Show answer

    Answer: A and B.

    Approving the pattern on the first document (A) prevents a 150-way mistake, and gating irreversible overwrites (B) protects the originals. Pausing on every document (C) defeats delegation, time-based pauses (D) do not align with risk, and no pauses (E) removes all steering on an irreversible bulk action.

  15. Q15D2 · Defining Objectives and TasksSelect one

    An objective contains many small unknowns you cannot pre-answer and there is no single fork to gate. What is the BEST instruction to the agent?

    • A. Silently choose interpretations and proceed quickly
    • B. State the assumptions it makes and surface them for review, so wrong ones can be caught
    • C. Refuse the task until every unknown is removed
    • D. Produce a separate full version for every possible interpretation
    Show answer

    Answer: B.

    When there are many small unknowns and no single fork to pause on, the best move is to have the agent surface its assumptions for review. Silent choices (A) hide wrong turns, refusing outright (C) is disproportionate, and enumerating every full version (D) is wasteful and impractical.

  16. Q16D2 · Defining Objectives and TasksSelect one

    A brief says 'summarise the incident' with no scope. The agent returns a two-line note when leadership expected a full timeline. What does this MOST directly reveal about the brief?

    • A. The model is too small
    • B. The definition of done lacked a scope, so the agent chose its own and stopped early
    • C. The reasoning effort was too low
    • D. The agent lacked connector access
    Show answer

    Answer: B.

    With no scope in the definition of done, the agent set its own bound and stopped, producing less than expected. Model size (A) and reasoning effort (C) do not define scope, and an access gap (D) would show as a stall or wrong data, not an under-scoped summary.

  17. Q17D2 · Defining Objectives and TasksSelect one

    Which statement BEST explains why 'move fast, let the agent assume' is a poor default for an ambiguous objective?

    • A. It is always literally slower up front
    • B. Speed bought by an unstated wrong assumption is repaid in rework, and worse if that assumption drove an irreversible action
    • C. Agents are incapable of making any assumptions
    • D. It merely uses more tokens than asking
    Show answer

    Answer: B.

    An unsurfaced wrong assumption compounds into rework and, on an irreversible step, into unrecoverable harm, so the apparent speed is a false economy. It is not always slower up front (A), agents can make assumptions (C), and token use (D) is not the substantive risk.

  18. Q18D3 · Context, Tools and PermissionsSelect one

    An agent drafting renewal emails cited an outdated policy AND pulled details from unrelated customers' files. Which pair of fixes addresses the TWO distinct causes?

    • A. Supply the current policy as the authoritative source, and scope the connector to only the relevant account records
    • B. Switch to a larger model, and remove all checkpoints to speed the run
    • C. Grant write access to every account, and paste in more documents
    • D. Lower the reasoning effort, and enable autonomous send
    Show answer

    Answer: A.

    The outdated policy is a context gap fixed by supplying the current policy, while cross-account leakage is an access/reach problem fixed by scoping the connector. A larger model (B) fixes neither cause, broad write and more documents (C) worsen risk and miss the context fix, and lower effort with autonomous send (D) addresses neither and adds danger.

  19. Q19D3 · Context, Tools and PermissionsSelect two

    Applying the read then write then send ladder for least privilege, which TWO postures are correct defaults?

    • A. Start at read and grant write only if the task must change data
    • B. Keep send gated behind human approval whenever it is enabled
    • C. Start at send and remove access only if problems appear
    • D. Grant write by default because most tasks need it
    • E. Grant send while withholding read, to limit exposure
    Show answer

    Answer: A and B.

    The safe ladder starts at read and adds write only to change data (A), and always gates send (B). Starting at send (C) exposes irreversible actions first, write-by-default (D) over-provisions, and send-without-read (E) is incoherent because the agent could act but not ground its action.

  20. Q20D3 · Context, Tools and PermissionsSelect one

    A team lead wants to widen an agent's access after it genuinely stalled on a real task for lack of reach. When is widening justified under least privilege?

    • A. Never; access should stay fixed forever
    • B. When a real run demonstrates a genuine need, granted as narrowly as possible
    • C. Whenever it would be convenient
    • D. Only for senior staff, regardless of the task
    Show answer

    Answer: B.

    Least privilege widens on evidence: a real run showing a genuine need justifies the narrowest additional grant that unblocks it. Fixed-forever access (A) ignores real needs, convenience (C) is the over-provisioning trap, and seniority (D) is not the basis for task-scoped access.

  21. Q21D3 · Context, Tools and PermissionsSelect one

    An agent produces work that is competent but consistently ignores a decision the team already made, while reporting no access issues. What is the MOST likely cause and fix?

    • A. Missing access; grant more connectors
    • B. Missing prior-decisions context; supply the record of what was already settled
    • C. A model defect; upgrade the model
    • D. Too little reasoning effort; raise it
    Show answer

    Answer: B.

    Ignoring a settled decision with no access complaint is a missing prior-decisions context gap, fixed by supplying that record. Granting connectors (A) addresses access, which is not the issue, and a bigger model (C) or more effort (D) will not surface a decision it was never given.

  22. Q22D3 · Context, Tools and PermissionsSelect one

    A workspace agent is given a connector to a channel to answer one question about a recent thread. What is the KEY question to ask before enabling it?

    • A. Which font the answer should use
    • B. Whether the task needs the whole channel history, or whether the connector can be scoped more narrowly
    • C. How many tokens the answer will consume
    • D. Whether the agent prefers that channel
    Show answer

    Answer: B.

    A channel connector exposes the whole history, so the key least-privilege question is whether the task truly needs that reach or can be scoped narrower. Font (A), token count (C) and an imagined preference (D) are irrelevant to the reach the connector grants.

  23. Q23D3 · Context, Tools and PermissionsSelect two

    An agent's output is competent but generic and it also stopped short saying it 'could not reach the pricing system'. Which TWO fixes correctly address the TWO different problems?

    • A. Supply the brand and product context it lacks
    • B. Grant a scoped read connector to the pricing system
    • C. Switch to a larger model to cover both issues at once
    • D. Paste every document you have into the task
    • E. Enable autonomous send so it can finish
    Show answer

    Answer: A and B.

    Generic output is a missing-context problem fixed by supplying brand and product context (A), while 'could not reach' is a missing-access problem fixed by a scoped connector (B). A larger model (C) addresses neither, pasting everything (D) buries signal, and autonomous send (E) is irrelevant and risky.

  24. Q24D3 · Context, Tools and PermissionsSelect one

    Why is 'to be safe, connect the agent to everything it might conceivably need' the WRONG default?

    • A. It is correct; broad access has no downside
    • B. Broad access is broad blast radius and broad reach; an autonomous agent may use any of it, so you start minimal and widen on evidence
    • C. It always makes the agent slower
    • D. It forces the agent onto a smaller model
    Show answer

    Answer: B.

    Every added connector is added reach an autonomous agent may exercise, so least privilege starts minimal and widens only on demonstrated need. Broad access has real downside (A), and it neither dictates speed (C) nor model size (D).

  25. Q25D3 · Context, Tools and PermissionsSelect one

    A task needs the agent to read from two internal systems and produce a draft you will send. Which provisioning is MOST aligned with least privilege?

    • A. Write access to both systems plus autonomous send
    • B. Scoped read connectors to the specific records in both systems, plus draft creation with send gated behind your approval
    • C. A single broad connector to the entire company drive
    • D. No connectors; have it guess from memory
    Show answer

    Answer: B.

    Scoped read to just the needed records, draft creation, and a gated send is the minimal set that lets the task succeed safely. Write plus autonomous send (A) over-provisions and removes the send gate, a broad drive connector (C) grants excess reach, and no connectors (D) forces guessing.

  26. Q26D3 · Context, Tools and PermissionsSelect one

    An agent quoted an internal margin note in a customer-facing draft. What is the ROOT cause and the strongest fix?

    • A. It hallucinated the note; lower the reasoning effort
    • B. It had reach into internal data it should never have accessed; scope its access so the note is unreachable
    • C. The definition of done was too strict; loosen it
    • D. The context window was too small; enlarge it
    Show answer

    Answer: B.

    The agent could quote the note because it had reach to it — an over-provisioning failure whose strongest fix is to make internal data unreachable. It did not hallucinate a real note it was given access to (A), and a strict definition of done (C) or window size (D) are unrelated to the leak.

  27. Q27D4 · Boundaries and GuardrailsSelect one

    An agent emailed several new external contacts despite a brief saying 'never contact anyone new without checking'. What is the correct fix?

    • A. Reword the instruction in stronger language
    • B. Move the boundary into the system: enforce an approval gate so it cannot send to a new contact without authorisation
    • C. Use a bigger model that follows instructions better
    • D. Remove the definition of done to simplify the run
    Show answer

    Answer: B.

    The failure is a respected boundary on an irreversible action; the fix is to enforce it as a system-level approval gate. Firmer wording (A) is still just a respected boundary, a bigger model (C) can still misread, and removing the definition of done (D) is unrelated and harmful.

  28. Q28D4 · Boundaries and GuardrailsSelect one

    A manager needs to bound the worst case of an overnight agent run over an unbounded queue. Which SINGLE control most directly guarantees the run cannot process more than a known number of items?

    • A. A note in the brief asking it to be efficient
    • B. An enforced scope cap on the number of items processed
    • C. A friendlier tone setting
    • D. Removing all checkpoints so it finishes faster
    Show answer

    Answer: B.

    An enforced scope cap directly limits how many items the run can touch, bounding the worst case. A polite note (A) is a respected boundary that will not stop a runaway, tone (C) is irrelevant, and removing checkpoints (D) increases risk.

  29. Q29D4 · Boundaries and GuardrailsSelect one

    Which describes the BEST-designed approval gate for an irreversible external send?

    • A. It pauses after the send, so the agent is not slowed
    • B. It pauses before the send and surfaces the full draft and recipients so the human can decide quickly
    • C. It asks 'proceed?' with no detail
    • D. It gates every step equally to maximise safety
    Show answer

    Answer: B.

    A good gate pauses before the irreversible step and shows enough context — the draft and recipients — for a real, cheap decision. Pausing after (A) controls nothing, a detail-free prompt (C) invites blind approval, and gating every step (D) causes rubber-stamp fatigue.

  30. Q30D4 · Boundaries and GuardrailsSelect one

    An agent 'reviewing partnership requests' interpreted 'clear yes' broadly and sent to companies it had never dealt with, one message quoting internal terms. Which combination of fixes is MOST complete?

    • A. Reword 'clear yes' more precisely and trust the agent
    • B. Enforce an approval gate on all external sends, require explicit sign-off for new contacts, and remove reach to internal terms
    • C. Use the most capable model and enable autonomous send
    • D. Add a spend cap and nothing else
    Show answer

    Answer: B.

    Both harms came from respected boundaries on irreversible/sensitive actions, so the complete fix enforces the send gate, adds a higher-stakes gate for new contacts, and removes reach to internal terms. Rewording and trusting (A) leaves a respected boundary, a bigger model with autonomous send (C) removes control, and a spend cap alone (D) addresses neither the send nor the leak.

  31. Q31D4 · Boundaries and GuardrailsSelect one

    For a low-stakes internal draft you will read before doing anything with it, which boundary posture is MOST appropriate?

    • A. Heavy enforced gates on every step
    • B. Light review after the run, with no gate, because nothing irreversible happens
    • C. A spend cap of zero so it cannot run at all
    • D. Disable all tools
    Show answer

    Answer: B.

    Boundaries should match stakes: a reversible internal draft you will review needs only light review. Heavy gates (A) waste attention, a zero cap (C) blocks the task, and disabling all tools (D) prevents harmless work.

  32. Q32D4 · Boundaries and GuardrailsSelect two

    Which TWO controls must be system-ENFORCED rather than merely requested when an agent runs unattended over paid tools with an external publish step?

    • A. An approval gate before the publish
    • B. A hard spend cap on the paid tools
    • C. A preference for a formal writing tone
    • D. A suggestion to wrap up before lunch
    • E. A preference for bullet points
    Show answer

    Answer: A and B.

    The irreversible publish needs an enforced approval gate (A) and the paid tools need an enforced spend cap (B) so the worst case is bounded and known. Tone (C), a timing wish (D) and formatting (E) are preferences for which respected boundaries suffice.

  33. Q33D4 · Boundaries and GuardrailsSelect one

    A leader argues 'the agent behaved perfectly in testing, so we can trust it to respect the do-not-send rule in production'. What is the strongest counter?

    • A. They are right; consistent testing proves it is safe
    • B. Usual behaviour is not a control; one misread on an unusual input can cross a merely-requested line irreversibly, so enforce the gate
    • C. Testing is irrelevant to agent behaviour
    • D. A larger model would remove the need for any gate
    Show answer

    Answer: B.

    A respected boundary can hold 99% of the time and still fail catastrophically once on an irreversible action, so enforcement is required regardless of test behaviour. Trusting test behaviour (A) mistakes usual for guaranteed, testing is not irrelevant (C), and no model (D) removes the need to gate an irreversible action.

  34. Q34D4 · Boundaries and GuardrailsSelect one

    Which is the STRONGEST way to ensure an agent cannot move a regulated dataset into an outbound message?

    • A. Instruct it in the brief not to include regulated data
    • B. Do not grant the agent reach to the regulated dataset in the first place
    • C. Ask it to summarise the dataset carefully
    • D. Increase the reasoning effort so it is more careful
    Show answer

    Answer: B.

    The strongest data boundary is architectural: an agent cannot move what it was never able to reach. Instruction (A) can be missed, careful summarisation (C) still touches the data, and higher effort (D) does not create a data boundary.

  35. Q35D5 · Reviewing and Verifying Agent WorkSelect one

    A stakeholder will rely on an agent-produced deliverable in one hour. Which SINGLE approach gives the most verification value in that limited time?

    • A. Re-read the agent's summary twice for reassurance
    • B. Verify against the definition of done, then risk-weight the spot-check and reproduce the single highest-stakes claim
    • C. Re-run the entire task from scratch
    • D. Ask the agent to grade its own work
    Show answer

    Answer: B.

    Ticking the acceptance criteria, risk-weighting the sample, and reproducing the top-stakes claim concentrate limited time where errors matter most. Re-reading the summary (A) and self-grading (D) verify nothing independent, and re-running everything (C) wastes the hour.

  36. Q36D5 · Reviewing and Verifying Agent WorkSelect one

    An agent claims 'reviewed all 90 pages' but the evidence trail shows it opened only 66. What has MOST likely occurred, and what must you do?

    • A. Nothing unusual; trust the claim
    • B. A silent partial completion — the unreviewed pages must be handled and the brief tightened to require a coverage flag
    • C. A hallucinated model needing replacement
    • D. A reasoning-effort error
    Show answer

    Answer: B.

    Reporting full coverage while the trail shows partial work is silent partial completion, caught by checking the trail against the claim; the fix handles the remainder and adds a coverage requirement. It is not normal (A), not a 'hallucinated model' (C), and not an effort setting (D) — it is a coverage-versus-claim mismatch.

  37. Q37D5 · Reviewing and Verifying Agent WorkSelect one

    Which statement BEST explains why the completion summary is the LEAST reliable part of an unwatched deliverable?

    • A. Summaries are always shorter than the work
    • B. It is the agent's own claim about its work, produced by the same reasoning, and fluency can mask silent partial completion
    • C. Summaries are written in a different language
    • D. Summaries use more tokens than artefacts
    Show answer

    Answer: B.

    The summary is a self-report from the same reasoning that did the work, and its fluency can conceal missing coverage, so it must be checked against artefacts and the definition of done. Length (A), language (C) and token use (D) are not why it is unreliable.

  38. Q38D5 · Reviewing and Verifying Agent WorkSelect one

    You verify a bulk agent output by sampling 8 of 80 items; two fail. What is the MOST defensible conclusion?

    • A. The other 72 are fine; a small failure rate is acceptable
    • B. The failures invalidate the assumption of uniform quality; widen the check substantially or reject and re-run the batch
    • C. Fix the two and ship the remaining 78 unchecked
    • D. Spot-checking is unreliable, so ignore the result
    Show answer

    Answer: B.

    Failures in the sample mean you can no longer assume the unchecked items are correct, so you widen the check or reject the batch. Assuming the rest are fine (A) and shipping 78 unchecked (C) ignore the signal, and spot-checking is a valid technique (D).

  39. Q39D5 · Reviewing and Verifying Agent WorkSelect two

    In limited time, which TWO checks best verify an agent's claim that it 'fixed every broken link' across a large knowledge base?

    • A. Test a risk-weighted sample of the links it says it fixed, prioritising customer-facing pages
    • B. Test some links it did not mention, in case it missed them
    • C. Re-read its summary more carefully
    • D. Ask it to confirm in the same chat
    • E. Assume completion because the report is detailed
    Show answer

    Answer: A and B.

    A risk-weighted sample of claimed fixes (A) and an adversarial check of links it did not mention (B) concentrate effort on likely and high-stakes failures. Re-reading the summary (C) and asking in-session (D) verify nothing independent, and detail (E) is not evidence of completion.

  40. Q40D5 · Reviewing and Verifying Agent WorkSelect one

    Why should you prefer to review the artefact rather than the agent's assertion about it?

    • A. The assertion is easier to read, so prefer it
    • B. The artefact — the actual change list or updated file — is checkable, whereas the assertion is only a claim about it
    • C. They contain identical information
    • D. Whichever is shorter is more accurate
    Show answer

    Answer: B.

    The artefact is the concrete, verifiable work product, while the assertion is merely a description that could be wrong. Preferring the assertion for ease (A) trusts the story, they are not identical (C), and length (D) does not determine accuracy.

  41. Q41D5 · Reviewing and Verifying Agent WorkSelect one

    An agent's report reads fluently, but you cannot verify one claim against any artefact or source. Applying the review discipline, what is the BEST interpretation?

    • A. The claim is correct because the report is fluent
    • B. The definition of done was too vague to check that claim; treat it as unverified and sharpen the brief next iteration
    • C. The model is broken and must be replaced
    • D. Agent work can never be verified
    Show answer

    Answer: B.

    An unverifiable claim usually signals an upstream brief defect — no checkable criterion — so treat it as unverified and fix the brief next time. Fluency is not evidence (A), it is not a broken model (C), and agent work is verifiable when the brief supports it (D).

  42. Q42D5 · Reviewing and Verifying Agent WorkSelect one

    Which sequence is the MOST efficient structured review of a forty-minute agent run before spending any time on the full transcript?

    • A. Transcript, then plan, then summary
    • B. Plan, then checkpoints, then verify against the definition of done, then inspect artefacts, then spot-check
    • C. Summary only, then close the task
    • D. Reproduce every claim before reading anything
    Show answer

    Answer: B.

    The structured order — plan, checkpoints, done-criteria, artefacts, spot-check — lets you stop early if an upstream step fails, reserving the transcript for last. Reading the transcript first (A) is the slow move, the summary alone (C) verifies nothing, and reproducing every claim (D) is impractical and skips cheaper checks.

  43. Q43D6 · Reliability and IterationSelect one

    A weekly triage agent shows three symptoms at once: wrong amounts, records missing while reporting 'all processed', and later entries drifting in tone. Which approach is MOST disciplined?

    • A. Conclude agents cannot do the task and switch to the biggest model
    • B. Recognise three distinct failure modes, instrument the run, then iterate one diagnosed change at a time, re-testing on varied inputs
    • C. Rewrite the entire brief at once and ship immediately
    • D. Trust the 'all processed' claim because it worked in the first test
    Show answer

    Answer: B.

    The symptoms map to a wrong source, silent partial completion and drift — three modes each with its own fix — so you instrument, then change one thing at a time and re-test on varied inputs. A bigger model (A) fixes no cause, rewriting everything at once (C) obscures what worked, and trusting the completion claim (D) ignores the partial completion.

  44. Q44D6 · Reliability and IterationSelect one

    In the failure taxonomy, which fix pairs correctly with 'the agent stalled and could not reach the ticketing system'?

    • A. Sharpen the objective
    • B. Grant a scoped connector to the ticketing system (a wrong-tool/access fix)
    • C. Add a coverage requirement to the definition of done
    • D. Switch to a smaller model
    Show answer

    Answer: B.

    A stall for lack of reach is a wrong-tool/access failure, fixed by provisioning the specific connector at least privilege. Sharpening the goal (A) addresses misunderstood objectives, a coverage requirement (C) addresses partial completion, and model size (D) is unrelated to access.

  45. Q45D6 · Reliability and IterationSelect two

    An agent 'did something adjacent to what you wanted'. Which TWO statements correctly diagnose and fix this?

    • A. It is a misunderstood objective — the goal was stated as a task or was ambiguous
    • B. Restate the goal as a concrete, checkable outcome before re-running
    • C. It is drift; the only fix is to raise the reasoning effort
    • D. It is silent partial completion; add a coverage check
    • E. It is missing access; grant more connectors
    Show answer

    Answer: A and B.

    Producing something near-but-not the goal is a misunderstood objective (A), fixed by restating it as a checkable outcome (B). Drift (C) is gradual degradation, silent partial completion (D) is a false completeness claim, and missing access (E) shows as a stall or wrong data — none matches 'adjacent to what you wanted'.

  46. Q46D6 · Reliability and IterationSelect one

    What BEST distinguishes a reliable delegation from a lucky one?

    • A. The reliable one used a more expensive model
    • B. The reliable one meets the definition of done across varied, harder inputs, not just the one easy case that happened to work
    • C. The lucky one had a longer brief
    • D. There is no difference in practice
    Show answer

    Answer: B.

    Reliability is defined by holding up across varied and edge-case inputs, whereas luck is a single easy input succeeding. Model cost (A) and brief length (C) do not determine reliability, and the distinction is very real (D).

  47. Q47D6 · Reliability and IterationSelect one

    You have observed a failure and want to iterate. Applying the LEARN loop, what is the correct order of the middle steps?

    • A. Revise a change, then examine instrumentation, then attribute a mode
    • B. Examine the instrumentation, attribute it to a taxonomy mode, then revise exactly one thing
    • C. Attribute a mode, revise everything at once, then examine the result
    • D. Revise the model first, then examine, then attribute
    Show answer

    Answer: B.

    LEARN examines the instrumentation, attributes the failure to a mode, then revises one thing before re-testing on a new input. Revising before examining or attributing (A, D) skips diagnosis, and revising everything at once (C) prevents you from knowing what worked.

  48. Q48D6 · Reliability and IterationSelect one

    Why is instrumentation a prerequisite for reliable iteration, rather than an optional extra?

    • A. It makes the agent run faster
    • B. You cannot classify or fix a failure you cannot see; the plan, evidence trail and outcome-versus-done reveal which mode you are in
    • C. It automatically upgrades the model
    • D. It is only relevant to developers writing code
    Show answer

    Answer: B.

    Instrumentation makes behaviour observable so you can classify the failure and target the fix rather than guess. It is not about speed (A) or model upgrades (C), and it is a delegation discipline for any knowledge worker, not only developers (D).

  49. Q49D6 · Reliability and IterationSelect two

    A delegation succeeded once on an easy input. Which TWO actions best establish whether it is genuinely reliable?

    • A. Re-test it on a harder input with edge cases
    • B. Re-test it on a different but realistic input and check it still meets the definition of done
    • C. Re-run the identical easy input several more times
    • D. Switch to a cheaper model and ship
    • E. Declare it reliable because it worked once
    Show answer

    Answer: A and B.

    Reliability is proven by re-testing on harder and on different realistic inputs while confirming the definition of done still holds (A, B). Re-running the same easy input (C) proves nothing new, a cheaper model (D) does not establish reliability, and one success (E) is only a hypothesis.

  50. Q50D6 · Reliability and IterationSelect one

    A drifting long-run agent is told more firmly to 'stay on task', yet it drifts again. What does this MOST clearly demonstrate?

    • A. The brief simply needs even firmer wording
    • B. Drift is a structural problem: a respected instruction will not hold on a long run, so add checkpoints or shorten the scope
    • C. The model is defective and must be replaced
    • D. The task is impossible to delegate
    Show answer

    Answer: B.

    Drift on a long unwatched run needs structural fixes — checkpoints that re-anchor or shorter scoped runs — because a respected instruction cannot hold across the duration. Firmer wording (A) is still just a respected boundary, the model is not necessarily defective (C), and the task is not impossible (D).

Last updated Sep 18, 2026