Applied AI Foundations
Applied AI Foundations · Mock Exam 1
A 50-item independent mock exam for the Applied AI Foundations track, weighted to the six domains, with full explanations and a readiness indicator.
This is Mock Exam 1 for the Applied AI Foundations track: 50 items across the six domains. It is an independent mock exam built from publicly available OpenAI learning objectives — not an official OpenAI assessment, not official exam questions, and not endorsed by OpenAI. Use it as your diagnostic: sit it first to find your two weakest domains, then study those domain pages before you attempt Mock Exam 2. For how the real credential works, see the assessment model.
Instructions
- Time: 60 minutes, matching the length of an Applied AI Foundations study session; you can also run it untimed on a first diagnostic pass.
- Items: 50, single-response and multiple-response. Each item states how many answers to select.
- Selection: for a select-one item choose exactly one option; for a select-two item you must choose both correct options and no incorrect one to score the item.
- No guessing penalty: answer every item — an unanswered item scores the same as a wrong one.
- Target: aim for ≥ 80% raw (≈ 40/50) before you take the real OpenAI Academy assessment, which itself passes at 80%.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | Finding and Scoping Opportunities | 16% | 8 |
| 2 | Decomposing Work into Steps | 18% | 9 |
| 3 | Inputs, Outputs and Contracts | 16% | 8 |
| 4 | Choosing the Right Capability | 20% | 10 |
| 5 | Review Points and Human Oversight | 16% | 8 |
| 6 | Repeatability and Improvement | 14% | 7 |
Total: 8 + 9 + 8 + 10 + 8 + 7 = 50 items.
Readiness interpretation
This is an independent readiness indicator, never an official score.
| Raw score (of 50) | Band | Interpretation |
|---|---|---|
| 40+ (80%+) | Assessment ready | You are tracking the 80% Academy badge threshold; take Mock Exam 2 to confirm |
| 35–39 (70–79%) | Building confidence | Close; revise your weakest one or two domains and re-test |
| under 70% | Keep learning | Work through the domain pages before re-attempting |
| 45+ (90%+) | Strong readiness | Comfortable margin across the domains |
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A team lead notices that onboarding-email personalisation happens about 30 times a week, takes 12 minutes each, tolerates small wording errors, and uses only public marketing copy. On the FTED screen, how should this task be classified?
Show answer
Answer: B.
Both value axes (frequency × time = 6 hours/week) and both safety axes (correctable errors, public data) score well, so it sits in the automate-now zone. Twelve minutes across 30 instances is not trivial (A). Marketing email is not inherently high-risk (C). Model price is a build detail that comes after the screen, not a blocker to classification (D).
Which statement best captures why importance is not the same as value when screening automation candidates?
Show answer
Answer: B.
A once-a-year showpiece may be highly important yet has almost no repeat frequency, so the workflow you build never amortises. Cost of building (A) is not what distinguishes the two concepts. They are explicitly not identical (C). Sensitivity is a separate axis unrelated to importance (D).
A workflow would draft press statements that are published verbatim to the company newsroom. Which single FTED axis most constrains how you build it?
Show answer
Answer: C.
Low error tolerance on an external, hard-to-retract artefact forces a mandatory human gate before publication. Frequency (A) affects payback, not how carefully you gate. These statements are not necessarily quick (B). FTED applies to any recurring task including writing (D).
You must pick one pilot. Task X saves 8 hours/week with a light weekly sample; task Y saves 9 hours/week but needs a human to approve every one of its irreversible outputs. Which do you build first and why?
Show answer
Answer: B.
Ranking uses risk-adjusted payback: Y's per-item approval of irreversible output erodes most of its 9 hours and carries higher risk, so X's net saving wins. Raw hours (A) ignore the review cost. X is clearly worth automating (C). Building both at once (D) contradicts the one-pilot constraint and dilutes focus.
A manager asks you to 'use AI for our whole procurement process'. What is the FIRST scoping move?
Show answer
Answer: B.
'The whole process' is a vague wish; scoping narrows it to a bounded problem you can actually build and test. Buying a model (A) is premature. An autonomous agent (C) skips scoping entirely. One giant prompt (D) is both premature and a decomposition anti-pattern.
Which element of a scoped opportunity most directly prevents scope creep during the build?
Show answer
Answer: C.
The boundary is what stops 'draft the reply' quietly growing into 'and send it' and 'and issue the refund'. The trigger (A) defines when it runs, not what it must avoid. Knowledge files (B) and model size (D) do not constrain scope.
Which TWO tasks are the strongest automation candidates on the FTED screen?
Show answer
Answer: A and C.
A and C are frequent, time-consuming, correctable and low-sensitivity — the automate-now profile. Loan disbursements (B) are irreversible and financial. The annual report (D) fails on frequency. Individual legal advice (E) is regulated and high-harm.
A task passes the value screen easily but touches regulated patient data that policy forbids leaving a sanctioned environment. What does the data-sensitivity axis tell you?
Show answer
Answer: C.
Sensitivity is a separate axis that can veto a high-value task; the answer is a sanctioned workspace with clearance, or manual handling. Sensitivity does not discount value (A). Summarising is not inherently safe (B). Model price (D) has nothing to do with data governance.
A single long prompt produces a report that 'works in the demo but occasionally drops a whole section in production and is hard to debug'. What is the BEST fix?
Show answer
Answer: B.
Dropped sections and undebuggable output are the mega-prompt signature; decomposition exposes each section as a step. Padding the prompt (A) worsens the blob. A bigger model (C) hides failures rather than exposing them. Temperature (D) does not address structure.
You must summarise fifteen independent regional reports into one national brief that reconciles totals and terminology. Which pattern fits and what is the critical step?
Show answer
Answer: B.
Independent pieces combined into one consistent output is fan-out/fan-in, and consistency is enforced at the fan-in, which must reconcile rather than concatenate. Sequential (A) misfits independent work. The other two (C, D) are the wrong top-level shape here.
Why is a separate critique step generally better than asking the model to check its own answer in the same reply?
Show answer
Answer: B.
Reviewing in the same pass carries the reasoning that produced the draft; a separate, checklist-driven critique is more objective. It is not about token count (A) or speed (D), and it does not remove human review (C) — often the critique step is where the gate lives.
A workflow reads receipts and outputs a reimbursement total, but a wrong total could be a misread amount OR a misapplied policy rule. Which decomposition makes the failure traceable?
Show answer
Answer: B.
Separating extraction from the rule-based transform makes each independently checkable, so a wrong total traces to reading or logic. Draft-then-critique (A) reviews prose, not this. A detailed single prompt (C) keeps errors blended. Fan-out (D) parallelises but still blends extract and transform per receipt.
When should you STOP splitting a task into more steps?
Show answer
Answer: B.
Splitting earns its keep only when it separates distinct jobs or adds a needed check; beyond that it adds latency and context loss. 'More is always better' (A) and a fixed count (C) are wrong, and prompt length (D) is not the criterion.
A proposal workflow keeps producing an executive summary that describes an earlier version of the body. What ordering fixes this?
Show answer
Answer: B.
A summary must describe finished content, so it is written last in the sequence. Writing it first (A) guarantees drift as the body changes. One prompt (C) is the mega-prompt trap. Skipping it (D) fails the requirement.
A launch-pack workflow produces five assets that sometimes contradict each other on price and key claims. What is the ROOT-CAUSE fix?
Show answer
Answer: B.
Contradictions come from each asset inventing its own facts; a shared extracted fact sheet gives one source of truth all assets draw from. An instruction to be consistent (A) is padding. Generating twice (C) doubles cost without fixing the cause. A bigger model (D) does not create a shared source of truth.
Which task is best kept as a single step rather than decomposed?
Show answer
Answer: B.
A single small job with one output does not benefit from splitting; decomposition would only add overhead. The multi-section report (A), the twelve-document brief (C) and the draft-then-critique task (D) all have distinct sub-goals that fail separately and warrant decomposition.
Which TWO signals in a task description point toward the fan-out/fan-in pattern?
Show answer
Answer: A and B.
Independent pieces combined with a consistency requirement is the fan-out/fan-in signature, where the fan-in reconciles. Acting on facts in one document (C) points to extract-then-transform. A review pass (D) points to draft-then-critique. Strict dependence (E) points to sequential.
A step 'sometimes returns a paragraph, sometimes bullets' and a colleague gets a different format than you. What is the ROOT problem?
Show answer
Answer: B.
Inconsistent format across inputs and users means the contract never specified an output shape; pinning a schema, template or fixed fields fixes it. Model size (A) and temperature (C) do not define structure, and the account (D) is irrelevant to format.
Which set best defines the five parts of a step contract?
Show answer
Answer: B.
A step contract pins what goes in, what must be present, what comes out, what good means, and what happens on missing input. Option A lists model settings, C lists observability metrics, and D lists project-management fields, none of which is a contract.
A categorisation step runs and its required category list is missing. What should it do?
Show answer
Answer: B.
With no safe default and correctness depending on the list, the step must fail loudly rather than guess. Inventing categories (A) fabricates undetectable data, a silent empty string (C) hides the failure, and an arbitrary pick (D) is a disguised guess.
Which is a strong, observable acceptance criterion for a support reply?
Show answer
Answer: C.
Acceptance criteria must be checkable conditions a person or step can verify. 'High quality' (A), 'reads professionally' (B) and 'comprehensive' (D) are all subjective and unverifiable.
You want a step's output parsed by a downstream system. What should the contract require?
Show answer
Answer: B.
Machine consumption needs a strict schema and only the JSON so parsing is reliable. Prose around the data (A) and a narrative (C) break parsers, and letting the model choose (D) reintroduces drift.
After you edit step 2, the workflow breaks at step 3. What is the MOST likely cause?
Show answer
Answer: B.
Editing a step commonly changes its output shape, breaking the consumer that relied on the old contract; re-check both edges. Model size (A) and temperature (C) are not implied, and step 1 (D) was not touched.
When is a marked DEFAULT the right missing-input strategy?
Show answer
Answer: B.
A default is appropriate only when it is safe and documented, and it must be marked so review sees the value was missing. If correctness depends on the value (A), fail instead; avoiding interruptions at any cost (C) leads to silent guessing; and 'always fail' (D) is too rigid when a safe default exists.
Which TWO contract rows are most often left blank in flaky workflows?
Show answer
Answer: A and C.
Undefined missing-input behaviour and an unchecked handoff are the classic gaps that make workflows flaky. Release date (B), word count (D) and colour (E) are not contract rows.
You repeatedly run the same task and it needs the same style guide and reference files every time. Which rung fits BEST?
Show answer
Answer: B.
Recurring context that must be present every time is the defining signal for a Project. Re-pasting (A) is the waste a Project removes, and an API app (C) or agent (D) over-reaches for a solo recurring task.
The decisive question separating a Project from a custom GPT is:
Show answer
Answer: B.
A Project is your recurring workspace; a custom GPT packages the task so others run it independently. Model (A), file count (C) and temperature (D) do not determine which of the two you need.
A task must run inside your company's web app, on demand, wired to your database. Which capability is required?
Show answer
Answer: D.
Embedding in another product, on demand, integrated with a database is exactly what an API application is for. Saved instructions (A), Projects (B) and custom GPTs (C) all live inside ChatGPT and cannot embed programmatically into your app.
A user needs the latest news about a named company from this week. Which in-conversation capability fits BEST?
Show answer
Answer: B.
A quick current-facts lookup is what search is for. Deep research (A) is heavier, for sourced multi-source reports; data analysis (C) is for computation over files; and Canvas (D) is for iterating on documents.
The task is 'produce a thoroughly sourced comparison of the five leading vendors, with citations'. Which capability fits BEST?
Show answer
Answer: C.
A multi-source, cited, comprehensive comparison is the deep-research use case. Plain chat (A) cannot source it reliably, a single search headline (B) under-delivers, and Canvas (D) is an editing surface, not a research tool.
You uploaded a sales spreadsheet and need outliers found and a trend plotted. What must be switched on?
Show answer
Answer: B.
Computation, outlier detection and plotting over a file require data analysis; uploads alone only let the model read the file (A). Search (C) is for current facts, and a custom GPT (D) is a packaging rung, not a compute capability.
Ten colleagues need to run the same configured assistant themselves, getting consistent results without your involvement. Which rung?
Show answer
Answer: B.
Others running the same task independently with consistent behaviour is the custom-GPT signal. A solo Project (A) does not hand the task to others, a one-off prompt (C) yields inconsistency, and an API app (D) over-reaches for in-ChatGPT self-serve.
Which TWO statements correctly justify reaching for the lightest capability rung that meets the need?
Show answer
Answer: A and C.
Climbing the ladder adds total cost of ownership, and under-reaching also has a cost, so you right-size to the actual need. Heavier rungs are not inherently slower (B) or less accurate at lighter rungs (D), and it is a design principle, not a policy rule (E).
A recurring deliverable is a document you refine over several turns in the same session. Which in-conversation capability best supports the iteration?
Show answer
Answer: C.
Canvas is the dedicated surface for iterating on a document or code across turns. Search (A) and deep research (B) gather information, and data analysis (D) computes over files — none is an editing surface.
A multi-step task can run semi-autonomously with review at checkpoints, executing several actions rather than answering a single question. Which rung fits?
Show answer
Answer: B.
Delegated multi-step execution under oversight is exactly a workspace agent. A saved instruction (A) and a one-off prompt (C) answer, they do not execute steps, and deep research (D) is a research capability, not an execution rung.
Which TWO practices make it possible to attribute a quality change to a specific prompt edit and roll it back if it hurts?
Show answer
Answer: A and C.
Changing one variable per version makes effects attributable, and keeping the prior version enables rollback. A bigger model (B) and a longer prompt (D) do not aid traceability, and reviewing every output (E) catches errors but does not attribute which edit caused a change.
A workflow sends customer emails automatically with no human check because 'the model is reliable'. What is the problem?
Show answer
Answer: B.
External, irreversible actions require a human gate regardless of model reliability. 'Reliability justifies it' (A) ignores the mandatory trigger, and model size (C) or temperature (D) do not address the missing gate.
A step runs 3,000 times a day, errors are correctable, and stakes per item are low. Which review mode fits BEST?
Show answer
Answer: B.
High-volume, correctable, low-stakes work is the textbook case for sampling with tracking. Gating all 3,000 (A) throttles and breeds rubber-stamping, no review (C) is reckless, and deep research (D) is unrelated to oversight.
Where should a review gate be placed in a draft-then-send workflow?
Show answer
Answer: B.
The gate belongs just before the irreversible action, reviewing the final content. Before drafting (A) reviews nothing meaningful, after sending (C) is not a gate at all, and placement absolutely matters (D).
Which of these is a MANDATORY review trigger?
Show answer
Answer: C.
Irreversible actions like issuing a payment mandate a human gate. Length (A), speed (B) and formatting (D) are irrelevant to whether a gate is required.
How should the sampling rate change for a brand-new automated step?
Show answer
Answer: B.
New processes lack an evidence base, so you sample heavily at first and relax as quality is demonstrated. Starting low and never changing (A) ignores drift, not sampling new steps (C) is backwards, and waiting for complaints (D) is reactive, not oversight.
Why does over-gating a high-volume, low-stakes step undermine the very safety it seeks?
Show answer
Answer: B.
Excessive gates dull attention and turn review into rubber-stamping, defeating the control. It does not affect hallucination (A) or inference speed (C), and token cost (D) is not the safety issue described.
Which TWO steps in a workflow REQUIRE a mandatory gate rather than sampling?
Show answer
Answer: A and C.
External publication and regulatory filing are irreversible/regulated/external — mandatory gates. Receipt categorisation (B) is high-volume correctable work suited to sampling, and brainstorming (D) and personal summaries (E) are low-stakes, needing at most a spot-check.
What makes a review gate an objective check rather than a reviewer guessing?
Show answer
Answer: B.
Acceptance criteria give the reviewer concrete, checkable conditions, turning review from subjective judgment into an objective check. Model size (A), prompt length (C) and volume (D) do not make a review objective.
You go on leave and your backup cannot run your workflow because it only exists in your head. What is the BEST fix?
Show answer
Answer: B.
A written runbook is what lets a colleague run the workflow end to end without asking questions. 'Figure it out' (A) and 'wait' (D) leave the workflow un-runnable, and a one-off video (C) rarely captures inputs, failure handling and the definition of done reliably.
After several undated prompt edits, a workflow got worse and no one knows which change caused it. What practice would have prevented this?
Show answer
Answer: B.
Dated single-change versions make effects attributable and enable rollback. A bigger model (A) and longer prompts (C) do not address traceability, and reviewing every output (D) catches errors but does not attribute which edit caused them.
A manager asks whether an improved workflow is actually better. Which answer reflects evidence-based iteration?
Show answer
Answer: B.
Improvement is shown against measured metrics on a fixed sample. 'Feels more thorough' (A) and 'I like it more' (C) are vibes, and prompt length (D) is not a quality measure.
Measuring cycle time on a workflow shows 8 of its 20 minutes are human review. What does this tell you about the next improvement?
Show answer
Answer: B.
Cycle-time measurement locates the bottleneck; with review dominating, review design is the lever, not prompting. Prompt length (A) and model price (C) do not touch the review time, and the measurement is highly useful (D).
Which TWO are legitimate data sources for measuring a workflow's ongoing quality?
Show answer
Answer: A and C.
Sampled items scored against acceptance criteria, and the rework/rejection rate at the gate, are objective quality measurements. The author's impression (B) is vibes, and prompt word count (D) and release notes (E) are not quality data.
Which TWO elements are essential in a runbook so a colleague can run the workflow unaided?
Show answer
Answer: A and C.
A definition of done tells the runner when to stop, and failure handling tells them what to do when input is missing — both essential for unaided execution. The author's opinion (B), the training cutoff (D) and a run count (E) do not help a colleague execute it.
Last updated Sep 18, 2026