AI Foundations
AI Foundations Mock Exam 2
A harder, timed 60-item independent mock exam for the OpenAI Academy AI Foundations track, used as a readiness gate before the real assessment.
This is the second full-length independent mock exam for the AI Foundations track, built from publicly available OpenAI learning objectives. It is not an official OpenAI assessment. It is deliberately harder than Mock Exam 1 – more multi-constraint stems, more FIRST / BEST / MOST cost-effective / TWO qualifiers and more scenario framing – so use it as your readiness gate, timed, once you have closed the gaps the diagnostic exposed. All 60 items are new and distinct from Mock Exam 1 and the domain-page questions.
Instructions
- Time: 75 minutes. Take this one timed to rehearse working under a clock.
- Items: 60, single-answer and multiple-response. Each item states how many answers to select.
- Selection: single-answer items use one choice; multiple-response items say Select two and you must pick exactly two – all correct and none incorrect.
- No guessing penalty: answer every question; an unanswered item simply scores zero.
- Target: aim for at least 80% raw (≈ 48/60) here before you take the real Academy assessment, because the Academy badge threshold is 80%.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| 1 | AI and Generative AI Fundamentals | 8 |
| 2 | How Language Models Behave | 10 |
| 3 | ChatGPT Surfaces and Features | 8 |
| 4 | Prompting and Instructions | 12 |
| 5 | Context, Files and Memory | 7 |
| 6 | Verifying and Evaluating Output | 9 |
| 7 | Responsible and Safe Use | 6 |
Total: 8 + 10 + 8 + 12 + 7 + 9 + 6 = 60 items.
Readiness interpretation
This is an independent readiness indicator, not a score and not a pass mark. Because this exam is harder, a score in the upper bands here is a strong signal of readiness.
| Raw score | Band | What it means |
|---|---|---|
| under 70% | Keep learning | The harder framing found real gaps; revisit those domains before the assessment. |
| 70–79% | Building confidence | Solid, but tighten the domains where multi-constraint items caught you. |
| 80–89% | Assessment ready | At or above the Academy badge threshold on the harder set; you are in good shape. |
| 90%+ | Strong readiness | Excellent on the tougher exam; you are well prepared for the Academy assessment. |
The 80% line matches the OpenAI Academy badge threshold. An Academy badge or pathway certificate of completion is not a certification and does not guarantee eligibility for a future OpenAI certification.
Take the mock exam
Two ways to use the questions below. The interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction. The review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
60 questions · one at a time · 75-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A stakeholder claims that because ChatGPT wrote correct Python, it must therefore understand mathematics the way a person does. What is the MOST accurate correction?
Show answer
Answer: B.
The model generates probable tokens; correct output demonstrates capability, not human-style understanding. A overclaims. C is backwards, since it is a language model. D assumes rote replay, which mischaracterises generative next-token prediction and is usually untrue for novel prompts.
You must estimate whether a 6,000-word briefing plus a 2,000-word appendix will fit alongside instructions in a GPT Instant window on ChatGPT Business. Using the standard rule of thumb, what is the BEST estimate and conclusion?
Show answer
Answer: B.
At roughly four characters per token, 8,000 words is about 10,700 tokens, well under 54K, leaving room for instructions. A underestimates by using far too few tokens per word. C wrongly claims 10,700 exceeds 54K. D both overestimates the count and misstates the limit.
A team asks for the SINGLE most reliable general explanation of why an LLM sometimes miscounts letters in a word. Which is it?
Show answer
Answer: A.
Models operate on tokens rather than characters, so exact letter counting is not a natural operation. B invents a feature. C confuses training recency with character handling. D misapplies context limits to a per-word counting task.
A knowledge worker wants the MOST cost-effective way to handle a repeatable, high-volume extraction job with clear rules, and asks which factor matters FIRST. What should drive the choice?
Show answer
Answer: B.
For clear, repeatable, high-volume tasks, a lightweight fast model is the cost-effective fit. A overspends. C uses the wrong modality. D optimises for recency, which is irrelevant to a rules-based extraction task.
Which statement BEST captures why fluency is not the same as accuracy in generated text?
Show answer
Answer: B.
Well-formed wording reflects probability, not truth, so fluent text can be false. A wrongly equates fluency with correctness. C ties accuracy to speed. D is false, since citations themselves can be fabricated.
A colleague asks you to name TWO facts that BEST explain why a precise, unsourced statistic about a small private company should be treated with caution.
Show answer
Answer: A and B.
The model can fabricate precise figures, and niche private data is often unavailable without search. C is false; there is no built-in verification. D is false; unsourced numbers are common. E is an overstatement that is not why caution is warranted.
A manager argues that adopting AI means the company should stop hiring analysts because the model computes everything reliably. What is the MOST accurate framing to offer FIRST?
Show answer
Answer: B.
The model assists but does not replace human judgment, and exact computation needs tools. A overclaims reliability. C understates its usefulness. D wrongly limits failure modes to recency, ignoring hallucination and arithmetic errors.
For a task that must reflect a regulatory change made three months ago, which combination is the MOST dependable FIRST choice?
Show answer
Answer: B.
Recent facts need retrieval plus verification against the authoritative source. A risks stale or fabricated content past the cutoff. C adds compute but cannot supply missing recent facts. D invites a confident guess with no grounding.
A user sets what they think is a temperature of zero and declares the model can no longer hallucinate or vary. What is the MOST accurate correction?
Show answer
Answer: B.
Reducing randomness affects variation, not factual grounding, and the chat product does not surface raw temperature. A conflates determinism with truth. C misstates what the setting does. D invents a lookup mode that does not exist.
A user wants to reliably retrieve an excellent draft they generated last week, but the chat is gone. What is the BEST practice going forward?
Show answer
Answer: B.
Because generation is non-deterministic, valuable outputs must be saved intentionally. A expects exact reproduction that will not happen. C does not make outputs reproducible. D wrongly assumes memory stores full drafts, which it does not.
On ChatGPT Enterprise, a user compares the GPT Instant and GPT Reasoning total context windows and asks which pairing is correct.
Show answer
Answer: A.
On Enterprise the GPT Instant total is 128K and Reasoning is 256K. B uses the Business Instant figure. C swaps the two values. D quotes an API-scale window not applicable to these ChatGPT plans.
A user needs BEST quality on a hard, open-ended strategy problem but is cost-conscious. Which selection reasoning is soundest?
Show answer
Answer: B.
Hard open-ended work needs a capable model, but effort should match difficulty rather than defaulting to maximum. A under-provisions for a hard task. C pairs high effort with an underpowered model. D uses the wrong modality entirely.
A model gives a fast, wrong answer to a chained arithmetic-and-logic problem. A user increases the reasoning effort and it improves but is still occasionally wrong. What is the MOST reliable next step?
Show answer
Answer: B.
Exact computation belongs in a tool, with the logic verified, since more effort alone does not guarantee arithmetic correctness. A skips verification on a task known to fail. C reduces rigour. D relies on repeated fallible generations rather than grounded computation.
A user notices the model enthusiastically endorses whichever option they frame favourably. To get the MOST balanced decision support, what should they do FIRST?
Show answer
Answer: B.
Neutral framing that forces balanced pros, cons and counterarguments counters sycophancy. A amplifies the bias. C adds capacity irrelevant to framing. D changes modality and does not address the reasoning need.
In a marathon planning chat, the model starts contradicting decisions made hours earlier. Which explanation and remedy is MOST accurate?
Show answer
Answer: B.
Long chats lose early context; carrying forward a summary or restarting preserves the decisions. A misreads normal behaviour. C confuses training cutoff with in-chat context. D is a billing concept unrelated to contradictions.
Select the TWO statements that correctly describe consequences of probabilistic, non-deterministic generation for everyday ChatGPT use.
Show answer
Answer: A and B.
Non-determinism explains varied wording and the need to save good outputs. C is false; preferences do not fix exact bytes. D is false; there is no built-in fact verification. E is false; variability and hallucination are separate phenomena.
A user reports that quality degraded after they pasted a dozen long, loosely related reports into one chat, though a token counter shows they are under the window. What is the MOST likely cause and fix?
Show answer
Answer: B.
Under the limit, relevance and volume still matter; pruning to relevant files fixes dilution. A contradicts the premise and prunes blindly. C invents a limitation. D adds compute without reducing the clutter causing the problem.
A user must confirm a competitor price announced this morning, then produce an exact side-by-side cost table from a 300-row sheet. Which pairing of features is correct, in order?
Show answer
Answer: B.
Search handles the fresh fact and data analysis computes exact figures. A misuses memory and image generation. C uses deep research for a one-line fact and wrongly expects canvas to compute. D is a tutoring feature that does neither task.
A team wants a shared, standing configuration of instructions and reference files that every member's chats can draw on for one client, without pasting anything. Which is the BEST fit?
Show answer
Answer: B.
A shared Project holds standing instructions and files across members' chats. A is per-person and small-scale. C automates timing, not shared context. D degrades as it grows and does not cleanly share configuration.
A user needs a well-sourced, multi-article synthesis on a niche standard AND wants each claim traceable to a source. Which feature and follow-up is BEST?
Show answer
Answer: B.
Deep research synthesises cited material, and confirming key citations makes claims traceable. A offers no sourcing. C and D are the wrong modalities and provide no traceable synthesis.
A manager insists every task run on the most powerful model to be safe. Which single point BEST pushes back on cost-effectiveness grounds?
Show answer
Answer: B.
Matching model and effort to task avoids paying frontier prices for simple work. A is false on cost. C is false; capable models handle simple tasks but wastefully. D denies the real cost and latency differences.
An analyst wants to learn a forecasting method deeply rather than be handed the final numbers, then later needs the numbers computed exactly. Which sequence is BEST?
Show answer
Answer: A.
Study mode teaches the method; data analysis computes exact numbers with code. B uses search for a concept and misuses memory for computation. C is the wrong modality. D skips learning and risks unreliable arithmetic.
Select the TWO correctly matched feature-to-need pairs.
Show answer
Answer: A and B.
Company knowledge answers from connected internal sources and canvas supports inline editing. C should use data analysis, not memory. D should use search, not image generation. E should use a file upload, not a scheduled task.
A user uploads a 60-page PDF and asks for a summary that states a specific compliance figure they will forward to legal. What is the correct posture BEFORE forwarding?
Show answer
Answer: B.
A high-stakes figure destined for legal must be checked against the source document. A over-trusts summarisation. C relies on self-reported confidence. D alters the figure without confirming it.
A task prompt says: you are a financial analyst; using the attached CSV, report the three largest cost categories as a table with columns Category and Total; success means the totals reconcile to the sheet. Which components are present?
Show answer
Answer: A.
The prompt names a role, a task, the CSV as context, a table format and a reconciliation success criterion. B, C and D each undercount the components clearly present in the stem, so they misidentify the well-formed prompt.
An output must feed an automated importer that expects fixed fields. The first attempt returned prose with the data embedded. What is the BEST single instruction to add?
Show answer
Answer: B.
A defined schema with no extra prose makes output importable. A shortens prose but keeps it unstructured. C adds compute without defining structure. D changes tone, which is irrelevant to machine import.
A prompt lists five prohibitions and no positive target; outputs keep missing the mark. Applying prompt-anatomy thinking, what is the MOST effective single fix?
Show answer
Answer: B.
A positive target with an example steers far better than stacked prohibitions. A adds more negatives without a target. C removes needed detail. D adds randomness, not direction.
A complex deliverable needs research, a recommendation, a slide outline and a summary email, each depending on the previous. What is the MOST reliable prompting structure?
Show answer
Answer: B.
Sequential, reviewed steps let errors be caught before they propagate through dependent artefacts. A makes review and error isolation hard. C skips required work. D introduces disorder without benefit.
A user reports that the model consistently ignores the single most important constraint. Which combination of fixes is MOST likely to work?
Show answer
Answer: B.
Prominence, labelling and turning it into a checkable criterion all raise the chance the constraint is honoured. A reduces prominence. C abandons the requirement. D drowns it in noise.
Two prompts differ only in that one adds a named audience and a concrete success criterion. The second consistently produces better output. What does this BEST illustrate?
Show answer
Answer: B.
The improvement comes from relevant specifics and a checkable criterion, not raw length. A confuses the mechanism with length. C denies the observed effect. D wrongly restricts the principle to one model class.
A user wants a screening email that invites a call and names one specific skill from the job post. In prompt anatomy, where do those two requirements belong?
Show answer
Answer: B.
Required actions and specific content are task and constraint elements. A sets a persona, not required content. C names who it is for, not what it must contain. D defines shape, not the required actions and skill mention.
A first draft is 90% right but uses the wrong currency and one wrong date. What is the MOST disciplined iteration move?
Show answer
Answer: B.
A precise, minimal correction preserves the 90% that works and fixes the two errors. A discards good work. C repeats the same errors. D changes the tool when a small edit suffices.
Select the TWO changes that would MOST improve a prompt whose output was on-topic but flat and aimed at the wrong reader.
Show answer
Answer: A and B.
A tone instruction lifts flat copy and naming the audience fixes the wrong-reader issue. C adds negatives without a target. D adds compute unrelated to tone or audience. E repeats the same shortfall.
You want consistent two-sentence summaries across 200 documents with identical structure. Which approach gives the MOST consistent shape at scale?
Show answer
Answer: B.
A template, example and explicit structure produce consistent shape across many items. A leaves consistency to chance. C guarantees variation. D adds cost without enforcing structure.
Which is the STRONGEST success criterion for a quarterly board update, judged by how checkable it is?
Show answer
Answer: B.
B is specific and verifiable across length, content and reconciliation. A, C and D are vague and cannot be checked, so the model cannot reliably satisfy them or know when it has.
A colleague pads prompts with adjectives believing it raises quality, yet results are no better. What is the MOST accurate explanation to give?
Show answer
Answer: B.
Substantive components, not decoration, drive quality. A is false. C ignores the well-established effect of prompt structure. D overcorrects into an absolute rule that is not the point.
Using an information-placement view, which TWO statements about a reference document reused across many chats in a single workstream are correct?
Show answer
Answer: A and B.
A reused reference document belongs in the Project, and pasting it repeatedly is inefficient and error-prone. C misuses memory for a document. D automates timing, not document access. E denies the reuse a Project explicitly provides.
A user wants a small standing preference applied to all future chats without re-typing it, and asks how memory behaves. Which statement is MOST accurate?
Show answer
Answer: B.
Memory holds small stable preferences across chats and is user-manageable. A overstates its scope. C confuses memory with training. D contradicts its cross-chat persistence.
A marathon chat kept open so nothing is lost is producing vaguer, slower answers. What is happening and the BEST fix?
Show answer
Answer: B.
A bloated, aging context degrades quality; summarising essentials and restarting restores focus. A gives up unnecessarily. C confuses training cutoff with context. D invents a memory-capacity cause.
A one-off constraint that applies only to the current deliverable keeps getting saved to memory, so it wrongly affects later unrelated chats. What is the fix and the principle?
Show answer
Answer: B.
One-off instructions belong in the prompt; only stable preferences belong in memory. A causes exactly the reported bleed-through. C would still over-apply it across the Project's chats. D is an overreaction that discards useful stable preferences.
A consultant handles three clients and needs each client's preferences and files kept strictly separate with the LEAST ongoing effort. What is the BEST structure?
Show answer
Answer: B.
A Project per client keeps contexts cleanly separated with little ongoing effort. A risks cross-contamination. C mixes clients in memory. D is high-effort and error-prone every session.
Select the TWO practices that BEST maintain clean context across a week of varied work.
Show answer
Answer: A and B.
Fresh chats per topic and using Projects and memory appropriately keep context clean. C clutters every chat. D causes drift and dilution. E stores sensitive data where it should not live.
A board pack is due in an hour; it reads fluently, but two subtotal rows do not match the headline total and one citation is unfindable. What should you do FIRST?
Show answer
Answer: B.
Internal-consistency and citation failures are blocking issues; fix the numbers and the source first. A ships known errors. C trusts the fallible source. D conceals the problem instead of resolving it.
An analyst validates a summary's claim that a contract's notice period is 60 days. The contract was uploaded. What MOST appropriately validates it?
Show answer
Answer: B.
Grounding the claim means reading the actual clause in the source document. A over-trusts a summary that can misread. C re-uses the fallible generator. D compares to an irrelevant document rather than the source.
A dashboard reports 92% accuracy on an AI-assisted tagging task and a manager wants to scale it. What matters MOST to check FIRST before expanding?
Show answer
Answer: B.
Aggregate accuracy can hide costly error patterns, so error distribution and impact matter most first. A ignores where the 8% falls. C is cosmetic. D relies on confidence, which is not an accuracy signal.
For a number the model COMPUTED from data you provided, which validation action correctly matches its provenance?
Show answer
Answer: B.
A computed figure is validated by independent recomputation, matching its provenance. A suits externally sourced facts, not computed ones. C applies to cited claims. D uses another fallible generator rather than recomputing.
A researcher has a summary with eight citations and limited time. What is the MOST effective triage to catch fabrication FIRST?
Show answer
Answer: B.
Checking the claims that matter most, then escalating if any is fabricated, is efficient triage. A assumes quantity implies reliability. C checks style, not truth. D relies on the source that produced them.
Select the TWO verification steps that are FREE and require NO external source.
Show answer
Answer: A and B.
Relevance and internal-consistency checks need only the output itself. C, D and E each require an external authoritative source and so are neither free nor source-independent.
Select the TWO practices that are weak checks and do NOT actually verify accuracy.
Show answer
Answer: A and B.
Self-reported confidence and inter-model agreement are not grounding and can be confidently wrong. C, D and E are genuine verification against a source or by independent computation, so they are strong checks, not weak ones.
A colleague argues that an internal-only forecast does not need verification because it will not be published. The forecast will set next year's budget. What is the BEST response?
Show answer
Answer: B.
Rigour scales with consequence, and a budget-setting forecast is high-stakes regardless of publication. A and C wrongly tie checking to audience or authority. D re-uses the same fallible generator instead of verifying.
An output you will act on is long and reads smoothly. Applying a cost-ordered verification approach, which check comes FIRST?
Show answer
Answer: B.
Start with the cheapest, fastest check that catches obvious errors, then escalate only as needed. A and C are heavyweight later steps. D is not a verification step at all and wastes effort before checking correctness.
A company policy names an Enterprise workspace as the ONLY sanctioned place for confidential data, but a personal account is faster today. An employee is under deadline pressure. What should they do?
Show answer
Answer: B.
Confidential data must stay in the sanctioned workspace regardless of deadline pressure. A violates policy and risks exposure. C still exposes data on an unsanctioned account. D delegates a governance decision to the model.
A vendor comparison the model produced lists many pros for Option A and mostly cons for Option B, matching how the user framed the request. What is the issue and the BEST fix?
Show answer
Answer: B.
The lopsided result reflects biased framing plus sycophancy; a neutral, balanced re-request fixes it. A ignores the bias. C addresses fabricated facts, not skew. D addresses long-chat drift, which is not the cause here.
When is disclosure of AI assistance MOST clearly expected?
Show answer
Answer: B.
Disclosure is expected where the context's rules or norms require it, such as academic or regulated settings. A is private and low-stakes. C ignores real disclosure norms. D ties disclosure to correctness, which is not the criterion.
A model output defaults to gendered role assumptions, casting the engineer as male and the assistant as female. What is this and the correct FIRST response?
Show answer
Answer: B.
Default gendered roles are a form of bias to correct and watch for. A dismisses a real fairness issue. C misclassifies bias as fabrication. D is unrelated to how the bias arises.
An educator drafts a quiz with ChatGPT to use with students next week. What responsible-use step is MOST essential before using it?
Show answer
Answer: B.
AI-drafted educational content must be reviewed for accuracy, bias and fit before use with students. A over-trusts a draft. C relies on the same fallible source. D exposes students to unreviewed errors first.
Select the TWO outputs that MOST clearly require a mandatory human review gate before acting.
Show answer
Answer: A and B.
A published legal disclaimer and a hiring rejection are high-stakes and consequential, needing human review. C, D and E are low-stakes and easily reversible, so they do not require a formal gate.
A user insists the model must be right because it answered instantly and never said it was unsure. What is the MOST accurate response?
Show answer
Answer: B.
Confidence and speed are not accuracy signals; the model can be fluently wrong without hedging. A treats surface cues as proof. C is false; models often do not flag uncertainty. D invents a verified-lookup source that does not exist.
A user needs a reusable team assistant that always applies a fixed review checklist AND can be shared, while a separate one-off question needs today's exchange rate. Which pairing is correct?
Show answer
Answer: A.
A custom GPT packages a reusable, shareable assistant and Search fetches a fresh rate. B misuses memory and image generation. C uses a scheduler and heavyweight research for the wrong needs. D is an editing surface for neither task.
A user needs a small, stable personal preference applied everywhere, a single 25-page report summarised once, and a client's standing brand rules shared with a team. Which mapping is correct?
Show answer
Answer: A.
Small stable preferences fit memory, a one-time document fits a file upload, and shared standing rules fit a shared Project. B misplaces each item. C overloads memory with a document and shared rules. D relies on a degrading single chat.
Last updated Sep 18, 2026