AI Foundations
AI Foundations Mock Exam 1
A 60-item, domain-weighted independent mock exam for the OpenAI Academy AI Foundations track, with full explanations and a readiness indicator.
This is a full-length, domain-weighted independent mock exam for the AI Foundations track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment – there is no public OpenAI exam to have questions from. Use it as your diagnostic: sit it first to find your two weakest domains before you enrol in and take the real OpenAI Academy assessment. All 60 items are new and do not repeat the domain-page questions or Mock Exam 2.
Instructions
- Time: 75 minutes, matching the length of the Academy AI Foundations course experience. You can also take it untimed the first time.
- Items: 60, single-answer and multiple-response. Each item states how many answers to select.
- Selection: single-answer items use one choice; multiple-response items say Select two and you must pick exactly two – all correct and none incorrect.
- No guessing penalty: answer every question; an unanswered item simply scores zero.
- Target: aim for at least 80% raw (≈ 48/60) before you take the real Academy assessment, because the Academy badge threshold is 80%.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| 1 | AI and Generative AI Fundamentals | 8 |
| 2 | How Language Models Behave | 10 |
| 3 | ChatGPT Surfaces and Features | 8 |
| 4 | Prompting and Instructions | 12 |
| 5 | Context, Files and Memory | 7 |
| 6 | Verifying and Evaluating Output | 9 |
| 7 | Responsible and Safe Use | 6 |
Total: 8 + 10 + 8 + 12 + 7 + 9 + 6 = 60 items.
Readiness interpretation
This is an independent readiness indicator, not a score and not a pass mark. Read your raw percentage against these bands.
| Raw score | Band | What it means |
|---|---|---|
| under 70% | Keep learning | Revisit the heaviest domains (Prompting, How Language Models Behave) before retrying. |
| 70–79% | Building confidence | Close to ready; target your weakest one or two domains. |
| 80–89% | Assessment ready | At or above the Academy badge threshold; review any weak domain and proceed. |
| 90%+ | Strong readiness | Consistent across domains; you are well prepared for the Academy assessment. |
The 80% line is deliberate: it matches the OpenAI Academy badge threshold. Remember that an Academy badge or pathway certificate of completion is not a certification.
Take the mock exam
Two ways to use the questions below. The interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction. The review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
60 questions · one at a time · 75-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
What is the relationship between artificial intelligence, machine learning and generative AI?
Show answer
Answer: C.
AI is the umbrella; machine learning is one approach to AI; generative AI is a category of machine learning focused on producing new content. B inverts the nesting. A denies the well-established containment. D conflates a broad field (ML) with one of its subcategories (generative AI).
A large language model produces its reply by which underlying process?
Show answer
Answer: B.
An LLM generates text by predicting the most probable next token given the context, one token at a time. A describes a lookup system, not a generative model. C describes rule-based software, not a neural network. D describes web search, which only happens when a search tool is explicitly enabled.
Roughly how many tokens does 750 words of typical English prose occupy?
Show answer
Answer: C.
A common rule of thumb is that a token is about three-quarters of a word, so 1,000 tokens is roughly 750 words. A and B are far too low; D is roughly ten times too high. Knowing the rough ratio helps estimate how much text fits in a context window.
Why can a model give a confidently wrong answer about an event that happened last week?
Show answer
Answer: A.
Each model is trained up to a knowledge cutoff date; events after it are unknown unless a search tool is enabled, and the model may still answer fluently. B invents a withholding policy. C is false; recency, not complexity, is the issue. D confuses context limits with the knowledge cutoff.
What does it mean that the latest OpenAI models are multimodal?
Show answer
Answer: B.
Multimodal means the model handles more than one modality of input, for example text plus images (vision), while producing text output. A is about deployment, not modality. C describes an orchestration pattern, not what multimodal means. D narrows multimodality to language fine-tuning, which is unrelated.
A user believes that the more their team chats with ChatGPT, the smarter their own copy of the model becomes over time. What is accurate?
Show answer
Answer: B.
The model weights are set during training and do not change as you chat; enterprise business data is not used to train models by default. A confuses use with training. C ties intelligence to billing rather than model capability. D confuses model learning with context size, which is plan-dependent and unrelated to usage volume.
Which task is the best fit for a probability-based text generator with no additional tools?
Show answer
Answer: B.
Tone rewriting is a pure language task that plays to the model's strength. A and D require live data the model cannot know without a search tool. C requires exact arithmetic best done by a calculation tool, not next-token prediction.
Which TWO statements correctly distinguish training from inference?
Show answer
Answer: A and B.
Training produces the fixed parameters ahead of time; inference is the run-time generation of a response. C is wrong because inference does not change weights. D is wrong because training is not live. E invents a re-download step that does not exist.
A user runs the identical prompt twice in two new chats and gets differently worded answers. What does this indicate?
Show answer
Answer: B.
Generation is probabilistic, so the same prompt can yield differently worded but often equivalent answers. A misreads normal behaviour as a fault. C is unfounded. D is impossible; the cutoff is a fixed property of the model, not something that shifts between two requests.
Why is a fast, fluent, confident answer not evidence that it is accurate?
Show answer
Answer: A.
The model optimises for probable, well-formed text; that says nothing about factual truth, so fluent text can be wrong. B is too absolute; speed does not determine accuracy. C is false, there is no built-in truth signal in confident phrasing. D is false; any model can produce fluent text.
On a ChatGPT Business plan, what is the total context window for a GPT Instant (fast) model versus a GPT Reasoning model?
Show answer
Answer: A.
On Business, the GPT Instant total window is 54K and the GPT Reasoning total is 256K. B matches Enterprise Instant, not Business. C is the Reasoning figure applied incorrectly to both. D is an API-scale window, not a ChatGPT plan figure.
A simple rephrasing task is being run at the highest reasoning effort and is slow and expensive. What is the best adjustment?
Show answer
Answer: B.
The guidance is to use the lowest reasoning effort that produces the needed result; a rephrasing task does not need deep reasoning. A wastes time and money. C uses an inappropriate specialised model. D adds cost and clutter without helping a simple language task.
A multi-step logic problem gets a fast, confident answer that skips steps and is wrong. What is the most appropriate fix?
Show answer
Answer: B.
Multi-step logic benefits from a reasoning model or explicit step-by-step working. A repeats the same shortcoming. C removes needed detail. D would add randomness, not rigour, and is not exposed for chat use in the way implied.
A user prompts the model to explain why their new slogan is obviously the best one and receives glowing agreement. What behaviour is this and the best counter?
Show answer
Answer: B.
The model tends to agree with the framing it is given; asking for weaknesses and counterarguments counters this sycophancy. A addresses fabricated facts, not agreeableness. C addresses wording variance. D addresses long-conversation drift, none of which is the issue here.
In a very long chat, the model appears to forget details stated at the start. What is the most likely cause?
Show answer
Answer: A.
As a conversation grows, older turns can drop out of the working context, so the model no longer sees them. B invents a deletion behaviour. C confuses training cutoff with in-chat context. D is a billing limit, unrelated to forgetting mid-conversation.
ChatGPT sums a 25-row expense table and returns a total. What is the most reliable way to trust the number?
Show answer
Answer: B.
Exact arithmetic should be computed with a tool or checked externally, not trusted to next-token prediction. A over-trusts the model. C relies on self-reported confidence, which is not an accuracy signal. D changes tone, not correctness.
A user pastes eight long documents into one chat and quality drops even though the total is under the window limit. What is the most likely explanation?
Show answer
Answer: B.
Even within the limit, large amounts of loosely related context can dilute the model's attention on what matters. A contradicts the premise that the total was under the limit. C and D invent hard rules that do not exist; the issue is relevance and volume, not a per-document cap.
Which TWO behaviours follow directly from the model predicting probable text rather than looking up facts?
Show answer
Answer: A and B.
Because it generates probable text, it can invent plausible citations and state wrong numbers confidently. C is false; models often answer rather than refuse when uncertain. D is false; generation is non-deterministic. E is false; browsing requires an explicit search tool.
A team wants to summarise anonymised aggregate figures rather than raw records containing customer identifiers. Why is this the more responsible choice?
Show answer
Answer: A.
Working from anonymised aggregates minimises personal-data exposure while meeting the need, which is good data hygiene. B confuses privacy with speed. C invents an automatic rejection that does not exist. D confuses privacy with accuracy, which are unrelated.
A high-volume task classifies thousands of short support tickets into fixed categories. Which model is most cost-effective?
Show answer
Answer: B.
Luna is designed for clear, repeatable, high-volume work such as classification, making it the cost-effective fit. A overspends on a simple task. C and D are specialised for domains and modalities unrelated to text classification.
A user needs a five-page, well-sourced briefing that synthesises many articles on a regulatory topic. Which feature fits best?
Show answer
Answer: B.
Deep research is built to gather and synthesise many sources into a longer, cited briefing. A cannot browse or synthesise at that depth. C and D are for images and speech, not multi-source research.
A team reuses the same brand guidelines and tone-of-voice instructions across dozens of chats each week. Where should this live?
Show answer
Answer: B.
A Project holds custom instructions and files that persist across its chats, ideal for reused brand context. A is error-prone and wasteful. C is impossible; you cannot add to training data. D schedules recurring runs, not standing context.
A user must total and chart a 400-row spreadsheet and needs the numbers to be exactly correct. Which feature should they use?
Show answer
Answer: B.
Data analysis runs actual code to compute totals and charts, giving reliable numbers. A relies on next-token estimation, which is unreliable for arithmetic. C stores preferences. D is a tutoring feature, neither of which computes spreadsheet totals.
An employee needs an answer that exists only in the company's connected internal document store. Which feature is correct?
Show answer
Answer: B.
Company knowledge connects to internal sources so answers can draw on them. A searches the public web, not internal documents. C produces images. D cannot contain private company documents added after training.
A user is repeatedly revising a single proposal document and wants to edit it inline rather than copy-paste between chat and a document. Which surface is best?
Show answer
Answer: A.
Canvas provides a side-by-side editable document surface for iterative drafting. B is speech-oriented. C automates recurring runs. D gathers sources; none supports inline document editing the way canvas does.
A user wants a reusable, shareable assistant preconfigured for a repeated team task. Which surface fits?
Show answer
Answer: B.
A custom GPT packages instructions and behaviour into a reusable, shareable assistant. A does not persist or share. C stores small preferences, not a configured assistant. D retrieves web results and is not a reusable assistant.
Which TWO tasks are correctly matched to a ChatGPT surface or feature?
Show answer
Answer: A and B.
Search handles fresh web facts and canvas supports inline document editing. C is wrong; exact totals need data analysis, not memory. D is wrong; image generation makes pictures, not news. E is wrong; a scheduled task automates runs, it does not hold a document for summarising.
A prompt reads write something about our new pricing and the output is generic. Which prompt components are most likely missing?
Show answer
Answer: B.
The prompt lacks audience, context and constraints, so the model has nothing to make the output specific. A ignores the obvious gaps. C and D address compute and capacity, not the missing instructions that cause generic output.
Which is the best statement of a success criterion for an executive summary?
Show answer
Answer: B.
A good success criterion is specific and checkable: length, required content and style. A, C and D are vague and cannot be verified, so the model cannot reliably meet them or know when it has.
A first output is close but too formal. What is the disciplined next step?
Show answer
Answer: B.
Disciplined iteration makes one targeted change rather than discarding a near-miss. A throws away progress. C repeats the same result. D changes the tool when only a small tone adjustment is needed.
A task requires researching a market, choosing a strategy, and drafting both a deck outline and a press release. What is the best prompting approach?
Show answer
Answer: B.
Multi-stage work is best decomposed into reviewable steps so errors are caught early. A overloads one prompt and makes review hard. C removes a necessary step. D changes randomness, not structure.
A user wants summaries in a consistent two-bullet style across many documents. What is the most reliable way to get it?
Show answer
Answer: B.
Specifying the format and giving an example makes the output shape reliable and repeatable. A leaves it to chance. C adds words without structure. D adds compute without defining the format.
A user keeps saying what they do not want and the output still misses. What is the best fix?
Show answer
Answer: B.
Positive, concrete instructions steer far better than a list of prohibitions. A piles on negatives without telling the model the target. C changes formatting, not substance. D is irrelevant to prompt quality.
Which output-shape request is most appropriate when the result will be imported into another system?
Show answer
Answer: B.
Downstream systems need predictable structure, so a defined table or JSON with named fields is best. A is hard to parse reliably. C and D are unsuitable formats for machine import.
A colleague believes a longer prompt is automatically a better prompt. What is the most accurate correction?
Show answer
Answer: B.
What helps is relevant, clear context; padding can dilute focus and reduce quality. A is false. C invents an arbitrary cap. D wrongly ties the principle to one model class.
A prompt buries its most important constraint in the last sentence and the model ignores it. What is the best fix?
Show answer
Answer: B.
Placing and labelling the key constraint prominently makes it more likely to be honoured. A abandons the requirement. C buries it further among more instructions. D reduces compute, which does not address placement.
After three identical re-sends of a vague prompt, a user concludes the model cannot do the task. What is the accurate diagnosis?
Show answer
Answer: B.
Repeating a vague prompt cannot fix vagueness; adding specifics usually can. A blames capability prematurely. C and D invent limits the scenario does not support; the root cause is prompt quality.
A marketing prompt produces on-topic but flat copy. Adding which single component would most likely lift the tone to match the brand?
Show answer
Answer: A.
Tone is controlled by an explicit voice instruction and an example, which directly addresses flat copy. B, C and D add capacity, compute or material but none specifies the tone the brand needs.
Which prompt component tells the model the shape of the output, such as a table with named columns?
Show answer
Answer: B.
The format specification defines the output shape, such as a table with named columns. The role sets a persona, the audience sets who it is for, and the success criterion sets when it is done; none of those defines shape.
Which TWO signals in an AI output should raise suspicion and prompt verification before you act on it?
Show answer
Answer: A and B.
Unsourced precise figures and internal-consistency failures are genuine red flags. C is not a warning sign, since fluency is expected and unrelated to accuracy. D is a good sign, not a concern. E, speed, says nothing about correctness.
Which TWO situations most clearly call for keeping a human accountable rather than letting the model decide?
Show answer
Answer: A and B.
Layoff decisions and public regulatory filings are consequential and require human accountability. C, D and E are low-stakes, easily reversible language tasks that do not need a decision gate, so they are the weaker choices.
A team pastes the same brand-voice guidelines into every new chat. Where should these live instead?
Show answer
Answer: A.
Project custom instructions persist across the Project's chats, removing repetitive pasting. B degrades as the chat grows. C is impossible for users. D keeps the inefficient status quo.
Which item is the best candidate for memory?
Show answer
Answer: B.
Memory suits small, stable, recurring preferences. A is a single-use instruction that belongs in the prompt. C belongs in a file upload. D is sensitive data that should not be stored in memory.
A user needs the model to summarise one specific 20-page report. Where should the report go?
Show answer
Answer: B.
A specific document to summarise belongs in a file upload the model can read. A stores small facts, not documents. C is for standing behaviour. D automates timing, not document ingestion.
Why can attaching many loosely related files degrade answers even below the context limit?
Show answer
Answer: B.
Loosely related material dilutes focus, so relevance matters even under the limit. A is false; files are used. C invents a one-file rule. D asserts an unrelated side effect that does not occur.
A user manages three clients in one long chat and the model mixes up their preferences. What is the best structural fix?
Show answer
Answer: B.
Separating clients into distinct Projects or chats prevents cross-contamination of context. A leaves the mixed context in place. C compounds the mixing in memory. D adds compute without fixing the structural overlap.
A one-off instruction applies only to the single email you are writing now. Which TWO statements about where it belongs are correct?
Show answer
Answer: A and B.
A single-use instruction goes in the immediate prompt and should be kept out of memory. C would wrongly persist a one-off across chats. D automates recurring runs, which does not apply. E is for shared internal reference sources, not a one-off email instruction.
Which TWO are good context-hygiene practices?
Show answer
Answer: A and B.
Fresh chats per topic and attaching only relevant files keep context clean. C clutters the window. D stores sensitive data inappropriately. E is inefficient when Projects or memory can hold standing context.
ChatGPT returns five precise, unsourced market statistics for a niche segment you will cite externally. What is the best next step?
Show answer
Answer: B.
Unsourced, specific statistics for external use must be verified against primary sources. A risks publishing fabrications. C relies on self-reported confidence. D disguises the problem without confirming truth.
Which check is free and should usually come first on an output you will act on?
Show answer
Answer: B.
A fast plausibility and consistency read costs nothing and catches many errors first. A and C are heavier, costlier steps for later. D adds another fallible model rather than a grounded check.
A citation has a plausible title, real-looking authors and a real journal name. How do you validate it?
Show answer
Answer: B.
Only finding the real source confirms a citation is genuine and supports the claim. A trusts plausible-looking fabrication. C asks the same fallible source. D checks style, not existence or content.
Two colleagues verify a high-stakes figure by asking two different AI tools; both agree, so they accept it. What is the flaw?
Show answer
Answer: B.
Two models can be confidently wrong in the same way; agreement between generators is not verification against a source. A is false. C just adds more fallible generators. D reduces independence further.
Which output most clearly requires a mandatory human review gate before acting?
Show answer
Answer: B.
A published, customer-facing legal notice is high-stakes and irreversible, so it needs human review. A, C and D are low-stakes and easily corrected, so they do not require a formal gate.
A research summary lists eight citations; the analyst opens two and one does not exist. What should they do?
Show answer
Answer: B.
A fabricated citation among the sample means the whole set is unreliable and each must be checked. A assumes the rest are fine. C keeps unverified sources. D asks the source that produced the fabrication.
Which TWO checks require no external source to perform?
Show answer
Answer: A and B.
Internal-consistency and relevance checks use only the output itself. C, D and E each require an external authoritative source to complete, so they are not internal checks.
An internal-only memo contains a revenue figure that will drive next year's hiring plan, and a colleague says skip the check because it is internal. What is the best response?
Show answer
Answer: B.
A figure driving hiring is high-stakes regardless of being internal, so it must be verified. A and C wrongly tie checking to audience rather than consequence. D relies on the same fallible source.
An employee wants to summarise a spreadsheet containing customer names and emails. What is the most responsible approach?
Show answer
Answer: B.
Personal data must be handled in a sanctioned tool under policy, with minimisation where possible. A ignores policy and risk. C does not address the AI-use question. D exposes personal data publicly.
A hiring manager wants ChatGPT to decide who to reject from a candidate pool. What is the correct posture?
Show answer
Answer: B.
Consequential decisions about people require human accountability and bias checks; AI can assist but not decide. A, C and D delegate a high-stakes, bias-prone decision to a fallible system without oversight.
Which data class should generally not be pasted into ChatGPT unless explicitly sanctioned?
Show answer
Answer: B.
Confidential or regulated personal data needs an explicitly sanctioned tool and handling. A, C and D are already public or non-sensitive, so they carry little exposure risk.
Which TWO practices protect sensitive data when using ChatGPT?
Show answer
Answer: A and B.
Using a sanctioned workspace and minimising identifiers both reduce exposure. C persists sensitive data where it should not live. D exposes it publicly. E lets urgency override policy, which is exactly the wrong trade-off.
Last updated Sep 18, 2026