AI Cert Prep
Type to search documentation.

AI Foundations

AI Foundations Mock Exam 1

A 60-item, domain-weighted independent mock exam for the OpenAI Academy AI Foundations track, with full explanations and a readiness indicator.

This is a full-length, domain-weighted independent mock exam for the AI Foundations track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment – there is no public OpenAI exam to have questions from. Use it as your diagnostic: sit it first to find your two weakest domains before you enrol in and take the real OpenAI Academy assessment. All 60 items are new and do not repeat the domain-page questions or Mock Exam 2.

Instructions

  • Time: 75 minutes, matching the length of the Academy AI Foundations course experience. You can also take it untimed the first time.
  • Items: 60, single-answer and multiple-response. Each item states how many answers to select.
  • Selection: single-answer items use one choice; multiple-response items say Select two and you must pick exactly two – all correct and none incorrect.
  • No guessing penalty: answer every question; an unanswered item simply scores zero.
  • Target: aim for at least 80% raw (≈ 48/60) before you take the real Academy assessment, because the Academy badge threshold is 80%.

Domain distribution

#DomainItems here
1AI and Generative AI Fundamentals8
2How Language Models Behave10
3ChatGPT Surfaces and Features8
4Prompting and Instructions12
5Context, Files and Memory7
6Verifying and Evaluating Output9
7Responsible and Safe Use6

Total: 8 + 10 + 8 + 12 + 7 + 9 + 6 = 60 items.

Readiness interpretation

This is an independent readiness indicator, not a score and not a pass mark. Read your raw percentage against these bands.

Raw scoreBandWhat it means
under 70%Keep learningRevisit the heaviest domains (Prompting, How Language Models Behave) before retrying.
70–79%Building confidenceClose to ready; target your weakest one or two domains.
80–89%Assessment readyAt or above the Academy badge threshold; review any weak domain and proceed.
90%+Strong readinessConsistent across domains; you are well prepared for the Academy assessment.

The 80% line is deliberate: it matches the OpenAI Academy badge threshold. Remember that an Academy badge or pathway certificate of completion is not a certification.

Take the mock exam

Two ways to use the questions below. The interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction. The review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

60 questions · one at a time · 75-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · AI and Generative AI FundamentalsSelect one

    What is the relationship between artificial intelligence, machine learning and generative AI?

    • A. They are three unrelated fields that occasionally overlap.
    • B. Generative AI is the broadest field; machine learning and AI are subsets of it.
    • C. AI is the broadest field, machine learning is a subset of AI, and generative AI is a subset of machine learning.
    • D. Machine learning and generative AI are the same thing under different names.
    Show answer

    Answer: C.

    AI is the umbrella; machine learning is one approach to AI; generative AI is a category of machine learning focused on producing new content. B inverts the nesting. A denies the well-established containment. D conflates a broad field (ML) with one of its subcategories (generative AI).

  2. Q2D1 · AI and Generative AI FundamentalsSelect one

    A large language model produces its reply by which underlying process?

    • A. Retrieving the closest matching answer from a stored database of prior questions.
    • B. Predicting the next token repeatedly, one piece at a time, based on the preceding text.
    • C. Executing a fixed decision tree written by its engineers.
    • D. Looking up the answer on the live web for every request.
    Show answer

    Answer: B.

    An LLM generates text by predicting the most probable next token given the context, one token at a time. A describes a lookup system, not a generative model. C describes rule-based software, not a neural network. D describes web search, which only happens when a search tool is explicitly enabled.

  3. Q3D1 · AI and Generative AI FundamentalsSelect one

    Roughly how many tokens does 750 words of typical English prose occupy?

    • A. About 75 tokens.
    • B. About 250 tokens.
    • C. About 1,000 tokens.
    • D. About 10,000 tokens.
    Show answer

    Answer: C.

    A common rule of thumb is that a token is about three-quarters of a word, so 1,000 tokens is roughly 750 words. A and B are far too low; D is roughly ten times too high. Knowing the rough ratio helps estimate how much text fits in a context window.

  4. Q4D1 · AI and Generative AI FundamentalsSelect one

    Why can a model give a confidently wrong answer about an event that happened last week?

    • A. The model's training data has a knowledge cutoff, and without a search tool it cannot know recent events.
    • B. The model deliberately withholds recent information for safety.
    • C. Recent events are always too complex for any model.
    • D. The model forgot the event because the chat was too long.
    Show answer

    Answer: A.

    Each model is trained up to a knowledge cutoff date; events after it are unknown unless a search tool is enabled, and the model may still answer fluently. B invents a withholding policy. C is false; recency, not complexity, is the issue. D confuses context limits with the knowledge cutoff.

  5. Q5D1 · AI and Generative AI FundamentalsSelect one

    What does it mean that the latest OpenAI models are multimodal?

    • A. They can run on multiple operating systems.
    • B. They accept more than one type of input, such as text and images, alongside text output.
    • C. They use multiple separate models chained together for every task.
    • D. They can be fine-tuned in multiple languages only.
    Show answer

    Answer: B.

    Multimodal means the model handles more than one modality of input, for example text plus images (vision), while producing text output. A is about deployment, not modality. C describes an orchestration pattern, not what multimodal means. D narrows multimodality to language fine-tuning, which is unrelated.

  6. Q6D1 · AI and Generative AI FundamentalsSelect one

    A user believes that the more their team chats with ChatGPT, the smarter their own copy of the model becomes over time. What is accurate?

    • A. Each conversation permanently retrains the shared model for everyone.
    • B. The base model's weights are fixed at inference; conversations do not retrain it, and enterprise data is not used for training by default.
    • C. The model gets smarter only if you upgrade your plan.
    • D. Chatting more increases the context window automatically.
    Show answer

    Answer: B.

    The model weights are set during training and do not change as you chat; enterprise business data is not used to train models by default. A confuses use with training. C ties intelligence to billing rather than model capability. D confuses model learning with context size, which is plan-dependent and unrelated to usage volume.

  7. Q7D1 · AI and Generative AI FundamentalsSelect one

    Which task is the best fit for a probability-based text generator with no additional tools?

    • A. Reporting today's exact stock price.
    • B. Rewriting a paragraph in a friendlier tone.
    • C. Calculating a large payroll spreadsheet to the cent.
    • D. Confirming a competitor's announcement from this morning.
    Show answer

    Answer: B.

    Tone rewriting is a pure language task that plays to the model's strength. A and D require live data the model cannot know without a search tool. C requires exact arithmetic best done by a calculation tool, not next-token prediction.

  8. Q8D1 · AI and Generative AI FundamentalsSelect two

    Which TWO statements correctly distinguish training from inference?

    • A. Training builds the model's parameters from large datasets; it happens once, ahead of time.
    • B. Inference is when the trained model generates a reply to your prompt.
    • C. Inference permanently updates the model's weights with each reply.
    • D. Training happens live during every conversation you have.
    • E. Inference requires re-downloading the training data each time.
    Show answer

    Answer: A and B.

    Training produces the fixed parameters ahead of time; inference is the run-time generation of a response. C is wrong because inference does not change weights. D is wrong because training is not live. E invents a re-download step that does not exist.

  9. Q9D2 · How Language Models BehaveSelect one

    A user runs the identical prompt twice in two new chats and gets differently worded answers. What does this indicate?

    • A. The model is malfunctioning and should be reported.
    • B. Text generation is non-deterministic, so wording can vary between runs.
    • C. The account has been compromised.
    • D. The knowledge cutoff changed between the two runs.
    Show answer

    Answer: B.

    Generation is probabilistic, so the same prompt can yield differently worded but often equivalent answers. A misreads normal behaviour as a fault. C is unfounded. D is impossible; the cutoff is a fixed property of the model, not something that shifts between two requests.

  10. Q10D2 · How Language Models BehaveSelect one

    Why is a fast, fluent, confident answer not evidence that it is accurate?

    • A. Fluency reflects how probable the wording is, not whether the underlying facts are true.
    • B. Fast answers are always less accurate than slow ones.
    • C. Confidence phrasing is a deliberate accuracy signal built into the model.
    • D. Fluent answers are only produced by reasoning models.
    Show answer

    Answer: A.

    The model optimises for probable, well-formed text; that says nothing about factual truth, so fluent text can be wrong. B is too absolute; speed does not determine accuracy. C is false, there is no built-in truth signal in confident phrasing. D is false; any model can produce fluent text.

  11. Q11D2 · How Language Models BehaveSelect one

    On a ChatGPT Business plan, what is the total context window for a GPT Instant (fast) model versus a GPT Reasoning model?

    • A. 54K for Instant and 256K for Reasoning.
    • B. 128K for Instant and 128K for Reasoning.
    • C. 256K for both.
    • D. 1M for both.
    Show answer

    Answer: A.

    On Business, the GPT Instant total window is 54K and the GPT Reasoning total is 256K. B matches Enterprise Instant, not Business. C is the Reasoning figure applied incorrectly to both. D is an API-scale window, not a ChatGPT plan figure.

  12. Q12D2 · How Language Models BehaveSelect one

    A simple rephrasing task is being run at the highest reasoning effort and is slow and expensive. What is the best adjustment?

    • A. Keep the highest effort to be safe.
    • B. Lower the reasoning effort to the lowest level that still gets the result.
    • C. Switch to a specialised security model.
    • D. Add more documents to the context.
    Show answer

    Answer: B.

    The guidance is to use the lowest reasoning effort that produces the needed result; a rephrasing task does not need deep reasoning. A wastes time and money. C uses an inappropriate specialised model. D adds cost and clutter without helping a simple language task.

  13. Q13D2 · How Language Models BehaveSelect one

    A multi-step logic problem gets a fast, confident answer that skips steps and is wrong. What is the most appropriate fix?

    • A. Re-send the identical prompt several times until it is right.
    • B. Switch to a reasoning model or ask the model to reason step by step.
    • C. Shorten the prompt to save tokens.
    • D. Increase the temperature setting.
    Show answer

    Answer: B.

    Multi-step logic benefits from a reasoning model or explicit step-by-step working. A repeats the same shortcoming. C removes needed detail. D would add randomness, not rigour, and is not exposed for chat use in the way implied.

  14. Q14D2 · How Language Models BehaveSelect one

    A user prompts the model to explain why their new slogan is obviously the best one and receives glowing agreement. What behaviour is this and the best counter?

    • A. Hallucination; add a citation requirement.
    • B. Sycophancy; ask for a balanced critique including weaknesses and counterarguments.
    • C. Non-determinism; run the prompt again.
    • D. Context overflow; start a new chat.
    Show answer

    Answer: B.

    The model tends to agree with the framing it is given; asking for weaknesses and counterarguments counters this sycophancy. A addresses fabricated facts, not agreeableness. C addresses wording variance. D addresses long-conversation drift, none of which is the issue here.

  15. Q15D2 · How Language Models BehaveSelect one

    In a very long chat, the model appears to forget details stated at the start. What is the most likely cause?

    • A. The earliest turns have fallen outside the effective context window.
    • B. The model deleted the information to protect privacy.
    • C. The knowledge cutoff removed it.
    • D. The account reached its monthly message limit.
    Show answer

    Answer: A.

    As a conversation grows, older turns can drop out of the working context, so the model no longer sees them. B invents a deletion behaviour. C confuses training cutoff with in-chat context. D is a billing limit, unrelated to forgetting mid-conversation.

  16. Q16D2 · How Language Models BehaveSelect one

    ChatGPT sums a 25-row expense table and returns a total. What is the most reliable way to trust the number?

    • A. Accept it because the model is good at arithmetic.
    • B. Have it computed with a data-analysis tool or verify it in a spreadsheet.
    • C. Ask the model whether it is confident in the total.
    • D. Re-ask in a friendlier tone.
    Show answer

    Answer: B.

    Exact arithmetic should be computed with a tool or checked externally, not trusted to next-token prediction. A over-trusts the model. C relies on self-reported confidence, which is not an accuracy signal. D changes tone, not correctness.

  17. Q17D2 · How Language Models BehaveSelect one

    A user pastes eight long documents into one chat and quality drops even though the total is under the window limit. What is the most likely explanation?

    • A. The window limit was exceeded after all.
    • B. Too much loosely related material dilutes attention and buries the relevant parts.
    • C. The model cannot read more than one document.
    • D. Documents must be under 1,000 words each.
    Show answer

    Answer: B.

    Even within the limit, large amounts of loosely related context can dilute the model's attention on what matters. A contradicts the premise that the total was under the limit. C and D invent hard rules that do not exist; the issue is relevance and volume, not a per-document cap.

  18. Q18D2 · How Language Models BehaveSelect two

    Which TWO behaviours follow directly from the model predicting probable text rather than looking up facts?

    • A. It can fabricate a plausible but non-existent citation.
    • B. It can state a wrong figure with full fluency and confidence.
    • C. It always refuses when unsure.
    • D. It guarantees identical wording on every run.
    • E. It can browse the live web without any tool.
    Show answer

    Answer: A and B.

    Because it generates probable text, it can invent plausible citations and state wrong numbers confidently. C is false; models often answer rather than refuse when uncertain. D is false; generation is non-deterministic. E is false; browsing requires an explicit search tool.

  19. Q19D7 · Responsible and Safe UseSelect one

    A team wants to summarise anonymised aggregate figures rather than raw records containing customer identifiers. Why is this the more responsible choice?

    • A. Aggregated, anonymised data reduces exposure of personal information while still serving the task.
    • B. Aggregation makes the model faster.
    • C. Raw records are always rejected by the tool.
    • D. Anonymised data cannot ever be wrong.
    Show answer

    Answer: A.

    Working from anonymised aggregates minimises personal-data exposure while meeting the need, which is good data hygiene. B confuses privacy with speed. C invents an automatic rejection that does not exist. D confuses privacy with accuracy, which are unrelated.

  20. Q20D3 · ChatGPT Surfaces and FeaturesSelect one

    A high-volume task classifies thousands of short support tickets into fixed categories. Which model is most cost-effective?

    • A. GPT-6 Pro for maximum quality.
    • B. GPT-5.6 Luna, suited to clear, repeatable, high-volume tasks.
    • C. A specialised life-sciences model.
    • D. A realtime voice model.
    Show answer

    Answer: B.

    Luna is designed for clear, repeatable, high-volume work such as classification, making it the cost-effective fit. A overspends on a simple task. C and D are specialised for domains and modalities unrelated to text classification.

  21. Q21D3 · ChatGPT Surfaces and FeaturesSelect one

    A user needs a five-page, well-sourced briefing that synthesises many articles on a regulatory topic. Which feature fits best?

    • A. A single fast reply from GPT Instant.
    • B. Deep research.
    • C. Image generation.
    • D. Voice mode.
    Show answer

    Answer: B.

    Deep research is built to gather and synthesise many sources into a longer, cited briefing. A cannot browse or synthesise at that depth. C and D are for images and speech, not multi-source research.

  22. Q22D3 · ChatGPT Surfaces and FeaturesSelect one

    A team reuses the same brand guidelines and tone-of-voice instructions across dozens of chats each week. Where should this live?

    • A. Pasted into each new chat by hand.
    • B. In a Project with custom instructions and shared files.
    • C. In the model's training data.
    • D. In a scheduled task.
    Show answer

    Answer: B.

    A Project holds custom instructions and files that persist across its chats, ideal for reused brand context. A is error-prone and wasteful. C is impossible; you cannot add to training data. D schedules recurring runs, not standing context.

  23. Q23D3 · ChatGPT Surfaces and FeaturesSelect one

    A user must total and chart a 400-row spreadsheet and needs the numbers to be exactly correct. Which feature should they use?

    • A. Ask the model to eyeball the totals.
    • B. Data analysis, which computes results with code rather than predicting them.
    • C. Memory.
    • D. Study mode.
    Show answer

    Answer: B.

    Data analysis runs actual code to compute totals and charts, giving reliable numbers. A relies on next-token estimation, which is unreliable for arithmetic. C stores preferences. D is a tutoring feature, neither of which computes spreadsheet totals.

  24. Q24D3 · ChatGPT Surfaces and FeaturesSelect one

    An employee needs an answer that exists only in the company's connected internal document store. Which feature is correct?

    • A. Web search.
    • B. Company knowledge, which draws on connected internal sources.
    • C. Image generation.
    • D. The model's training data.
    Show answer

    Answer: B.

    Company knowledge connects to internal sources so answers can draw on them. A searches the public web, not internal documents. C produces images. D cannot contain private company documents added after training.

  25. Q25D3 · ChatGPT Surfaces and FeaturesSelect one

    A user is repeatedly revising a single proposal document and wants to edit it inline rather than copy-paste between chat and a document. Which surface is best?

    • A. Canvas.
    • B. Voice mode.
    • C. A scheduled task.
    • D. Deep research.
    Show answer

    Answer: A.

    Canvas provides a side-by-side editable document surface for iterative drafting. B is speech-oriented. C automates recurring runs. D gathers sources; none supports inline document editing the way canvas does.

  26. Q26D3 · ChatGPT Surfaces and FeaturesSelect one

    A user wants a reusable, shareable assistant preconfigured for a repeated team task. Which surface fits?

    • A. A one-off chat.
    • B. A custom GPT.
    • C. Memory.
    • D. Search.
    Show answer

    Answer: B.

    A custom GPT packages instructions and behaviour into a reusable, shareable assistant. A does not persist or share. C stores small preferences, not a configured assistant. D retrieves web results and is not a reusable assistant.

  27. Q27D3 · ChatGPT Surfaces and FeaturesSelect two

    Which TWO tasks are correctly matched to a ChatGPT surface or feature?

    • A. Checking a price announced this morning uses Search.
    • B. Iteratively editing one long document inline uses canvas.
    • C. Guaranteeing exact spreadsheet totals uses memory.
    • D. Learning today's news headlines uses image generation.
    • E. Storing a 30-page report for one summary uses a scheduled task.
    Show answer

    Answer: A and B.

    Search handles fresh web facts and canvas supports inline document editing. C is wrong; exact totals need data analysis, not memory. D is wrong; image generation makes pictures, not news. E is wrong; a scheduled task automates runs, it does not hold a document for summarising.

  28. Q28D4 · Prompting and InstructionsSelect one

    A prompt reads write something about our new pricing and the output is generic. Which prompt components are most likely missing?

    • A. Nothing; the prompt is complete.
    • B. Audience, context and specific constraints such as format and length.
    • C. A higher reasoning effort.
    • D. A larger context window.
    Show answer

    Answer: B.

    The prompt lacks audience, context and constraints, so the model has nothing to make the output specific. A ignores the obvious gaps. C and D address compute and capacity, not the missing instructions that cause generic output.

  29. Q29D4 · Prompting and InstructionsSelect one

    Which is the best statement of a success criterion for an executive summary?

    • A. Make it good.
    • B. One page, three key risks named, each with a recommended action, no jargon.
    • C. Write professionally.
    • D. Be thorough and detailed.
    Show answer

    Answer: B.

    A good success criterion is specific and checkable: length, required content and style. A, C and D are vague and cannot be verified, so the model cannot reliably meet them or know when it has.

  30. Q30D4 · Prompting and InstructionsSelect one

    A first output is close but too formal. What is the disciplined next step?

    • A. Discard everything and start a brand-new prompt from scratch.
    • B. Iterate: keep what works and give one targeted instruction to relax the tone.
    • C. Send the identical prompt again.
    • D. Switch to a specialised model.
    Show answer

    Answer: B.

    Disciplined iteration makes one targeted change rather than discarding a near-miss. A throws away progress. C repeats the same result. D changes the tool when only a small tone adjustment is needed.

  31. Q31D4 · Prompting and InstructionsSelect one

    A task requires researching a market, choosing a strategy, and drafting both a deck outline and a press release. What is the best prompting approach?

    • A. One giant prompt asking for everything at once.
    • B. Break it into sequential steps, reviewing each before moving on.
    • C. Ask the model to skip the research.
    • D. Increase the temperature.
    Show answer

    Answer: B.

    Multi-stage work is best decomposed into reviewable steps so errors are caught early. A overloads one prompt and makes review hard. C removes a necessary step. D changes randomness, not structure.

  32. Q32D4 · Prompting and InstructionsSelect one

    A user wants summaries in a consistent two-bullet style across many documents. What is the most reliable way to get it?

    • A. Hope the model remembers the style.
    • B. State the exact format and show a short example of the desired output.
    • C. Use a longer prompt with more adjectives.
    • D. Raise the reasoning effort.
    Show answer

    Answer: B.

    Specifying the format and giving an example makes the output shape reliable and repeatable. A leaves it to chance. C adds words without structure. D adds compute without defining the format.

  33. Q33D4 · Prompting and InstructionsSelect one

    A user keeps saying what they do not want and the output still misses. What is the best fix?

    • A. Add more negative instructions.
    • B. State positively what the output should be, with concrete criteria.
    • C. Repeat the prompt louder in capitals.
    • D. Switch accounts.
    Show answer

    Answer: B.

    Positive, concrete instructions steer far better than a list of prohibitions. A piles on negatives without telling the model the target. C changes formatting, not substance. D is irrelevant to prompt quality.

  34. Q34D4 · Prompting and InstructionsSelect one

    Which output-shape request is most appropriate when the result will be imported into another system?

    • A. A free-form paragraph.
    • B. A structured format such as a defined table or JSON with named fields.
    • C. A poem.
    • D. A voice recording.
    Show answer

    Answer: B.

    Downstream systems need predictable structure, so a defined table or JSON with named fields is best. A is hard to parse reliably. C and D are unsuitable formats for machine import.

  35. Q35D4 · Prompting and InstructionsSelect one

    A colleague believes a longer prompt is automatically a better prompt. What is the most accurate correction?

    • A. Length always helps; add as much as possible.
    • B. Clarity and relevant context matter more than length; irrelevant text can hurt.
    • C. Prompts must be under 50 words.
    • D. Length only matters for reasoning models.
    Show answer

    Answer: B.

    What helps is relevant, clear context; padding can dilute focus and reduce quality. A is false. C invents an arbitrary cap. D wrongly ties the principle to one model class.

  36. Q36D4 · Prompting and InstructionsSelect one

    A prompt buries its most important constraint in the last sentence and the model ignores it. What is the best fix?

    • A. Delete the constraint.
    • B. Move the critical constraint to a prominent position and label it clearly.
    • C. Add three more constraints.
    • D. Lower the reasoning effort.
    Show answer

    Answer: B.

    Placing and labelling the key constraint prominently makes it more likely to be honoured. A abandons the requirement. C buries it further among more instructions. D reduces compute, which does not address placement.

  37. Q37D4 · Prompting and InstructionsSelect one

    After three identical re-sends of a vague prompt, a user concludes the model cannot do the task. What is the accurate diagnosis?

    • A. The model is incapable of the task.
    • B. The prompt lacks the specificity the model needs; refine it rather than repeat it.
    • C. The account is rate-limited.
    • D. The context window is too small.
    Show answer

    Answer: B.

    Repeating a vague prompt cannot fix vagueness; adding specifics usually can. A blames capability prematurely. C and D invent limits the scenario does not support; the root cause is prompt quality.

  38. Q38D4 · Prompting and InstructionsSelect one

    A marketing prompt produces on-topic but flat copy. Adding which single component would most likely lift the tone to match the brand?

    • A. A tone-and-voice instruction with an example of the desired style.
    • B. A larger context window.
    • C. A higher reasoning effort.
    • D. More documents.
    Show answer

    Answer: A.

    Tone is controlled by an explicit voice instruction and an example, which directly addresses flat copy. B, C and D add capacity, compute or material but none specifies the tone the brand needs.

  39. Q39D4 · Prompting and InstructionsSelect one

    Which prompt component tells the model the shape of the output, such as a table with named columns?

    • A. The role.
    • B. The format specification.
    • C. The audience.
    • D. The success criterion.
    Show answer

    Answer: B.

    The format specification defines the output shape, such as a table with named columns. The role sets a persona, the audience sets who it is for, and the success criterion sets when it is done; none of those defines shape.

  40. Q40D6 · Verifying and Evaluating OutputSelect two

    Which TWO signals in an AI output should raise suspicion and prompt verification before you act on it?

    • A. A precise statistic offered with no source.
    • B. Subtotals that do not add up to the stated headline total.
    • C. The output being written in fluent, confident prose.
    • D. The answer directly addressing the question asked.
    • E. The response arriving quickly.
    Show answer

    Answer: A and B.

    Unsourced precise figures and internal-consistency failures are genuine red flags. C is not a warning sign, since fluency is expected and unrelated to accuracy. D is a good sign, not a concern. E, speed, says nothing about correctness.

  41. Q41D7 · Responsible and Safe UseSelect two

    Which TWO situations most clearly call for keeping a human accountable rather than letting the model decide?

    • A. Deciding which employees to lay off.
    • B. Approving a public regulatory filing.
    • C. Suggesting synonyms for a headline.
    • D. Brainstorming names for an internal project.
    • E. Rewording a sentence to be clearer.
    Show answer

    Answer: A and B.

    Layoff decisions and public regulatory filings are consequential and require human accountability. C, D and E are low-stakes, easily reversible language tasks that do not need a decision gate, so they are the weaker choices.

  42. Q42D5 · Context, Files and MemorySelect one

    A team pastes the same brand-voice guidelines into every new chat. Where should these live instead?

    • A. In a Project's custom instructions so they apply across its chats.
    • B. In a single very long chat kept open forever.
    • C. In the model's training data.
    • D. Nowhere; keep pasting them.
    Show answer

    Answer: A.

    Project custom instructions persist across the Project's chats, removing repetitive pasting. B degrades as the chat grows. C is impossible for users. D keeps the inefficient status quo.

  43. Q43D5 · Context, Files and MemorySelect one

    Which item is the best candidate for memory?

    • A. A one-off instruction for a single email.
    • B. A stable standing preference, such as always reply in British English.
    • C. A 40-page report needed for one summary.
    • D. A confidential customer list.
    Show answer

    Answer: B.

    Memory suits small, stable, recurring preferences. A is a single-use instruction that belongs in the prompt. C belongs in a file upload. D is sensitive data that should not be stored in memory.

  44. Q44D5 · Context, Files and MemorySelect one

    A user needs the model to summarise one specific 20-page report. Where should the report go?

    • A. Into memory.
    • B. Uploaded as a file for that chat or Project.
    • C. Into custom instructions.
    • D. Into a scheduled task.
    Show answer

    Answer: B.

    A specific document to summarise belongs in a file upload the model can read. A stores small facts, not documents. C is for standing behaviour. D automates timing, not document ingestion.

  45. Q45D5 · Context, Files and MemorySelect one

    Why can attaching many loosely related files degrade answers even below the context limit?

    • A. Files are never read by the model.
    • B. Irrelevant material competes for attention and can bury the parts that matter.
    • C. The model can only read one file ever.
    • D. Attachments disable memory.
    Show answer

    Answer: B.

    Loosely related material dilutes focus, so relevance matters even under the limit. A is false; files are used. C invents a one-file rule. D asserts an unrelated side effect that does not occur.

  46. Q46D5 · Context, Files and MemorySelect one

    A user manages three clients in one long chat and the model mixes up their preferences. What is the best structural fix?

    • A. Keep going in the same chat but ask more carefully.
    • B. Use a separate Project or chat per client to keep contexts clean.
    • C. Store all three clients in memory together.
    • D. Increase the reasoning effort.
    Show answer

    Answer: B.

    Separating clients into distinct Projects or chats prevents cross-contamination of context. A leaves the mixed context in place. C compounds the mixing in memory. D adds compute without fixing the structural overlap.

  47. Q47D5 · Context, Files and MemorySelect two

    A one-off instruction applies only to the single email you are writing now. Which TWO statements about where it belongs are correct?

    • A. It belongs in the current prompt for this email.
    • B. It should not go into memory, which is for stable, recurring preferences.
    • C. It belongs in a Project's custom instructions so it applies to every chat.
    • D. It should be saved as a scheduled task.
    • E. It should be added to company knowledge.
    Show answer

    Answer: A and B.

    A single-use instruction goes in the immediate prompt and should be kept out of memory. C would wrongly persist a one-off across chats. D automates recurring runs, which does not apply. E is for shared internal reference sources, not a one-off email instruction.

  48. Q48D5 · Context, Files and MemorySelect two

    Which TWO are good context-hygiene practices?

    • A. Start a fresh chat for a new topic rather than extending a marathon thread.
    • B. Attach only the files relevant to the current task.
    • C. Keep every document from every project attached at all times.
    • D. Store sensitive customer data in memory for convenience.
    • E. Paste all standing instructions into every prompt manually.
    Show answer

    Answer: A and B.

    Fresh chats per topic and attaching only relevant files keep context clean. C clutters the window. D stores sensitive data inappropriately. E is inefficient when Projects or memory can hold standing context.

  49. Q49D6 · Verifying and Evaluating OutputSelect one

    ChatGPT returns five precise, unsourced market statistics for a niche segment you will cite externally. What is the best next step?

    • A. Publish them; the numbers look precise.
    • B. Verify each figure against a primary source before using it.
    • C. Ask the model if it is sure.
    • D. Round the numbers to look less specific.
    Show answer

    Answer: B.

    Unsourced, specific statistics for external use must be verified against primary sources. A risks publishing fabrications. C relies on self-reported confidence. D disguises the problem without confirming truth.

  50. Q50D6 · Verifying and Evaluating OutputSelect one

    Which check is free and should usually come first on an output you will act on?

    • A. Commissioning an external audit.
    • B. A quick plausibility and internal-consistency read of the output.
    • C. Running a paid data-vendor comparison.
    • D. Asking a second AI tool.
    Show answer

    Answer: B.

    A fast plausibility and consistency read costs nothing and catches many errors first. A and C are heavier, costlier steps for later. D adds another fallible model rather than a grounded check.

  51. Q51D6 · Verifying and Evaluating OutputSelect one

    A citation has a plausible title, real-looking authors and a real journal name. How do you validate it?

    • A. Accept it; the details look real.
    • B. Locate the actual source and confirm it exists and says what is claimed.
    • C. Ask the model to confirm the citation.
    • D. Check that the formatting is correct.
    Show answer

    Answer: B.

    Only finding the real source confirms a citation is genuine and supports the claim. A trusts plausible-looking fabrication. C asks the same fallible source. D checks style, not existence or content.

  52. Q52D6 · Verifying and Evaluating OutputSelect one

    Two colleagues verify a high-stakes figure by asking two different AI tools; both agree, so they accept it. What is the flaw?

    • A. There is no flaw; agreement proves accuracy.
    • B. Both are probabilistic generators and can share the same error; agreement is not grounding.
    • C. They should have asked three tools.
    • D. They should have used the same tool twice.
    Show answer

    Answer: B.

    Two models can be confidently wrong in the same way; agreement between generators is not verification against a source. A is false. C just adds more fallible generators. D reduces independence further.

  53. Q53D6 · Verifying and Evaluating OutputSelect one

    Which output most clearly requires a mandatory human review gate before acting?

    • A. A private brainstorming list for your own use.
    • B. A customer-facing legal notice to be published.
    • C. A draft title for a personal blog post.
    • D. A list of synonyms.
    Show answer

    Answer: B.

    A published, customer-facing legal notice is high-stakes and irreversible, so it needs human review. A, C and D are low-stakes and easily corrected, so they do not require a formal gate.

  54. Q54D6 · Verifying and Evaluating OutputSelect one

    A research summary lists eight citations; the analyst opens two and one does not exist. What should they do?

    • A. Trust the remaining six by default.
    • B. Treat all citations as suspect and verify every one before use.
    • C. Remove only the missing citation and publish the rest.
    • D. Ask the model to fix the citations.
    Show answer

    Answer: B.

    A fabricated citation among the sample means the whole set is unreliable and each must be checked. A assumes the rest are fine. C keeps unverified sources. D asks the source that produced the fabrication.

  55. Q55D6 · Verifying and Evaluating OutputSelect two

    Which TWO checks require no external source to perform?

    • A. Confirming the output's internal figures add up to its own stated total.
    • B. Checking the answer actually addresses the question asked.
    • C. Verifying a quoted statute against the official register.
    • D. Confirming a citation exists in a library database.
    • E. Comparing a market size to an industry report.
    Show answer

    Answer: A and B.

    Internal-consistency and relevance checks use only the output itself. C, D and E each require an external authoritative source to complete, so they are not internal checks.

  56. Q56D6 · Verifying and Evaluating OutputSelect one

    An internal-only memo contains a revenue figure that will drive next year's hiring plan, and a colleague says skip the check because it is internal. What is the best response?

    • A. Agree; internal outputs do not need checking.
    • B. Verify the figure, because stakes, not audience, determine the level of checking.
    • C. Check only if it will be published.
    • D. Ask the model to double-check itself.
    Show answer

    Answer: B.

    A figure driving hiring is high-stakes regardless of being internal, so it must be verified. A and C wrongly tie checking to audience rather than consequence. D relies on the same fallible source.

  57. Q57D7 · Responsible and Safe UseSelect one

    An employee wants to summarise a spreadsheet containing customer names and emails. What is the most responsible approach?

    • A. Paste it into any available account to move fast.
    • B. Use only a sanctioned workspace and follow data-handling policy, minimising or removing personal data where possible.
    • C. Email it to a colleague first.
    • D. Post it in a public GPT.
    Show answer

    Answer: B.

    Personal data must be handled in a sanctioned tool under policy, with minimisation where possible. A ignores policy and risk. C does not address the AI-use question. D exposes personal data publicly.

  58. Q58D7 · Responsible and Safe UseSelect one

    A hiring manager wants ChatGPT to decide who to reject from a candidate pool. What is the correct posture?

    • A. Let the model make the final rejection decisions to save time.
    • B. Use the model to assist analysis, but keep a human accountable for consequential decisions and check for bias.
    • C. Trust the model because it is neutral.
    • D. Automate rejections with no review.
    Show answer

    Answer: B.

    Consequential decisions about people require human accountability and bias checks; AI can assist but not decide. A, C and D delegate a high-stakes, bias-prone decision to a fallible system without oversight.

  59. Q59D7 · Responsible and Safe UseSelect one

    Which data class should generally not be pasted into ChatGPT unless explicitly sanctioned?

    • A. A published press release.
    • B. Confidential or regulated personal data.
    • C. A public product FAQ.
    • D. A blog draft you wrote.
    Show answer

    Answer: B.

    Confidential or regulated personal data needs an explicitly sanctioned tool and handling. A, C and D are already public or non-sensitive, so they carry little exposure risk.

  60. Q60D7 · Responsible and Safe UseSelect two

    Which TWO practices protect sensitive data when using ChatGPT?

    • A. Use only the workspace sanctioned for that data class.
    • B. Remove or mask personal identifiers you do not need before pasting.
    • C. Store confidential data in memory for reuse.
    • D. Share sensitive chats in a public GPT.
    • E. Assume any account is fine if the task is urgent.
    Show answer

    Answer: A and B.

    Using a sanctioned workspace and minimising identifiers both reduce exposure. C persists sensitive data where it should not live. D exposes it publicly. E lets urgency override policy, which is exactly the wrong trade-off.

Last updated Sep 18, 2026