AI Cert Prep
Type to search documentation.

AI Foundations

D1 · AI and Generative AI Fundamentals

What AI, machine learning, generative AI and large language models are, how tokens and probability-based generation work, the difference between training and inference, knowledge cutoffs and multimodality.

This domain is worth 14% of the mock – roughly 8 of 60 items. It is the vocabulary layer for everything else. It tests whether you can place ChatGPT correctly in the family tree of AI, explain in plain terms why it produces text the way it does, and reason from that mechanism to its strengths and its failure modes. Almost every later domain is an application of one idea introduced here: the model predicts likely text, it does not look facts up.

What you need to know

Artificial intelligence is the broad field of getting machines to do things that normally need human intelligence. Machine learning is the subset that learns patterns from data instead of following hand-written rules. Generative AI is the slice of machine learning that produces new content – text, images, audio, code. A large language model (LLM) is a generative model trained on enormous amounts of text to predict the next chunk of text. ChatGPT is a product built on top of OpenAI’s LLMs. Understanding that stack, and that the model generates by predicting probable continuations one token at a time from what it learned during training, explains its fluency, its speed, its knowledge cutoff and its tendency to sound confident even when wrong.

Learning objectives

By the end of this page you should be able to:

  1. Distinguish AI, machine learning, generative AI and LLMs, and place ChatGPT within them.
  2. Explain tokens and why the model works in tokens rather than words or letters.
  3. Describe probability-based generation and why it makes the model fluent but not factual.
  4. Separate training from inference and explain what a knowledge cutoff is and is not.
  5. Explain multimodality – what today’s models can take in and produce.
  6. Reason from the generation mechanism to a concrete strength or limitation.

1.1 The family tree: AI, ML, generative AI, LLM

These four terms are nested, not synonyms. The single most common D1 item asks you to order or relate them.

text
┌─────────────────────────────────────────────────────────┐
│ Artificial Intelligence │
│ machines doing tasks that need human-like intelligence │
│ ┌──────────────────────────────────────────────────┐ │
│ │ Machine Learning │ │
│ │ systems that learn patterns from data │ │
│ │ ┌────────────────────────────────────────────┐ │ │
│ │ │ Generative AI │ │ │
│ │ │ creates new content (text, image, audio) │ │ │
│ │ │ ┌──────────────────────────────────────┐ │ │ │
│ │ │ │ Large Language Models (LLMs) │ │ │ │
│ │ │ │ predict likely text; ChatGPT is a │ │ │ │
│ │ │ │ product built on OpenAI's LLMs │ │ │ │
│ │ │ └──────────────────────────────────────┘ │ │ │
│ │ └────────────────────────────────────────────┘ │ │
│ └──────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────┘
TermOne-line definitionEveryday example that is not generative
Artificial intelligenceMachines performing tasks that need human-like intelligenceA chess engine; a spam filter rule set
Machine learningAI that learns patterns from data rather than fixed rulesA fraud-detection model scoring transactions
Generative AIML that produces new contentAn image generator; an LLM
Large language modelA generative model trained to predict likely textThe model behind ChatGPT

Assessment signal

Stems that say “which term is the broadest”, “ChatGPT is an example of”, or “generative AI differs from traditional machine learning because” are testing this nesting. The discriminator: traditional ML classifies or predicts a label; generative AI produces new content.

1.2 Tokens: the units the model actually reads and writes

A model does not process words or characters directly. Text is broken into tokens – common chunks that are often a word, sometimes a word-piece, sometimes punctuation. A rough English rule of thumb is one token ≈ 4 characters ≈ ¾ of a word, so 1,000 tokens is about 750 words.

text
"Generation is probabilistic."
▼ tokenizer
["Generation", " is", " probabil", "istic", "."]
one token one word-piece word-pc punct

Why this matters in practice:

Consequence of tokenisationWhat you observe
The model sees chunks, not lettersIt miscounts letters (“how many r’s in strawberry”) and struggles with anagrams
Cost and limits are measured in tokensContext windows and API prices are quoted per token, not per word
Rare words split into many tokensUnusual names or code identifiers cost more and are handled less reliably
Whitespace and case are part of the token” the” and “The” are different tokens

Assessment signal

“Why does the model get letter-counting or spelling-manipulation tasks wrong?” points here: it reads tokens, not characters. “How is usage measured/priced?” also points here: tokens, not words.

1.3 Probability-based generation

An LLM generates text one token at a time. At each step it produces a probability distribution over all possible next tokens and picks from it, then feeds the result back in and repeats. It is, in effect, a very sophisticated next-token predictor.

text
prompt ─► [model] ─► distribution over next tokens
"Paris" 0.71 · "the" 0.09 · "France" 0.04 · …
pick ─► append ─► feed back ─► repeat

Two facts fall straight out of this mechanism and drive most of Domain 2:

  • Fluency is a property of the mechanism, not a truth signal. The model is optimised to produce probable text, which usually reads well, regardless of whether the content is correct.
  • The same prompt can give different wording because sampling introduces variation (see determinism in D2).

Predicting likely continuations is exactly right for drafting, rephrasing, summarising, translating, brainstorming, explaining and coding-by-pattern – tasks where a fluent, plausible continuation is the goal.

1.4 Training versus inference

These are two entirely separate phases, and confusing them causes real mistakes (like believing your chat “teaches” the base model).

PhaseWhenWhat happensDoes it change the model?
TrainingBefore release, at OpenAI, on huge datasetsThe model’s parameters are learned from dataYes – this is where the weights are set
InferenceEvery time you send a promptThe trained model predicts tokens for your inputNo – the weights are fixed; nothing you type retrains the base model

Consequences to remember:

  • Your prompts and uploaded files inform the model only for that conversation (plus memory/features you enable). They do not update the underlying model for everyone.
  • Whether your content is used to improve future models is a data-control and plan setting, separate from inference (Enterprise does not train on business data by default). Do not conflate “the model used my message this turn” with “the model was retrained on my message”.
text
TRAINING (once, before you ever see it) INFERENCE (every prompt)
huge text corpus ─► learn parameters your prompt ─► fixed model ─► tokens
weights are set and frozen weights never change here

1.5 Knowledge cutoff

Because knowledge is baked in during training, each model has a knowledge cutoff: the point after which it has no built-in information about the world. Ask about an event after the cutoff and, without a tool, the model will either say it does not know or – worse – guess fluently.

Model (September 2026)Knowledge cutoff
GPT-6 Astra (gpt-6-astra)30 April 2026
GPT-5.6 Sol / Terra / Luna16 February 2026

How to work with a cutoff:

  • For anything time-sensitive (prices, news, releases, “latest”), use a tool that reaches current information – Search or deep research in ChatGPT – rather than the model’s memory.
  • Treat unsourced claims about recent events as suspect by default.
  • A later cutoff is not “smarter”, only more recently informed.

Assessment signal

“The model gave outdated information / did not know about a recent event” points to the knowledge cutoff, and the fix is a retrieval tool (Search / deep research), not a bigger or more expensive model.

1.6 Multimodality

Modern OpenAI models are multimodal: they can take in more than one type of content. In September 2026 the latest models accept text and image input and produce text output, and are multilingual with vision. In ChatGPT you also get image generation, voice, and file/data handling as product features layered on top.

CapabilityWhat it means in ChatGPTEveryday use
Text in, text outThe core chat interactionDrafting, Q&A, analysis
Image input (vision)Upload a photo, screenshot or diagram and ask about it“What does this error screen say?”
Image generationProduce an image from a descriptionA slide illustration, a concept sketch
VoiceSpeak to ChatGPT and hear replies; voice with videoHands-free brainstorming
Files and dataUpload documents/spreadsheets; run data analysisSummarise a PDF; chart a CSV

The exam-relevant point: “multimodal” describes the kinds of input and output, and knowing that today’s latest models take text + image input and return text tells you when you need a specialised feature (image generation, voice) rather than the base chat.

Decision framework

Use the AI-fit triage to decide, in seconds, whether a task suits a probability-based text generator at all before you even open a chat.

AskIf yesIf no
Is a fluent, plausible result the goal (draft, summary, rephrase, explain, ideate)?Strong fit for an LLMReconsider – you may need a calculator, database or search
Does it need current information past the cutoff?Add a retrieval tool (Search / deep research)Base model knowledge is fine
Does it need exact arithmetic, counting or precise data lookup?Use a tool (data analysis / code) or verifyThe model alone is acceptable
Would a plausible-but-wrong answer cause harm?Verify before use; add a review gateLower verification is acceptable

The framework’s value: it stops you from asking a next-token predictor to be a database or a calculator, which is the root cause of most beginner disappointment.

Common mistakes

MistakeWhy it happensWhat to do instead
Treating ChatGPT as a search engineIt answers in the same conversational shape as searchUse Search/deep research for current facts; treat model knowledge as a lead
Thinking “AI”, “ML” and “generative AI” are the same thingThe terms are used loosely in the pressKeep the nesting: generative AI creates content; older ML classifies or predicts labels
Believing your chats retrain the model“It learned from me” feels intuitiveInference is fixed-weight; training is separate; data use is a plan/setting
Expecting exact letter counts or arithmeticThe output is fluent, so it looks authoritativeRemember it reads tokens, not characters; use a tool for exact work
Assuming a later model “knows everything now”Newer sounds betterEvery model has a cutoff; time-sensitive facts need a tool
Confusing multimodal with “does anything”The word sounds all-encompassingMultimodal = specific input/output types; check the actual capability
Thinking higher price means more accuracyCost correlates with capability, not truthVerification, not spend, is what makes output trustworthy

Scenario challenge

Scenario. Priya, a new analyst, is delighted with ChatGPT after it drafted a crisp market summary. Encouraged, she asks it for “the exact number of stores our competitor opened last quarter” and “how many times the letter s appears in our tagline”, then pastes both answers into a board slide. She also tells a colleague, “the more we use it, the smarter our copy of it gets”. The competitor number looks precise (it cites “47 new stores”); the letter count is off by two; and she is about to present both.

Expert reasoning trace.

  1. Locate each task on the AI-fit triage. The market summary is a fluent-continuation task – a strong fit, done well. The exact competitor count needs current, specific data past nothing the model was trained to know reliably; that is a factual-lookup task, not a generation task. The letter count is a character-level operation the model performs poorly because it reads tokens, not letters.
  2. Diagnose the competitor figure. A precise unsourced number on a niche, recent topic is the hallucination signature. The fix is a retrieval tool (Search / deep research) or an authoritative source, then verification – not trusting the fluent “47”.
  3. Diagnose the letter count. Tokenisation explains the error directly; the reliable fix is a tool that counts characters, or doing it by hand.
  4. Correct the “it gets smarter as we use it” belief. Inference does not retrain the model; her chats do not improve the base model, and whether her data trains future models is a separate data-control setting she should check, not assume.
  5. Set the review gate. Both figures are going on an external-facing board slide, so they need verification before use regardless of how confident the text sounds.

Exam-correct outcome: keep the summary, replace the competitor number with a verified sourced figure (via a retrieval tool or authoritative data), recompute the letter count with a tool or by hand, and drop the mistaken belief that usage retrains the model.

Assessment traps

TrapWhy it is temptingThe discriminator
“ChatGPT looks things up when you ask”It answers like a search boxIt predicts likely text from training; it only looks up when a retrieval tool is invoked
“Generative AI and machine learning are the same”Both are “AI”Generative AI creates content; it is a subset of ML
“My prompts train the model”“It learned from me” feels rightTraining is a separate, pre-release phase; inference is fixed-weight
“A newer model has no knowledge cutoff”Newer implies up-to-dateEvery model has a cutoff; only tools reach beyond it
“It counts letters accurately because it reads text”It clearly reads textIt reads tokens, not characters, so character-level tasks fail
“Multimodal means it can do any task”The word sounds totalIt names specific input/output types (e.g. text + image in, text out)
“The pricier model is more truthful”Cost feels like qualityPrice tracks capability/latency, not factual guarantees; verification does that

Practice questions

Q1 · Which term is the BROADEST of the following? (Select one)

A. Large language model B. Generative AI C. Artificial intelligence D. Machine learning

Answer: C. Artificial intelligence is the outermost category; machine learning is a subset of it, generative AI a subset of ML, and an LLM a specific kind of generative model. B, D and A are all nested inside AI, so none is the broadest.

Q2 · A colleague says generative AI is 'just machine learning'. What is the MOST accurate correction? (Select one)

A. They are unrelated fields. B. Generative AI is a subset of machine learning that creates new content, whereas traditional ML often just classifies or predicts a label. C. Machine learning is a subset of generative AI. D. Generative AI does not use data.

Answer: B. Generative AI sits inside machine learning but is distinguished by producing new content rather than only scoring or labelling. A is wrong because they are nested, not separate. C reverses the nesting. D is wrong – generative models are trained on data.

Q3 · Why does ChatGPT sometimes miscount the letters in a word? (Select one)

A. It is not connected to the internet. B. It processes text as tokens (chunks), not individual characters. C. Its knowledge cutoff has passed. D. The temperature is set too high.

Answer: B. The model reads and writes tokens, which are word-pieces, so character-level operations like letter counting are unreliable. A is irrelevant to counting. C concerns time-sensitive facts, not spelling. D affects variation, not character awareness.

Q4 · Roughly how many words is 1,000 tokens of English text? (Select one)

A. About 250 words. B. About 500 words. C. About 750 words. D. About 2,000 words.

Answer: C. A common rule of thumb is one token ≈ ¾ of a word, so 1,000 tokens is roughly 750 words. The other figures do not match the ≈4-characters-per-token heuristic.

Q5 · In one sentence, how does an LLM generate a reply? (Select one)

A. It retrieves the closest matching answer from a database. B. It predicts the next token repeatedly from a probability distribution, appending each to the growing text. C. It runs a fixed decision tree over the prompt. D. It searches the web and summarises the top results.

Answer: B. Generation is next-token prediction sampled from a probability distribution, one token at a time. A and C describe non-generative systems. D describes a retrieval tool, which the base model uses only when invoked.

Q6 · A user is convinced that the more their team chats with ChatGPT, the smarter their copy of the model becomes. What is the accurate picture? (Select one)

A. Correct – every chat retrains the model for that account. B. The model’s weights are fixed at inference; chats do not retrain the base model, and whether data improves future models is a separate data-control setting. C. Only paid plans retrain the model on your chats. D. Chats retrain the model but only overnight.

Answer: B. Training and inference are separate phases; sending prompts does not update the weights, and data-use for future training is governed by plan settings, not by chatting. A, C and D all wrongly equate using the model with retraining it.

Q7 · A user asks about an event that happened two weeks ago and gets a confidently wrong answer with no source. What is the FIRST thing to suspect and the fix? (Select one)

A. A bug in the app; restart it. B. The event is past the model’s knowledge cutoff; use a retrieval tool such as Search or deep research. C. The temperature is too low; raise it. D. The prompt was too short; make it longer.

Answer: B. Recent events fall past the training knowledge cutoff, so the base model has no grounded information and may guess fluently; the fix is a tool that reaches current information. A, C and D do not address the missing-current-knowledge cause.

Q8 · What does it mean that the latest OpenAI models are 'multimodal' in September 2026? (Select one)

A. They can perform any task a human can. B. They accept more than one type of input – text and images – and are multilingual with vision, while producing text output. C. They only work with audio. D. They automatically browse the web.

Answer: B. Multimodal names the input/output types – text and image input, text output for the latest models. A overstates it. C is too narrow. D confuses multimodality with a retrieval tool.

Q9 · Which task is the BEST fit for a probability-based text generator with no additional tools? (Select one)

A. Computing an exact 12-column financial total. B. Drafting a first-pass rewrite of a paragraph in a warmer tone. C. Reporting today’s stock price. D. Counting the exact characters in a code file.

Answer: B. Rephrasing is a fluent-continuation task, exactly what next-token prediction does well. A and D need exact computation/counting that the model does unreliably. C needs current data past the cutoff.

Q10 · Which TWO statements about training versus inference are correct? (Select two)

A. Training sets the model’s parameters before release. B. Inference updates the model’s parameters with each prompt. C. At inference the model’s weights are fixed and it predicts tokens for your input. D. Inference happens once, before the model is released. E. Training happens every time you press send.

Answer: A and C. Training is the pre-release phase that fixes the weights; inference is the per-prompt phase where those fixed weights generate tokens. B and E wrongly claim inference/pressing send retrains the model. D swaps the definitions of the two phases.

Q11 · A precise, unsourced statistic about a small company appears in an answer. Which TWO facts explain why you should be cautious? (Select two)

A. The model generates the most probable text, not the most verified. B. Long-tail, niche facts are exactly where unsupported generation is riskiest. C. Precise numbers are always accurate. D. The model always cites its sources automatically. E. Statistics are outside what LLMs can ever produce.

Answer: A and B. Because generation optimises for plausibility and niche facts are sparse in training data, a precise unsourced figure on an obscure topic is a classic hallucination signal. C and D are false. E is wrong – models produce statistics readily, which is precisely the problem.

Q12 · Which statement best captures why 'fluent' does not mean 'accurate'? (Select one)

A. Fluency is deliberately degraded to save cost. B. The model is optimised to produce probable, natural-sounding text, which is independent of whether the content is true. C. Accurate answers are always shorter. D. Fluency only appears in paid plans.

Answer: B. The generation objective rewards likely, well-formed text, so fluency is a property of the mechanism rather than a signal of truth. A, C and D are all false and unrelated to the fluency-versus-accuracy distinction.

Key takeaways

  • AI ⊃ machine learning ⊃ generative AI ⊃ LLMs; ChatGPT is a product on top of OpenAI’s LLMs.
  • Models work in tokens (≈ ¾ of a word), which is why they price by token and stumble on letter-level tasks.
  • Generation is next-token prediction from a probability distribution – fluent by design, not factual by design.
  • Training sets fixed weights before release; inference uses those fixed weights per prompt and does not retrain the model.
  • Every model has a knowledge cutoff; time-sensitive facts need a retrieval tool, not a bigger model.
  • Multimodal names input/output types (today: text + image in, text out), not omnipotence.
  • Use the AI-fit triage: fluent-continuation tasks fit the model; exact, current or high-harm tasks need tools and verification.

Last updated Sep 18, 2026