# D1 · AI and Generative AI Fundamentals

What AI, machine learning, generative AI and large language models are, how tokens and probability-based generation work, the difference between training and inference, knowledge cutoffs and multimodality.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain is worth **14% of the mock – roughly 8 of 60 items**. It is the vocabulary layer for everything else. It tests whether you can place ChatGPT correctly in the family tree of AI, explain in plain terms *why* it produces text the way it does, and reason from that mechanism to its strengths and its failure modes. Almost every later domain is an application of one idea introduced here: the model predicts likely text, it does not look facts up.

## What you need to know

Artificial intelligence is the broad field of getting machines to do things that normally need human intelligence. Machine learning is the subset that learns patterns from data instead of following hand-written rules. Generative AI is the slice of machine learning that produces new content – text, images, audio, code. A large language model (LLM) is a generative model trained on enormous amounts of text to predict the next chunk of text. ChatGPT is a product built on top of OpenAI's LLMs. Understanding that stack, and that the model generates by predicting probable continuations one token at a time from what it learned during training, explains its fluency, its speed, its knowledge cutoff and its tendency to sound confident even when wrong.

## Learning objectives

By the end of this page you should be able to:

1. Distinguish **AI, machine learning, generative AI and LLMs**, and place ChatGPT within them.
2. Explain **tokens** and why the model works in tokens rather than words or letters.
3. Describe **probability-based generation** and why it makes the model fluent but not factual.
4. Separate **training from inference** and explain what a **knowledge cutoff** is and is not.
5. Explain **multimodality** – what today's models can take in and produce.
6. Reason from the generation mechanism to a concrete strength or limitation.

---

## 1.1 The family tree: AI, ML, generative AI, LLM

These four terms are nested, not synonyms. The single most common D1 item asks you to order or relate them.

```text
┌─────────────────────────────────────────────────────────┐
│ Artificial Intelligence                                   │
│  machines doing tasks that need human-like intelligence   │
│  ┌──────────────────────────────────────────────────┐    │
│  │ Machine Learning                                    │   │
│  │  systems that learn patterns from data              │   │
│  │  ┌────────────────────────────────────────────┐    │   │
│  │  │ Generative AI                                 │   │  │
│  │  │  creates new content (text, image, audio)     │   │  │
│  │  │  ┌──────────────────────────────────────┐    │   │  │
│  │  │  │ Large Language Models (LLMs)           │   │  │  │
│  │  │  │  predict likely text; ChatGPT is a     │   │  │  │
│  │  │  │  product built on OpenAI's LLMs        │   │  │  │
│  │  │  └──────────────────────────────────────┘    │   │  │
│  │  └────────────────────────────────────────────┘    │   │
│  └──────────────────────────────────────────────────┘    │
└─────────────────────────────────────────────────────────┘
```

| Term | One-line definition | Everyday example that is *not* generative |
| --- | --- | --- |
| Artificial intelligence | Machines performing tasks that need human-like intelligence | A chess engine; a spam filter rule set |
| Machine learning | AI that learns patterns from data rather than fixed rules | A fraud-detection model scoring transactions |
| Generative AI | ML that produces new content | An image generator; an LLM |
| Large language model | A generative model trained to predict likely text | The model behind ChatGPT |

:::tip[Assessment signal]
Stems that say "which term is the *broadest*", "ChatGPT is an example of", or "generative AI differs from traditional machine learning because" are testing this nesting. The discriminator: traditional ML *classifies or predicts a label*; generative AI *produces new content*.
:::

## 1.2 Tokens: the units the model actually reads and writes

A model does not process words or characters directly. Text is broken into **tokens** – common chunks that are often a word, sometimes a word-piece, sometimes punctuation. A rough English rule of thumb is **one token ≈ 4 characters ≈ ¾ of a word**, so 1,000 tokens is about 750 words.

```text
"Generation is probabilistic."
  ▼ tokenizer
["Generation", " is", " probabil", "istic", "."]
   one token    one   word-piece  word-pc  punct
```

Why this matters in practice:

| Consequence of tokenisation | What you observe |
| --- | --- |
| The model sees chunks, not letters | It miscounts letters ("how many r's in strawberry") and struggles with anagrams |
| Cost and limits are measured in tokens | Context windows and API prices are quoted per token, not per word |
| Rare words split into many tokens | Unusual names or code identifiers cost more and are handled less reliably |
| Whitespace and case are part of the token | " the" and "The" are different tokens |

:::tip[Assessment signal]
"Why does the model get letter-counting or spelling-manipulation tasks wrong?" points here: it reads tokens, not characters. "How is usage measured/priced?" also points here: tokens, not words.
:::

## 1.3 Probability-based generation

An LLM generates text one token at a time. At each step it produces a probability distribution over all possible next tokens and picks from it, then feeds the result back in and repeats. It is, in effect, a very sophisticated next-token predictor.

```text
prompt ─► [model] ─► distribution over next tokens
                       "Paris" 0.71 · "the" 0.09 · "France" 0.04 · …
                       pick ─► append ─► feed back ─► repeat
```

Two facts fall straight out of this mechanism and drive most of Domain 2:

- **Fluency is a property of the mechanism, not a truth signal.** The model is optimised to produce *probable* text, which usually reads well, regardless of whether the content is correct.
- **The same prompt can give different wording** because sampling introduces variation (see determinism in D2).

<Tabs>
  <TabItem label="Where it shines">
    Predicting likely continuations is exactly right for drafting, rephrasing, summarising, translating, brainstorming, explaining and coding-by-pattern – tasks where a fluent, plausible continuation *is* the goal.
  </TabItem>
  <TabItem label="Where it fails">
    It is a poor fit for tasks that need guaranteed correctness with no verification: exact arithmetic, precise counts, current facts beyond the cutoff, or anything where a plausible-but-wrong answer is dangerous. Those need tools, retrieval, or human checking.
  </TabItem>
</Tabs>

## 1.4 Training versus inference

These are two entirely separate phases, and confusing them causes real mistakes (like believing your chat "teaches" the base model).

| Phase | When | What happens | Does it change the model? |
| --- | --- | --- | --- |
| **Training** | Before release, at OpenAI, on huge datasets | The model's parameters are learned from data | Yes – this is where the weights are set |
| **Inference** | Every time you send a prompt | The trained model predicts tokens for your input | No – the weights are fixed; nothing you type retrains the base model |

Consequences to remember:

- Your prompts and uploaded files inform the model **only for that conversation** (plus memory/features you enable). They do not update the underlying model for everyone.
- Whether your content is used to *improve future models* is a **data-control and plan setting**, separate from inference (Enterprise does not train on business data by default). Do not conflate "the model used my message this turn" with "the model was retrained on my message".

```text
TRAINING (once, before you ever see it)      INFERENCE (every prompt)
huge text corpus ─► learn parameters         your prompt ─► fixed model ─► tokens
        weights are set and frozen                 weights never change here
```

## 1.5 Knowledge cutoff

Because knowledge is baked in during training, each model has a **knowledge cutoff**: the point after which it has no built-in information about the world. Ask about an event after the cutoff and, without a tool, the model will either say it does not know or – worse – guess fluently.

| Model (September 2026) | Knowledge cutoff |
| --- | --- |
| GPT-6 Astra (`gpt-6-astra`) | 30 April 2026 |
| GPT-5.6 Sol / Terra / Luna | 16 February 2026 |

How to work with a cutoff:

- For anything time-sensitive (prices, news, releases, "latest"), use a tool that reaches current information – **Search** or **deep research** in ChatGPT – rather than the model's memory.
- Treat unsourced claims about recent events as suspect by default.
- A later cutoff is not "smarter", only more recently informed.

:::tip[Assessment signal]
"The model gave outdated information / did not know about a recent event" points to the knowledge cutoff, and the fix is a retrieval tool (Search / deep research), not a bigger or more expensive model.
:::

## 1.6 Multimodality

Modern OpenAI models are **multimodal**: they can take in more than one type of content. In September 2026 the latest models accept **text and image input** and produce **text output**, and are multilingual with vision. In ChatGPT you also get image generation, voice, and file/data handling as product features layered on top.

| Capability | What it means in ChatGPT | Everyday use |
| --- | --- | --- |
| Text in, text out | The core chat interaction | Drafting, Q&A, analysis |
| Image input (vision) | Upload a photo, screenshot or diagram and ask about it | "What does this error screen say?" |
| Image generation | Produce an image from a description | A slide illustration, a concept sketch |
| Voice | Speak to ChatGPT and hear replies; voice with video | Hands-free brainstorming |
| Files and data | Upload documents/spreadsheets; run data analysis | Summarise a PDF; chart a CSV |

The exam-relevant point: "multimodal" describes the *kinds of input and output*, and knowing that today's latest models take text + image input and return text tells you when you need a specialised feature (image generation, voice) rather than the base chat.

## Decision framework

Use **the AI-fit triage** to decide, in seconds, whether a task suits a probability-based text generator at all before you even open a chat.

| Ask | If yes | If no |
| --- | --- | --- |
| Is a *fluent, plausible* result the goal (draft, summary, rephrase, explain, ideate)? | Strong fit for an LLM | Reconsider – you may need a calculator, database or search |
| Does it need **current** information past the cutoff? | Add a retrieval tool (Search / deep research) | Base model knowledge is fine |
| Does it need **exact** arithmetic, counting or precise data lookup? | Use a tool (data analysis / code) or verify | The model alone is acceptable |
| Would a plausible-but-wrong answer cause harm? | Verify before use; add a review gate | Lower verification is acceptable |

The framework's value: it stops you from asking a next-token predictor to *be* a database or a calculator, which is the root cause of most beginner disappointment.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Treating ChatGPT as a search engine | It answers in the same conversational shape as search | Use Search/deep research for current facts; treat model knowledge as a lead |
| Thinking "AI", "ML" and "generative AI" are the same thing | The terms are used loosely in the press | Keep the nesting: generative AI creates content; older ML classifies or predicts labels |
| Believing your chats retrain the model | "It learned from me" feels intuitive | Inference is fixed-weight; training is separate; data use is a plan/setting |
| Expecting exact letter counts or arithmetic | The output is fluent, so it looks authoritative | Remember it reads tokens, not characters; use a tool for exact work |
| Assuming a later model "knows everything now" | Newer sounds better | Every model has a cutoff; time-sensitive facts need a tool |
| Confusing multimodal with "does anything" | The word sounds all-encompassing | Multimodal = specific input/output types; check the actual capability |
| Thinking higher price means more accuracy | Cost correlates with capability, not truth | Verification, not spend, is what makes output trustworthy |

## Scenario challenge

**Scenario.** Priya, a new analyst, is delighted with ChatGPT after it drafted a crisp market summary. Encouraged, she asks it for "the exact number of stores our competitor opened last quarter" and "how many times the letter *s* appears in our tagline", then pastes both answers into a board slide. She also tells a colleague, "the more we use it, the smarter our copy of it gets". The competitor number looks precise (it cites "47 new stores"); the letter count is off by two; and she is about to present both.

**Expert reasoning trace.**

1. **Locate each task on the AI-fit triage.** The market *summary* is a fluent-continuation task – a strong fit, done well. The *exact competitor count* needs current, specific data past nothing the model was trained to know reliably; that is a factual-lookup task, not a generation task. The *letter count* is a character-level operation the model performs poorly because it reads tokens, not letters.
2. **Diagnose the competitor figure.** A precise unsourced number on a niche, recent topic is the hallucination signature. The fix is a retrieval tool (Search / deep research) or an authoritative source, then verification – not trusting the fluent "47".
3. **Diagnose the letter count.** Tokenisation explains the error directly; the reliable fix is a tool that counts characters, or doing it by hand.
4. **Correct the "it gets smarter as we use it" belief.** Inference does not retrain the model; her chats do not improve the base model, and whether her data trains future models is a separate data-control setting she should check, not assume.
5. **Set the review gate.** Both figures are going on an external-facing board slide, so they need verification before use regardless of how confident the text sounds.

**Exam-correct outcome:** keep the summary, replace the competitor number with a verified sourced figure (via a retrieval tool or authoritative data), recompute the letter count with a tool or by hand, and drop the mistaken belief that usage retrains the model.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "ChatGPT looks things up when you ask" | It answers like a search box | It predicts likely text from training; it only *looks up* when a retrieval tool is invoked |
| "Generative AI and machine learning are the same" | Both are "AI" | Generative AI creates content; it is a *subset* of ML |
| "My prompts train the model" | "It learned from me" feels right | Training is a separate, pre-release phase; inference is fixed-weight |
| "A newer model has no knowledge cutoff" | Newer implies up-to-date | Every model has a cutoff; only tools reach beyond it |
| "It counts letters accurately because it reads text" | It clearly reads text | It reads *tokens*, not characters, so character-level tasks fail |
| "Multimodal means it can do any task" | The word sounds total | It names specific input/output types (e.g. text + image in, text out) |
| "The pricier model is more truthful" | Cost feels like quality | Price tracks capability/latency, not factual guarantees; verification does that |

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · Which term is the BROADEST of the following? (Select one)">
    A. Large language model
    B. Generative AI
    C. Artificial intelligence
    D. Machine learning

    **Answer: C.** Artificial intelligence is the outermost category; machine learning is a subset of it, generative AI a subset of ML, and an LLM a specific kind of generative model. B, D and A are all nested inside AI, so none is the broadest.
  </AccordionItem>

  <AccordionItem title="Q2 · A colleague says generative AI is 'just machine learning'. What is the MOST accurate correction? (Select one)">
    A. They are unrelated fields.
    B. Generative AI is a subset of machine learning that creates new content, whereas traditional ML often just classifies or predicts a label.
    C. Machine learning is a subset of generative AI.
    D. Generative AI does not use data.

    **Answer: B.** Generative AI sits inside machine learning but is distinguished by producing new content rather than only scoring or labelling. A is wrong because they are nested, not separate. C reverses the nesting. D is wrong – generative models are trained on data.
  </AccordionItem>

  <AccordionItem title="Q3 · Why does ChatGPT sometimes miscount the letters in a word? (Select one)">
    A. It is not connected to the internet.
    B. It processes text as tokens (chunks), not individual characters.
    C. Its knowledge cutoff has passed.
    D. The temperature is set too high.

    **Answer: B.** The model reads and writes tokens, which are word-pieces, so character-level operations like letter counting are unreliable. A is irrelevant to counting. C concerns time-sensitive facts, not spelling. D affects variation, not character awareness.
  </AccordionItem>

  <AccordionItem title="Q4 · Roughly how many words is 1,000 tokens of English text? (Select one)">
    A. About 250 words.
    B. About 500 words.
    C. About 750 words.
    D. About 2,000 words.

    **Answer: C.** A common rule of thumb is one token ≈ ¾ of a word, so 1,000 tokens is roughly 750 words. The other figures do not match the ≈4-characters-per-token heuristic.
  </AccordionItem>

  <AccordionItem title="Q5 · In one sentence, how does an LLM generate a reply? (Select one)">
    A. It retrieves the closest matching answer from a database.
    B. It predicts the next token repeatedly from a probability distribution, appending each to the growing text.
    C. It runs a fixed decision tree over the prompt.
    D. It searches the web and summarises the top results.

    **Answer: B.** Generation is next-token prediction sampled from a probability distribution, one token at a time. A and C describe non-generative systems. D describes a retrieval tool, which the base model uses only when invoked.
  </AccordionItem>

  <AccordionItem title="Q6 · A user is convinced that the more their team chats with ChatGPT, the smarter their copy of the model becomes. What is the accurate picture? (Select one)">
    A. Correct – every chat retrains the model for that account.
    B. The model's weights are fixed at inference; chats do not retrain the base model, and whether data improves future models is a separate data-control setting.
    C. Only paid plans retrain the model on your chats.
    D. Chats retrain the model but only overnight.

    **Answer: B.** Training and inference are separate phases; sending prompts does not update the weights, and data-use for future training is governed by plan settings, not by chatting. A, C and D all wrongly equate using the model with retraining it.
  </AccordionItem>

  <AccordionItem title="Q7 · A user asks about an event that happened two weeks ago and gets a confidently wrong answer with no source. What is the FIRST thing to suspect and the fix? (Select one)">
    A. A bug in the app; restart it.
    B. The event is past the model's knowledge cutoff; use a retrieval tool such as Search or deep research.
    C. The temperature is too low; raise it.
    D. The prompt was too short; make it longer.

    **Answer: B.** Recent events fall past the training knowledge cutoff, so the base model has no grounded information and may guess fluently; the fix is a tool that reaches current information. A, C and D do not address the missing-current-knowledge cause.
  </AccordionItem>

  <AccordionItem title="Q8 · What does it mean that the latest OpenAI models are 'multimodal' in September 2026? (Select one)">
    A. They can perform any task a human can.
    B. They accept more than one type of input – text and images – and are multilingual with vision, while producing text output.
    C. They only work with audio.
    D. They automatically browse the web.

    **Answer: B.** Multimodal names the input/output types – text and image input, text output for the latest models. A overstates it. C is too narrow. D confuses multimodality with a retrieval tool.
  </AccordionItem>

  <AccordionItem title="Q9 · Which task is the BEST fit for a probability-based text generator with no additional tools? (Select one)">
    A. Computing an exact 12-column financial total.
    B. Drafting a first-pass rewrite of a paragraph in a warmer tone.
    C. Reporting today's stock price.
    D. Counting the exact characters in a code file.

    **Answer: B.** Rephrasing is a fluent-continuation task, exactly what next-token prediction does well. A and D need exact computation/counting that the model does unreliably. C needs current data past the cutoff.
  </AccordionItem>

  <AccordionItem title="Q10 · Which TWO statements about training versus inference are correct? (Select two)">
    A. Training sets the model's parameters before release.
    B. Inference updates the model's parameters with each prompt.
    C. At inference the model's weights are fixed and it predicts tokens for your input.
    D. Inference happens once, before the model is released.
    E. Training happens every time you press send.

    **Answer: A and C.** Training is the pre-release phase that fixes the weights; inference is the per-prompt phase where those fixed weights generate tokens. B and E wrongly claim inference/pressing send retrains the model. D swaps the definitions of the two phases.
  </AccordionItem>

  <AccordionItem title="Q11 · A precise, unsourced statistic about a small company appears in an answer. Which TWO facts explain why you should be cautious? (Select two)">
    A. The model generates the most probable text, not the most verified.
    B. Long-tail, niche facts are exactly where unsupported generation is riskiest.
    C. Precise numbers are always accurate.
    D. The model always cites its sources automatically.
    E. Statistics are outside what LLMs can ever produce.

    **Answer: A and B.** Because generation optimises for plausibility and niche facts are sparse in training data, a precise unsourced figure on an obscure topic is a classic hallucination signal. C and D are false. E is wrong – models produce statistics readily, which is precisely the problem.
  </AccordionItem>

  <AccordionItem title="Q12 · Which statement best captures why 'fluent' does not mean 'accurate'? (Select one)">
    A. Fluency is deliberately degraded to save cost.
    B. The model is optimised to produce probable, natural-sounding text, which is independent of whether the content is true.
    C. Accurate answers are always shorter.
    D. Fluency only appears in paid plans.

    **Answer: B.** The generation objective rewards likely, well-formed text, so fluency is a property of the mechanism rather than a signal of truth. A, C and D are all false and unrelated to the fluency-versus-accuracy distinction.
  </AccordionItem>
</Accordions>

## Key takeaways

- AI ⊃ machine learning ⊃ generative AI ⊃ LLMs; ChatGPT is a product on top of OpenAI's LLMs.
- Models work in **tokens** (≈ ¾ of a word), which is why they price by token and stumble on letter-level tasks.
- Generation is **next-token prediction from a probability distribution** – fluent by design, not factual by design.
- **Training** sets fixed weights before release; **inference** uses those fixed weights per prompt and does not retrain the model.
- Every model has a **knowledge cutoff**; time-sensitive facts need a retrieval tool, not a bigger model.
- **Multimodal** names input/output types (today: text + image in, text out), not omnipotence.
- Use the AI-fit triage: fluent-continuation tasks fit the model; exact, current or high-harm tasks need tools and verification.
