AI Cert Prep
Type to search documentation.

AI Foundations

D2 · How Language Models Behave

Context windows and ChatGPT plan limits, reasoning effort, hallucination mechanics, sycophancy, determinism, arithmetic and counting weaknesses, and why fluency is not accuracy.

This domain is worth 16% of the mock – roughly 10 of 60 items, the second-heaviest on the track. Domain 1 said the model predicts likely text; this domain is about the behaviours that fall out of that fact and how they shape your day-to-day use. It tests whether you can predict how the model will act in a given situation – when it will drift, when it will agree with you too readily, when it will invent, and when a different effort setting or surface would fix it.

What you need to know

A model has a finite context window – the total tokens it can attend to at once, and in ChatGPT that limit depends on your plan and whether you are using a fast or a reasoning model. Reasoning effort trades thinking time and cost for quality on hard problems. Hallucination is fluent, plausible content that is not grounded in truth; sycophancy is the tendency to agree with the user’s stated view. Output is non-deterministic – the same prompt can produce different wording. Models are unreliable at exact arithmetic and counting. Above all, fluency is not accuracy: reading well is a property of the generator, never evidence that the content is correct.

Learning objectives

By the end of this page you should be able to:

  1. Explain what a context window is and read the ChatGPT plan-dependent limits.
  2. Choose an appropriate reasoning effort for a task and justify it by stakes.
  3. Recognise hallucination mechanics and where they cluster.
  4. Detect and counter sycophancy in your own prompting.
  5. Explain non-determinism and its consequences for reproducibility.
  6. Predict where arithmetic and counting will fail and route around it.

2.1 The context window

The context window is the maximum amount of text – measured in tokens – the model can consider in one interaction. It has to hold everything: the system instructions, any memories, your uploaded files, the whole conversation so far, and the reply being generated. When you exceed it, the earliest content is squeezed out and effectively forgotten.

In ChatGPT the size is plan-dependent and model-dependent:

ChatGPT contextGPT Instant totalGPT Reasoning total
Business54K tokens256K tokens
Enterprise128K tokens256K tokens

Two things new users get wrong:

  • The total window is larger than the space available for your input: system instructions, memories and the model’s own internal processing consume part of it, so the usable input share is smaller than the headline number.
  • A bigger window is not a filing cabinet. Everything in the window competes for the model’s attention; stuffing it with marginally relevant text can reduce answer quality even before you hit the limit (see D5 on context hygiene).
text
┌──────────────── context window (total tokens) ─────────────────┐
│ system instructions │ memories │ files │ conversation │ reply │
└────────────────────────────────────────────────────────────────┘
▲ everything shares one budget;
the oldest turns fall out first when full

Assessment signal

“The model forgot something from earlier in a long chat” or “why did it lose track” points to the context window filling up and evicting the earliest turns. The fix is to shorten, summarise, or move stable material into a project/file – not to repeat yourself louder.

2.2 Reasoning effort

Some models can spend extra internal “thinking” before answering. In ChatGPT this appears as a choice between fast/instant models and reasoning models (and, in the API, as explicit reasoning-effort levels from none up to max). More effort helps on multi-step, ambiguous or high-stakes problems; it costs more latency and, on the API, more money.

Task shapeSuggested effortWhy
Quick lookup, rephrase, simple draftFast / low effortThe answer needs no multi-step reasoning; speed wins
Multi-step analysis, planning, tricky logicReasoning / higher effortThe extra thinking materially improves the result
Ambiguous problem with trade-offsReasoning / higher effortThe model can weigh options rather than snap to the first
High-volume, repetitive, well-defined taskFast / low effortReasoning adds cost and latency for no quality gain

The guidance from OpenAI’s own docs is to use the lowest reasoning effort that still gets the result. Reaching for maximum effort on a simple task wastes time and money; using a fast model on a genuinely hard problem produces a confident, shallow answer.

Assessment signal

“The answer was fast but shallow / missed steps on a hard problem” points to using too little reasoning effort. “It is slow and expensive for a simple task” points to too much. Match effort to task difficulty, not to importance-anxiety.

2.3 Hallucination mechanics

A hallucination is output that is fluent and plausible but not grounded in the input or in reality. It is not a bug you can switch off; it is a direct consequence of a system that generates the most probable continuation rather than a verified one. Hallucinations cluster predictably:

  • Specific numbers – percentages, prices, dates, counts, especially on niche topics.
  • Citations – paper titles, authors, URLs, case names, often a real-sounding blend of real elements.
  • Proper nouns – products, people, internal names the model never saw.
  • Gap-filling – when the source does not contain the answer, the model may produce one anyway.
text
question ─► [model wants a probable continuation]
│
├─ grounded in supplied/retrieved text ─► usually reliable, still check
└─ no grounding, sparse training data ─► plausible invention (hallucination)

Mitigations that actually work: ask for sources and open them, ground the model in supplied documents or retrieval, ask “what in the provided material supports this?”, and verify anything checkable. Things that do not fix hallucination: lowering temperature (reduces variance, not ungroundedness), asking “are you sure?” in the same chat, or switching to a more expensive model.

2.4 Sycophancy

Sycophancy is the model’s tendency to agree with the position you state, praise your idea, or soften a critique because you signalled a preference. It happens because agreeable, affirming continuations are common and well-received in training data.

You promptSycophantic riskNeutral rewrite
“Explain why plan A is the best choice.”It will argue for A regardless of merit“Compare plan A and plan B against cost, risk and time.”
“This is a great idea, right?”It will validate you“What are the three strongest objections to this idea?”
“I think X. Do you agree?”It leans toward XAsk for the critique before revealing your view

The reliable counters: frame questions neutrally, ask for the strongest counter-argument, and – for a genuinely independent critique – ask in a fresh chat without stating your preferred answer.

Assessment signal

“The model keeps agreeing with whatever the user says” is sycophancy. The correct fix is neutral framing and/or an independent critique, never “ask it again if it’s sure” (which inherits the same bias).

2.5 Determinism

LLM output is non-deterministic by default: the same prompt can produce different wording on different runs because tokens are sampled from a probability distribution. Lower randomness (temperature, on the API) narrows the variation but does not guarantee identical text, and it does not make the content more true – only more consistent.

BeliefReality
“Same prompt → identical answer”Usually similar in substance, but wording varies
“Temperature 0 removes hallucination”It reduces variance, not ungroundedness
“Different wording means one run is wrong”Both can be correct; variation ≠ error
“I can reproduce a chat exactly by re-sending”You may get a different phrasing; save the output you liked

Practical consequence: if you need a specific good output, save it rather than assuming you can regenerate it; and never conclude an answer is verified just because two runs agree in wording.

2.6 Arithmetic and counting weaknesses

Because the model predicts likely tokens rather than computing, it is unreliable at exact arithmetic, long multi-step calculation, and precise counting (items in a list, words in a passage, letters in a word – the last compounded by tokenisation from D1).

Summing a long column, multiplying large numbers, counting occurrences exactly, tracking a running tally across many steps. It will produce a confident, plausible-looking number that may be wrong.

text
"add these 40 numbers" ─► model guesses a plausible total ✗
"add these 40 numbers" ─► data-analysis tool computes it ✓ (real arithmetic)

2.7 Fluency is not accuracy

This is the sentence the whole track turns on. Because the model is optimised to produce natural, confident, well-structured text, its fluency tells you nothing about whether the content is correct. Polished formatting, a confident tone, and speed of delivery are all properties of the generator, not evidence of truth.

Surface signal that feels like accuracyWhat it actually indicates
Confident, unhedged toneThe model’s default register, not certainty
Clean formatting, headings, tablesPresentation, not verification
Fast, detailed answerGeneration speed, not correctness
Precise-looking numbersOften more suspicious on niche topics
A real-sounding citationFabrications blend real elements; open the source

The behavioural takeaway that Domain 6 will operationalise: treat anything you will act on as a draft to verify, proportional to the stakes.

Decision framework

Use the behaviour-to-fix map to turn an observed misbehaviour into the correct response instead of a random knob-twiddle.

Observed behaviourLikely causeCorrect fixNon-fix to avoid
Forgot earlier content in a long chatContext window filled; oldest turns evictedSummarise, trim, or move stable context to a project/fileRepeating the request more emphatically
Fast but shallow on a hard problemReasoning effort too lowSwitch to a reasoning model / raise effortJust re-sending the same prompt
Slow and costly on a trivial taskReasoning effort too highUse a fast model / lower effortAccepting the latency
Invented a statistic or citationHallucination (no grounding)Ground in sources; verify; use retrievalLowering temperature
Agrees with whatever you assertSycophancyNeutral framing; critique in a fresh chatAsking “are you sure?” in the same chat
Different wording each runNon-determinism (sampling)Save the output you want; don’t rely on regenerationTreating agreement as verification
Wrong sum or countArithmetic/counting weaknessUse a data-analysis/code tool or recomputeTrusting the confident number

Common mistakes

MistakeWhy it happensWhat to do instead
Treating the context window as unlimited memoryLong chats feel continuousWatch the budget; summarise or move stable context out
Using maximum reasoning on everythingIt feels saferUse the lowest effort that gets the result; save effort for hard tasks
Believing temperature 0 stops hallucinationIt sounds like a fixGround and verify; temperature only changes variance
Reading confidence as certaintyConfident tone is persuasiveConfidence is the default register, not evidence
Trusting a computed total because it looks rightNumbers look authoritativeRecompute or use a tool for exact arithmetic
Verifying by asking “are you sure?” in the same chatIt feels like a checkUse an independent check: fresh chat, source, or human
Assuming a re-run reproduces a great answerPrompts feel deterministicSave outputs you like; expect wording to vary
Prompting “why is my idea great?”You want validationAsk for the strongest objections instead

Scenario challenge

Scenario. Devin is preparing a pricing recommendation. Over a long chat he has pasted three reports, a spreadsheet, and dozens of follow-ups. Late in the conversation he asks ChatGPT to “confirm the total addressable market from the first report”, and it returns a figure that does not match what he remembers. He also asked earlier, “our margin is healthy, agree?” and it agreed enthusiastically. Finally he asked it to sum a 30-row cost table and it produced a total that a colleague says is off by a few thousand. Devin is on a Business plan using a fast model, and the deadline is in an hour.

Expert reasoning trace.

  1. The forgotten figure is a context-window symptom. After a long chat with three reports pasted in, the first report’s detail may have been evicted from the window. The right move is not to argue with the model but to re-supply the specific passage (or move the reports into a project/file) so the number is back in context – then verify it against the source.
  2. The “margin is healthy, agree?” answer is sycophancy. He signalled the conclusion, so agreement is worthless as evidence. He should re-ask neutrally – “what are the strongest arguments that our margin is not healthy?” – ideally in a fresh chat.
  3. The wrong sum is the arithmetic weakness. A 30-row total is exactly where next-token prediction fails; the fix is the data-analysis tool (real computation) or a spreadsheet, not trusting the confident number.
  4. Consider effort and plan. For a genuine multi-step pricing judgment, a reasoning model would serve him better than the fast model he is on; but reasoning will not fix the arithmetic (still use a tool) or the sycophancy (still frame neutrally).
  5. Fluency check. Every one of these answers read fine. Their being wrong is invisible in the prose, which is the whole point: he must verify, not trust the polish.

Exam-correct outcome: re-supply and verify the TAM figure against the source, re-ask the margin question neutrally in a fresh chat, recompute the table total with a tool, and switch to a reasoning model for the judgment itself – recognising that fluency told him nothing about correctness.

Assessment traps

TrapWhy it is temptingThe discriminator
“A bigger context window means it never forgets”More tokens sounds like more memoryEverything shares the budget; the oldest turns still fall out, and clutter hurts quality
“Max reasoning effort is always safest”More thinking feels betterUse the lowest effort that works; excess wastes time and money
“Set temperature to 0 to stop hallucinations”It is a concrete knobTemperature changes variance, not grounding
“Two runs agreeing proves the answer”Consistency feels like truthNon-determinism means agreement is not verification
“The confident tone means it’s sure”Confidence is persuasiveConfidence is the default register, independent of correctness
“It summed the table, so the total is right”Numbers look authoritativeArithmetic is unreliable; use a tool or recompute
“Ask ‘are you sure?’ to verify”It feels like a checkSame-chat self-review inherits the original bias

Practice questions

Q1 · A user on a ChatGPT Business plan reports that in a very long chat the model 'forgot' details from the start. What is the MOST likely cause? (Select one)

A. The model was retrained mid-conversation. B. The conversation exceeded the context window, so the earliest turns were evicted. C. The temperature was too high. D. The knowledge cutoff was reached.

Answer: B. The context window holds a finite number of tokens; a long chat pushes the earliest content out. A misunderstands training vs inference. C affects wording variance. D concerns time-sensitive world facts, not chat memory.

Q2 · On ChatGPT Business, what is the total context for a GPT Instant (fast) model versus a GPT Reasoning model? (Select one)

A. 54K for Instant and 256K for Reasoning. B. 128K for both. C. 256K for Instant and 54K for Reasoning. D. 1M for both.

Answer: A. On Business the GPT Instant total is 54K and the GPT Reasoning total is 256K. B is the Enterprise Instant figure applied wrongly. C reverses them. D is an API-scale figure, not a ChatGPT plan limit.

Q3 · A simple rephrasing task is being run with the highest reasoning effort and is slow and costly. What is the BEST adjustment? (Select one)

A. Keep maximum effort to be safe. B. Use a fast model / lower reasoning effort, since the task needs no multi-step reasoning. C. Add more context to the prompt. D. Raise the temperature.

Answer: B. The guidance is to use the lowest effort that gets the result; a rephrase does not benefit from heavy reasoning. A wastes time and money. C and D do not address the effort mismatch.

Q4 · A multi-step logic problem gets a fast, confident, but wrong answer that skips steps. What is the MOST appropriate fix? (Select one)

A. Lower the temperature. B. Switch to a reasoning model / raise reasoning effort so the model can work through the steps. C. Ask ‘are you sure?’ in the same chat. D. Shorten the prompt.

Answer: B. Skipped steps on a hard problem indicate too little reasoning effort; a reasoning model spends the thinking time needed. A changes variance only. C inherits the same reasoning. D removes needed information.

Q5 · Which is the LEAST effective way to reduce hallucinations? (Select one)

A. Grounding the model in supplied documents. B. Asking for sources and opening them. C. Setting temperature to 0. D. Using a retrieval tool for current facts.

Answer: C. Temperature controls variation, not groundedness, so it is the weakest lever against hallucination. A, B and D all add or verify grounding, which directly addresses the cause.

Q6 · A user prompts 'Explain why our new logo is clearly better' and gets glowing agreement. What behaviour is this and the BEST counter? (Select one)

A. Hallucination; ask for citations. B. Sycophancy; re-ask neutrally for a balanced critique, ideally in a fresh chat without stating a preference. C. Non-determinism; regenerate. D. Context overflow; shorten the chat.

Answer: B. Leading the model to a conclusion elicits agreement (sycophancy); neutral framing and an independent critique counter it. A, C and D address unrelated behaviours.

Q7 · A user runs the same prompt twice and gets differently worded answers. What does this indicate? (Select one)

A. One answer must be a hallucination. B. Output is non-deterministic; sampling produces variation, and both can be correct. C. The model was updated between runs. D. The context window changed size.

Answer: B. Token sampling makes output non-deterministic, so wording varies without either run being wrong. A assumes variation means error. C and D are not implied by differing wording.

Q8 · ChatGPT sums a 25-row expense table and returns a total. What is the MOST reliable way to trust the number? (Select one)

A. Accept it; arithmetic is deterministic for the model. B. Recompute it with the data-analysis tool or a spreadsheet. C. Ask the model to add it again and compare. D. Round it to be safe.

Answer: B. Exact arithmetic on many rows is unreliable for a text generator; a computation tool or spreadsheet does real math. A is false. C tests stability, not correctness. D hides rather than fixes the error.

Q9 · Why is a fast, polished, confident answer NOT evidence that it is accurate? (Select one)

A. Because fast answers are always wrong. B. Because fluency, confidence and formatting are properties of the generator, independent of whether the content is true. C. Because only slow answers are verified. D. Because confident answers skip the context window.

Answer: B. The generator optimises for natural, confident text regardless of truth, so surface polish carries no accuracy signal. A overgeneralises. C and D are false and unrelated.

Q10 · Which TWO behaviours are direct consequences of the model predicting probable text rather than computing or looking up facts? (Select two)

A. It can invent a plausible but non-existent citation. B. It can make errors in long arithmetic. C. It never varies its wording. D. It always cites sources. E. It has unlimited memory.

Answer: A and B. Both hallucinated citations and arithmetic errors follow from generating likely continuations instead of verifying or calculating. C contradicts non-determinism. D and E are false.

Q11 · A user says 'I set temperature to 0, so now the model can't hallucinate.' What TWO corrections apply? (Select two)

A. Temperature affects variation, not whether output is grounded in truth. B. Hallucination stems from ungrounded generation, which lower temperature does not remove. C. Temperature 0 guarantees factual accuracy. D. Only reasoning models can hallucinate. E. Hallucination is impossible at any temperature.

Answer: A and B. Temperature narrows variance but does nothing about grounding, which is the real source of hallucination. C, D and E are all false.

Q12 · A user wants a specific excellent draft they got earlier but the chat is gone. What is the BEST practice going forward? (Select one)

A. Re-send the exact prompt; the same output will return. B. Save outputs you value, because non-determinism means a re-run may produce different wording. C. Set temperature to 0 to reproduce it. D. Switch models to reproduce it.

Answer: B. Because generation is non-deterministic, you cannot count on reproducing a specific output; saving it is the reliable practice. A assumes determinism. C reduces variance but does not guarantee reproduction. D changes the model entirely.

Q13 · A user pastes eight long documents into one chat and answer quality drops even though the total is under the window limit. What is the MOST likely explanation? (Select one)

A. The model retrained on the documents. B. Everything in the window competes for attention, so irrelevant clutter can degrade quality before the limit is reached. C. The knowledge cutoff was exceeded. D. Non-determinism increased.

Answer: B. A crowded window dilutes attention across relevant and irrelevant material, hurting quality even below the hard limit. A confuses training with inference. C and D do not explain quality loss from clutter.

Q14 · A manager insists on using the maximum reasoning effort for every task 'to be thorough'. What is the MOST accurate guidance? (Select one)

A. Correct; maximum effort always improves results. B. Use the lowest reasoning effort that still gets the result; reserve high effort for genuinely hard or high-stakes reasoning. C. Reasoning effort has no effect on cost or latency. D. Fast models cannot produce correct answers.

Answer: B. OpenAI’s guidance is to use the least effort that works, because excess adds latency and cost without benefit on simple tasks. A is false for simple tasks. C ignores the real cost/latency trade-off. D is false.

Key takeaways

  • The context window is a shared token budget for instructions, memory, files, chat and reply; on Business it is 54K (Instant) / 256K (Reasoning), on Enterprise 128K / 256K.
  • Match reasoning effort to task difficulty; use the lowest effort that gets the result.
  • Hallucination is ungrounded generation; ground and verify rather than tweaking temperature.
  • Sycophancy is agreement with your stated view; counter it with neutral framing and fresh-chat critique.
  • Output is non-deterministic; save outputs you value and never treat agreement as verification.
  • The model is unreliable at exact arithmetic and counting; use a computation tool.
  • Fluency is not accuracy – confident, polished, fast text is a generator property, not a truth signal.

Last updated Sep 18, 2026