Start with OpenAI
Assessment Model and Readiness Scoring
How OpenAI Academy assessments work, how our independent mock exams differ, how question anatomy is designed, and how the readiness indicator should be read.
Read this before your first mock exam. It explains what the real assessments do, what ours do differently and why, and how to interpret a score without fooling yourself.
How an Academy assessment works
| Property | Value |
|---|---|
| When | At the end of every Academy course; taking it is optional, passing it is required for the badge |
| Length | 10–20 questions per attempt |
| Source | Randomised selection from a 50-question bank per assessment |
| Pass mark | 80% or higher |
| Retakes | Allowed, with a newly randomised selection |
| Result | An OpenAI Academy badge issued through Accredible |
| Pathway effect | Every course in a pathway must be passed at ≥ 80% for the pathway certificate of completion |
Two consequences follow from the bank being small and the selection random.
First, coverage beats cramming. With 10–20 items drawn from 50, any single narrow topic may or may not appear, but across retakes everything appears. Studying “what is likely on the test” is not a strategy when the sample changes each attempt.
Second, 80% is unforgiving on short forms. On a 10-item form you can afford two wrong answers. On a 15-item form, three. That is why our mocks are longer and our internal target is higher than 80%: a long mock at 85% is a much better predictor of a short assessment at 80% than a short quiz at 80% is.
How our mock exams differ
| Academy assessment | Our mock exams | |
|---|---|---|
| Items per sitting | 10–20 | 50–60 |
| Bank | 50 items per assessment | 100–120 items per track, across two exams, plus 60–150 in-page questions |
| Timing | Untimed in practice | Timed, defaulting to the track’s recommended duration; 30/60/90-minute drills and untimed mode available |
| Scoring | 80% to earn a badge | Independent readiness bands, with the same 80% line called out |
| Feedback | Pass or fail | Per-domain breakdown, full correction with distractor analysis, retry-incorrect-only |
| Status | Official | Independent practice. Not official OpenAI questions. |
Our items are written from published learning objectives and product documentation, not from anyone’s recollection of a real assessment. That is both an ethical and a practical choice: an item written from the objective tests the skill, whereas an item written from a leak tests memory of a leak.
Anatomy of a well-built item
Every question in this section carries the same structure, whether it lives on a domain page or in a mock bank:
| Component | Purpose |
|---|---|
| Scenario | A short, realistic work situation – who you are, what you have, what is at stake |
| Stem | One precise question, usually with a qualifier: FIRST, BEST, MOST cost-effective, TWO |
| Four options | One correct, three plausible-but-weaker; all similar in length and register |
| Correct answer | Marked, with the reasoning that makes it correct rather than merely acceptable |
| Distractor analysis | Why each wrong option is wrong – the part that actually teaches |
| Domain and objective | So a wrong answer maps to a page you can re-read |
How to read a qualifier
FIRST asks for sequence, not quality: several options may be good, only one comes first. BEST asks for a trade-off judgment against the stated constraint. MOST cost-effective means the cheapest option that still meets the requirement – not the cheapest option. TWO means an item is scored all-or-nothing: both right, no extras.
Readiness bands
Our results screen reports an independent readiness indicator. It is not an OpenAI score, and it has no standing with OpenAI.
| Score | Band | Interpretation | Next action |
|---|---|---|---|
| under 70% | Keep learning | Gaps are structural, not incidental | Re-read your two weakest domains, redo their in-page questions |
| 70–79% | Building confidence | You know the material but not reliably | Drill weak domains; retry incorrect items only |
| 80–89% | Assessment ready | At or above the Academy badge threshold | Take the Academy course, then its assessment |
| 90%+ | Strong readiness | Comfortable margin on a short randomised form | Move to the next track |
Per-domain scores matter more than the total. A 3 of 8 in one domain and 95% overall still means a real chance of failing a short form that happens to sample that domain twice.
Practice modes and how to use them
| Mode | Length | Use it for |
|---|---|---|
| Domain questions | 10–24 per page | Learning while reading; immediate feedback |
| Short drill | 30-minute timed run through a mock | Warm-up, or a focused re-test after revision |
| Full mock 1 | 50–60 items, recommended duration | Diagnostic before you study hard |
| Full mock 2 | 50–60 harder items, timed | Readiness gate before the real assessment |
| Retry incorrect only | Varies | Closing a specific gap; available after any submission |
Progress is saved in your browser, so a closed tab offers to resume the attempt.
Common self-deception patterns
Reading the explanation before committing to an answer
The explanation is the highest-value part of the item, and it is worthless if you read it while your answer is still undecided – recognition feels like knowledge. Commit, then expand. If you were unsure, mark the item and revisit it a day later.
Taking the same mock twice and calling the second score progress
The second sitting of the same bank measures recall of items, not competence. Use mock 1 for diagnosis and mock 2 for readiness, and put at least one revision cycle between them.
Studying only the heaviest domains
Domain weights tell you where the marks are, not where your gaps are. Weight your revision by weight × your error rate, not by weight alone.
Reading about hands-on tasks instead of doing them
The certification programme is explicitly built around demonstrating skills, with ChatGPT acting as tutor, practice environment and feedback loop. Every track here has a hands-on checklist for that reason. An hour of doing beats three hours of reading about doing.
Honest limits of this material
- Item counts, durations and weights on our mocks are our design choices, not published OpenAI parameters. The published parameters are exactly two: 10–20 items from a 50-item bank, and 80% to pass.
- Product facts drift. Model names, prices and feature availability were verified in September 2026; the appendix pages tell you where to re-check each one.
- Nothing here guarantees a badge, a certificate, or eligibility for any future OpenAI certification.
Last updated Sep 18, 2026