AI Leadership
AI Leadership · Mock Exam 1
A 50-item, domain-weighted independent mock exam for the OpenAI Academy AI Leadership course, with full explanations and a readiness indicator.
This is a full-length, domain-weighted independent mock exam for the AI Leadership track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment or practice exam. Use it as your diagnostic: sit it first to find your two weakest domains. All 50 items are new and do not repeat the domain-page questions.
Instructions
- Time: 60 minutes. This is our own suggested pacing for practice, not an official OpenAI limit.
- Items: 50, executive judgement scenarios. Each item states how many answers to select — most are single-answer, some are select-two.
- Selection: for single-answer items pick the one best option; for select-two items you must select both correct options and no others to score the item.
- No guessing penalty: answer every question. There is no deduction for a wrong answer.
- Target: aim for ≥ 80% raw before taking the real OpenAI Academy assessment — the 80% line matches the Academy badge threshold. Remember that Academy badges and pathway certificates of completion are not certifications.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| D1 | Opportunity and Value | 9 |
| D2 | Strategy and Roadmap | 9 |
| D3 | Governance and Risk | 9 |
| D4 | Adoption and Change Management | 8 |
| D5 | Workforce and Enablement | 7 |
| D6 | Measurement and ROI | 8 |
Total: 9 + 9 + 9 + 8 + 7 + 8 = 50 items.
Readiness interpretation
This is an independent readiness indicator, never an official score.
| Raw score | Band | What it means |
|---|---|---|
| under 70% | Keep learning | Revisit the heaviest domains (Opportunity, Strategy, Governance) before re-attempting |
| 70–79% | Building confidence | Close the gaps in your two weakest domains, then re-sit |
| 80–89% | Assessment ready | You are at or above the Academy badge threshold; tidy up remaining weak spots |
| 90%+ | Strong readiness | Broad command of the material; proceed to the Academy assessment with confidence |
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it. See the assessment model page for how the real Academy assessment differs.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A COO lists twelve candidate AI use cases and asks which to fund first. Nine promise 'efficiency', two promise 'innovation', one names a specific stated priority: cut invoice-processing cycle time by 30%. What should the leader anchor the shortlist on?
Show answer
Answer: B.
Value is judged by contribution to a stated priority, and only the invoice case names one with a measurable target, so it anchors the shortlist. A chases novelty over traceability. C spreads scarce attention across unqualified ideas and produces no learning. D optimises cost, not value, and would fund cheap-but-irrelevant work.
A sponsor claims an assistant will save 40,000 agent-hours a year, but the contact centre is already running below capacity and cannot reduce headcount this year. How should the benefit be characterised to the board?
Show answer
Answer: B.
Time saved in an under-utilised team is capacity, not cash, until it is deliberately redeployed; reporting it honestly protects credibility. A books a saving the P&L will never show. C promises cuts the plan explicitly rules out. D relabels capacity as revenue with no product or demand behind it.
Two initiatives compete for one slot: a customer-facing chatbot that could lift NPS but touches regulated advice, and an internal meeting-summary tool with modest but certain time savings and low risk. The organisation has no AI track record yet. Which is the stronger FIRST bet?
Show answer
Answer: B.
With no track record, sequencing a low-risk, quickly provable win earns the credibility and capability needed for regulated work later. A leads with the highest-risk option before the organisation can govern it. C freezes all progress pending perfection. D ignores capacity limits and dilutes the first proof point.
A leader is screening a proposed use case for whether it is a strong EARLY AI candidate. Which two attributes most strengthen the case? (Select two)
Show answer
Answer: A and B.
High volume makes benefit measurable and material, and an existing human review point makes early error rates safe to absorb. C removes the safeguard that makes early adoption defensible. D removes the baseline needed to prove benefit. E introduces a compliance blocker that disqualifies the case rather than strengthening it.
A business-unit head insists their idea 'will obviously pay for itself' and refuses to state assumptions. The leader must decide whether to advance it to a funding gate. What is the most defensible action?
Show answer
Answer: B.
A fundable case needs a transparent, driver-based estimate with assumptions and a range so it can be challenged and tracked; that is the gate's whole purpose. A funds on conviction alone. C discards a potentially good idea rather than pricing it. D fabricates evidence, which corrodes the entire portfolio's credibility.
An AI opportunity portfolio contains only large, multi-year transformation bets. A board member worries about the shape. What is the most balanced correction?
Show answer
Answer: A.
A healthy portfolio balances near-term provable wins against longer strategic bets, so it delivers learning and momentum without betting everything on slow maturation. B abandons strategic ambition. C worsens the concentration risk. D ignores the legitimate imbalance the board flagged.
During opportunity sizing, a sponsor's estimate assumes AI handles 100% of cases with zero review. Experienced staff say realistically 60% could be automated with review on the rest. How should the leader treat the gap?
Show answer
Answer: B.
Credible sizing uses the realistic automation share and accounts for residual human review, expressed as a range. A inflates the case on an unachievable assumption. C picks a number with no basis. D waits for a level of autonomy that may never be appropriate for the task.
A functional leader wants to pursue a flashy generative use case that competitors are publicising, but it maps to none of the company's four stated priorities. What should the AI leader recommend?
Show answer
Answer: B.
Discipline means initiatives earn a place by advancing a stated priority; one that traces to nothing should be parked. A lets competitor PR set the agenda. C reverse-engineers strategy to justify a pet project. D just hides the same misalignment under a different budget line.
A team reports that its AI pilot 'saved 20 minutes per report'. The leader wants to know whether that is real value. Which follow-up question matters MOST first?
Show answer
Answer: B.
Without a measured baseline, a per-report saving is an assertion, so the measurement method and pre-AI baseline come first. A matters, but a benefit with no baseline cannot be netted against cost meaningfully. C measures sentiment, not value. D is an implementation detail irrelevant to whether the benefit is real.
The executive team asks for an AI roadmap that reaches enterprise-wide scale in one quarter. Only one pilot has run and its benefit is not yet baselined. What sequencing should the leader defend?
Show answer
Answer: B.
Scale should be gated on proven, baselined benefit, so prove then scale then embed is the defensible sequence. A promises a timeline the evidence cannot support. C scales on sentiment. D embeds unproven capability everywhere, multiplying risk before any evidence exists.
Four functions each want to buy a different AI point tool. The CIO warns of overlapping spend, fragmented governance and duplicated integration. What is the BEST strategic response?
Show answer
Answer: B.
A platform-first assessment captures shared governance and integration economics while still allowing justified point tools. A produces the exact fragmentation the CIO flagged. C blocks all value while a monolith is chosen. D lets size, not fit, dictate the standard and may starve smaller functions of what they need.
A leader is deciding whether to BUILD a bespoke capability rather than buy a product. Which two conditions most justify building? (Select two)
Show answer
Answer: A and B.
Build when the capability is differentiating and must be owned, or when no product fits a durable need and the org can maintain it. C ignores total cost of ownership over the lifetime. D lets curiosity drive a strategic decision. E is a temporary procurement friction, not a reason to take on permanent build-and-maintain cost.
An AI strategy draft lists many initiatives but never states how they connect to the company's priorities or to each other. A director calls it 'a shopping list'. What is the most important fix?
Show answer
Answer: B.
A strategy, not a list, requires each initiative to trace to a priority and to sit in a deliberate sequence. A adds volume, not coherence. C is cosmetic. D deletes the very anchor that would turn the list into a strategy.
A CFO will fund AI only if the leader picks one funding model. The programme spans several years, needs disciplined stage-gates, and must not let unproven bets run unchecked. Which model fits best?
Show answer
Answer: B.
Stage-gated funding matches capital release to proven progress, giving the discipline the CFO requires. A hands over all capital before any proof. C removes the control entirely. D is an accounting trick that hides, rather than manages, cost.
Halfway through the year, a scaled initiative is missing its benefit target and its assumptions have proven wrong. The sponsor wants to press on 'because we've already invested'. What should the leader advise?
Show answer
Answer: B.
Good portfolio discipline judges the go-forward case on current evidence and ignores sunk cost. A is the sunk-cost fallacy. C throws more capital at a failing set of assumptions. D conceals information the organisation needs to reallocate.
A leader is asked to produce an initial AI strategy draft after the AI Leadership course. Which content most belongs in that draft?
Show answer
Answer: B.
An initial strategy draft aligns priorities, opportunities, sequencing, governance and measurement — exactly the leadership artefact the course targets. A and C are downstream execution detail, not strategy. D is background noise, not a plan.
Two departments propose nearly identical AI initiatives with separate budgets and vendors. Before funding, the leader wants to consolidate them well. Which two conditions should the consolidated initiative satisfy? (Select two)
Show answer
Answer: A and B.
Consolidation works when there is one owner and shared budget plus common governance and measurement, capturing shared learning and avoiding duplicate spend. C reintroduces the duplication being removed. D fragments the metrics that should now be shared. E inflates cost rather than consolidating it.
A leader must decide whether to build on a general-purpose AI platform or buy a narrow point solution for a single, unusual regulated workflow that the platform cannot handle. Everything else in the org is well served by the platform. What is the best call?
Show answer
Answer: B.
Platform-for-the-many plus point-tool-for-the-genuinely-distinct is the balanced pattern. A bends a critical regulated need to a tool that cannot meet it. C throws away platform economics. D abandons a required workflow to protect tidiness.
A programme dashboard shows only monthly licence counts and total prompts. The board asks whether these prove the programme is delivering value. What is the leader's most accurate response?
Show answer
Answer: B.
Licences and prompt counts measure access and activity, not value; value needs baselined benefit, quality-paired outcomes and net cost. A and C mistake access and volume for delivered value. D treats a rising vanity metric as sufficient evidence of return.
Staff report that summarising a public press release inside the sanctioned workspace still requires monthly risk-committee sign-off, so some have started pasting content into personal ChatGPT accounts. What is the ROOT problem a leader should fix?
Show answer
Answer: B.
Uniform heavy controls on low-risk tasks create shadow AI; the fix is tiering controls to risk so trivial tasks are frictionless. A punishes a symptom of bad design. C misdiagnoses speed for a policy problem. D bans a plainly safe activity.
A firm must ensure that when an employee leaves, all access to the AI workspace is removed automatically from the corporate identity system. Which enterprise control most directly delivers this?
Show answer
Answer: B.
SCIM automates provisioning and deprovisioning through the identity provider, removing access the moment the identity is disabled. A is a policy without enforcement. C protects stored data but not access lifecycle. D catches lapses late and manually, not automatically.
Legal fears that customer data entered into ChatGPT will be used to train OpenAI's models. In a properly configured Business or Enterprise workspace, what is the accurate position a leader should state?
Show answer
Answer: B.
Business and Enterprise workspaces do not train on business data by default, which is precisely why sanctioned workspaces beat personal accounts. A is false for those plans. C misstates the default as opt-out per session. D is simply incorrect for the governed plans.
A leader is drafting an AI risk taxonomy. Which two categories most belong in a working taxonomy for a large enterprise? (Select two)
Show answer
Answer: A and B.
Data/privacy and output-quality are core risk categories any taxonomy must cover. C is not an enterprise risk category. D is a competitive-strategy consideration, not a governance risk. E is not a risk at all.
A global firm must ensure EU customers' regulated data is processed within the EU. The leader is choosing an approach. Which capability makes this achievable in a governed workspace?
Show answer
Answer: B.
Regional data residency lets an organisation keep processing within a specified region such as the EU. A is a performance/cost feature. C confuses working hours with data location. D is unrelated to where data is processed.
A newly formed AI governance committee insists on reviewing every individual prompt use case before approval, creating a growing backlog. Adoption is stalling. How should the leader reshape governance?
Show answer
Answer: B.
Governance should enable by tiering decision rights and pre-approving low-risk patterns, reserving scrutiny for genuine risk. A scales a broken process. C removes needed controls. D abandons risk management altogether.
A regulated organisation wants staff to use AI but must enforce that only approved user groups can access certain sensitive projects. Which control set most directly enforces this?
Show answer
Answer: A.
RBAC plus SSO ties access to verified identity and role, enforcing who can reach sensitive projects. B is an insecure anti-pattern. C is guidance without enforcement. D affects model behaviour, not access.
The audit team asks how the organisation will demonstrate, months later, exactly who accessed which AI workspace data and when. Which enterprise capability answers this?
Show answer
Answer: B.
Compliance and audit logs provide the durable, reviewable record auditors require. A shows usage counts, not access events. C is a model property. D is anecdotal and not a control.
A vendor pitches an AI agent capability but the leader learns it runs only with US data residency and no zero-data-retention option. The workload involves EU-regulated data that must not leave the region. What should the leader conclude?
Show answer
Answer: B.
Data-residency and retention constraints are hard fit-checks; a workload that must stay in-region cannot use a US-only, no-ZDR service. A treats a compliance boundary as negotiable. C lets capability override compliance. D is unrealistic and misplaces the responsibility.
Six months after a company-wide launch email, AI usage is high in the analytics team and near zero everywhere else. The sponsor is surprised. What does this pattern indicate?
Show answer
Answer: B.
Early enthusiasts adopt from an announcement, but the mainstream needs role-specific enablement to cross the chasm. A blames the tool for a change-management gap. C invents a motive with no evidence. D is a trivial explanation that does not fit high enthusiast usage.
Usage has plateaued after a rollout. The leader can invest in only one lever. Which has the STRONGEST effect on getting the mainstream majority to adopt?
Show answer
Answer: B.
Manager modelling and embedding into routines is the strongest driver of mainstream adoption. A repeats a message that already failed to move the majority. C buys access, not behaviour. D rewards volume, which invites gaming rather than valuable use.
A programme reports success as '6,000 licences issued' to the board. The leader is uneasy. Why is this metric misleading as a measure of success?
Show answer
Answer: B.
Licences count access; adoption is measured by active, valuable use, and unused seats inflate the picture. A is about cost, not the metric's validity. C misreads scale. D invents a licensing rule irrelevant to the point.
A leader is designing a champion network to drive adoption. Which two design choices most improve its effectiveness? (Select two)
Show answer
Answer: A and B.
Embedded, credible champions plus real time, recognition and support make a network effective. C removes the peer credibility that makes champions work. D hides the very people colleagues should turn to. E destroys the context champions need to help.
A staff survey finds many people say 'I don't know what I'd use it for in my role'. Usage is low. What is the BEST response?
Show answer
Answer: B.
Concrete, role-specific examples remove the 'what would I use it for' barrier directly. A repeats generic messaging that already failed. C forces activity without relevance, inviting gaming. D leaves the barrier in place.
An enthusiastic early cohort loves the tool, but the leader notices the same people are producing all the reported wins while most licence-holders never log in. What is the most useful next move?
Show answer
Answer: B.
Value at scale comes from moving the majority, so enablement should target the un-engaged. A mistakes early-adopter success for organisational adoption. C punishes the people succeeding. D deepens the divide rather than closing it.
A manager says AI adoption is 'not my job — that's IT's programme'. Team usage is the lowest in the company. How should the leader address this?
Show answer
Answer: B.
Adoption is a line-management responsibility; making it an explicit, supported expectation aligns the strongest lever. A treats a leadership issue as a tooling issue. C works around the manager rather than engaging the key influence. D gives up on a solvable problem.
The board wants a single number to track whether the AI rollout is 'working'. The leader must recommend something better than licences issued. Which is the strongest single adoption signal?
Show answer
Answer: B.
Recurring, work-embedded active use across the target population reflects real adoption. A counts access. C rewards volume that may be noise. D measures a one-off event, not sustained use.
A company plans identical two-hour AI training for all 2,500 staff regardless of role. The leader questions the design. What is the BEST improvement?
Show answer
Answer: B.
Role- and level-segmented enablement teaches each group what its work actually needs, which drives real capability. A enforces a one-size-fits-none design. C trims duration without fixing relevance. D relies on unmanaged cascade that usually decays.
A manager wants to announce the team is 'OpenAI certified' after everyone earned the Academy AI Foundations badge. What is the accurate position the leader must insist on?
Show answer
Answer: B.
OpenAI states plainly that Academy badges and pathway certificates of completion are not certifications; the claim would be inaccurate. A, C and D all misrepresent a badge as a certification in different ways, which the leader must correct.
A leader is choosing training formats to produce durable behaviour change, not just awareness. Which two formats are strongest? (Select two)
Show answer
Answer: A and B.
Applied practice on real tasks and sustained peer reinforcement change behaviour. C, D and E are all one-off, passive awareness activities that rarely alter how people work.
As AI drafts routine documents faster, a leader must redesign an analyst role. Which capability should the redesigned role most emphasise?
Show answer
Answer: B.
Role redesign shifts humans toward review, judgement and higher-value exception work. A optimises a task AI now does. C manufactures redundant work. D refuses the change entirely.
Morale has dropped and staff are asking 'are they going to replace us?'. Leadership must respond. What is the MOST effective response?
Show answer
Answer: B.
Honest framing of role change plus concrete reskilling support builds trust and engagement. A makes a promise that will be broken. C leaves a vacuum that rumour fills. D is corrosive and drives out talent.
A leader wants to route employees onto appropriate OpenAI Academy learning. A team of API developers and a team of business analysts need different paths. What is the best routing principle?
Show answer
Answer: B.
Matching learning paths to role and product context is the point of a capability-based enablement plan. A wastes time on irrelevant depth. C produces uneven, unaligned capability. D misroutes technical staff into a strategy course.
Finance asks the leader to justify a training budget when 'the tools are intuitive'. What is the strongest justification?
Show answer
Answer: B.
Enablement is what converts licence spend into realised value; skipping it wastes the larger investment. A invents a compliance mandate. C argues by imitation. D trivialises training as filler.
A team deployed an AI assist and now wants to prove it saved time, but no measurement was taken before launch. What is the core problem?
Show answer
Answer: B.
A benefit claim needs a baseline; without a before-measurement there is nothing to compare against. A is a minor detail. C is tooling, not the measurement gap. D measures sentiment, not the time saving in question.
A leader wants leading indicators to watch early, before financial results arrive. Which two are LEADING indicators for an AI initiative? (Select two)
Show answer
Answer: A and B.
Recurring use and task-time reduction show up early and predict later results, so they are leading indicators. C, D and E are lagging outcomes reported long after and shaped by many factors beyond the initiative.
A support team's average handle time dropped 30%, but the rate of reopened tickets rose sharply over the same period. How should the leader judge the result?
Show answer
Answer: B.
Cycle-time gains must be paired with quality; rising reopens can cancel or reverse the apparent benefit. A reports a partial metric as a win. C dismisses a valid quality signal. D doubles down on the metric that may be causing the harm.
A pilot 'was basically free' on a trial, but at production scale the bill grew. Which two cost drivers must the leader model before scaling? (Select two)
Show answer
Answer: A and B.
Usage-based consumption and per-seat licensing are the drivers that scale the bill with volume and headcount. C, D and E are irrelevant to unit economics.
A region grew revenue 12% in the same period it adopted AI, but it also hired two senior reps and the whole market rose. The leader must report to the board. What is the MOST honest attribution approach?
Show answer
Answer: B.
Honest attribution isolates confounders and reports a defensible range with assumptions. A overclaims by ignoring hiring and market lift. C under-reports and hides real benefit. D forfeits the ability to steer the programme with evidence.
The board asks for a single credible line on an AI initiative's return. Which board line best reflects disciplined measurement?
Show answer
Answer: B.
A credible board line cites a baseline, a range, conservative attribution and net run cost. A is a hero claim with no evidence. C substitutes sentiment for value. D counts access, not return.
A leader reports gross benefit from an AI programme but omits the licence, usage and enablement costs. The CFO objects. What is the correct fix?
Show answer
Answer: B.
ROI must net the full run cost against gross benefit to be honest. A overstates return by hiding cost. C is not a real option and would misstate accounts. D swings to the opposite distortion and hides the benefit.
Last updated Sep 18, 2026