AI Leadership
AI Leadership · Mock Exam 2
A harder, 50-item domain-weighted independent mock exam for the OpenAI Academy AI Leadership course, used as a readiness gate before the Academy assessment.
This is the harder of the two independent mock exams for the AI Leadership track — more multi-constraint stems and more FIRST / BEST / MOST cost-effective / TWO qualifiers. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment or practice exam. Use it as your readiness gate: reach 80%+ here before you book the real OpenAI Academy assessment. All 50 items are new and are distinct from Mock Exam 1 and the domain-page questions.
Instructions
- Time: 60 minutes, our own suggested pacing for practice, not an official OpenAI limit.
- Items: 50 executive judgement scenarios, deliberately more demanding than Mock Exam 1. Each item states how many answers to select.
- Selection: for single-answer items pick the one best option; for select-two items you must select both correct options and no others to score the item.
- No guessing penalty: answer every question. There is no deduction for a wrong answer.
- Target: aim for ≥ 80% raw before taking the real OpenAI Academy assessment; the 80% line matches the Academy badge threshold. Academy badges and pathway certificates of completion are not certifications.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| D1 | Opportunity and Value | 9 |
| D2 | Strategy and Roadmap | 9 |
| D3 | Governance and Risk | 9 |
| D4 | Adoption and Change Management | 8 |
| D5 | Workforce and Enablement | 7 |
| D6 | Measurement and ROI | 8 |
Total: 9 + 9 + 9 + 8 + 7 + 8 = 50 items.
Readiness interpretation
This is an independent readiness indicator, never an official score.
| Raw score | Band | What it means |
|---|---|---|
| under 70% | Keep learning | The harder framing has exposed real gaps; revisit the weak domains before re-attempting |
| 70–79% | Building confidence | Close your two weakest domains, then re-sit this harder set |
| 80–89% | Assessment ready | At or above the Academy badge threshold on the harder mock; you are in good shape |
| 90%+ | Strong readiness | Strong command under multi-constraint pressure; proceed to the Academy assessment |
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it. If you have not sat Mock Exam 1 yet, start there as your diagnostic.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A CEO gives the leader one funding slot and three candidates: a high-value regulated use case with no review capacity yet, a medium-value low-risk case with a clear baseline, and a low-value flashy demo competitors are running. Constraints: build credibility, prove measurable value, stay within risk appetite. Which should be funded FIRST?
Show answer
Answer: B.
Only the medium, low-risk, baselined case satisfies value, risk appetite and credibility-building simultaneously, which is what FIRST demands under all three constraints. A breaks the risk-appetite and review-capacity constraints. C fails the measurable-value test. D forfeits all near-term learning while waiting.
Two sponsors present benefit cases. Sponsor A claims £5m from full automation with no assumptions listed. Sponsor B claims £1.2m–£1.8m with a stated automation rate, a counterfactual and residual review costs. The board asks which case the leader trusts MORE for funding. What is the best answer?
Show answer
Answer: B.
Credibility, not magnitude, decides trust; B's transparent range with counterfactual can be challenged and tracked. A is a hero number with no basis. C rejects the legitimate practice of estimating. D blends an unsupported figure into the credible one, contaminating it.
A leader must separate cashable savings from freed capacity in a benefit case. Which two statements are TRUE about how to treat each? (Select two)
Show answer
Answer: A and B.
Cashable savings hit the budget once realised; freed capacity only becomes value on redeployment. C books capacity as cash the P&L never sees. D conflates two fundamentally different benefit types. E fabricates a multiplier with no basis.
A leader is pressured to fund the initiative with the single largest headline number. It concentrates the whole portfolio in one multi-year bet with unproven assumptions. Which reasoning should MOST guide the decision?
Show answer
Answer: B.
Risk-adjusted value and portfolio balance beat a raw headline; an all-in unproven bet is fragile. A ignores probability and risk. C removes the learning and momentum that diversification provides. D decides by seniority, not merit.
An opportunity looks strong on paper but the only data that would make it work is data the organisation is contractually barred from using for AI. What is the FIRST thing the leader should do?
Show answer
Answer: B.
A binding data restriction is a disqualifying constraint that must be resolved before value is even discussed. A invites contractual and reputational harm. C bets the case on an uncertain future renegotiation. D risks breaching the restriction with unverified anonymisation.
A leader must decide the MOST cost-effective way to test whether a high-uncertainty opportunity is worth pursuing at all. Time and budget are tight. What is the best approach?
Show answer
Answer: B.
A small, baselined, time-boxed pilot is the cheapest way to generate a go/no-go decision on an uncertain idea. A spends enterprise cost before proving anything. C is the most expensive path for a high-uncertainty bet. D discards evidence entirely.
The value of a proposed initiative is real but small, while a lower-scoring alternative would unlock learning that many future initiatives depend on. Capacity allows only one this quarter. What consideration should tip the decision?
Show answer
Answer: B.
Strategic value includes option value — learning that unlocks a downstream pipeline can outweigh a small standalone benefit. A ignores that pipeline effect. C decides by politics. D under-funds both to the point neither proves anything.
A leader reviews a benefit estimate that quietly assumes the freed hours convert to revenue at the team's full billing rate, though the team is salaried and not billable. What is the MOST accurate correction?
Show answer
Answer: B.
Freed salaried time is valued at cost or redeployment value, not an unearned billing rate. A applies a rate the team never charges, inflating the case. C compounds the error. D over-corrects and hides a genuine, if modest, benefit.
A leader must present the single most defensible sizing artefact to a sceptical CFO for a mid-sized initiative. Which artefact best withstands challenge?
Show answer
Answer: B.
A transparent driver-based model with assumptions, counterfactual and range is the most challengeable and therefore most defensible artefact. A hides all logic in one number. C imports external data with no local validity. D inherits a vendor's optimistic assumptions.
The board wants scale next quarter; the leader has one pilot with strong user enthusiasm but no baselined benefit, and a second pilot with a modest but baselined and positive benefit. Which is the correct sequencing decision under a prove-scale-embed model?
Show answer
Answer: B.
Scale is gated on proven, baselined benefit, so only the baselined-positive pilot qualifies now. A scales on enthusiasm. C scales an unproven pilot to hit a date. D embeds unmeasured work everywhere before any evidence exists.
A leader is arbitrating platform versus point-tool decisions across the enterprise. Which two conditions most favour buying a separate POINT tool rather than extending the platform? (Select two)
Show answer
Answer: A and B.
Point tools are justified for genuine, durable, distinct needs the platform cannot serve without harm. C is a preference, not a requirement. D duplicates a capability that already exists. E adds fragmentation on principle rather than need.
A three-year custom-build has consumed significant budget, has no production use, and now traces to none of the current stated priorities after a strategy refresh. The sponsor cites the money already spent. What should the leader recommend?
Show answer
Answer: B.
The go-forward decision ignores sunk cost, and work tracing to no priority should not be funded further. A is the sunk-cost fallacy. C escalates commitment to a misaligned bet. D corrupts the strategy to protect the project.
A CFO offers two funding structures for a multi-year AI programme with uncertain later phases: (1) full multi-year budget approved now, or (2) stage-gated release tied to proven benefit. The leader wants both discipline and momentum. Which is the BEST choice and why?
Show answer
Answer: B.
Stage-gated funding keeps discipline over uncertain later phases while still funding earned momentum. A commits all capital before uncertain phases prove out. C is an accounting deferral, not a funding model. D removes the very tracking that would justify continued spend.
A strategy draft is strong on opportunities and roadmap but silent on governance and measurement. A director says it will fail review. Which combination of additions MOST completes it as a leadership artefact?
Show answer
Answer: B.
Governance and measurement are the missing pillars of a complete strategy; adding both ties the draft together. A adds volume to a section already strong. C is execution detail. D is communications, not strategy substance.
Two business units each demand to own the enterprise AI platform decision. Delay is now blocking every downstream initiative. What is the BEST governance response for the strategy?
Show answer
Answer: B.
Clear decision rights with a named owner unblock the enterprise; strategy needs someone accountable to decide. A leaves the blockage unresolved. C creates the fragmentation the enterprise choice was meant to avoid. D overloads the board and slows everything.
A leader must decide whether to build, buy or partner for a capability that is important but not differentiating, where a mature product exists and no internal team can maintain a bespoke system. What is the BEST call?
Show answer
Answer: B.
Buy when the capability is non-differentiating, a product fits, and the org cannot maintain a build. A takes on maintenance the org cannot sustain for no differentiation. C over-engineers a commodity. D blocks value waiting to build what should be bought.
The roadmap embeds AI deeply into a core process before the pilot benefit is baselined, on the sponsor's confidence. A risk officer objects. Which principle most supports the objection?
Show answer
Answer: B.
Embed comes last because deep integration multiplies risk as well as value and should follow proof at scale. A ignores the risk multiplier of premature embedding. C substitutes authority for evidence. D delays measurement past the point it can guide the decision.
A leader is told the strategy must fix the exact models and prices for the next three years to give teams certainty. Given how fast the model lineup changes, what is the BEST way to handle this in the strategy?
Show answer
Answer: B.
Durable strategy states selection principles and a review cadence rather than binding to fast-changing model specifics. A locks in facts that will be stale within months. C omits guidance teams need. D wastes cost and still hard-codes a choice.
After tightening controls, shadow use of personal AI accounts rises, not falls. Investigation shows every task — even trivial public-data summaries — now needs committee approval. What is the MOST effective governance change?
Show answer
Answer: B.
Shadow AI grows when governance over-controls low-risk work; tiering controls removes the incentive to bypass it. A treats symptoms with surveillance while the cause remains. C scales a mis-designed process. D swings to no control, which is unacceptable for higher-risk work.
A leader must map two requirements to enterprise controls: (1) automatically remove access when staff leave, and (2) keep EU-regulated data processed in the EU. Which two control choices correctly satisfy them respectively? (Select two)
Show answer
Answer: A and B.
SCIM automates deprovisioning and regional residency keeps processing in-region — each maps to its requirement. C is a performance feature unrelated to access. D confuses model capacity with data location. E is unenforced policy, not an automatic control.
A leader must decide whether a proposed control set is 'enabling governance' or 'governance theatre'. The set requires a full risk-committee review for every prompt, produces long queues, and has never rejected a case. Which judgement is MOST accurate?
Show answer
Answer: B.
Reviewing everything with no triage and no rejections adds friction without managing risk — the definition of theatre. A mistakes activity for control. C tunes cadence, not design. D reframes a bug (slow adoption) as a feature.
A vendor offers a durable cloud-agent capability but only with US data residency and no zero-data-retention option. The organisation has one workload that must stay in the EU and another with no residency constraint. What is the BEST governance decision?
Show answer
Answer: B.
Match the capability to workloads whose constraints it satisfies: the unconstrained one, not the EU-residency one. A puts EU-regulated data into a US-only, no-ZDR service. C directly violates the residency requirement. D forgoes value where the capability is perfectly compliant.
A leader is assembling a risk taxonomy and wants the category that most directly covers the danger of a confident but wrong AI output driving a business decision. Which category is it?
Show answer
Answer: B.
A confident-but-wrong output reaching a decision is squarely output-quality (accuracy/hallucination) risk. A concerns where data is processed. C concerns supplier dependence. D concerns spend, not correctness.
Legal, worried about training on inputs, proposes banning AI outright. The leader knows the sanctioned Business/Enterprise workspace does not train on business data by default and offers SSO, RBAC and audit logs. What is the MOST balanced recommendation?
Show answer
Answer: B.
The governed workspace's default no-training plus SSO, RBAC and audit logs directly address the concern while enabling value. A drives shadow use and forfeits benefit. C is the ungoverned path legal fears. D removes a control the organisation needs for assurance.
An organisation must both restrict sensitive projects to approved groups and be able to prove after the fact who accessed what. Which pairing of controls MOST directly meets both needs together?
Show answer
Answer: B.
RBAC/SSO enforces access and audit logs provide the after-the-fact record — together they meet both needs. A is unrelated to access or audit. C is insecure and unauditable. D addresses neither enforcement nor evidence.
A leader is asked to approve an agentic workload that will act with limited human oversight on customer records. Which condition should the leader treat as a hard prerequisite before approval?
Show answer
Answer: B.
Higher-autonomy work on sensitive records requires data-handling, residency/retention and risk-appropriate oversight as prerequisites. A is a model choice, not a control. C is a benefit, not a safeguard. D is not evidence of governance readiness.
The governance model routes every AI decision — trivial or critical — to a single central committee, and it is now the bottleneck for the whole programme. What redesign BEST preserves control while removing the bottleneck?
Show answer
Answer: B.
Tiered delegation — local owners and pre-approved patterns for low risk, central review for high risk — preserves control and removes the bottleneck. A speeds a centralised choke point without fixing it. C removes control. D introduces inconsistent, unaccountable decisions.
A rollout crossed the early enthusiasts but stalled at roughly a third of the target population. Two options remain within budget: a second broadcast launch email, or manager-led embedding with role-specific use cases. Which is MOST likely to move the mainstream, and why?
Show answer
Answer: B.
The chasm is crossed by manager modelling and role relevance, not by repeating a broadcast that already failed. A repeats an approach the mainstream ignored. C concedes defeat prematurely. D rewards volume and invites gaming rather than valuable use.
A leader is choosing metrics to judge whether adoption is genuinely progressing across the organisation. Which two are the STRONGEST signals? (Select two)
Show answer
Answer: A and B.
Regular use in core workflows and breadth across functions are the strongest adoption signals. C counts access. D counts a one-off open. E rewards volume that may be meaningless activity.
A programme celebrates '6,000 licences issued' and '2 million prompts submitted' as proof of success. The CEO asks the leader whether the programme is truly working. What is the MOST accurate answer?
Show answer
Answer: B.
Access and raw activity counts cannot confirm valuable, recurring, work-embedded use. A and C mistake access and volume for adoption. D misjudges scale; the issue is the metric type, not the magnitude.
A leader inherits a champion network that is all senior executives, meets quarterly, and has no time allocated. Adoption in the field is flat. What single change would MOST improve it?
Show answer
Answer: B.
Effective champions are embedded peers with real time and recognition, close to the daily work. A deepens the wrong composition. C changes cadence, not credibility or capacity. D is a relabel with no substance.
Managers in one division treat AI adoption as optional and never use it themselves; their teams have the lowest usage. Another division's managers use it daily and expect it in routines; their teams lead adoption. What does this MOST strongly imply for the leader's plan?
Show answer
Answer: B.
The contrast points to manager modelling and expectation as the primary lever, so it should anchor the plan. A ignores a clear pattern. C attributes to talent what the evidence attributes to leadership behaviour. D treats a management issue as a tooling mandate.
A survey shows two barriers: 'I don't know what to use it for' (majority) and 'I don't trust the output' (a vocal minority). Budget allows one primary intervention this quarter. Where should the leader focus FIRST for the biggest adoption gain?
Show answer
Answer: B.
The majority barrier — 'what would I use it for' — gates mainstream adoption, so a role-specific use-case library yields the biggest gain first. A serves the loudest minority, not the largest blocker. C delays action. D adds access nobody is using.
A leader must design enablement for 2,500 staff with very different roles, under a fixed budget. Which design gives the BEST capability return per pound spent?
Show answer
Answer: B.
Segmented, applied paths concentrate spend on the capability each role actually needs, maximising return. A spends on irrelevance for most. C wastes time on unneeded depth. D leaves capability to chance.
A director wants marketing to say the workforce is 'OpenAI certified' and 'certification-ready' after completing Academy pathways. Which two statements must the leader insist are TRUE and used instead? (Select two)
Show answer
Answer: A and B.
OpenAI states badges and pathway certificates are not certifications; the accurate claim is course/assessment completion earning badges. C contradicts OpenAI's explicit wording. D still calls a badge a certification. E overstates a completion badge as a validated exam.
As AI automates routine drafting, a leader redesigns a knowledge-worker role. Which redesign MOST protects value and morale together?
Show answer
Answer: B.
Redesigning toward review, exceptions and judgement, backed by reskilling, protects both value and morale. A cuts capability without redesign. C refuses the change and forgoes value. D manufactures make-work that fools no one.
Anxiety is spreading: staff hear rumours of layoffs 'because of AI'. Leadership has not decided any cuts. What is the MOST effective communication approach?
Show answer
Answer: B.
Honest, early communication with concrete reskilling and a realistic picture builds trust without over-promising. A lets rumour fill the silence. C makes an unkeepable promise. D dismisses legitimate concern and erodes trust.
A leader must route three groups to appropriate OpenAI Academy learning: API developers, workplace users, and Codex-using engineers. Which routing is MOST role-appropriate?
Show answer
Answer: B.
Matching each group to the product-relevant Academy track is the core of capability-based routing. A misroutes technical staff into a strategy course. C produces unaligned capability. D optimises time over relevance.
Finance will fund enablement only if the leader shows it is not 'nice-to-have'. Which argument is the STRONGEST link between enablement and realised value?
Show answer
Answer: B.
Enablement is the mechanism that turns licence spend into the adoption the business case depends on, protecting the larger investment. A is a soft benefit, not the value link. C argues by convention. D trivialises it as administration.
A team wants to claim a 25% time saving from an AI tool but only started measuring after launch, comparing this month to last month. Multiple process changes also happened. What is the MOST defensible way to salvage a credible claim?
Show answer
Answer: B.
A credible claim needs a proper baseline, controls for confounders, and a conservative range. A ignores confounding process changes. C over-corrects and discards recoverable evidence. D inflates an already-suspect number.
A leader must pick the two metrics that best show, early and honestly, whether an AI initiative is on track before annual financials land. Which two? (Select two)
Show answer
Answer: A and B.
Recurring use and quality-paired cycle-time are early, honest signals of progress. C is a lagging annual outcome. D is sentiment, not measurement. E rewards volume, not value or quality.
An AI assist cut average handling time 30%, but customer satisfaction and reopened-ticket rates worsened. Finance still wants to book the full time-saving as benefit. What should the leader insist on?
Show answer
Answer: B.
Speed and quality must be judged together; worsening quality can erase or reverse a time-saving. A books a gross figure that quality losses may cancel. C dismisses a real quality signal. D worsens the quality problem it should address.
A pilot ran on a free trial. Before approving scale, the leader must model the true run cost. Which two cost drivers MOST determine the at-scale bill? (Select two)
Show answer
Answer: A and B.
Consumption charges and per-seat licensing scale with volume and headcount, driving the at-scale bill. C is a sentiment count. D and E are irrelevant to unit economics.
A region using AI grew 15%, but it also launched a new product, gained two competitors' customers, and rode a rising market. The leader must brief the board on AI's contribution. What is the MOST honest framing?
Show answer
Answer: B.
Honest attribution decomposes growth, names confounders, and reports a conservative, assumption-backed range. A overclaims by ignoring the product launch and market lift. C hides genuine benefit. D offers a vague claim that cannot be defended or steered by.
The board wants a single quarterly line that captures return without overclaiming. Which construction is BEST?
Show answer
Answer: B.
A defensible board line states baseline, net range, conservative attribution and run cost. A is a hero claim. C is a competitive boast, not a return figure. D infers return from adoption without measuring benefit or cost.
A leader must justify continuing an adoption-support budget when a finance partner argues the tool 'sells itself'. Usage sits at a third of the target population and has plateaued. Which argument is MOST compelling?
Show answer
Answer: B.
A plateau at a third signals the mainstream has not crossed the chasm, and active enablement is what moves them — the budget's purpose. A argues by convention. C blames the tool for a change gap. D abandons the majority to a level that will not rise on its own.
A sponsor presents only leading indicators after two quarters and asks the board to declare victory. Before any success claim, which two additional evidence types should the leader require? (Select two)
Show answer
Answer: A and B.
A proven success needs lagging outcomes tied to the priority and net benefit after cost with conservative attribution. C is a raw activity count. D is sentiment, not outcome evidence. E measures access, not delivered value.
Two initiatives report the same gross benefit. Initiative X uses a high-cost model at high volume; initiative Y uses a low-cost model for the same task quality. The leader must advise on portfolio value. What matters MOST?
Show answer
Answer: B.
With equal gross benefit and quality, the lower run cost makes Y the higher net-value choice — unit economics decide. A ignores cost. C equates model capability with value even when quality is matched at lower cost. D denies the clear cost difference.
A leader launches AI across ten country offices at once. Three months later, usage is strong in two offices whose leaders actively use and expect the tool, and near zero in the rest. What is the FIRST corrective move?
Show answer
Answer: B.
The success/failure split tracks whether local leaders model and expect use, so engaging low-usage office leaders is the first fix. A repeats a broadcast the mainstream ignores. C abandons offices without addressing the leadership gap. D rewards the already-adopting sites and ignores the laggards.
A leader must build a capability map before routing training. A team's members range from AI-sceptical novices to a few power users already building workflows. What is the BEST enablement design for this single team?
Show answer
Answer: B.
Capability mapping tailors enablement to current levels and can turn power users into peer leaders. A bores and under-serves the advanced members. C loses the novices. D relies on unstructured cascade that usually decays and leaves novices unsupported.
Last updated Sep 18, 2026