AI Leadership
D1 · Opportunity and Value
Finding credible AI opportunities across the value chain, sizing benefit honestly, distinguishing cost-out from revenue and risk plays, balancing a portfolio, and killing weak ideas early.
This domain carries 18% of the mock — roughly 9 of 50 items — and it is where most real AI programmes go wrong before a single tool is bought. It tests whether you can look across a value chain, name where generative AI actually creates value, size that value in numbers a CFO will not laugh at, and — just as important — decide which ideas to stop. The standalone Academy AI Leadership course (180 minutes, the longest in the catalogue) frames leadership as setting priorities, aligning stakeholders and producing an initial AI strategy draft; opportunity identification is the raw material for all three.
What you need to know
Opportunity work is triage, not enthusiasm. Every function will bring you a wish list; your job is to convert vague “AI could help here” into a small set of framed bets with an owner, a benefit type, a rough size and a kill condition. Value comes in three flavours — cost-out (do the same work with less time or fewer errors), revenue (win, retain or grow customers) and risk (reduce loss, fraud, downtime or compliance exposure) — and each is measured and defended differently. A credible portfolio mixes quick cost-out wins that build belief with a few higher-variance revenue and risk bets. The discipline that separates leaders from tinkerers is killing bad ideas cheaply and early, before they consume a year of a team’s calendar.
Learning objectives
By the end of this page you should be able to:
- Map where generative AI creates value across a value chain, distinguishing high-volume/high-variance work from work AI cannot yet own.
- Classify an opportunity as a cost-out, revenue or risk play and choose the right benefit measure for each.
- Size a benefit credibly using a driver-based estimate with explicit assumptions and a confidence band.
- Balance a portfolio across effort, value, time-to-impact and risk rather than chasing one big bet.
- Kill weak opportunities early using a written set of disqualifiers, and defend that decision to their sponsors.
- Frame an opportunity on a single page a steering group can approve or reject.
1.1 Where AI actually creates value
Generative AI does not create value uniformly. It concentrates in work that is language-shaped, high-volume, and tolerant of a review step. Walk the value chain function by function and ask three questions of each activity: is the input mostly text, code or images? does it happen thousands of times? can a human check the output before it matters?
| Value-chain area | Where AI lands well | Where it does not (yet) |
|---|---|---|
| Marketing & sales | Drafting, personalisation, research synthesis, call summaries | Final brand voice sign-off, high-stakes contract terms |
| Customer service | Deflection, draft replies, knowledge search, agent assist | Regulated advice, emotionally sensitive escalations without a human |
| Operations & supply | Document extraction, exception triage, SOP drafting | Physical process control, safety-critical decisions |
| Finance | Narrative reporting, reconciliation prep, policy Q&A | Sign-off on filed numbers, audit judgement |
| Legal & compliance | First-pass review, clause search, policy drafting | Legal opinions, regulatory positions of record |
| Engineering | Code drafting, review assist, migration, docs (see the Codex track) | Architecture decisions of record, production deploys without gates |
| HR & enablement | JD drafting, training content, policy Q&A | Hiring/firing decisions, anything touching protected characteristics |
Value density of an activity high │ ████ contracts drafting ████ support draft replies │ ████ research synthesis ████ document extraction │ ──────────────────────────────────────────────────────── low │ ░░ physical control ░░ audit sign-off ░░ legal opinions └───────────────────────────────────────────────────────────► high volume · text-shaped · reviewable → AI-friendly low volume · physical · irreversible → keep human-ownedAssessment signal
Stems that describe “high volume”, “repetitive”, “text-heavy”, “drafts a human then checks” point to a strong AI opportunity. Stems with “irreversible”, “safety-critical”, “regulated advice of record”, “physical” point away from full automation and toward assist-with-review or no play at all.
1.2 The three kinds of value: cost-out, revenue, risk
Every opportunity resolves to one dominant benefit type. Naming it wrong is the most common sizing error, because you then measure the wrong thing and cannot defend the number.
| Benefit type | The claim | Primary measure | How finance stress-tests it |
|---|---|---|---|
| Cost-out | Same output, less time or fewer errors | Hours saved × loaded rate; error-rework cost avoided | “Do those hours leave the P&L or just move?” |
| Revenue | More won, retained or expanded | Incremental conversion, retention, deal size | “What is the counterfactual — would this have happened anyway?” |
| Risk | Less loss, fraud, downtime, exposure | Expected loss avoided = probability × impact | “Is the risk real and are you double-counting insurance?” |
The hardest honesty test is cost-out: time saved is not money saved unless the time is redeployed or removed. Five minutes saved per ticket across an under-utilised team is capacity, not cash, until you either raise volume or reduce headcount — and the second is a workforce decision (see D5), not an automatic saving.
1.3 Sizing a benefit credibly
Board members have seen inflated AI business cases and discount them reflexively. Credibility comes from a driver-based estimate — a small chain of observable quantities — plus stated assumptions and a range, never a single hero number.
Claim: AI-assisted first-draft replies in customer support.
Volume : 40,000 tickets / monthCoverage : 60% are draftable → 24,000Time saved : 4 min → 2 min per ticket = 2 min savedRate : £25 / hour loadedGross saving: 24,000 × 2/60 × £25 = £20,000 / monthAdjust : 30% of saved time redeployed to backlog (real), 70% is slackCredible : ~£6,000 / month cashable, ~£14,000 capacityRange : ±40% pending a 6-week pilotClaim: faster, personalised proposal drafting shortens sales cycle.
Baseline win rate : 22%Hypothesis : +2pp from faster, tailored proposalsDeals / quarter : 300 → +6 winsAvg deal : £18,000Gross uplift : £108,000 / quarterCounterfactual : assume 40% would have closed anyway → net +£65,000Confidence : low until an A/B test on 2 regions runsThe move that earns trust is the counterfactual line and the confidence band. A leader who volunteers “40% would have closed anyway” is believed on the remaining 60%; a leader who claims the full number is discounted entirely.
Assessment signal
When a stem gives you a benefit number with no assumptions, no range and no counterfactual, the correct answer is almost always to demand the driver-based breakdown before funding — not to approve or reject on the headline.
1.4 Portfolio balance
One big bet is a career risk and a programme risk. Leaders assemble a portfolio balanced on effort, value, time-to-impact and uncertainty, so that early wins fund belief while a few larger bets carry the upside.
VALUE ▲ high │ Major bets Transformative │ (revenue, risk; (rare; long, │ scale carefully) board-sponsored) │ │ Quick wins Fill-ins low │ (cost-out; ship (nice-to-have; │ first, build trust) defer) └──────────────────────────────────────► EFFORT low highA healthy first-year portfolio is roughly 60% quick cost-out wins, 30% scaled revenue/risk plays, 10% exploratory — enough momentum to keep sponsors funding, enough ambition to matter. The strategy and roadmap domain turns this portfolio into a sequenced plan.
1.5 Killing bad ideas early
The most valuable leadership behaviour in this domain is stopping work. Every opportunity should carry disqualifiers written before enthusiasm sets in, so the kill decision is a policy, not a personality clash.
| Disqualifier | Kill signal |
|---|---|
| No owner | Nobody in the business will be accountable for the outcome |
| No baseline | You cannot measure today, so you can never prove improvement |
| Benefit is capacity only, forever | Time saved will never be redeployed or removed |
| Requires data you cannot lawfully use | Privacy, residency or contract blocks the input |
| Value below the cost to run | Token, seat and oversight cost exceeds the benefit |
| Needs accuracy the tool cannot reach | High-stakes, no viable human-review design |
Kill cheaply: a two-week spike, a back-of-envelope size, and a documented “no” preserve credibility far better than a year-long project that dies quietly.
Decision framework
The V-A-L-U-E opportunity screen
Run every candidate through five gates in order. The first “no” stops the screen — you do not size an opportunity you cannot own or measure.
| Gate | Question | If no |
|---|---|---|
| Value-chain fit | Is the work text-shaped, high-volume and reviewable? | Reject or shrink scope |
| Accountable owner | Is there a named business owner, not just IT? | Reject until sponsored |
| Lawful data | Can we use the required data (privacy, residency, contracts)? | Route to governance first |
| Unit economics | Does credible benefit exceed run cost with margin? | Reject or redesign |
| Evidence path | Can we baseline now and prove impact later? | Fix measurement before funding |
Only candidates that pass all five earn a one-page frame and a slot in the portfolio. The screen is deliberately cheap to run — minutes per idea — so it can be applied to a whole backlog in an afternoon.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| Counting time saved as cash | It is the easy number and looks big | Split cashable vs capacity; only bank redeployed or removed time |
| Funding on a single hero number | It is persuasive in a deck | Require a driver-based estimate with a ±band and a counterfactual |
| One big transformational bet | Ambition feels like leadership | Balance a portfolio; let quick wins fund belief |
| Technology-led opportunity list | Vendors and engineers push capabilities | Start from business pain and value-chain volume, not features |
| Never killing anything | Sunk cost and sponsor politics | Write disqualifiers up front; make the kill a policy |
| Ignoring run cost | Pilots feel free on a trial | Include tokens, seats and human review in unit economics |
| No baseline before launch | Everyone wants to start building | Measure the current state first (D6) |
| Claiming revenue with no counterfactual | It inflates the case | State what would have happened anyway and net it out |
Scenario challenge
Scenario. You are the transformation lead at a 4,000-person insurer. Three functions bring you AI ideas in the same week. Claims operations wants to auto-draft settlement letters (120,000 letters a year, currently 12 minutes each, heavily templated, always reviewed by a handler). Marketing wants a “GPT-powered” campaign personalisation engine with a promised “40% lift” and no baseline. The chief underwriter wants AI to make certain low-value coverage decisions automatically to cut cycle time. The CFO has given you funding for two initiatives this quarter and wants a one-page rationale for each choice, including what you are declining and why.
Expert reasoning trace.
-
Screen before sizing. Run V-A-L-U-E on all three. Claims letters: text-shaped, high-volume, reviewable, owner is the head of claims, data is internal, evidence is easy (time per letter is logged) — passes all gates. Marketing: fails Evidence (no baseline) and its number fails the credibility test (a bare 40% with no driver chain). Underwriting auto-decisions: fails Value-chain fit and Lawful/accountable — an irreversible, regulated decision of record with no human review is not an assist opportunity, it is a governance problem.
-
Classify benefit type honestly. Claims is cost-out: 120,000 × ~6 minutes saved × loaded rate — but I split cashable from capacity, because handlers are not going to be reduced this year, so most of it is capacity plus a quality/turnaround gain worth quantifying separately. Marketing would be a revenue play, but I cannot size a revenue play with no baseline and no counterfactual.
-
Fix, don’t reject, the salvageable one. Marketing is not a “no” forever — it is a “not yet”. I fund a two-week baseline-and-A/B design as a cheap spike, not a build. That respects the sponsor while refusing to fund on a hero number.
-
Decline the dangerous one clearly. Underwriting auto-decisions get a written “no” with the reason: regulated decisions of record need a human-review design and a governance review first. I offer the underwriter the assist version (AI drafts the rationale, a human decides) as the fundable alternative.
-
Compose the portfolio. Fund the claims quick win (builds belief, real turnaround gain) and the marketing baseline spike (keeps the higher-upside revenue bet alive at low cost). Document the underwriting decline and its safe alternative.
Board-ready outcome: two funded initiatives (one cashable/capacity cost-out win, one cheap revenue-hypothesis spike), one documented decline with a safer path offered, and every number carrying assumptions and a range. The CFO can defend all of it.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| Approve the “40% lift” campaign because the upside is largest | Big numbers dominate a deck | No baseline and no counterfactual → unsizeable → fund a spike, not the build |
| Treat all time saved as savings | Simple multiplication | Only redeployed or removed time is cashable; the rest is capacity |
| Automate the regulated decision to cut cycle time | Speed is the stated goal | Irreversible regulated decisions of record need human review; offer assist instead |
| Fund the single transformational bet | It feels like real leadership | Portfolio balance beats one high-variance bet; quick wins fund belief |
| Rank ideas by technical excitement | Engineers and vendors push capability | Start from business pain and volume, not features |
| Skip the kill decision to avoid politics | Sponsors resist “no” | Written disqualifiers make the kill a policy, defensible to the sponsor |
Practice questions
Each item states how many responses to select. Commit before revealing.
Q1 · A support team drafts 30,000 templated replies a month, each reviewed by an agent. Which benefit type BEST describes AI-assisted drafting here? (Select one)
A. Revenue — it will win new customers. B. Cost-out — same output produced with less handler time, subject to how the time is used. C. Risk — it reduces compliance exposure. D. Transformational — it changes the business model.
Answer: B. High-volume, templated, reviewed drafting is a classic cost-out play measured in handler time. It is not primarily a revenue driver (A) or a risk-loss reducer (C), and drafting assistance is incremental, not a business-model change (D).
Q2 · A sponsor presents a single '£2m annual saving' figure with no assumptions. What is the BEST leadership response before funding? (Select one)
A. Approve it; the number is large enough to justify itself. B. Reject it; large AI savings claims are never real. C. Ask for a driver-based breakdown with assumptions, a counterfactual and a range. D. Halve the number as a conservative rule of thumb.
Answer: C. Credibility comes from the driver chain, assumptions and a band, not from the headline. Approving (A) funds a hero number; rejecting outright (B) is as unjustified as approving; halving (D) is arbitrary and hides the real drivers.
Q3 · An underwriting lead wants AI to make certain regulated coverage decisions automatically to cut cycle time. What is the MOST appropriate framing? (Select one)
A. Fund full automation; cycle time is the stated business goal. B. Decline full automation of the regulated decision of record; offer an assist design where AI drafts and a human decides. C. Fund it but add a monthly audit after launch. D. Send it to marketing for a personalisation pilot.
Answer: B. A regulated, largely irreversible decision of record needs a human-review design; the fundable version is assist-with-decision-by-human. Full automation (A) ignores the risk; a retrospective audit (C) does not prevent the harm; (D) is unrelated.
Q4 · Which of the following make an opportunity a strong early AI candidate? (Select two)
A. The work is high-volume and text-shaped. B. The output is reviewed by a human before it matters. C. The decision is irreversible and safety-critical. D. It is the most technically novel idea in the backlog. E. It requires customer data your contracts forbid you to process.
Answer: A and B. Volume plus a natural review step is the sweet spot for generative AI. Irreversible safety-critical work (C) is a reason to avoid full automation; technical novelty (D) is not a value criterion; forbidden data (E) is a hard disqualifier.
Q5 · A pilot saves five minutes per case across a team that is currently under-utilised. How should the benefit be reported? (Select one)
A. As a direct cash saving equal to the hours times the loaded rate. B. As capacity created, distinct from cash, unless the time is redeployed to more volume or removed. C. As revenue, since faster cases can serve more customers. D. As risk reduction.
Answer: B. Time saved on an under-utilised team is capacity, not cash, until it is redeployed or removed. Reporting it as cash (A) overstates the case; it is not inherently revenue (C) or risk (D).
Q6 · Your first-year AI portfolio contains one large, high-variance revenue bet and nothing else. What is the MAIN weakness? (Select one)
A. It is too cheap to matter. B. It has no quick wins to build sponsor belief and no diversification if the bet fails. C. It uses too many models. D. It ignores cost-out entirely, which is illegal.
Answer: B. A single high-variance bet risks the whole programme’s credibility; quick cost-out wins fund belief and diversify. It is not about cost (A) or model count (C), and there is nothing illegal about the mix (D).
Q7 · A proposed initiative has no named business owner, only an interested IT manager. Which V-A-L-U-E gate does it fail, and what follows? (Select one)
A. Value-chain fit; shrink the scope. B. Accountable owner; do not size or fund it until the business sponsors it. C. Unit economics; cut the model cost. D. Evidence path; add a dashboard.
Answer: B. Without an accountable business owner the initiative fails the Accountable-owner gate and should not proceed. It is not a value-chain (A), economics (C) or evidence (D) failure — those gates are not even reached.
Q8 · A revenue business case claims +£1m from faster proposals. What single addition MOST improves its credibility? (Select one)
A. A larger font on the headline number. B. A counterfactual estimate of how many deals would have closed anyway, netted out. C. A promise to revisit it next year. D. Switching to a cheaper model.
Answer: B. Netting out the counterfactual is what makes a revenue claim believable to finance. Presentation (A) is irrelevant; deferring (C) does not improve the estimate; model cost (D) is a run-cost question, not a benefit-credibility one.
Q9 · Which items belong in an opportunity's written disqualifiers, set before enthusiasm builds? (Select two)
A. There is no baseline, so improvement can never be proven. B. The required data cannot be used lawfully. C. The idea came from marketing rather than engineering. D. The vendor’s demo was impressive. E. The initiative would need more than one model.
Answer: A and B. No baseline and unlawful data are legitimate up-front kill conditions. The originating function (C) and demo quality (D) are irrelevant to disqualification, and needing multiple models (E) is an implementation detail, not a disqualifier.
Q10 · A CFO asks why you are declining a popular AI idea. What is the STRONGEST way to defend the decision? (Select one)
A. Say the team is too busy this quarter. B. Point to the written disqualifier it triggered, agreed before the idea gained momentum. C. Explain that you personally dislike the sponsor’s approach. D. Note that a competitor tried something similar and failed.
Answer: B. A pre-agreed, written disqualifier makes the kill a policy decision rather than a personal one and is defensible to any sponsor. Capacity excuses (A), personal preference (C) and anecdote (D) are all weak and political.
Q11 · An operations exec lists ten AI ideas ranked by 'how advanced the technology is'. What is the FIRST correction to make? (Select one)
A. Re-rank by business pain and value-chain volume, not by technical novelty. B. Pick the top three by novelty and start immediately. C. Buy the most advanced platform to cover all ten. D. Ask each vendor to demo before deciding.
Answer: A. Opportunity value follows business pain and volume, not technical excitement; the ranking axis is wrong. Acting on novelty (B), buying a platform first (C) or leading with vendor demos (D) all repeat the technology-led mistake.
Q12 · A cost-out pilot shows a clear quality gain (fewer errors) but little cashable time saving. How should this be positioned in the portfolio? (Select one)
A. Kill it, because there is no cash saving. B. Keep it, sizing the avoided rework and quality/turnaround benefit as a risk-and-quality play rather than a pure cash cost-out. C. Reclassify it as a transformational bet. D. Report the quality gain as revenue.
Answer: B. Error reduction is real value — measure it as rework avoided and quality/turnaround improvement even when cash time saving is small. Killing it (A) discards genuine benefit; it is not transformational (C) or revenue (D) simply because the cost-out number is thin.
Key takeaways
- AI value concentrates in high-volume, text-shaped, reviewable work; irreversible, safety-critical and regulated-of-record work is not a full-automation opportunity.
- Every opportunity is dominated by one benefit type — cost-out, revenue or risk — and each is measured and defended differently.
- Time saved is not cash unless it is redeployed or removed; separate cashable from capacity.
- Size with a driver-based estimate, explicit assumptions, a counterfactual and a range — never a hero number.
- Balance a portfolio: quick cost-out wins to build belief, a few scaled revenue/risk bets, a little exploration.
- Kill early and in writing. Pre-agreed disqualifiers turn a hard “no” into a defensible policy.
- Run the V-A-L-U-E screen before sizing: value-chain fit, accountable owner, lawful data, unit economics, evidence path.
- The portfolio you assemble here becomes the input to the strategy and roadmap domain.
Last updated Sep 18, 2026