AI Cert Prep
Type to search documentation.

AI Leadership

D6 · Measurement and ROI

Leading vs lagging indicators, baselining before deployment, quality and cycle-time metrics, the unit economics of tokens and seats, attribution problems, and reporting to a board without overclaiming.

This domain carries 16% of the mock — roughly 8 of 50 items. It tests whether a leader can prove value honestly: choosing leading and lagging indicators, baselining before deployment (the most-skipped and most-costly discipline), measuring quality and cycle time, understanding the unit economics of tokens and seats, confronting the attribution problem, and reporting to a board without the overclaiming that eventually destroys credibility. It closes the loop opened by opportunity sizing: the number you promised there is the number you must now measure.

What you need to know

Measurement fails at the start, not the end: if you did not baseline the current state before deploying, you can never prove improvement — so baselining is the first act, not an afterthought. Distinguish leading indicators (active use, cycle time, adoption — visible in weeks) from lagging indicators (cost, revenue, retention — visible in quarters). Measure quality alongside speed, or you will optimise throughput while degrading outcomes. Understand unit economics: cost is tokens (usage-priced per million) plus seats (per user), and value must clear that cost with margin. Confront attribution honestly — AI is rarely the only variable — and report to the board with baselines, ranges and counterfactuals rather than hero numbers, because one exposed overclaim discounts everything that follows.

Learning objectives

By the end of this page you should be able to:

  1. Baseline the current state before deployment so improvement is provable.
  2. Choose leading and lagging indicators appropriate to each initiative and time horizon.
  3. Measure quality and cycle time together, not throughput alone.
  4. Model unit economics from token and seat costs against realised value.
  5. Address the attribution problem with controlled comparison where possible.
  6. Report to a board credibly, with ranges, counterfactuals and no overclaiming.

6.1 Baseline before deployment

The single most common measurement failure is deploying first and asking “did it help?” afterwards — by which point there is no clean before-picture to compare against. Baseline first.

Without a baselineWith a baseline
“It feels faster” — unprovable“Handle time fell from 6.2 to 4.1 min” — provable
Benefit contested by scepticsBenefit defensible to finance
Attribution impossibleChange measured against a known start
Board discounts the claimBoard trusts the method
text
RIGHT ORDER WRONG ORDER
─────────── ───────────
1. define the metric 1. deploy the tool
2. measure the baseline 2. announce success
3. deploy 3. try to prove it later
4. measure the change 4. discover there is no baseline
5. attribute honestly 5. claim vibes as ROI

Assessment signal

Stems saying “we deployed and now want to prove ROI”, “how do we know it helped”, “leadership wants the business case validated” almost always hinge on whether a baseline exists. The correct answer establishes or reconstructs a baseline before claiming benefit; distractors report anecdote or activity as ROI.

6.2 Leading vs lagging indicators

You need both, on different clocks. Leading indicators tell you early whether the initiative is on track; lagging indicators confirm the financial outcome later.

TypeExamplesHorizonAnswers
LeadingActive use, adoption rate, cycle time, quality score, rework rateWeeks“Is this working and being used?”
LaggingCost-to-serve, revenue, retention, marginQuarters“Did it pay off financially?”

Reporting only lagging indicators leaves you blind for a quarter; reporting only leading indicators lets you mistake activity for value. Boards want the lagging outcome; the programme is steered by the leading ones. A healthy report shows leading indicators trending now and states when the lagging outcome will be measurable.

6.3 Quality and cycle time together

Speed without quality is a trap: AI can make people faster at producing worse work, and a cycle-time win that raises error or rework rates is a net loss. Always pair a throughput metric with a quality metric.

InitiativeCycle-time metricPaired quality metric
Support draft repliesHandle timeCSAT, reopen rate, escalation rate
Proposal draftingTime to first draftWin rate, error/rework rate
Document reviewTime per documentMiss rate on material issues
Code assistanceTime to changeDefect/incident rate, review findings

The leadership discipline: define the quality guardrail before celebrating the speed gain, and treat a speed improvement that breaches the guardrail as a failure, not a success.

6.4 Unit economics: tokens and seats

A leader must understand the two cost drivers, because “the pilot was free” on a trial becomes a real bill at scale.

  • Tokens — usage-based cost for API/agent workloads, priced per million input and output tokens. Heavier reasoning, longer context and more output all raise it; capability tiers differ widely in price (a cheaper, lighter model can be an order of magnitude cheaper per token than a top reasoning model).
  • Seats — per-user subscription cost for ChatGPT Business/Enterprise. Cost scales with provisioned seats, which is exactly why the licence graveyard is a financial problem, not just an adoption one: you pay for seats whether or not they are used.

Two leadership levers follow directly. First, match the model tier to the task: routing high-volume, simple work to a lighter, cheaper model and reserving the top reasoning tier for genuinely hard work is often the largest single cost lever. Second, reclaim unused seats: paying for provisioned-but-idle seats is pure waste that a usage report exposes.

Assessment signal

Stems about “the pilot was cheap but the bill grew”, “cost per user”, “we’re paying for seats nobody uses”, “which model to run at volume” are unit-economics items. Correct answers right-size the model tier, reclaim idle seats, or check value against total run cost; distractors assume trial pricing or ignore oversight cost.

6.5 The attribution problem

AI is rarely the only thing that changed. Sales also hired, the market shifted, a process was tidied at the same time. Overclaiming attribution is the fastest route to losing board trust when a sceptic pulls the thread.

Attribution tacticStrength
A/B or holdout (some teams use AI, some do not)Strongest; isolates the AI effect
Before/after with a stable baselineGood if no other big change occurred
Cohort comparison (similar teams)Moderate; controls for some confounds
Correlation with usageWeak; suggestive, not causal
“Revenue went up and we launched AI”None; pure coincidence claim

Where the stakes justify it, design an A/B or holdout from the start. Where you cannot, state the confounds openly and claim a range rather than the whole delta. Honesty about attribution is itself credibility: a leader who says “AI contributed an estimated 30–50% of this improvement, the rest is the process change” is believed; one who claims 100% is not.

6.6 Reporting to a board without overclaiming

The board report is where every earlier discipline is cashed in or squandered. One exposed overclaim discounts every future number you present.

DoDon’t
Show the baseline and the change against itPresent a raw after-number with no before
Give ranges and state confidenceGive a single hero figure
Separate cashable from capacity benefitBank all time saved as cash
State the counterfactual / confoundsClaim the whole delta as AI
Pair speed gains with quality guardrailsReport throughput while quality slips
Report leading indicators now, lagging when readyPromise lagging outcomes prematurely
text
CREDIBLE BOARD LINE
"Support handle time fell 34% (baseline 6.2 → 4.1 min) on the
60% of tickets in scope. ~30% of the saved time is redeployed
(cashable), the rest is capacity. Leading indicators are strong;
the cost-to-serve outcome will be measurable next quarter.
An A/B against two control regions attributes ~80% of the drop
to the assist. Net run cost is below benefit with margin."

That is what “reporting without overclaiming” looks like: specific, bounded, attributed, and therefore trusted.

Decision framework

The P-R-O-V-E measurement plan

Attach this to every funded initiative before it deploys, so measurement is designed in, not retrofitted.

LetterStepFailure if skipped
Pre-baselineMeasure the current state before deployingNo provable improvement
Right indicatorsChoose leading + lagging, with a quality guardrailActivity mistaken for value; quality slips
Own the economicsModel tokens + seats + oversight vs valueScaling a loss-making pilot
Verify attributionDesign A/B or holdout, or state confoundsOverclaimed, later discredited
Evidence to the boardReport baseline, range, counterfactual honestlyCredibility lost on the first exposed claim

Common mistakes

MistakeWhy it happensWhat to do instead
Deploying before baseliningEveryone wants to startBaseline the metric first; improvement is otherwise unprovable
Reporting only lagging (or only leading) indicatorsOne clock feels simplerReport both; steer on leading, prove on lagging
Celebrating speed while quality dropsSpeed is visible; quality is notPair every throughput metric with a quality guardrail
Assuming trial pricing at scalePilots feel freeModel tokens + seats + oversight before scaling
Running everything on the top modelBest model feels safestRight-size the tier to the task; it is the biggest cost lever
Paying for idle seatsNobody reviews provisioningReclaim unused seats; a usage report exposes them
Claiming 100% attribution to AIThe whole delta is temptingUse A/B/holdout or state confounds and claim a range
A single hero number to the boardIt is persuasive onceRanges, baselines and counterfactuals earn lasting trust

Scenario challenge

Scenario. Your CEO wants to tell the board that “AI delivered £4m of value this year”. The claim rests on: a support team whose handle time “feels much faster” since ChatGPT assist launched (no before-measurement was taken); a sales region that grew 12% in the same period it adopted AI proposal drafting (and also hired two senior reps and entered a rising market); and a company-wide “£4m” figure produced by multiplying every user’s estimated minutes saved by their loaded rate — including 40% of seats that show no active use. The CFO is sceptical and the audit committee will ask questions.

Expert reasoning trace.

  1. Refuse the composite hero number. “£4m” is exactly the claim that collapses under one sharp question, and collapsing it would discredit the entire AI programme. My job is to make the number smaller and defensible rather than large and fragile.

  2. Fix the support claim — or downgrade it. “Feels faster” with no baseline is unprovable. I check whether handle-time logs predate the launch; if a baseline can be reconstructed, I report the measured delta on the in-scope tickets; if not, I report it as a leading-indicator trend to be measured properly next quarter, not as banked value. No baseline, no ROI claim.

  3. Confront the sales attribution. A 12% regional growth coinciding with AI adoption is not evidence AI caused it — two senior hires and a rising market are large confounds. The honest move is a cohort or holdout comparison against similar regions; absent that, I claim only a range of the uplift and state the confounds. Claiming the full 12% would be the overclaim the audit committee is built to catch.

  4. Strip the seat inflation. Multiplying estimated minutes by loaded rate across all seats — including 40% idle — is double nonsense: it counts capacity as cash and counts non-users as beneficiaries. I remove idle seats entirely (and flag the wasted seat spend as a cost to reclaim, per D4), and split the remaining benefit into cashable vs capacity.

  5. Rebuild the report credibly. Present the measured support delta (or a labelled trend), a ranged and attributed sales contribution, and a defensible productivity figure net of idle seats and split cashable/capacity — with the run cost (tokens + seats + oversight) subtracted. The total will be well below £4m and utterly defensible.

  6. Give the board the method, not just the number. Baselines, ranges, counterfactuals and the reclaimed-seat action signal a leader in control — which is worth more than a large fragile figure.

Board-ready outcome: the £4m composite is declined; the support claim is baselined or downgraded to a trend; the sales uplift is attributed with a holdout/cohort and reported as a range; idle seats are stripped and flagged for reclaim; the final number is smaller, attributed, cost-net and trusted — and the programme’s credibility survives the audit committee.

Assessment traps

TrapWhy it is temptingThe discriminator
Report “it feels faster” as ROIThe team is convincedNo baseline means no provable benefit; measure or downgrade to a trend
Attribute all of a coincident revenue rise to AIThe whole delta looks greatConfounds (hires, market) demand A/B/holdout or a stated range
Multiply minutes saved across all seatsProduces a big headlineExcludes idle seats; counts capacity as cash — strip and split
Run every workload on the top modelBest model feels safestRight-sizing the tier is often the biggest cost lever
Report only the lagging financial outcomeIt is what the board wantsYou are then blind for a quarter; pair with leading indicators
Celebrate a cycle-time win while quality slipsSpeed is the visible metricA speed gain that breaches the quality guardrail is a net loss

Practice questions

Each item states how many responses to select. Commit before revealing.

Q1 · A team deployed an AI assist and now wants to prove it saved time, but no before-measurement was taken. What is the core problem? (Select one)

A. The model is too slow. B. Without a baseline, improvement cannot be proven; the measurement should have preceded deployment. C. The team used the wrong model tier. D. The report is too short.

Answer: B. No baseline means no provable improvement — baselining must precede deployment. Model speed (A), tier choice (C) and report length (D) are not the fundamental measurement failure.

Q2 · Which are LEADING indicators for an AI initiative? (Select two)

A. Weekly active use and adoption rate. B. Cycle time on the target task. C. Quarterly cost-to-serve. D. Annual revenue. E. Full-year margin.

Answer: A and B. Active use and cycle time move in weeks and steer the programme. Cost-to-serve (C), revenue (D) and margin (E) are lagging financial outcomes visible over quarters.

Q3 · A support team's handle time dropped 30% but reopened-ticket rate rose sharply. How should this be judged? (Select one)

A. A clear success; speed improved. B. Not a success as-is; the speed gain breached the quality guardrail, which may be a net loss. C. Irrelevant; only handle time matters. D. A reason to remove human review.

Answer: B. A throughput gain that degrades quality can be a net loss; speed and quality must be judged together. Speed alone (A, C) ignores quality, and removing review (D) would worsen it.

Q4 · A pilot 'was basically free' on a trial, but at scale the bill grew. Which cost drivers must be modelled? (Select two)

A. Token usage for API/agent workloads. B. Per-user seat costs for provisioned licences. C. The colour scheme of the interface. D. The number of board meetings held. E. The vendor’s marketing spend.

Answer: A and B. Tokens and seats are the two real cost drivers at scale. UI colour (C), meeting counts (D) and vendor marketing (E) are not part of your unit economics.

Q5 · A region grew 12% in the same period it adopted AI, but it also hired two senior reps and the market rose. What is the MOST honest attribution approach? (Select one)

A. Attribute the full 12% to AI. B. Use a holdout or cohort comparison, or claim only a range and state the confounds. C. Attribute nothing to AI ever. D. Attribute it to the market only.

Answer: B. With large confounds, a controlled comparison or a stated range is the honest claim. Claiming all of it (A), none of it (C) or the market alone (D) are all unsupported extremes.

Q6 · 40% of provisioned ChatGPT seats show no active use. What is the correct measurement-and-cost response? (Select one)

A. Include those seats’ estimated savings in the ROI figure. B. Exclude idle seats from benefit claims and reclaim them to cut cost. C. Buy more seats to improve the average. D. Report the seats as active to protect the programme.

Answer: B. Idle seats deliver no benefit and cost money; exclude them from claims and reclaim them. Counting their savings (A) inflates ROI, buying more (C) worsens waste, and misreporting (D) is dishonest.

Q7 · A CEO wants to present a single '£4m of AI value' figure built from estimated minutes saved across all users. What should a leader advise? (Select one)

A. Present it as-is; a big number impresses the board. B. Rebuild it with baselines, exclude idle seats, split cashable from capacity, attribute honestly, and report a defensible range. C. Double it to account for hidden benefits. D. Refuse to report any value at all.

Answer: B. A defensible, bounded, attributed number preserves credibility; a fragile hero figure risks the whole programme. Presenting as-is (A) or doubling (C) invites collapse under scrutiny; refusing entirely (D) forgoes legitimate reporting.

Q8 · Why is reporting ONLY lagging financial indicators risky mid-programme? (Select one)

A. Lagging indicators are always wrong. B. They appear only after quarters, leaving the programme blind to whether it is working now. C. Boards never ask for financial outcomes. D. They cost too much to compute.

Answer: B. Lagging indicators arrive late, so leading indicators are needed to steer in the meantime. They are not inherently wrong (A), boards do want them (C), and cost (D) is not the issue.

Q9 · A high-volume, simple classification workload is running on the most expensive top-tier reasoning model. What is the STRONGEST cost lever? (Select one)

A. Reduce the number of users. B. Right-size to a lighter, cheaper model tier suited to the simple task. C. Turn off logging. D. Increase the reasoning effort for accuracy.

Answer: B. Matching model tier to task difficulty is often the single largest cost lever; a simple, high-volume task rarely needs the top tier. Cutting users (A) reduces value, turning off logging (C) harms governance, and raising effort (D) increases cost.

Q10 · For a proposal-drafting initiative, which metric PAIR correctly combines cycle time with quality? (Select one)

A. Time to first draft and number of drafts produced. B. Time to first draft and win rate (with error/rework rate). C. Number of logins and seats issued. D. Model latency and token count.

Answer: B. Time to first draft (speed) paired with win rate and rework (quality) measures the outcome, not just throughput. Draft counts (A), logins/seats (C) and latency/tokens (D) are activity or cost metrics, not quality.

Q11 · Following P-R-O-V-E, what MUST happen before an initiative deploys? (Select one)

A. The board must approve the final ROI number. B. The current-state baseline must be measured. C. All seats must be provisioned. D. The most advanced model must be selected.

Answer: B. Pre-baselining is the first P-R-O-V-E step and must precede deployment. A final ROI number (A) cannot exist yet, full seat provisioning (C) is premature, and model choice (D) is not a measurement prerequisite.

Q12 · Which TWO practices make a board report on AI value credible rather than fragile? (Select two)

A. Showing the baseline and the change measured against it. B. Stating a range with the counterfactual or confounds. C. Presenting one large hero figure with no method. D. Banking all time saved as cash. E. Omitting run costs to keep the number high.

Answer: A and B. Baselines and ranged, attributed claims withstand scrutiny. A hero figure (C), banking all time as cash (D) and omitting run costs (E) all make the number fragile and eventually discrediting.

Key takeaways

  • Baseline before deployment; without a before-picture, improvement is unprovable and the board discounts the claim.
  • Report leading indicators (use, cycle time, quality — weeks) and lagging indicators (cost, revenue, margin — quarters) together.
  • Pair speed with quality; a cycle-time win that breaches the quality guardrail is a net loss.
  • Model unit economics — tokens plus seats plus oversight — and right-size the model tier; reclaim idle seats.
  • Confront attribution honestly with A/B or holdout comparisons, or state the confounds and claim a range.
  • Report without overclaiming: baselines, ranges, counterfactuals, cashable-vs-capacity — one exposed overclaim discredits everything.
  • Attach P-R-O-V-E to every initiative before it deploys: pre-baseline, right indicators, own the economics, verify attribution, evidence to the board.

Last updated Sep 18, 2026