# D6 · Measurement and ROI

Leading vs lagging indicators, baselining before deployment, quality and cycle-time metrics, the unit economics of tokens and seats, attribution problems, and reporting to a board without overclaiming.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain carries **16%** of the mock — roughly **8 of 50 items**. It tests whether a leader can prove value honestly: choosing leading and lagging indicators, **baselining before deployment** (the most-skipped and most-costly discipline), measuring quality and cycle time, understanding the unit economics of tokens and seats, confronting the attribution problem, and reporting to a board without the overclaiming that eventually destroys credibility. It closes the loop opened by [opportunity sizing](/openai/leadership/domains/d1-opportunity-and-value/): the number you promised there is the number you must now measure.

## What you need to know

Measurement fails at the start, not the end: if you did not **baseline** the current state before deploying, you can never prove improvement — so baselining is the first act, not an afterthought. Distinguish **leading indicators** (active use, cycle time, adoption — visible in weeks) from **lagging indicators** (cost, revenue, retention — visible in quarters). Measure **quality** alongside speed, or you will optimise throughput while degrading outcomes. Understand **unit economics**: cost is tokens (usage-priced per million) plus seats (per user), and value must clear that cost with margin. Confront **attribution** honestly — AI is rarely the only variable — and report to the board with baselines, ranges and counterfactuals rather than hero numbers, because one exposed overclaim discounts everything that follows.

## Learning objectives

By the end of this page you should be able to:

1. **Baseline** the current state before deployment so improvement is provable.
2. **Choose** leading and lagging indicators appropriate to each initiative and time horizon.
3. **Measure** quality and cycle time together, not throughput alone.
4. **Model** unit economics from token and seat costs against realised value.
5. **Address** the attribution problem with controlled comparison where possible.
6. **Report** to a board credibly, with ranges, counterfactuals and no overclaiming.

---

## 6.1 Baseline before deployment

The single most common measurement failure is deploying first and asking "did it help?" afterwards — by which point there is no clean before-picture to compare against. Baseline first.

| Without a baseline | With a baseline |
| --- | --- |
| "It feels faster" — unprovable | "Handle time fell from 6.2 to 4.1 min" — provable |
| Benefit contested by sceptics | Benefit defensible to finance |
| Attribution impossible | Change measured against a known start |
| Board discounts the claim | Board trusts the method |

```text
  RIGHT ORDER                     WRONG ORDER
  ───────────                     ───────────
  1. define the metric            1. deploy the tool
  2. measure the baseline         2. announce success
  3. deploy                       3. try to prove it later
  4. measure the change           4. discover there is no baseline
  5. attribute honestly           5. claim vibes as ROI
```

:::tip[Assessment signal]
Stems saying **"we deployed and now want to prove ROI", "how do we know it helped", "leadership wants the business case validated"** almost always hinge on whether a baseline exists. The correct answer establishes or reconstructs a baseline before claiming benefit; distractors report anecdote or activity as ROI.
:::

## 6.2 Leading vs lagging indicators

You need both, on different clocks. Leading indicators tell you early whether the initiative is on track; lagging indicators confirm the financial outcome later.

| Type | Examples | Horizon | Answers |
| --- | --- | --- | --- |
| **Leading** | Active use, adoption rate, cycle time, quality score, rework rate | Weeks | "Is this working and being used?" |
| **Lagging** | Cost-to-serve, revenue, retention, margin | Quarters | "Did it pay off financially?" |

Reporting only lagging indicators leaves you blind for a quarter; reporting only leading indicators lets you mistake activity for value. Boards want the lagging outcome; the programme is steered by the leading ones. A healthy report shows leading indicators trending now and states when the lagging outcome will be measurable.

## 6.3 Quality and cycle time together

Speed without quality is a trap: AI can make people faster at producing worse work, and a cycle-time win that raises error or rework rates is a net loss. Always pair a throughput metric with a quality metric.

| Initiative | Cycle-time metric | Paired quality metric |
| --- | --- | --- |
| Support draft replies | Handle time | CSAT, reopen rate, escalation rate |
| Proposal drafting | Time to first draft | Win rate, error/rework rate |
| Document review | Time per document | Miss rate on material issues |
| Code assistance | Time to change | Defect/incident rate, review findings |

The leadership discipline: define the quality guardrail *before* celebrating the speed gain, and treat a speed improvement that breaches the guardrail as a failure, not a success.

## 6.4 Unit economics: tokens and seats

A leader must understand the two cost drivers, because "the pilot was free" on a trial becomes a real bill at scale.

<Tabs>
<TabItem label="The two cost drivers">

- **Tokens** — usage-based cost for API/agent workloads, priced per million input and output tokens. Heavier reasoning, longer context and more output all raise it; capability tiers differ widely in price (a cheaper, lighter model can be an order of magnitude cheaper per token than a top reasoning model).
- **Seats** — per-user subscription cost for ChatGPT Business/Enterprise. Cost scales with *provisioned* seats, which is exactly why the [licence graveyard](/openai/leadership/domains/d4-adoption-and-change-management/) is a financial problem, not just an adoption one: you pay for seats whether or not they are used.

</TabItem>
<TabItem label="The economics that must hold">

```text
  Value per period  >  (seats × seat price)
                     +  (tokens used × token price)
                     +  human review / oversight cost
                     +  run + support overhead

  If not, redesign: right-size the model tier,
  reclaim unused seats, tighten context length,
  or narrow scope — before scaling.
```

</TabItem>
</Tabs>

Two leadership levers follow directly. First, **match the model tier to the task**: routing high-volume, simple work to a lighter, cheaper model and reserving the top reasoning tier for genuinely hard work is often the largest single cost lever. Second, **reclaim unused seats**: paying for provisioned-but-idle seats is pure waste that a usage report exposes.

:::tip[Assessment signal]
Stems about **"the pilot was cheap but the bill grew", "cost per user", "we're paying for seats nobody uses", "which model to run at volume"** are unit-economics items. Correct answers right-size the model tier, reclaim idle seats, or check value against total run cost; distractors assume trial pricing or ignore oversight cost.
:::

## 6.5 The attribution problem

AI is rarely the only thing that changed. Sales also hired, the market shifted, a process was tidied at the same time. Overclaiming attribution is the fastest route to losing board trust when a sceptic pulls the thread.

| Attribution tactic | Strength |
| --- | --- |
| A/B or holdout (some teams use AI, some do not) | Strongest; isolates the AI effect |
| Before/after with a stable baseline | Good if no other big change occurred |
| Cohort comparison (similar teams) | Moderate; controls for some confounds |
| Correlation with usage | Weak; suggestive, not causal |
| "Revenue went up and we launched AI" | None; pure coincidence claim |

Where the stakes justify it, design an **A/B or holdout** from the start. Where you cannot, state the confounds openly and claim a *range* rather than the whole delta. Honesty about attribution is itself credibility: a leader who says "AI contributed an estimated 30–50% of this improvement, the rest is the process change" is believed; one who claims 100% is not.

## 6.6 Reporting to a board without overclaiming

The board report is where every earlier discipline is cashed in or squandered. One exposed overclaim discounts every future number you present.

| Do | Don't |
| --- | --- |
| Show the baseline and the change against it | Present a raw after-number with no before |
| Give ranges and state confidence | Give a single hero figure |
| Separate cashable from capacity benefit | Bank all time saved as cash |
| State the counterfactual / confounds | Claim the whole delta as AI |
| Pair speed gains with quality guardrails | Report throughput while quality slips |
| Report leading indicators now, lagging when ready | Promise lagging outcomes prematurely |

```text
  CREDIBLE BOARD LINE
  "Support handle time fell 34% (baseline 6.2 → 4.1 min) on the
   60% of tickets in scope. ~30% of the saved time is redeployed
   (cashable), the rest is capacity. Leading indicators are strong;
   the cost-to-serve outcome will be measurable next quarter.
   An A/B against two control regions attributes ~80% of the drop
   to the assist. Net run cost is below benefit with margin."
```

That is what "reporting without overclaiming" looks like: specific, bounded, attributed, and therefore trusted.

## Decision framework

### The P-R-O-V-E measurement plan

Attach this to every funded initiative before it deploys, so measurement is designed in, not retrofitted.

| Letter | Step | Failure if skipped |
| --- | --- | --- |
| **P**re-baseline | Measure the current state before deploying | No provable improvement |
| **R**ight indicators | Choose leading + lagging, with a quality guardrail | Activity mistaken for value; quality slips |
| **O**wn the economics | Model tokens + seats + oversight vs value | Scaling a loss-making pilot |
| **V**erify attribution | Design A/B or holdout, or state confounds | Overclaimed, later discredited |
| **E**vidence to the board | Report baseline, range, counterfactual honestly | Credibility lost on the first exposed claim |

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Deploying before baselining | Everyone wants to start | Baseline the metric first; improvement is otherwise unprovable |
| Reporting only lagging (or only leading) indicators | One clock feels simpler | Report both; steer on leading, prove on lagging |
| Celebrating speed while quality drops | Speed is visible; quality is not | Pair every throughput metric with a quality guardrail |
| Assuming trial pricing at scale | Pilots feel free | Model tokens + seats + oversight before scaling |
| Running everything on the top model | Best model feels safest | Right-size the tier to the task; it is the biggest cost lever |
| Paying for idle seats | Nobody reviews provisioning | Reclaim unused seats; a usage report exposes them |
| Claiming 100% attribution to AI | The whole delta is tempting | Use A/B/holdout or state confounds and claim a range |
| A single hero number to the board | It is persuasive once | Ranges, baselines and counterfactuals earn lasting trust |

## Scenario challenge

**Scenario.** Your CEO wants to tell the board that "AI delivered £4m of value this year". The claim rests on: a support team whose handle time "feels much faster" since ChatGPT assist launched (no before-measurement was taken); a sales region that grew 12% in the same period it adopted AI proposal drafting (and also hired two senior reps and entered a rising market); and a company-wide "£4m" figure produced by multiplying every user's estimated minutes saved by their loaded rate — including 40% of seats that show no active use. The CFO is sceptical and the audit committee will ask questions.

**Expert reasoning trace.**

1. **Refuse the composite hero number.** "£4m" is exactly the claim that collapses under one sharp question, and collapsing it would discredit the entire AI programme. My job is to make the number *smaller and defensible* rather than large and fragile.

2. **Fix the support claim — or downgrade it.** "Feels faster" with no baseline is unprovable. I check whether handle-time logs predate the launch; if a baseline can be reconstructed, I report the measured delta on the in-scope tickets; if not, I report it as a leading-indicator trend to be measured properly next quarter, not as banked value. No baseline, no ROI claim.

3. **Confront the sales attribution.** A 12% regional growth coinciding with AI adoption is not evidence AI caused it — two senior hires and a rising market are large confounds. The honest move is a cohort or holdout comparison against similar regions; absent that, I claim only a *range* of the uplift and state the confounds. Claiming the full 12% would be the overclaim the audit committee is built to catch.

4. **Strip the seat inflation.** Multiplying estimated minutes by loaded rate across *all* seats — including 40% idle — is double nonsense: it counts capacity as cash and counts non-users as beneficiaries. I remove idle seats entirely (and flag the wasted seat spend as a cost to reclaim, per D4), and split the remaining benefit into cashable vs capacity.

5. **Rebuild the report credibly.** Present the measured support delta (or a labelled trend), a ranged and attributed sales contribution, and a defensible productivity figure net of idle seats and split cashable/capacity — with the run cost (tokens + seats + oversight) subtracted. The total will be well below £4m and utterly defensible.

6. **Give the board the method, not just the number.** Baselines, ranges, counterfactuals and the reclaimed-seat action signal a leader in control — which is worth more than a large fragile figure.

**Board-ready outcome:** the £4m composite is declined; the support claim is baselined or downgraded to a trend; the sales uplift is attributed with a holdout/cohort and reported as a range; idle seats are stripped and flagged for reclaim; the final number is smaller, attributed, cost-net and trusted — and the programme's credibility survives the audit committee.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| Report "it feels faster" as ROI | The team is convinced | No baseline means no provable benefit; measure or downgrade to a trend |
| Attribute all of a coincident revenue rise to AI | The whole delta looks great | Confounds (hires, market) demand A/B/holdout or a stated range |
| Multiply minutes saved across all seats | Produces a big headline | Excludes idle seats; counts capacity as cash — strip and split |
| Run every workload on the top model | Best model feels safest | Right-sizing the tier is often the biggest cost lever |
| Report only the lagging financial outcome | It is what the board wants | You are then blind for a quarter; pair with leading indicators |
| Celebrate a cycle-time win while quality slips | Speed is the visible metric | A speed gain that breaches the quality guardrail is a net loss |

## Practice questions

Each item states how many responses to select. Commit before revealing.

<Accordions>
  <AccordionItem title="Q1 · A team deployed an AI assist and now wants to prove it saved time, but no before-measurement was taken. What is the core problem? (Select one)">
    A. The model is too slow.
    B. Without a baseline, improvement cannot be proven; the measurement should have preceded deployment.
    C. The team used the wrong model tier.
    D. The report is too short.

    **Answer: B.** No baseline means no provable improvement — baselining must precede deployment. Model speed (A), tier choice (C) and report length (D) are not the fundamental measurement failure.
  </AccordionItem>

  <AccordionItem title="Q2 · Which are LEADING indicators for an AI initiative? (Select two)">
    A. Weekly active use and adoption rate.
    B. Cycle time on the target task.
    C. Quarterly cost-to-serve.
    D. Annual revenue.
    E. Full-year margin.

    **Answer: A and B.** Active use and cycle time move in weeks and steer the programme. Cost-to-serve (C), revenue (D) and margin (E) are lagging financial outcomes visible over quarters.
  </AccordionItem>

  <AccordionItem title="Q3 · A support team's handle time dropped 30% but reopened-ticket rate rose sharply. How should this be judged? (Select one)">
    A. A clear success; speed improved.
    B. Not a success as-is; the speed gain breached the quality guardrail, which may be a net loss.
    C. Irrelevant; only handle time matters.
    D. A reason to remove human review.

    **Answer: B.** A throughput gain that degrades quality can be a net loss; speed and quality must be judged together. Speed alone (A, C) ignores quality, and removing review (D) would worsen it.
  </AccordionItem>

  <AccordionItem title="Q4 · A pilot 'was basically free' on a trial, but at scale the bill grew. Which cost drivers must be modelled? (Select two)">
    A. Token usage for API/agent workloads.
    B. Per-user seat costs for provisioned licences.
    C. The colour scheme of the interface.
    D. The number of board meetings held.
    E. The vendor's marketing spend.

    **Answer: A and B.** Tokens and seats are the two real cost drivers at scale. UI colour (C), meeting counts (D) and vendor marketing (E) are not part of your unit economics.
  </AccordionItem>

  <AccordionItem title="Q5 · A region grew 12% in the same period it adopted AI, but it also hired two senior reps and the market rose. What is the MOST honest attribution approach? (Select one)">
    A. Attribute the full 12% to AI.
    B. Use a holdout or cohort comparison, or claim only a range and state the confounds.
    C. Attribute nothing to AI ever.
    D. Attribute it to the market only.

    **Answer: B.** With large confounds, a controlled comparison or a stated range is the honest claim. Claiming all of it (A), none of it (C) or the market alone (D) are all unsupported extremes.
  </AccordionItem>

  <AccordionItem title="Q6 · 40% of provisioned ChatGPT seats show no active use. What is the correct measurement-and-cost response? (Select one)">
    A. Include those seats' estimated savings in the ROI figure.
    B. Exclude idle seats from benefit claims and reclaim them to cut cost.
    C. Buy more seats to improve the average.
    D. Report the seats as active to protect the programme.

    **Answer: B.** Idle seats deliver no benefit and cost money; exclude them from claims and reclaim them. Counting their savings (A) inflates ROI, buying more (C) worsens waste, and misreporting (D) is dishonest.
  </AccordionItem>

  <AccordionItem title="Q7 · A CEO wants to present a single '£4m of AI value' figure built from estimated minutes saved across all users. What should a leader advise? (Select one)">
    A. Present it as-is; a big number impresses the board.
    B. Rebuild it with baselines, exclude idle seats, split cashable from capacity, attribute honestly, and report a defensible range.
    C. Double it to account for hidden benefits.
    D. Refuse to report any value at all.

    **Answer: B.** A defensible, bounded, attributed number preserves credibility; a fragile hero figure risks the whole programme. Presenting as-is (A) or doubling (C) invites collapse under scrutiny; refusing entirely (D) forgoes legitimate reporting.
  </AccordionItem>

  <AccordionItem title="Q8 · Why is reporting ONLY lagging financial indicators risky mid-programme? (Select one)">
    A. Lagging indicators are always wrong.
    B. They appear only after quarters, leaving the programme blind to whether it is working now.
    C. Boards never ask for financial outcomes.
    D. They cost too much to compute.

    **Answer: B.** Lagging indicators arrive late, so leading indicators are needed to steer in the meantime. They are not inherently wrong (A), boards do want them (C), and cost (D) is not the issue.
  </AccordionItem>

  <AccordionItem title="Q9 · A high-volume, simple classification workload is running on the most expensive top-tier reasoning model. What is the STRONGEST cost lever? (Select one)">
    A. Reduce the number of users.
    B. Right-size to a lighter, cheaper model tier suited to the simple task.
    C. Turn off logging.
    D. Increase the reasoning effort for accuracy.

    **Answer: B.** Matching model tier to task difficulty is often the single largest cost lever; a simple, high-volume task rarely needs the top tier. Cutting users (A) reduces value, turning off logging (C) harms governance, and raising effort (D) increases cost.
  </AccordionItem>

  <AccordionItem title="Q10 · For a proposal-drafting initiative, which metric PAIR correctly combines cycle time with quality? (Select one)">
    A. Time to first draft and number of drafts produced.
    B. Time to first draft and win rate (with error/rework rate).
    C. Number of logins and seats issued.
    D. Model latency and token count.

    **Answer: B.** Time to first draft (speed) paired with win rate and rework (quality) measures the outcome, not just throughput. Draft counts (A), logins/seats (C) and latency/tokens (D) are activity or cost metrics, not quality.
  </AccordionItem>

  <AccordionItem title="Q11 · Following P-R-O-V-E, what MUST happen before an initiative deploys? (Select one)">
    A. The board must approve the final ROI number.
    B. The current-state baseline must be measured.
    C. All seats must be provisioned.
    D. The most advanced model must be selected.

    **Answer: B.** Pre-baselining is the first P-R-O-V-E step and must precede deployment. A final ROI number (A) cannot exist yet, full seat provisioning (C) is premature, and model choice (D) is not a measurement prerequisite.
  </AccordionItem>

  <AccordionItem title="Q12 · Which TWO practices make a board report on AI value credible rather than fragile? (Select two)">
    A. Showing the baseline and the change measured against it.
    B. Stating a range with the counterfactual or confounds.
    C. Presenting one large hero figure with no method.
    D. Banking all time saved as cash.
    E. Omitting run costs to keep the number high.

    **Answer: A and B.** Baselines and ranged, attributed claims withstand scrutiny. A hero figure (C), banking all time as cash (D) and omitting run costs (E) all make the number fragile and eventually discrediting.
  </AccordionItem>
</Accordions>

## Key takeaways

- **Baseline before deployment**; without a before-picture, improvement is unprovable and the board discounts the claim.
- Report **leading indicators** (use, cycle time, quality — weeks) and **lagging indicators** (cost, revenue, margin — quarters) together.
- **Pair speed with quality**; a cycle-time win that breaches the quality guardrail is a net loss.
- Model **unit economics** — tokens plus seats plus oversight — and right-size the model tier; reclaim idle seats.
- Confront **attribution** honestly with A/B or holdout comparisons, or state the confounds and claim a range.
- **Report without overclaiming**: baselines, ranges, counterfactuals, cashable-vs-capacity — one exposed overclaim discredits everything.
- Attach **P-R-O-V-E** to every initiative before it deploys: pre-baseline, right indicators, own the economics, verify attribution, evidence to the board.
