Appendix · AWS
KPI and Metrics Library
A catalogue of measurable AI KPIs by business function with formulas, baselines and gaming risks, plus tangible vs intangible benefits, leading vs lagging pairs, adoption metrics, a KPI tree, baseline-construction methods and vanity metrics to avoid — for AIB-C01.
This is a working catalogue for the measurement work Domain 2 tests (tasks 2.2.1–2.2.4): KPIs by business function, tangible versus intangible benefits, leading versus lagging indicators, adoption and health metrics, a KPI tree, and how to build a baseline when none exists. It is independent preparation. The AIB-C01 beta exam does not ask you to instrument anything; it asks you to pick the right metric, define a defensible baseline, and spot a metric that measures activity instead of value. Every formula here stays at that decision level.
Baseline before deployment
The single most-repeated Domain 2 discriminator: value you did not baseline is value you cannot claim. When a stem says a team wants to prove ROI after launch and never measured the before-state, the correct answer reconstructs a baseline (§9) — an after-only number is not evidence.
How the metrics fit together
BUSINESS OBJECTIVE (e.g. cut cost-to-serve 15%) │ ▼ VALUE KPI (lagging) ── cost per contact, ROI, revenue │ ▼ OPERATIONAL KPI ──── handle time, containment, rework rate │ ▼ ADOPTION / HEALTH (leading) ── activation, weekly active use, │ task completion, quality score ▼ INSTRUMENTED SIGNAL ── events you can actually countRead it top-down to design, bottom-up to explain. Executives care about the objective; the leading indicators near the bottom are what tell you three weeks in whether the lagging value at the top will ever arrive.
1. KPIs by business function
Each KPI below carries a definition, a formula, where the baseline comes from, how to set a target, and the way it gets gamed. Pick the smallest set that ties to the objective; more metrics is not more insight.
| KPI | Definition / formula | Baseline source | Target guidance | Gaming risk |
|---|---|---|---|---|
| Containment rate | Contacts resolved without a human / total contacts | Current self-service logs | Set below 100%; over-containment traps customers | Deflecting hard cases into a dead end to boost the ratio |
| Average handle time (AHT) | Total handle time / contacts handled | Historical CRM timestamps | Realistic minutes saved, not zero | Rushing to close, hurting resolution quality |
| First-contact resolution | Resolved on first contact / total | QA sampling | Improve, not maximise at cost of accuracy | Marking reopened issues as new |
| CSAT / CES | Survey score after AI-handled contact | Pre-AI survey mean | Match or beat human baseline | Cherry-picking who gets surveyed |
| KPI | Definition / formula | Baseline source | Target guidance | Gaming risk |
|---|---|---|---|---|
| Lead conversion | Won / qualified leads | CRM historical cohort | Lift over matched cohort | Re-labelling weak leads as unqualified |
| Content throughput | Assets produced per period | Prior quarter output | Volume with a quality gate | Volume of low-value drafts |
| Campaign ROI | (Revenue − cost) / cost | Comparable prior campaign | Beat the control campaign | Attributing organic revenue to AI |
| Cycle time to publish | Brief-to-publish days | Historical average | Days saved | Skipping review to shorten cycle |
| KPI | Definition / formula | Baseline source | Target guidance | Gaming risk |
|---|---|---|---|---|
| Time-to-insight | Days from question to validated finding | Historical project logs | Reduce, with validity check | Counting unvalidated outputs |
| Experiments per quarter | Completed experiments per period | Prior year cadence | Increase throughput | Splitting one experiment into many |
| Literature-review time | Hours per review | Time-and-motion sample | Hours saved | Shallow reviews to save hours |
| KPI | Definition / formula | Baseline source | Target guidance | Gaming risk |
|---|---|---|---|---|
| Cycle time | Commit-to-deploy time | DORA baseline | Reduce without defect rise | Merging unreviewed code |
| Change failure rate | Failed changes / total | Incident history | Hold or reduce | Reclassifying failures |
| Review throughput | PRs reviewed per period | Prior sprint average | Increase with quality gate | Rubber-stamping reviews |
| Defect escape rate | Defects found in prod / total | QA history | Reduce | Under-reporting prod defects |
| KPI | Definition / formula | Baseline source | Target guidance | Gaming risk |
|---|---|---|---|---|
| Cost per transaction | Processing cost / transactions | GL and headcount data | Reduce unit cost | Shifting cost to another line |
| Close cycle time | Days to close the books | Historical close logs | Reduce days | Deferring reconciliations |
| Forecast accuracy | 1 − |actual − forecast| / actual | Prior forecast error | Improve accuracy | Padding forecasts to look accurate |
| Invoice exception rate | Exceptions / invoices | AP system history | Reduce exceptions | Auto-approving to cut exceptions |
| KPI | Definition / formula | Baseline source | Target guidance | Gaming risk |
|---|---|---|---|---|
| Time-to-hire | Days from req to accept | ATS history | Reduce, watch quality-of-hire | Lowering the bar to fill fast |
| Onboarding time-to-productive | Days to ramp | Manager assessment baseline | Reduce | Declaring ramp early |
| Forecast error (demand) | MAPE vs actuals | Historical demand data | Reduce MAPE | Smoothing to hide misses |
| On-time-in-full (OTIF) | Orders delivered on time and complete / total | ERP history | Improve | Loosening the on-time window |
Every table shares one rule: the baseline source is a before number, and every KPI has a gaming risk. A KPI with no gaming risk column is a KPI you have not thought hard enough about.
2. Tangible vs intangible benefits
Task 2.2.1 splits benefits explicitly. Tangible benefits convert to currency directly; intangible ones do not, which is why they get dropped from business cases — and why the mistake is not to make them defensible.
| Type | Examples | How to defend it |
|---|---|---|
| Tangible | Cost reduction, revenue growth, hours saved × loaded rate, error-cost avoided | Tie to a GL line or a headcount rate; show before and after |
| Intangible | Customer satisfaction, employee productivity, brand trust, decision speed, risk reduction | Proxy it: CSAT delta, retention lift, survey-measured time saved, incidents avoided × expected cost |
The technique for intangibles is to attach a measurable proxy and, where possible, a conservative monetary bridge. “Employees are happier” is not defensible; “eNPS rose 8 points and regretted attrition fell 3 points, worth roughly one avoided backfill per quarter at £X” is. Always mark such figures as indicative.
| Intangible | Defensible proxy | Monetary bridge (conservative) |
|---|---|---|
| Employee productivity | Self-reported + sampled hours saved per week | Hours × loaded hourly rate × active users |
| Customer satisfaction | CSAT / CES / NPS delta on AI-handled interactions | Retention lift × customer lifetime value |
| Decision quality/speed | Time-to-decision; rework rate | Value of faster cycle; cost of avoided rework |
| Risk reduction | Incidents avoided vs baseline rate | Expected loss avoided × probability |
3. Leading vs lagging indicators
Task 2.2.4 asks you to identify leading indicators that predict success. Lagging indicators confirm value after the fact; leading indicators tell you early whether it is coming. Pair them.
| Programme goal | Leading indicator (early, predictive) | Lagging indicator (confirms value) |
|---|---|---|
| Support cost reduction | Weekly active agents using the assistant; task-completion rate | Cost per contact; AHT; CSAT |
| Sales lift | Reps adopting the tool; assisted opportunities created | Conversion rate; revenue per rep |
| Developer velocity | Suggestions accepted; PRs assisted | Cycle time; change-failure rate |
| Content programme | Drafts started with AI; review pass rate | Publish cycle time; campaign ROI |
If a leading indicator is flat three weeks in, the lagging value will not arrive — that is the point of watching it. A programme that reports only lagging metrics finds out too late.
4. Adoption and health metrics
Value requires use. These metrics tell you whether the tool is actually being adopted well, and they are the leading indicators most programmes forget to instrument.
| Metric | Definition | What it warns you about |
|---|---|---|
| Activation | Users who reached first successful use / provisioned | Onboarding friction; licences bought but unused |
| Weekly active use | Distinct users using it in a rolling week / target users | Pilot enthusiasm fading; no habit forming |
| Task completion rate | Tasks finished with AI / tasks started | The tool fails on real work, not demos |
| Containment rate | Handled without human handoff / total | Over- or under-automation |
| Escalation rate | Handoffs to a human / total | Where the tool hits its limits |
| Rework rate | Outputs needing correction / total | Quality problem masquerading as productivity |
| Quality sampling score | Human-graded sample against a rubric | The number CSAT can hide |
Containment and escalation are two sides of one coin: read them together, because a high containment rate with a rising rework rate means you are trapping users, not helping them.
5. A worked KPI tree
Start from the objective and decompose until you reach something you can count. This is the artefact that connects an executive goal to an instrumented signal, and it is the shape the exam rewards.
OBJECTIVE: Reduce cost-to-serve in support by 15% in 12 months │ ├── Value KPI: Cost per contact (lagging) │ │ │ ├── Driver: Contacts handled without a human → Containment rate │ ├── Driver: Time per human-handled contact → AHT │ └── Guardrail: Customer satisfaction → CSAT (must not fall) │ └── Leading indicators (weeks 1–6): ├── Weekly active agents using the assistant ├── Task-completion rate on real tickets └── Escalation rate trend (falling = tool coping)Worked arithmetic: 500,000 contacts/year at £6.00 each = £3.0m. A 30% containment rate at £0.40 per contained contact, with the remaining 70% at a reduced £5.40 AHT-adjusted cost, gives 500,000 × (0.30 × 0.40 + 0.70 × 5.40) = 500,000 × (0.12 + 3.78) = £1.95m — a 35% cost reduction if CSAT holds. The guardrail metric is why the tree includes CSAT: a cost win that drops satisfaction is not a win.
6. Baselines: constructing one when none exists
Task 2.2.2 requires a baseline before implementation, but the common real-world blocker is that no baseline was ever recorded. There are four defensible ways to build one; pick by data availability and cost.
| Method | How | When to use | Weakness |
|---|---|---|---|
| Time-and-motion sample | Observe/measure a representative sample of the current process | No historical logs exist | Sampling bias; observer effect |
| Historical proxy | Use existing system logs (CRM, ERP, ATS) as the before-state | Timestamps/records already captured | Proxy may not match the exact metric |
| Control group | Run AI for one group, hold another unchanged, compare | You can split fairly and ethically | Contamination; group comparability |
| Staged rollout | Compare cohorts before and after phased enablement | Big-bang launch is risky anyway | Time trends confound the comparison |
The exam’s preferred answer to “we have no baseline” is build one now by the cheapest credible method — never “measure after launch and assume the difference is the AI”. A control group or staged rollout is the strongest because it isolates the AI effect from background change.
7. Metrics that look good and mean nothing
Vanity metrics feel like progress and predict nothing. Recognising them is a direct exam skill.
| Vanity metric | Why it is tempting | Replace with |
|---|---|---|
| Number of prompts / queries run | Big number, easy to pull | Task-completion rate; value KPI |
| Licences purchased | Looks like adoption | Weekly active use; activation |
| Model accuracy in isolation | Technical and impressive | Business outcome the accuracy serves |
| Total outputs generated | Shows the tool is “busy” | Outputs that passed review; rework rate |
| Pilot NPS from volunteers | Enthusiastic early adopters | Sampled quality score across all users |
The tell of a vanity metric: it can rise while the business outcome is flat or falling. If a metric can go up while cost-per-contact and CSAT do nothing, it is measuring activity, not value.
Key takeaways
- Choose the smallest KPI set that ties to the objective; every KPI needs a formula, a before-baseline and an acknowledged gaming risk.
- Make intangible benefits defensible with a measurable proxy and a conservative, clearly-labelled monetary bridge — do not drop them.
- Pair a leading indicator (adoption, task completion) with each lagging value KPI so you learn early, not late.
- Read containment against escalation and rework; a high containment rate with rising rework means you are trapping users.
- Build the KPI tree from objective to instrumented signal, and keep a guardrail metric (like CSAT) so a cost win cannot quietly destroy quality.
- With no baseline, construct one now — time-and-motion, historical proxy, control group or staged rollout — never measure after launch and assume the delta is the AI.
- A metric that can rise while the business outcome is flat is a vanity metric; replace it.
Last updated Sep 18, 2026