Applied AI Foundations
D6 · Repeatability and Improvement
Documenting a workflow so a colleague can run it, versioning prompts, measuring quality and cycle time, and iterating on evidence rather than vibes.
This domain is worth 14% of the mock — roughly 7 of 50 items. It is where “a thing I do” becomes “a thing our team does”: a workflow that is documented so someone else can run it, versioned so changes are traceable, measured so you know whether it is any good, and improved on evidence rather than on how it feels. It is the payoff of the whole track — decomposition, contracts, capability choice and oversight only compound if the workflow is repeatable and gets better over time.
What you need to know
A repeatable workflow is one a colleague can run end to end without asking you a question, because it is documented (a runbook: trigger, steps, inputs, outputs, review points, what to do when it fails). Prompts and configurations are versioned — dated, with a change note — so you can tell what changed and roll back if a change makes things worse. Quality and cycle time are measured against the acceptance criteria and the clock, so improvement is judged on numbers, not vibes. Iteration is evidence-based: change one thing, measure, keep it only if the metric improved. The failure mode is the workflow that lives only in the author’s head, changes silently, and is “improved” by intuition with no way to tell if it actually got better.
Learning objectives
By the end of this page you should be able to:
- Document a workflow as a runbook a colleague can execute without you.
- Version prompts and configurations so changes are traceable and reversible.
- Measure a workflow on quality (against acceptance criteria) and cycle time.
- Iterate on evidence: change one variable, measure, keep or revert.
- Recognise vibes-based “improvement” and the single-author bottleneck as anti-patterns.
6.1 The runbook — documenting for handover
A workflow that only you can run is a liability, not an asset. The test is blunt: could a colleague run this end to end without asking you a single question? If not, it is not repeatable. A runbook makes it so.
RUNBOOK: Weekly deal-risk digest TRIGGER Monday 09:00, or on demand INPUTS this week's rep notes; CRM export; (Project holds style guide + template) STEPS 1. Extract slipped deals (data analysis on the export) 2. Pull latest news per top-5 account (search) 3. Draft digest in the house template 4. REVIEW gate: check figures reconcile, tone on-brand OUTPUT one-page digest posted to #sales DONE digest posted; no unreconciled figures ON FAILURE if CRM export missing → request it, do not estimate OWNER Sales Ops; backup: Ops lead| Runbook element | Why it must be written |
|---|---|
| Trigger & inputs | So the runner knows when and with what |
| Steps (with capability) | So the runner reproduces the workflow, not improvises |
| Review points | So oversight survives the handover |
| Definition of done | So the runner knows when to stop |
| On failure | So the runner doesn’t guess when something is missing |
| Owner & backup | So the workflow has a maintainer |
Assessment signal
“A colleague needs to run this while you’re on leave” or “the workflow only works when you do it” is asking for documentation/runbook. The right answer writes down the steps, inputs, review points and failure handling — not “record a video of me doing it once” or “they’ll figure it out”.
6.2 Versioning prompts and configuration
Prompts and Project/GPT configurations drift as you tweak them. Without versioning you cannot answer two essential questions: what changed? and can we go back?
| Practice | What it buys you |
|---|---|
| Date each prompt version | You know which version produced which results |
| Write a one-line change note | You know why it changed |
| Keep the previous version | You can roll back a regression |
| Change one thing per version | You can attribute the effect to the change |
Worked example. A support-reply prompt is edited to “be more concise.” Replies get shorter — and the customer-satisfaction sample dips because they now omit a needed next step. Because the change was a single dated version with a note, you roll back to the prior version and instead change only the length guidance while keeping the next-step instruction. Without versioning you would be guessing which of several edits caused the dip.
Assessment signal
“After several edits the workflow got worse and no one knows why” is a versioning failure. The right answer keeps dated versions with change notes and changes one variable at a time so effects are attributable — not “revert everything and start over.”
6.3 Measuring quality
You cannot improve what you do not measure, and “it seems better” is not a measurement. Quality is measured against the step’s acceptance criteria from Domain 3.
| Quality signal | How to measure it |
|---|---|
| Correctness | Sampled error rate against acceptance criteria |
| Completeness | Fraction of outputs meeting the full checklist |
| Rework rate | How often output is edited/rejected at the review gate |
| Downstream complaints | Issues traced back to the workflow |
The review gate and sampling from Domain 5 are your data source: every sampled item is a quality measurement. A workflow with a sampling regime already has a quality metric — the sampled error rate — and just needs it tracked over time.
6.4 Measuring cycle time
Quality is half the picture; the other half is cycle time — how long the workflow takes end to end, and how much of that is human vs automated. Measuring it tells you where the value (and the bottleneck) actually is.
Manual baseline: ██████████████████████ 120 min/run (all human)After workflow: ████ 20 min/run ── of which ── ██ 12 min automated ██ 8 min human reviewTwo uses:
- Justify the workflow. 120 → 20 minutes is the payback you screened for in Domain 1.
- Find the next improvement. If 8 of the 20 minutes are human review, the biggest remaining lever is the review design (Domain 5), not a better prompt. Measure before you optimise, or you will optimise the wrong step.
6.5 Iterating on evidence, not vibes
Improvement is a loop: hypothesis → change one thing → measure against the metric → keep if better, revert if not. “Vibes” improvement — changing several things because the output “feels” better — cannot tell you what helped, cannot be defended, and often makes things quietly worse.
EVIDENCE-BASED VIBES-BASED (anti-pattern) 1 form a hypothesis 1 tweak several things at once 2 change ONE variable 2 read a couple of outputs 3 measure against the metric 3 "feels better", ship it 4 keep or revert on the number 4 no metric, no attribution, no rollback| Evidence-based | Vibes-based | |
|---|---|---|
| Variables changed | One at a time | Several at once |
| Decision basis | The measured metric | Impression from a few outputs |
| Reversible? | Yes — versioned | No — untracked |
| Learns over time? | Yes | No |
Assessment signal
“We changed the prompt and it feels better” versus “we A/B’d one change and the error rate dropped from 9% to 4%” — the exam rewards the measured, single-variable, attributable answer every time.
6.6 Closing the loop with the other domains
Repeatability is where the track’s pieces connect: acceptance criteria (D3) are the quality metric; the review sampling (D5) is the data source; the runbook documents the decomposition (D2) and capability choices (D4); and the measured payback validates the opportunity you scoped (D1). A workflow that is documented, versioned, measured and iterated is the finished product the whole Applied AI course is teaching you to build — and it is what a colleague can run, a manager can trust, and you can keep improving without guessing.
Decision framework
The DRIVE loop — the maintenance cycle for a live workflow.
| Letter | Step | What you do |
|---|---|---|
| D — Document | Runbook | Write it so a colleague runs it without asking you |
| R — Record versions | Version control | Date each prompt/config change with a one-line note |
| I — Instrument | Measure | Track quality (vs acceptance criteria) and cycle time |
| V — Vary one thing | Iterate | Change a single variable per experiment |
| E — Evaluate | Decide | Keep it if the metric improved; revert if not |
Run DRIVE whenever the workflow is live and being improved. The discipline is in V and E: one change, measured, kept or reverted on the number — never a bundle of tweaks judged by feel.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| The workflow lives only in the author’s head | It works when they run it | Write a runbook a colleague can execute unaided |
| No prompt/config versioning | Editing in place is faster | Date versions, note the change, keep the prior version |
| Changing several things then judging by feel | It seems efficient | Change one variable, measure against the metric |
| “It feels better” as the improvement test | Vibes are easy | Measure error rate / rework / cycle time; decide on the number |
| No cycle-time measurement | Only quality feels important | Measure time end to end to find the real bottleneck |
| Optimising the prompt when review is the bottleneck | Prompting is the familiar lever | Measure first; optimise the step that actually costs time |
| No definition of done in the runbook | The author “just knows” | Write the done condition so the runner knows when to stop |
| No rollback path after a bad change | Versioning was skipped | Keep prior versions so a regression can be reverted |
Scenario challenge
Scenario. Sofia built a contract-review workflow that flags risky clauses and drafts a summary memo. It saves her about 90 minutes per contract and she is proud of it. Three things are now going wrong. First, she is going on leave and her backup cannot run it — it exists as a set of prompts Sofia has “in her head” and pastes from a scratch doc. Second, over the last month she has edited the flagging prompt “a bunch of times” to catch more clause types, and lately it flags far more false positives, but she can’t tell which edit caused it. Third, her manager asks “is it actually better than before?” and Sofia can only say it “feels more thorough.” She wants to keep tuning the prompt until it feels right.
Expert reasoning trace.
- Diagnose all three as repeatability failures, not prompt-quality failures. The backup can’t run it (documentation gap), the false-positive regression is untraceable (versioning gap), and “feels more thorough” (measurement gap). None is fixed by more prompt tuning; tuning-by-feel is precisely what created the regression.
- Fix handover with a runbook. Write trigger, inputs, the ordered steps with their capabilities, the review gate, the definition of done, and what to do when a contract is missing a section — so the backup runs it unaided. “Record a video” or “they’ll figure it out” fails the no-questions test.
- Introduce versioning to isolate the regression. The false positives appeared after “a bunch of” undated edits, so there is no way to attribute or roll back. Adopt dated versions with one-line change notes, restore the last version known to have a low false-positive rate, then re-apply the individual clause-type additions one at a time, measuring after each, to find the edit that introduced the noise.
- Instrument quality and cycle time. Define the acceptance criteria (which clause types must be flagged; acceptable false-positive rate) and measure the sampled error/false-positive rate and the end-to-end time. Now “is it better?” has an answer: a number, before and after, against a baseline.
- Switch to evidence-based iteration. Replace “tune until it feels right” with the DRIVE loop: hypothesise, change one clause-type rule, measure the false-positive rate on a fixed sample, keep or revert. This both fixes the current regression and prevents the next one.
- Answer the manager properly. Not “it feels more thorough” but “against a 40-contract sample it flags 96% of the target clause types (baseline was manual) at an 8% false-positive rate, down from 19% last week after we reverted the bad edit, and cuts review time from ~110 to ~20 minutes.”
The decision: treat the problems as documentation, versioning and measurement gaps — write the runbook, adopt dated one-change-at-a-time versioning to isolate and revert the regression, and instrument quality and cycle time so iteration is evidence-based — not “keep tuning until it feels right,” which is the vibes anti-pattern that caused the regression in the first place.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| “Record a video of me doing it” as documentation | It captures the steps loosely | A runbook (steps, inputs, review, done, failure handling) is what lets a colleague run it unaided |
| “Keep tuning the prompt until it feels right” | Iteration feels like progress | Vibes iteration can’t attribute or measure; change one thing and measure |
| “It feels more thorough” as proof of improvement | Impressions are immediate | Improvement must be shown against a metric (error rate, rework, cycle time) |
| “Revert everything and start over” after a regression | It seems like a clean slate | Versioning lets you roll back to the last good version and re-apply changes one at a time |
| “Optimise the prompt” when review is the bottleneck | Prompting is the familiar lever | Measure cycle time first; optimise the step that actually costs the time |
| “The author just knows when it’s done” | It works for the author | Without a written definition of done the workflow can’t be handed over |
Practice questions
Each item states how many responses to select. Commit before revealing.
Q1 · You go on leave and your backup cannot run your workflow because it only exists in your head. What is the BEST fix? (Select one)
A. Tell them to figure it out from the chat history B. Write a runbook: trigger, inputs, ordered steps with capabilities, review points, definition of done, and failure handling C. Record a quick video once and hope it covers everything D. Wait until you’re back
Answer: B. A written runbook is what lets a colleague run the workflow end to end without asking questions. ‘Figure it out’ (A) and ‘wait’ (D) leave the workflow un-runnable. A one-off video (C) rarely captures inputs, failure handling and the definition of done reliably.
Q2 · After several undated prompt edits, a workflow got worse and no one knows which change caused it. What practice would have prevented this? (Select one)
A. Using a bigger model B. Versioning: dated versions with a one-line change note, changing one variable at a time C. Longer prompts D. Reviewing every output
Answer: B. Dated single-change versions make effects attributable and enable rollback. A bigger model (A) and longer prompts (C) don’t address traceability. Reviewing every output (D) catches errors but doesn’t attribute which edit caused them.
Q3 · A manager asks whether an improved workflow is actually better. Which answer reflects evidence-based iteration? (Select one)
A. “It feels more thorough now” B. “On a fixed 40-item sample the error rate fell from 19% to 8% and cycle time from 110 to 20 minutes” C. “I changed a lot of things and I like the output more” D. “The new prompt is longer”
Answer: B. Improvement is shown against measured metrics on a fixed sample. ‘Feels more thorough’ (A) and ‘I like it more’ (C) are vibes. Prompt length (D) is not a quality measure.
Q4 · What is the correct way to iterate on a workflow? (Select one)
A. Change several things at once so it improves faster B. Change one variable, measure against the metric, keep it if better or revert if not C. Change things until the output feels right D. Never change a working workflow
Answer: B. Single-variable, measured, keep-or-revert iteration is what makes improvement attributable and reversible. Changing several at once (A) and tuning by feel (C) prevent attribution. ‘Never change’ (D) forgoes improvement entirely.
Q5 · Measuring cycle time on a workflow shows 8 of its 20 minutes are human review. What does this tell you about the next improvement? (Select one)
A. Make the prompt longer B. The biggest remaining lever is the review design, not the prompt C. Switch to a cheaper model D. Nothing useful
Answer: B. Cycle-time measurement locates the bottleneck; with review dominating, review design is the lever, not prompting. Prompt length (A) and model price (C) don’t touch the review time. The measurement is highly useful (D).
Q6 · Which is the best data source for measuring a workflow's ongoing quality? (Select one)
A. The author’s overall impression B. The sampled items from the review regime, scored against the acceptance criteria C. The number of prompt words D. The model’s release notes
Answer: B. The sampling regime already generates quality measurements when scored against acceptance criteria. The author’s impression (A) is vibes. Prompt length (C) and release notes (D) aren’t quality data.
Q7 · A support-reply prompt was edited to 'be more concise' and satisfaction dropped because replies now omit the next step. With versioning in place, what is the BEST response? (Select one)
A. Rewrite the whole workflow from scratch B. Roll back to the prior version, then change only the length guidance while keeping the next-step instruction C. Keep the concise version; satisfaction will recover D. Add more knowledge files
Answer: B. Versioning lets you revert the regression and re-apply a single, isolated change. Starting over (A) discards working history. Keeping a version that measurably hurt satisfaction (C) is not evidence-based. Knowledge files (D) don’t address the omitted next step.
Q8 · Which TWO elements are essential in a runbook so a colleague can run the workflow unaided? (Select two)
A. A definition of done B. The author’s personal opinion of the output C. What to do when a required input is missing D. The model’s training cutoff date E. The number of times the author has run it
Answer: A and C. A definition of done tells the runner when to stop, and failure handling tells them what to do when input is missing — both essential for unaided execution. The author’s opinion (B), the training cutoff (D) and a run count (E) don’t help a colleague execute it.
Q9 · A workflow cut a task from 120 to 20 minutes per run. What does this measurement primarily support? (Select one)
A. Nothing — time is irrelevant B. It quantifies the payback that justified automating the opportunity and provides a baseline for further improvement C. It proves the output quality is high D. It sets the model temperature
Answer: B. Cycle-time savings quantify the payback screened in Domain 1 and give a baseline to improve against. Time is very relevant (A). It does not by itself prove quality (C) — that needs the acceptance-criteria metric. It has nothing to do with temperature (D).
Q10 · A team keeps 'improving' a workflow by tweaking multiple settings whenever output feels off, with no metric. What TWO problems does this create? (Select two)
A. Improvements can’t be attributed to any specific change B. There is no way to tell if the workflow actually got better C. The model becomes physically slower D. Token pricing increases E. The workflow automatically documents itself
Answer: A and B. Multi-variable, metric-free tweaking prevents attribution and can’t demonstrate real improvement — the vibes anti-pattern. It doesn’t change model speed (C) or token pricing (D), and it certainly doesn’t self-document (E).
Q11 · How do the earlier domains feed repeatability? (Select one)
A. They don’t; repeatability is independent B. Acceptance criteria (D3) are the quality metric, the review sampling (D5) is the data source, and the runbook documents the decomposition (D2) and capability choices (D4) C. Only the model choice matters D. Only the prompt wording matters
Answer: B. Repeatability connects the track: criteria supply the metric, sampling supplies the data, and the runbook records the decomposition and capabilities. It is not independent (A). Model choice (C) and prompt wording (D) alone don’t make a workflow measurable and handover-ready.
Q12 · A false-positive regression appeared after many untracked edits. What is the BEST recovery, given you now adopt versioning? (Select one)
A. Delete the workflow and rebuild from memory B. Restore the last version with a low false-positive rate, then re-apply the individual changes one at a time, measuring after each C. Keep all edits and lower the temperature D. Add more clause types to overwhelm the false positives
Answer: B. Rolling back to a known-good version and re-applying changes one at a time with measurement isolates the offending edit and restores quality. Rebuilding from memory (A) loses the good history. Keeping the edits and changing temperature (C) doesn’t isolate the cause. Adding more rules (D) compounds the noise.
Key takeaways
- A repeatable workflow is documented as a runbook — trigger, steps with capabilities, review points, definition of done, failure handling, owner — so a colleague runs it without asking you a question.
- Version prompts and configurations: date them, note the change, keep the prior version, change one thing at a time — so effects are attributable and regressions are reversible.
- Measure quality against acceptance criteria (using the sampling regime as your data source) and measure cycle time to find the real bottleneck.
- Iterate on evidence, not vibes: hypothesis → one change → measure → keep or revert. “It feels better” is not a measurement.
- Measure before you optimise, or you will tune the prompt when the review step was the bottleneck.
- Repeatability closes the loop on the track: D3 criteria are the metric, D5 sampling is the data, the runbook records D2 and D4, and cycle-time payback validates the D1 opportunity.
- The DRIVE loop — Document, Record versions, Instrument, Vary one thing, Evaluate — is the maintenance cycle for a live workflow.
Last updated Sep 18, 2026