Appendix · AWS
AI Governance Toolkit
Ready-to-adapt governance artefacts for AIB-C01 — an operating model, RACI, intake form, risk register, risk-tier control matrix, tool-classification policy, launch and incident checklists, vendor due diligence, and how AWS controls map onto them.
This is a set of ready-to-adapt artefacts for the governance work Domain 3 tests: an operating model, a RACI, a use-case intake form, a risk register, a risk-tier to control matrix, a tool-classification policy, launch and incident checklists, and a vendor due-diligence set. It is independent preparation, not an official AWS compliance framework — adapt every artefact to your own regulatory obligations before you rely on it. The AIB-C01 beta exam does not ask you to configure any of this; it asks you to recognise which structure, control or accountability fits a scenario. Everything here stays at that decision level.
Governance by design
The recurring Domain 3 discriminator is when governance enters. The tempting answer reviews a system just before launch; the correct answer builds classification, oversight mode and monitoring into project planning from the start (task 3.1.3). When a stem describes a control added late, the better option almost always moved it earlier.
The toolkit at a glance
Use case ──▶ INTAKE FORM ──▶ RISK CLASSIFICATION ──▶ CONTROL SET │ │ │ │ ▼ ▼ │ RISK REGISTER HUMAN OVERSIGHT │ (owned rows) + MONITORING ▼ │ OPERATING MODEL (who decides) ◀── RACI ────────┤ │ ▼ ▼ LAUNCH READINESS TOOL CLASSIFICATION │ (approved/blocked/eval) ▼ INCIDENT RESPONSEEach artefact answers a different governance question. Read down the diagram: intake captures the request, classification sets the tier, the tier drives the control set, the operating model and RACI say who decides, and launch and incident procedures close the loop.
1. AI governance operating model
Governance fails in two opposite ways: a single committee that becomes a bottleneck, or no forum at all so decisions happen by accident. The operating model names a small number of bodies, gives each clear decision rights, and sets a cadence. It maps directly to task 3.2.1 (cross-functional representation and clear accountability).
| Body | Membership | May decide | May not decide | Cadence |
|---|---|---|---|---|
| AI steering committee | Executive sponsor, CIO/CTO, line-of-business heads, finance | Portfolio priorities, budget, risk appetite, scale/pause/terminate | Individual prompt wording, technical architecture | Quarterly |
| AI governance council | Legal, compliance, security, data owner, responsible-AI lead, business rep | Risk-tier definitions, policy, high-tier launch approval, escalation outcomes | Which vendor a team likes best without a risk view | Monthly + on escalation |
| Responsible-AI review | Responsible-AI lead, data scientist, domain expert, affected-user advocate | Fairness/oversight adequacy, mitigation sign-off for medium+ tiers | Business ROI targets | Per high/medium use case |
| Use-case working group | Product owner, technical lead, business analyst | Day-to-day delivery, low-tier launch within policy | Anything outside its tier’s delegated authority | Weekly |
The design rule is delegated authority by tier: low-risk use cases clear at the working-group level so governance does not throttle volume, while high-risk use cases rise to the council. A body that cannot say no is not a governance body; each row’s “may not decide” column is deliberate.
2. RACI across the AI lifecycle
A RACI stops the two classic accountability failures: everyone assumes someone else owns a control, or one person is accountable for everything and therefore for nothing. Responsible does the work, Accountable owns the outcome (exactly one per row), Consulted gives input before, Informed hears after.
| Activity | Business owner | Data owner | Responsible-AI lead | Legal / compliance | Security | Governance council |
|---|---|---|---|---|---|---|
| Use-case intake | A/R | C | C | I | I | I |
| Risk classification | C | C | R | C | C | A |
| Data approval | C | A/R | C | C | C | I |
| Model / build-buy-partner choice | A/R | C | C | C | C | I |
| Launch approval (high tier) | C | C | C | C | C | A/R |
| Production monitoring | A | R | C | I | C | I |
| Incident response | C | C | C | C | R | A |
The exam signal: when a stem says “no one was clearly responsible for X after launch”, the fix is a named A in the monitoring or incident row — not a new tool. Note that data approval is accountable to the data owner, not the business owner who wants to use it; that separation is what prevents a team from waving its own data through.
3. Use-case intake form
Intake is the cheapest place to say no. A structured form forces a proposal to state its outcome, its data and its risk before anyone spends money, and it feeds prioritisation (task 2.1.3) and classification (task 3.2.4).
- Sponsor and business owner — named individual, not a team.
- Problem and current baseline — what the process costs or takes today, measured (link to the KPI Library baseline methods).
- Proposed outcome and target — the specific metric that should move, and by how much.
- Why AI, not rules — the deterministic-vs-AI test (task 1.2.1); if the logic is fixed, this is a rule engine.
- Data required — sources, sensitivity, ownership, whether it may leave the org, quality state.
- Consequence of a wrong output — reversible or not, who is affected, regulated or not.
- Human-in-the-loop point — where a person reviews before an action is taken.
- Build / buy / partner leaning — with a rough cost and timeline.
- Success and kill criteria — the numbers that would make you scale, pause or terminate.
A proposal that cannot fill “consequence of a wrong output” and “kill criteria” is not ready for a budget decision. That is the intake gate doing its job.
4. Risk register template with worked rows
The register is the living record the exam expects to exist before production, not after an incident. Each row names an owner, a likelihood and impact, the current control and a residual rating. These worked rows span the risk categories the guide lists (accuracy, privacy, security, IP, regulatory, reputational, workforce).
| ID | Risk category | Description | Likelihood | Impact | Owner | Mitigation / control | Residual |
|---|---|---|---|---|---|---|---|
| R-01 | Accuracy / reliability | Support assistant gives wrong policy answers (hallucination) | Medium | High | Support lead | Contextual grounding check against policy KB; escalation on low confidence | Medium |
| R-02 | Privacy | Customer PII pasted into prompts and retained | Medium | High | Data owner | PII redaction filter; retention set to minimum; staff training | Low |
| R-03 | Security | Prompt-injection via user-supplied text drives unintended action | Low | High | Security | Guardrails denied topics; no autonomous write actions; human approval | Low |
| R-04 | Intellectual property | Generated marketing copy reproduces third-party protected content | Medium | Medium | Legal | Output review for regulated claims; vendor IP indemnity confirmed | Medium |
| R-05 | Regulatory / compliance | Automated eligibility decision falls under regulated decisioning | Low | High | Compliance | Risk-tier = high; human decision-maker retained; audit log | Low |
| R-06 | Reputational | Public-facing bot produces an offensive or off-brand response | Medium | High | Brand / comms | Content filters; red-team before launch; kill switch and holding message | Medium |
| R-07 | Workforce | Staff fear role loss and disengage from the rollout | High | Medium | Change lead | Transparent comms on role change to oversight; reskilling path | Medium |
The teaching point is that likelihood-times-impact drives the tier, the tier drives the control, and the residual rating is what the governance council actually signs off — not the raw risk. A row with no named owner is not a managed risk.
5. Risk-tier to control-set matrix
This is the artefact that turns “apply a risk classification framework” (task 3.2.4) into something a reader can use tomorrow. Classify the use case by consequence and reach, then read the control set across the row. It mirrors the tiered logic of the EU AI Act and the NIST AI RMF without citing article numbers, which is exactly the level the exam tests.
| Tier | Typical use case | Approvals | Human oversight mode | Monitoring cadence | Documentation | Escalation |
|---|---|---|---|---|---|---|
| Minimal | Internal draft assistance, brainstorming | Working group | Human-in-the-loop optional; author reviews own output | Spot-check quarterly | Use-case entry only | To council if scope grows |
| Limited | Internal knowledge search, summarisation | Working group + responsible-AI note | Human reviews before external use | Monthly sampling | Intake + register rows | To council on incident |
| High | Customer-facing decisions, content, or advice | Governance council | Human-in-the-loop on each consequential action | Weekly metrics + drift watch | Full register, evaluation evidence, AI Service Card review | Immediate to council + sponsor |
| Unacceptable | Manipulative, unlawful, or safety-critical without recourse | Do not proceed | N/A | N/A | Rejection recorded | Terminate at intake |
The oversight column carries the exam’s core distinction: human-in-the-loop means a person approves before the action, human-on-the-loop means a person monitors and can intervene, and human-in-command means a person can override and shut down. Higher consequence pushes you up that ladder.
6. AI tool classification policy — the answer to shadow AI
Task 1.2.4 is explicit: a transparent classification of tools as approved / blocked / under evaluation is the concrete answer to shadow AI. Banning everything drives usage underground; approving everything abandons control. The policy names states and the criteria that move a tool between them.
A tool is approved for a stated set of use cases when it has cleared due diligence: data-use terms acceptable, no training on your data by default, residency and retention meet policy, security review passed, and an owner is named. Approval is scoped — approved for internal drafting is not approved for customer PII. Movement out: a term change, an incident, or a failed re-review moves it to under evaluation or blocked.
A tool is under evaluation while due diligence is in progress. Staff may pilot it only within a sandbox and with non-sensitive data. It moves to approved when the vendor question set (§9) is answered acceptably and an owner accepts the residual risk; it moves to blocked if a hard criterion fails (e.g. trains on your data with no opt-out, unacceptable residency).
A tool is blocked when a hard criterion fails or an incident makes it unsafe. Blocking is only credible if there is an approved alternative for the same job — otherwise staff route around it. The policy therefore pairs every block with a sanctioned substitute and a fast path to request evaluation of something new.
request ──▶ UNDER EVALUATION ──▶ APPROVED (scoped) │ │ hard criterion term change / fails incident ▼ ▼ BLOCKED ◀────────────┘ │ re-request ──▶ UNDER EVALUATIONThe exam signal for shadow AI is always visibility plus a sanctioned path, never a blanket ban.
7. Launch readiness checklist
Before a use case goes live, the sponsor confirms each item. This operationalises “envision → launch” from AWS CAF and the transition to production-grade from task 4.4.5.
- Risk tier assigned and control set applied.
- Baseline metric captured before go-live (you cannot claim value later without it).
- Human-oversight point defined and staffed for the chosen tier.
- Guardrails / content filters configured and tested against known-bad inputs.
- Monitoring and drift-detection in place with named owner and alert thresholds.
- Kill switch and rollback path tested; a holding message ready.
- Incident runbook written and the response owner (RACI A) briefed.
- Data terms, residency and retention confirmed; vendor indemnity on file.
- Success and kill criteria agreed by the steering committee.
- Workforce communication sent; affected staff know what changes.
8. Incident response outline for AI failures
AI incidents differ from classic IT outages: the system stays “up” while producing harmful, biased or wrong output. The outline names phases and owners.
Sources: monitoring alert (accuracy drop, drift, bias-drift), user or customer complaint, red-team finding, or a regulatory query. Log time, use case, tier and suspected category (accuracy, bias, privacy, security, IP, reputational).
For high-tier systems, invoke the kill switch or fall back to the human process; post a holding message. Containment authority sits with the incident A in the RACI, not with the delivery team, so containment is not delayed by a debate.
Classify severity by consequence and reach. Notify the governance council; notify legal and compliance early if PII, a regulated decision or IP is involved, because notification clocks may run. Preserve logs and prompts as evidence.
Fix root cause (data, prompt, guardrail, oversight gap — not just the symptom), re-test, and stage a controlled restart. Add a register row and a control so the same class of failure is caught earlier next time. Feed the lesson into intake and the launch checklist.
9. Vendor due-diligence question set
Build-buy-partner (task 2.1.2) usually lands on buy or partner, which makes vendor scrutiny a governance control, not a procurement formality. Ask every vendor:
| Area | Question | A weak answer looks like |
|---|---|---|
| Data use | How is our data used, and is it used to improve or train your models? | “It may be used to improve the service” with no opt-out |
| Training | Is training on our data off by default, and can we contractually prohibit it? | Off “on request” only, or unclear |
| Retention | How long is our content retained, and can we set it to zero / minimum? | Fixed long retention with no control |
| Residency | In which regions is our data stored and processed? | Cannot commit to a region we require |
| Indemnity | Do you indemnify us against third-party IP claims on generated output? | No IP indemnity, or heavily capped |
| Evaluation evidence | What evaluation, bias and safety evidence can you share? | Marketing claims, no methodology or data |
| Roadmap risk | What is your model deprecation and change policy, and notice period? | Models change with little notice; lock-in risk |
| Sub-processors | Who are your sub-processors and where do they operate? | Undisclosed or non-committal |
The discriminator on a vendor-selection item is rarely price; it is usually a data-use, indemnity or residency answer that a cheaper vendor cannot give.
10. How AWS’s own controls map onto these artefacts
The exam does not require you to configure any AWS control, but it expects you to recognise which AWS capability supports which governance artefact. Keep this at the mapping level.
| Governance need | AWS capability (strategic view) | Where it maps |
|---|---|---|
| Stop harmful, off-topic or unsafe output | Amazon Bedrock Guardrails — content filters, denied topics, PII redaction, contextual grounding (hallucination) checks | Risk register controls; launch checklist; incident containment |
| Detect and explain bias | Amazon SageMaker Clarify — bias detection across data prep, after training and in the deployed model, plus explainability | Responsible-AI review; register bias rows; tier control set |
| Catch drift and inaccurate predictions in production | Amazon SageMaker Model Monitor — alerts on inaccurate predictions from deployed models | Monitoring cadence rows; incident detection |
| Transparency on intended use and limits | AWS AI Service Cards — intended use cases, limitations, responsible-AI design choices | Vendor evidence; high-tier documentation |
| Who is accountable for what | AWS shared responsibility model for AI — AWS secures the cloud; the customer owns data, access, use-case appropriateness and whether output is fit for purpose | Operating model and RACI accountability |
The single most-tested idea in this table: no managed service removes the customer’s accountability for what the system is used for and whether its output is fit for purpose. Guardrails and Clarify are tools that support your controls; they are not a substitute for the operating model, the tiering and the human oversight that you own.
Key takeaways
- Governance is a set of bodies with decision rights and a cadence, not one committee; delegate low tiers so volume is not throttled.
- Every managed risk and every lifecycle activity has exactly one accountable owner — the fix for “no one owned it” is a name, not a tool.
- Classify first: tier drives approvals, oversight mode, monitoring cadence, documentation and escalation.
- Shadow AI is answered by a transparent approved / blocked / under-evaluation policy with a sanctioned path and a substitute for anything blocked.
- Capture the baseline before launch and write the incident runbook before you need it.
- Vendor due diligence — data use, training, retention, residency, indemnity, evidence, roadmap — is a governance control, not paperwork.
- AWS Guardrails, Clarify, Model Monitor and AI Service Cards support these artefacts; shared responsibility keeps use-case and fitness accountability with you.
Last updated Sep 18, 2026