AI Cert Prep
Type to search documentation.

Agents and Workflows

D4 · Boundaries and Guardrails

Setting the limits an agent operates inside — irreversible actions, approval gates, spend and time limits, data boundaries — and the difference between a boundary the agent respects and one the system enforces.

Worth 16% — about 8 of 50 items. This domain tests whether you can bound an agent so that when it misunderstands the task (and eventually it will), the damage is contained. The single most important idea here is the distinction between a boundary the agent is asked to respect and one the system enforces. On an irreversible action, a request is not a control; enforcement is. Get that distinction right and most of the domain follows.

What you need to know

Boundaries are the limits you place around an agent’s autonomy. The highest-stakes boundaries protect against irreversible actions — sending, publishing, paying, deleting, changing a system of record. The primary control for these is an approval gate: the agent pauses and a human authorises before the action fires. Other boundaries cap spend and time so a runaway or looping agent can’t burn budget or run forever, and enforce data boundaries so it can’t move sensitive data where it shouldn’t. The crucial distinction is between a respected boundary (the agent is instructed not to cross it and usually won’t) and an enforced boundary (the system makes crossing it impossible). For anything irreversible or sensitive, you need enforcement — a literal, autonomous agent that misreads the goal can cross a merely-requested line without knowing it did anything wrong.

Learning objectives

By the end of this page you should be able to:

  1. Identify irreversible actions and place approval gates before them.
  2. Distinguish a boundary the agent respects from one the system enforces, and know when only enforcement will do.
  3. Set spend and time limits appropriate to a task’s stakes.
  4. Define data boundaries that stop sensitive data crossing lines it shouldn’t.
  5. Design an approval gate that pauses at the right point with enough context to decide.
  6. Choose enforcement over instruction whenever an action is irreversible or a boundary protects against real harm.

4.1 Irreversible actions are the whole game

Every boundary decision starts by asking: what can this agent do that I can’t take back? Reversible actions are forgiving — a bad draft is deleted, a wrong analysis re-run. Irreversible actions are not, and they are exactly where an agent’s autonomy is most dangerous.

ActionReversible?Boundary needed
Draft, summarise, analyse in the workspaceYesLight review after
Save/overwrite a fileSometimes (overwrite may lose the original)Gate overwrites of originals
Send an email, post a messageNo — the recipient has itApproval gate before send
Publish externallyNo — it’s publicApproval gate + review
Pay, transfer, purchaseNo — money movedEnforced gate + sign-off
Delete, or change a system of recordNo / costlyEnforced gate; often block entirely
text
reversibility ◄──────────────────────────────►
fully reversible irreversible
│ │
draft / analyse overwrite send / publish / pay / delete
light review gate overwrites APPROVAL GATE (enforced)

Assessment signal

Any stem containing “send”, “publish”, “pay”, “delete”, “cannot be undone” or “external” is pointing at an irreversible action, and the correct answer will place an approval gate before it — never “trust the agent to be careful”.

4.2 Respected boundaries versus enforced boundaries

This is the distinction the domain is built around, and it is the most common trap.

A respected boundary lives in the brief: “do not email anyone outside the company”, “don’t spend more than an hour on this”. The agent reads it and, most of the time, honours it. But it is instruction, not control — a misread objective or an unexpected path can lead the agent across the line without malice, and you find out afterward.

An enforced boundary lives in the system: the agent literally cannot send without your click; the connector has no external-send permission; a spend cap stops the run at a hard number. Crossing it is impossible, not merely discouraged.

Respected boundaryEnforced boundary
Where it livesThe brief / instructionsThe system / configuration
What it doesAsks the agent not to crossMakes crossing impossible
Fails whenThe agent misreads or takes an odd pathIt doesn’t — that’s the point
Use forPreferences, style, low-stakes scopeIrreversible actions, sensitive data, spend
text
"please don't send externally" vs send tool disabled / gated
───────────────────────────── ──────────────────────────
agent usually complies agent cannot send, full stop
depends on correct understanding independent of understanding
OK for low-stakes preferences REQUIRED for irreversible / sensitive

The rule: if crossing the boundary would cause real, hard-to-reverse harm, it must be enforced, not merely requested. A well-behaved agent that respects a requested boundary 99% of the time is still an unacceptable control on an action that can’t be undone.

Assessment signal

When a stem offers “tell the agent not to…” versus “configure the system so it can’t…”, and the action is irreversible or sensitive, the enforced option is correct. “Instruct it clearly” is the tempting-but-wrong answer for high-stakes boundaries.

4.3 Approval gates: pausing at the right point

An approval gate is the enforced pause before an irreversible action. A good gate does three things: it pauses before the action (not after), it surfaces enough context to decide, and it makes the decision cheap for the human.

Bad gateGood gate
Pauses after the email is sentPauses before send, showing the full draft and recipients
Asks “proceed?” with no detailShows exactly what will happen and to whom
Gates everything, so humans rubber-stampGates the irreversible/high-stakes steps only
Gates nothing, trusts the agentGates every irreversible action

The design tension is real: gate too much and humans stop reading and click “approve” reflexively (gate fatigue); gate too little and something irreversible slips through. Resolve it by gating on reversibility and stakes, not on every step — the “one before many” checkpoint from D2 is a form of this.

4.4 Spend and time limits

Autonomy plus a loop equals a runaway. An agent that misunderstands “keep improving it” can iterate indefinitely, and one with paid tools can accumulate cost. Spend and time limits are enforced boundaries against this.

LimitProtects againstSet it by
Time / step limitEndless loops, drift on long runsThe task’s realistic duration, plus margin
Spend capRunaway paid-tool or model costThe value of the task; stop well below “surprising”
Scope cap (items processed)One→many mistakes compoundingThe batch you’ve verified the pattern on

These are enforced, not requested: “try not to spend too much” is a respected boundary and will not stop a looping agent. A hard cap will. The point is not to be stingy but to make the worst case bounded and known before you start the run.

4.5 Data boundaries

Data boundaries stop an agent moving sensitive information where it shouldn’t go — pasting confidential data into an external tool, including PII in an outbound message, or copying regulated data across a line. They connect directly to D3’s reach: the tightest data boundary is not granting reach to the sensitive data in the first place.

Data-boundary controlRespected or enforcedNotes
“Don’t include account numbers in emails”RespectedInstruction; can be missed
Redaction/DLP that blocks the sendEnforcedCatches what instruction misses
Not connecting the agent to the sensitive storeEnforcedStrongest — it can’t move what it can’t reach
Scoping a connector to non-sensitive dataEnforcedLeast-privilege applied to data

The strongest data boundary is architectural: an agent cannot leak data it was never able to touch. Where reach is unavoidable, enforce the outbound control rather than trusting the instruction.

4.6 Matching the boundary to the stakes

Not every task needs heavy boundaries. Over-gating a low-stakes drafting task wastes attention; under-gating an irreversible one invites harm. Match the control to the reversibility and sensitivity.

text
stakes / irreversibility
low ──────────────────────────────────────► high
internal draft shared doc external send / pay / delete
│ │ │
light review spot-check + ENFORCED approval gate +
no gate needed gate overwrites spend/time caps + sign-off

The judgment mirrors the whole track: reversible and low-stakes → keep it light; irreversible or sensitive → enforce. The mistake in both directions is treating all tasks the same.

Decision framework

Use the GATE test on every action an agent can take: Gauge reversibility, Authorise irreversible steps, Threshold spend and time, Enforce, don’t just ask.

StepQuestionAction
G — GaugeCan this action be undone?If no, it needs a gate — full stop
A — AuthoriseWho approves before it fires?Insert a human approval checkpoint before the irreversible step
T — ThresholdWhat’s the worst-case cost/time?Set an enforced spend cap and time/step limit
E — EnforceIs this a request or a control?For irreversible/sensitive actions, make it system-enforced, not brief-requested

Apply it tomorrow: list what the agent can do, run each through GATE, and turn every irreversible or sensitive action from a respected boundary into an enforced one.

Common mistakes

MistakeWhy it happensWhat to do instead
Trusting the agent to “be careful” with irreversible actionsIt usually behavesEnforce a gate; usual behaviour isn’t a control
Writing the boundary only in the briefInstruction feels sufficientEnforce high-stakes boundaries in the system
Gating after the action instead of beforeThe pause was added lateGate before the irreversible step fires
Gating every stepWanting maximum safetyGate by reversibility/stakes; over-gating causes rubber-stamping
No spend or time cap on a looping-capable task“It won’t loop”Set enforced caps sized to the task’s value
“Don’t paste sensitive data” as the only data controlInstruction is easy to writeEnforce with redaction, or don’t grant reach at all
Approval prompt with no contextAdded as an afterthoughtShow what will happen and to whom, so the human can decide
Same boundaries for every taskSimplicityMatch the boundary to the stakes and reversibility

Scenario challenge

Scenario. Ravi runs partnerships on ChatGPT Work. He delegates an agent to “review the inbound partnership requests in our shared inbox, draft polite responses, and send the clear yes/no cases so I only handle the maybe pile”. To speed things up, he adds to the brief: “Only send to companies we already have a relationship with; never send to anyone new without checking.” The agent processes forty requests overnight. In the morning, Ravi finds it emailed twelve external companies — including three it had never dealt with — because it interpreted “clear yes” broadly. One message quoted internal deal terms.

Expert reasoning trace.

  1. Name the irreversible action. Sending external emails is irreversible — once sent, they can’t be recalled. That alone means the send should never have been a merely-respected boundary.
  2. See the respected-vs-enforced failure. “Never send to anyone new without checking” is an instruction in the brief — a respected boundary. The agent misread “clear yes” and crossed it, exactly the failure mode of a respected boundary on an irreversible action. It didn’t disobey maliciously; it interpreted the goal and acted.
  3. The correct control is enforcement. The send should have been behind an enforced approval gate: the agent drafts, pauses, and Ravi authorises each send (or at least each send to a new company) before it fires. With enforcement, the agent’s broad interpretation of “clear yes” produces drafts awaiting approval, not sent emails.
  4. Add the data boundary. The message quoting internal deal terms is a data-boundary failure. The agent had reach to internal deal information and no enforced control stopped it leaving in an outbound message. Fix: don’t grant reach to internal deal terms for this task (D3), and/or enforce an outbound check — not merely instruct “don’t include internal terms”.
  5. Right-size the gating. Gating every draft would create rubber-stamp fatigue across forty items. Gate by stakes: auto-draft everything, enforce approval on all external sends (the irreversible step), and require explicit sign-off for any new company (the highest-stakes subset). Time/scope caps bound the overnight run so it can’t process an unbounded pile.
  6. Reframe the lesson. Ravi treated a control problem as a wording problem. No amount of clearer instruction makes a respected boundary safe for an irreversible action — the fix is to move the boundary from the brief into the system.

The point. The harm came from relying on a respected boundary where an enforced one was required. Sending is irreversible; internal deal terms are sensitive; both needed system-level controls (approval gate on send, no reach to internal terms), not a more strongly-worded brief.

Assessment traps

TrapWhy it is temptingThe discriminator
“Just instruct the agent clearly not to…”Clear instructions feel like controlOn irreversible/sensitive actions, only enforcement is a control
“It behaved fine in testing, so trust it”Usual behaviour reassuresA misread on the wrong run is unrecoverable; enforce the gate
“Gate every action to be safe”Maximum caution seems safestOver-gating causes rubber-stamping; gate by reversibility/stakes
“Add ‘don’t overspend’ to the brief”Easy to writeA respected cap won’t stop a loop; set an enforced spend limit
“The draft mentioned internal terms — reword the brief”Looks like a wording fixIt’s a data-boundary/reach failure; enforce or remove reach
“Approve after sending to save a step”Feels efficientA gate after an irreversible action controls nothing

Practice questions

Each item states how many responses to select. Attempt before revealing.

Q1 · Which action MOST requires an enforced approval gate before it fires? (Select one)

A. Drafting a summary in the workspace. B. Sending an email to an external customer. C. Analysing a spreadsheet you supplied. D. Rewriting a paragraph you’ll read next.

Answer: B. Sending an external email is irreversible, so it needs an enforced gate before it fires. Drafting (A), analysing (C) and rewriting (D) are reversible and safe with light review — nothing has left the workspace.

Q2 · What is the key difference between a respected boundary and an enforced boundary? (Select one)

A. Respected boundaries are written in a nicer tone. B. A respected boundary asks the agent not to cross a line; an enforced boundary makes crossing it impossible. C. Enforced boundaries only apply to large models. D. There is no real difference.

Answer: B. A respected boundary is instruction the agent usually honours; an enforced boundary is a system control that removes the possibility of crossing. Tone (A) is irrelevant, enforcement isn’t model-specific (C), and the difference is central, not cosmetic (D).

Q3 · For an irreversible action, why is 'instruct the agent clearly not to do it' insufficient? (Select one)

A. It isn’t insufficient; clear instructions always work. B. A misread objective or unexpected path can lead the agent across a merely-requested line, and the action can’t be undone. C. Instructions cost more tokens. D. Agents ignore all instructions.

Answer: B. Instruction depends on correct understanding, and an autonomous agent can cross a requested line by misreading the goal — unacceptable when the action is irreversible. Clear instructions don’t always hold (A). Token cost (C) is irrelevant, and agents don’t ignore all instructions (D) — they can simply misinterpret them.

Q4 · A good approval gate should… (Select one)

A. Pause after the action, to avoid slowing the agent. B. Pause before the irreversible action and show what will happen and to whom. C. Ask ‘proceed?’ with no detail. D. Gate every single step equally.

Answer: B. A gate must pause before the irreversible step and surface enough context — the draft and recipients — for a real decision. Pausing after (A) controls nothing. A detail-free prompt (C) invites blind approval. Gating every step (D) causes rubber-stamp fatigue.

Q5 · An agent capable of iterating and using paid tools is told 'don't spend too much'. What is the problem? (Select one)

A. Nothing; the instruction is enough. B. ‘Don’t spend too much’ is a respected boundary that won’t stop a looping run; set an enforced spend cap. C. Paid tools can’t loop. D. Spend limits slow the model down.

Answer: B. A vague requested limit won’t halt a runaway; an enforced spend cap makes the worst case bounded and known. The instruction alone is not enough (A). Loops can occur with paid tools (C), and an enforced cap doesn’t slow the model — it stops a runaway (D).

Q6 · What is the STRONGEST data boundary against an agent leaking a sensitive internal document? (Select one)

A. Instructing it not to share the document. B. Not granting the agent reach to the sensitive document at all. C. Asking it to summarise the document carefully. D. Using a larger model.

Answer: B. The strongest data boundary is architectural: an agent cannot leak what it was never able to reach. Instruction (A) can be missed. Careful summarisation (C) still touches the data. Model size (D) doesn’t create a data boundary.

Q7 · Why can gating every action be as harmful as gating none? (Select one)

A. It isn’t; more gates are always better. B. Excessive gates cause humans to approve reflexively, so a real risk slips through unnoticed. C. Gates disable the agent’s tools. D. Gates change the model.

Answer: B. Over-gating produces rubber-stamp fatigue, defeating the purpose of the gate when it matters. More gates aren’t always better (A). Gates don’t disable tools (C) or change the model (D).

Q8 · Which TWO boundaries should be system-enforced rather than merely written in the brief? (Select two)

A. A preference for bullet points over prose. B. An approval gate before any external send. C. A hard spend cap on a task that uses paid tools. D. A suggestion to keep the tone friendly. E. A preference to finish before lunch.

Answer: B and C. An approval gate on irreversible sends (B) and a hard spend cap on paid-tool use (C) protect against real, hard-to-reverse harm and must be enforced. Style (A), tone (D) and a soft timing wish (E) are preferences — respected boundaries are fine.

Q9 · An agent overwrites an original file while 'improving' it, losing the prior version. Which boundary would have prevented this? (Select one)

A. A friendlier tone. B. An enforced gate on overwriting originals, or saving to a new version instead. C. A larger context window. D. A web-search tool.

Answer: B. Overwriting can be irreversible, so gating overwrites or forcing versioned saves preserves the original. Tone (A), context window (C) and web search (D) do nothing to protect the file.

Q10 · A task is a low-stakes internal draft you'll read before doing anything with it. What boundary posture fits? (Select one)

A. Heavy enforced gates on every step. B. Light review after the run; no gate needed because nothing irreversible happens. C. A spend cap of zero so it can’t run. D. Disable all tools.

Answer: B. Match boundaries to stakes: a reversible internal draft you’ll review needs only light review, not heavy gating. Heavy gates (A) waste attention. A zero cap (C) or disabling tools (D) blocks a harmless task.

Q11 · An agent emailed new external contacts despite a brief saying 'never contact anyone new without checking'. What is the correct fix? (Select one)

A. Reword the instruction more firmly. B. Move the boundary into the system: enforce an approval gate on sends to new contacts so the agent cannot send without authorisation. C. Use a bigger model that follows instructions better. D. Remove the definition of done.

Answer: B. The failure is a respected boundary on an irreversible action; the fix is to enforce it as a system-level approval gate. Firmer wording (A) is still just a respected boundary. A bigger model (C) can still misread. Removing the definition of done (D) is unrelated and harmful.

Q12 · A manager wants to bound the worst case of an overnight agent run over an unbounded queue. Which TWO enforced limits fit best? (Select two)

A. A time or step limit sized to the realistic run. B. A scope cap on the number of items processed. C. A note in the brief asking it to be efficient. D. A friendlier tone setting. E. Removing all checkpoints so it finishes faster.

Answer: A and B. An enforced time/step limit (A) and a scope cap on items (B) make the worst case bounded and known for an unattended overnight run. A polite note (C) is a respected boundary that won’t stop a runaway. Tone (D) is irrelevant, and removing checkpoints (E) increases risk.

Key takeaways

  • Start every boundary decision from reversibility: what can this agent do that I can’t take back?
  • The core distinction is respected (instruction in the brief) versus enforced (system makes crossing impossible); irreversible or sensitive actions require enforcement.
  • Put an approval gate before any irreversible action, surfacing enough context to decide — never after, and never on every step.
  • Set enforced spend and time/step limits so a looping or runaway agent has a bounded, known worst case.
  • The strongest data boundary is architectural: an agent cannot leak what it was never granted reach to (links to D3).
  • Match the boundary to the stakes: light review for reversible low-stakes work, enforced gates and caps for irreversible or sensitive work.
  • A clearer instruction never fixes a control problem — use GATE to turn irreversible/sensitive actions from requests into controls.

Last updated Sep 18, 2026