AI Cert Prep
Type to search documentation.

Agents and Workflows

D2 · Defining Objectives and Tasks

Specifying a delegable objective — goal, definition of done, constraints, sources of truth and checkpoints — and handling ambiguity before an agent starts work.

This is one of the two heaviest domains — 18%, about 9 of 50 items on our mock. It tests whether you can turn a wish into a brief an agent can actually execute and you can actually verify. Most failed delegations are decided here: an agent given a vague objective produces vague work, and you cannot tell whether it succeeded because you never said what success was. The skill this domain rewards is writing the objective, not writing the prompt.

What you need to know

A delegable objective has five parts: a goal (the outcome, stated as a result not a task list), a definition of done (the acceptance test you’ll check against), constraints (what it must and must not do), sources of truth (which inputs are authoritative), and checkpoints (where it pauses for you). The goal says what; the definition of done says how you’ll know; constraints and sources bound how; checkpoints control when you get to steer. When an objective is ambiguous, the correct move is almost never to let the agent guess — it is to sharpen the brief, or to instruct the agent to ask before proceeding on the ambiguous point. A good brief is one you could hand to a capable new hire who will not ask clarifying questions.

Learning objectives

By the end of this page you should be able to:

  1. Write an objective as a result with a testable definition of done.
  2. Separate constraints from the goal, and hard constraints from preferences.
  3. Name the sources of truth an agent should treat as authoritative, and rank them when they conflict.
  4. Place checkpoints where the risk or uncertainty is highest.
  5. Handle an ambiguous objective without letting the agent silently guess.
  6. Diagnose a failed delegation as a brief defect before blaming the model.

2.1 Goal: state the outcome, not the keystrokes

The most common briefing error is describing the steps you’d take instead of the result you want. Steps constrain an agent to your path and stop it adapting; an outcome lets it plan while still being checkable.

Weak (steps / vague)Strong (outcome)
“Look through the tickets and see what’s going on”“Produce a list of the five most common support issues this month, with a count and one example ticket each”
“Help with the newsletter”“Draft a 300-word customer newsletter announcing the March release, in our house tone, ready for review”
“Analyse the sales data”“Identify which two regions missed their Q1 target and the single largest cause for each, citing the figures”

Assessment signal

Stems that say the objective is “to help with”, “to look into” or “to work on” something are flagging a goal defect — the answer is to restate the objective as a concrete, checkable outcome, not to add tools or a bigger model.

2.2 Definition of done: the acceptance test

The definition of done is what separates a delegable task from a wish. It is the checklist you will run against the result — so it must be observable, not a feeling.

text
GOAL: Produce a competitor pricing summary for the leadership deck.
DEFINITION OF DONE (checkable):
[ ] Covers all four named competitors
[ ] One price point per product tier, sourced to a public page with a link
[ ] A one-line "what changed since last quarter" per competitor
[ ] Fits on one slide (≤ 6 bullets)
[ ] Flags any figure it could not verify rather than guessing

A good definition of done does three things: it tells the agent when to stop, it tells you what to check, and — crucially — it tells the agent what to do when it can’t meet a criterion (flag it, don’t fabricate it). Compare that to “make a pricing summary”, where neither party knows what “done” means.

Property of a good definition of doneWhy it matters
ObservableYou can tick it without judgment calls
CompleteCovers every part of the goal, so nothing is silently dropped
BoundedStates a size/scope so the agent stops at the right place
Failure-awareSays what to do when a criterion can’t be met (flag, don’t fabricate)

2.3 Constraints: hard rules versus preferences

Constraints bound how the work is done. The critical distinction is between hard constraints (violating them makes the output unusable or unsafe) and preferences (nice to have). An agent that treats a preference as hard wastes effort; one that treats a hard constraint as a preference produces unusable or unsafe work.

Constraint typeExampleConsequence of violating
Hard — safety/legal“Never include a customer’s full account number”Unsafe; potential breach
Hard — factual“Only use figures from the attached report”Wrong output; can’t be trusted
Hard — scope“Do not contact any supplier directly”Irreversible external action
Preference — style“Prefer bullet points over prose”Suboptimal, still usable
Preference — length“Aim for about one page”Longer is acceptable if needed

Write hard constraints as prohibitions the agent must respect (must not, only, never) and mark preferences as such. D4 develops the further point that a hard constraint on an irreversible action often needs to be enforced by the system, not merely stated in the brief.

2.4 Sources of truth: what is authoritative

An agent will happily blend an authoritative figure with a stale one from the open web unless you tell it which source wins. Naming sources of truth — and ranking them — is what keeps a delegation grounded.

The brief says…The agent will…
Nothing about sourcesUse whatever it finds, including stale or wrong data
“Use the attached Q1 report for all figures”Ground numbers in the authoritative document
“Prefer the shared pricing sheet; if it’s silent, use the public site and flag it”Rank sources and surface uncertainty
“Company knowledge for policy, the ticket system for facts about this case”Route each fact to the right source
text
conflict resolution order (example)
┌────────────────────────────────────────┐
1 ► │ the document I attached to this task │ most authoritative
2 ► │ approved company knowledge / policy │
3 ► │ the live system of record (tickets, CRM) │
4 ► │ the public web │ least — and flag it
└────────────────────────────────────────┘

Assessment signal

When a stem describes an agent mixing correct and incorrect figures, or using an out-of-date number, the root cause is usually a missing source of truth, and the fix is to name and rank the authoritative sources — not to lower temperature or change models.

2.5 Checkpoints: where to pause for a human

Checkpoints are the steering points you build into the objective. They belong where uncertainty or risk is highest — not evenly spaced, and not only at the end. A checkpoint at the end is a review; a checkpoint mid-run is a chance to correct course before wasted work compounds.

Put a checkpoint…Because
Before any irreversible actionYou cannot undo a send, publish or payment
After the plan, before executionCheapest place to catch a misread objective
At a fork where the agent had to guessConfirm the assumption before it propagates
Before scaling from one item to manyVerify the pattern on one before it repeats fifty times

The “one before many” checkpoint is the highest-value one on this exam: have the agent do the first item, pause, let you approve the approach, then run the rest. It converts a fifty-way mistake into a one-way one.

2.6 Handling an ambiguous objective

When the objective is ambiguous, letting the agent silently pick an interpretation is the wrong answer nearly every time. There are three legitimate moves, in order of preference:

  1. Resolve it yourself before delegating — the cheapest fix when you know the answer.
  2. Instruct the agent to ask on the specific ambiguous point before proceeding (a checkpoint on the fork).
  3. State an explicit assumption the agent should make and surface, so you can catch a wrong one at review.
text
Objective is ambiguous
│
├─ Do I know the right interpretation? ── yes ─► resolve it in the brief
├─ Is it a single, identifiable fork? ── yes ─► tell the agent to pause and ask
└─ Many small unknowns I can't pre-answer? ─► tell it to state its assumptions
and flag them for review

The trap answer is “let the agent decide and move fast”. Speed bought by an unstated wrong assumption is a false economy — you pay it back in rework, and worse if the assumption drove an irreversible action.

2.7 A failed delegation is a brief defect first

When an agent returns the wrong thing, the disciplined response is to interrogate the brief before the model. Most failures map to a missing brief element.

SymptomLikely brief defect
Did something adjacent to what you wantedGoal stated as a task, not an outcome
“Finished” but half the work is missingNo definition of done to check against
Used a stale or wrong numberNo source of truth named
Took an action you didn’t wantNo hard constraint, or no checkpoint
Guessed and got it wrongAmbiguity left for the agent to resolve silently

This mindset — brief first, model second — is developed fully in D6. Here the point is that a strong objective prevents most failures before they happen.

Decision framework

Use the GDCSC brief — Goal, Done, Constraints, Sources, Checkpoints — as the five-part template for every delegation. If any row is blank, the task isn’t ready to hand off.

ElementThe question it answersA weak answer looks likeA strong answer looks like
GoalWhat outcome do I want?“Help with X”“Produce Y that does Z, ready for review”
DoneHow will I know it’s finished and correct?“When it’s good”A checkable list, size-bounded, failure-aware
ConstraintsWhat must it always/never do?(unstated)Hard rules marked must/never; preferences marked as such
SourcesWhich inputs are authoritative?(unstated)Named and ranked, with conflict order
CheckpointsWhere should it pause for me?(none, or only at the end)Before irreversible steps and before scaling one→many

Apply it tomorrow: take any task you’d delegate, fill all five rows, and hand off only when none is blank. A blank row is exactly where the delegation will fail.

Common mistakes

MistakeWhy it happensWhat to do instead
Briefing steps instead of the outcomeYou describe how you’d do itState the result; let the agent plan the path
No definition of doneThe goal felt obviousWrite the acceptance checklist before delegating
Mixing hard constraints with preferencesEverything feels importantMark must/never vs “prefer”; agents need the difference
Not naming a source of truthYou assumed it would use the right dataName and rank authoritative sources explicitly
Checkpoints only at the very endEnd review feels sufficientPut a checkpoint before irreversible steps and before one→many scaling
Letting the agent resolve ambiguity silently“Move fast” biasResolve it, or tell the agent to pause and ask
Telling it to fill gaps rather than flag themYou want a complete-looking answerInstruct: flag what you can’t verify, don’t fabricate
Blaming the model for a vague briefThe output was wrong, so the model must beInterrogate the brief first — most failures are brief defects

Scenario challenge

Scenario. Devon, a marketing manager on ChatGPT Work, delegates: “Put together a competitive summary of our top competitors’ pricing for the board deck next week.” An hour later the agent returns a two-page document. It covers three competitors (Devon has four), mixes current prices with figures that turn out to be a year old, invents a price for one tier it couldn’t find, and is written as flowing prose that won’t fit on a slide. Devon’s first instinct is that the agent “isn’t good enough” and asks whether a more capable model would fix it.

Expert reasoning trace.

  1. Interrogate the brief before the model. Every failure here maps to a missing GDCSC element, so a bigger model would produce the same shaped errors faster. This is a brief defect, not a capability defect.
  2. Goal defect. “Put together a competitive summary” is a task, not an outcome — there’s no statement of what the finished artefact must contain, so the agent chose its own scope and stopped at three competitors.
  3. Definition-of-done defect. No acceptance test meant nothing checked coverage of all four competitors, one price per tier, or slide-fit. “Done” was left to the agent’s judgment.
  4. Sources-of-truth defect. No authoritative source was named, so the agent blended current and year-old figures freely — the classic missing-source symptom.
  5. Failure-handling defect. Nothing told the agent what to do when it couldn’t find a price, so it fabricated one instead of flagging the gap.
  6. Constraint/format defect. No format constraint (“fits one slide, ≤ 6 bullets”) meant it produced two pages of prose.
  7. The fix. Rewrite the objective with all five GDCSC rows: goal = a one-slide competitor pricing summary; done = all four competitors, one sourced price per tier, a change-since-last-quarter line, ≤ 6 bullets, gaps flagged not filled; sources = the shared pricing sheet first, then the competitor’s public page, flag anything unverifiable; constraints = only public pricing, never internal deal terms; checkpoint = pause after the first competitor so Devon can approve the pattern before the agent does the other three.

The point. A more capable model was the wrong lever. The delegation failed because the brief was missing four of its five parts; fixing the brief fixes the output, and the “one before many” checkpoint would have caught the scope and format problems after one competitor instead of after all of them.

Assessment traps

TrapWhy it is temptingThe discriminator
“Use a bigger model to fix a wrong result”Capability feels like the missing ingredientMost wrong results are brief defects; fix the objective first
“The goal was obvious, no need to define done”It was obvious to youThe agent can’t check a definition of done that doesn’t exist
“Let the agent decide the ambiguous point to save time”Speed biasSilent wrong assumptions cost more in rework and risk
“Tell it to fill in missing data so the answer is complete”Complete-looking answers feel finishedInstruct it to flag gaps, not fabricate them
“One end-of-run review is enough”End review feels thoroughPut a checkpoint before irreversible steps and before one→many scaling
“List every preference as a firm requirement”Everything seems importantMarking preferences as hard constraints wastes effort and blocks good work

Practice questions

Each item states how many responses to select. Attempt before revealing.

Q1 · Which objective is MOST delegable to an agent? (Select one)

A. “Help with the quarterly report.” B. “Look into the sales numbers and see what’s interesting.” C. “Produce a one-page summary of which two regions missed their Q1 target and the single largest cause for each, citing figures from the attached report.” D. “Do something about the low sales.”

Answer: C. It states a concrete outcome, a checkable scope, and a named source of truth — everything an agent needs to execute and you need to verify. A, B and D are wishes, not objectives: none defines what the finished artefact contains or how you’d know it’s done.

Q2 · What is a 'definition of done' for a delegated task? (Select one)

A. The model and temperature to use. B. The observable acceptance test you’ll check the result against. C. The time limit for the run. D. The list of tools the agent may use.

Answer: B. The definition of done is the checkable criteria that tell the agent when to stop and tell you what to verify. Model settings (A), time limits (C) and tool lists (D) are other parts of a brief but none is the acceptance test.

Q3 · An agent returns work that looks finished but silently omits part of the task. Which brief element was MOST likely missing? (Select one)

A. A larger model. B. A definition of done covering every part of the goal. C. A friendlier tone instruction. D. A longer time limit.

Answer: B. Silent omission happens when there’s no acceptance test to check coverage against, so the agent stops when it feels done. A bigger model (A) or more time (D) doesn’t add a completeness check. Tone (C) is unrelated to coverage.

Q4 · The brief says 'never include a customer's full account number' and 'prefer bullet points'. How should the agent treat these? (Select one)

A. Both are suggestions it can ignore under time pressure. B. The account-number rule is a hard constraint it must not violate; the bullet-point rule is a preference. C. Both are hard constraints. D. Both are preferences.

Answer: B. A safety/privacy rule stated as ‘never’ is a hard constraint; a stylistic ‘prefer’ is a preference the agent can override when needed. Treating the account-number rule as optional (A, D) is unsafe; treating the bullet preference as hard (C) can block otherwise good work.

Q5 · An agent mixes a current price with a year-old figure in its output. What is the MOST likely root cause and fix? (Select one)

A. Temperature too high; lower it. B. No source of truth named; specify and rank the authoritative sources. C. The model is too small; upgrade it. D. The prompt was too short; make it longer.

Answer: B. Blending stale and current data is the classic missing-source-of-truth symptom; the fix is to name which source is authoritative and how to resolve conflicts. Temperature (A) governs variance, not which source is used. Model size (C) and prompt length (D) don’t tell the agent which data to trust.

Q6 · Where does the highest-value checkpoint usually belong in a task that repeats an action over fifty items? (Select one)

A. Only after all fifty are complete. B. After the first item, to approve the pattern before the other forty-nine run. C. Every ten seconds regardless of progress. D. Never; checkpoints slow the agent down.

Answer: B. The ‘one before many’ checkpoint converts a fifty-way mistake into a one-way one by verifying the approach on the first item before scaling. Reviewing only at the end (A) means the error already repeated fifty times. Time-based pauses (C) don’t align with risk. No checkpoints (D) removes your steering.

Q7 · An objective is ambiguous on one specific, identifiable point. What is the BEST instruction to the agent? (Select one)

A. Decide the point yourself and keep moving. B. Pause and ask about that point before proceeding. C. Ignore the ambiguous part entirely. D. Produce two full versions covering both interpretations.

Answer: B. A single identifiable fork is best handled by a checkpoint: have the agent pause and ask before the assumption propagates. Deciding silently (A) risks a wrong path; ignoring it (C) leaves the task incomplete; producing two full versions (D) is wasteful when a one-line question resolves it.

Q8 · An agent couldn't find a required data point, so it produced a plausible value anyway. What should the brief have instructed? (Select one)

A. To always fill in missing values so the output looks complete. B. To flag any value it cannot verify rather than fabricating one. C. To stop the entire task at the first missing value. D. To use a bigger model next time.

Answer: B. A failure-aware definition of done tells the agent to surface gaps instead of inventing data, so you can see what’s unverified. Filling gaps (A) hides fabrication. Halting the whole task (C) is disproportionate when one gap can be flagged. Model size (D) doesn’t change fabrication behaviour.

Q9 · A delegation fails. Which sequence reflects the disciplined diagnosis? (Select one)

A. Upgrade the model, then the tools, then the brief. B. Interrogate the brief first (goal, done, constraints, sources, checkpoints), then consider the model. C. Re-run the same brief several times and average the results. D. Assume the task is impossible.

Answer: B. Most failures are brief defects, so the objective and its five elements are the first thing to check before touching the model. Model/tool changes first (A) waste effort on the wrong cause. Re-running the same brief (C) reproduces the same defect. Declaring impossibility (D) skips diagnosis.

Q10 · Which pair belongs in a strong definition of done for a slide-ready summary? (Select two)

A. Covers all four named competitors. B. Uses the newest available model. C. Fits on one slide with no more than six bullets. D. Runs in under five minutes. E. Uses a warm, friendly tone throughout.

Answer: A and C. Coverage of all four competitors (A) and a size bound of one slide / six bullets (C) are observable acceptance criteria you can tick. Model choice (B) and run time (D) are operational, not acceptance tests. Tone (E) is a preference, not a done-criterion for a data summary.

Q11 · A manager delegates 'improve our onboarding docs' with no further detail. What is the FIRST thing to fix before an agent runs? (Select one)

A. Choose a faster model. B. Turn the vague goal into a concrete, checkable outcome with a definition of done. C. Grant the agent access to every connector available. D. Set the time limit to one hour.

Answer: B. ‘Improve’ is a wish with no outcome or acceptance test, so the first fix is to specify what ‘improved’ means and how you’ll check it. Model speed (A) and time limits (D) don’t define the outcome. Broad access (C) adds blast radius without clarifying the goal.

Q12 · When several sources give conflicting figures, what should the brief provide? (Select one)

A. Nothing — let the agent average them. B. A ranked conflict-resolution order stating which source wins and when to flag uncertainty. C. Instructions to always trust the public web. D. A larger context window.

Answer: B. Ranking the sources and telling the agent when to flag uncertainty is how you keep the output grounded when data disagrees. Averaging (A) blends good and bad data. Always trusting the web (C) is often the least authoritative choice. A bigger context window (D) doesn’t resolve which source to believe.

Q13 · A project lead wants an agent to migrate formatting across 120 documents. They write a clear outcome and definition of done but no checkpoints. Which TWO checkpoints add the most safety? (Select two)

A. Pause after the first document so the pattern can be approved before the rest run. B. Pause after every single document for individual approval. C. Pause before any step that would overwrite an original file irreversibly. D. Pause every thirty seconds regardless of progress. E. Never pause, to finish faster.

Answer: A and C. Approving the pattern on the first document (A) prevents a 120-way mistake, and gating irreversible overwrites (C) protects the originals. Pausing on every document (B) defeats the point of delegating. Time-based pauses (D) don’t align with risk. No pauses (E) removes all steering on an irreversible bulk action.

Q14 · Which statement best captures why 'move fast, let the agent assume' is a poor default for an ambiguous objective? (Select one)

A. It’s always slower in practice. B. Speed bought by an unstated wrong assumption is repaid in rework, and worse if the assumption drove an irreversible action. C. Agents cannot make assumptions at all. D. It uses more tokens than asking.

Answer: B. An unsurfaced wrong assumption compounds into rework and, on an irreversible step, into unrecoverable harm — so the apparent speed is a false economy. It isn’t always literally slower up front (A). Agents can make assumptions (C). Token use (D) is not the substantive risk.

Key takeaways

  • A delegable objective has five parts — Goal, Done, Constraints, Sources, Checkpoints (GDCSC); a blank row is where the delegation will fail.
  • State the outcome, not the keystrokes; steps constrain the agent and hide the acceptance test.
  • The definition of done is an observable, bounded, failure-aware checklist — it tells the agent when to stop and you what to verify.
  • Distinguish hard constraints (must/never) from preferences; agents need to know which they can override.
  • Name and rank sources of truth; blended stale and current data is a missing-source symptom, not a model problem.
  • Put checkpoints where risk is highest — before irreversible steps and before scaling one item to many.
  • Handle ambiguity by resolving it, telling the agent to pause and ask, or having it surface assumptions — never by letting it guess silently.
  • Diagnose a failed delegation as a brief defect first; a bigger model rarely fixes a vague objective.

Last updated Sep 18, 2026