# D2 · Defining Objectives and Tasks

Specifying a delegable objective — goal, definition of done, constraints, sources of truth and checkpoints — and handling ambiguity before an agent starts work.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This is one of the two heaviest domains — **18%**, about **9 of 50 items** on our mock. It tests whether you can turn a wish into a brief an agent can actually execute and you can actually verify. Most failed delegations are decided here: an agent given a vague objective produces vague work, and you cannot tell whether it succeeded because you never said what success was. The skill this domain rewards is writing the objective, not writing the prompt.

## What you need to know

A delegable objective has five parts: a **goal** (the outcome, stated as a result not a task list), a **definition of done** (the acceptance test you'll check against), **constraints** (what it must and must not do), **sources of truth** (which inputs are authoritative), and **checkpoints** (where it pauses for you). The goal says *what*; the definition of done says *how you'll know*; constraints and sources bound *how*; checkpoints control *when you get to steer*. When an objective is ambiguous, the correct move is almost never to let the agent guess — it is to sharpen the brief, or to instruct the agent to ask before proceeding on the ambiguous point. A good brief is one you could hand to a capable new hire who will not ask clarifying questions.

## Learning objectives

By the end of this page you should be able to:

1. **Write** an objective as a result with a testable definition of done.
2. **Separate** constraints from the goal, and hard constraints from preferences.
3. **Name** the sources of truth an agent should treat as authoritative, and rank them when they conflict.
4. **Place** checkpoints where the risk or uncertainty is highest.
5. **Handle** an ambiguous objective without letting the agent silently guess.
6. **Diagnose** a failed delegation as a brief defect before blaming the model.

---

## 2.1 Goal: state the outcome, not the keystrokes

The most common briefing error is describing the steps you'd take instead of the result you want. Steps constrain an agent to your path and stop it adapting; an outcome lets it plan while still being checkable.

| Weak (steps / vague) | Strong (outcome) |
| --- | --- |
| "Look through the tickets and see what's going on" | "Produce a list of the five most common support issues this month, with a count and one example ticket each" |
| "Help with the newsletter" | "Draft a 300-word customer newsletter announcing the March release, in our house tone, ready for review" |
| "Analyse the sales data" | "Identify which two regions missed their Q1 target and the single largest cause for each, citing the figures" |

:::tip[Assessment signal]
Stems that say the objective is "to help with", "to look into" or "to work on" something are flagging a **goal defect** — the answer is to restate the objective as a concrete, checkable outcome, not to add tools or a bigger model.
:::

## 2.2 Definition of done: the acceptance test

The definition of done is what separates a delegable task from a wish. It is the checklist you will run against the result — so it must be observable, not a feeling.

```text
GOAL:  Produce a competitor pricing summary for the leadership deck.

DEFINITION OF DONE (checkable):
  [ ] Covers all four named competitors
  [ ] One price point per product tier, sourced to a public page with a link
  [ ] A one-line "what changed since last quarter" per competitor
  [ ] Fits on one slide (≤ 6 bullets)
  [ ] Flags any figure it could not verify rather than guessing
```

A good definition of done does three things: it tells the agent when to stop, it tells you what to check, and — crucially — it tells the agent what to do when it *can't* meet a criterion (flag it, don't fabricate it). Compare that to "make a pricing summary", where neither party knows what "done" means.

| Property of a good definition of done | Why it matters |
| --- | --- |
| Observable | You can tick it without judgment calls |
| Complete | Covers every part of the goal, so nothing is silently dropped |
| Bounded | States a size/scope so the agent stops at the right place |
| Failure-aware | Says what to do when a criterion can't be met (flag, don't fabricate) |

## 2.3 Constraints: hard rules versus preferences

Constraints bound *how* the work is done. The critical distinction is between **hard constraints** (violating them makes the output unusable or unsafe) and **preferences** (nice to have). An agent that treats a preference as hard wastes effort; one that treats a hard constraint as a preference produces unusable or unsafe work.

| Constraint type | Example | Consequence of violating |
| --- | --- | --- |
| Hard — safety/legal | "Never include a customer's full account number" | Unsafe; potential breach |
| Hard — factual | "Only use figures from the attached report" | Wrong output; can't be trusted |
| Hard — scope | "Do not contact any supplier directly" | Irreversible external action |
| Preference — style | "Prefer bullet points over prose" | Suboptimal, still usable |
| Preference — length | "Aim for about one page" | Longer is acceptable if needed |

Write hard constraints as prohibitions the agent must respect (`must not`, `only`, `never`) and mark preferences as such. D4 develops the further point that a *hard* constraint on an irreversible action often needs to be **enforced by the system**, not merely stated in the brief.

## 2.4 Sources of truth: what is authoritative

An agent will happily blend an authoritative figure with a stale one from the open web unless you tell it which source wins. Naming sources of truth — and ranking them — is what keeps a delegation grounded.

| The brief says… | The agent will… |
| --- | --- |
| Nothing about sources | Use whatever it finds, including stale or wrong data |
| "Use the attached Q1 report for all figures" | Ground numbers in the authoritative document |
| "Prefer the shared pricing sheet; if it's silent, use the public site and flag it" | Rank sources and surface uncertainty |
| "Company knowledge for policy, the ticket system for facts about this case" | Route each fact to the right source |

```text
        conflict resolution order (example)
        ┌────────────────────────────────────────┐
   1 ►  │ the document I attached to this task     │  most authoritative
   2 ►  │ approved company knowledge / policy      │
   3 ►  │ the live system of record (tickets, CRM) │
   4 ►  │ the public web                           │  least — and flag it
        └────────────────────────────────────────┘
```

:::tip[Assessment signal]
When a stem describes an agent mixing correct and incorrect figures, or using an out-of-date number, the root cause is usually a **missing source of truth**, and the fix is to name and rank the authoritative sources — not to lower temperature or change models.
:::

## 2.5 Checkpoints: where to pause for a human

Checkpoints are the steering points you build into the objective. They belong where **uncertainty or risk is highest** — not evenly spaced, and not only at the end. A checkpoint at the end is a review; a checkpoint mid-run is a chance to correct course before wasted work compounds.

| Put a checkpoint… | Because |
| --- | --- |
| Before any irreversible action | You cannot undo a send, publish or payment |
| After the plan, before execution | Cheapest place to catch a misread objective |
| At a fork where the agent had to guess | Confirm the assumption before it propagates |
| Before scaling from one item to many | Verify the pattern on one before it repeats fifty times |

The "one before many" checkpoint is the highest-value one on this exam: have the agent do the first item, pause, let you approve the approach, *then* run the rest. It converts a fifty-way mistake into a one-way one.

## 2.6 Handling an ambiguous objective

When the objective is ambiguous, letting the agent silently pick an interpretation is the wrong answer nearly every time. There are three legitimate moves, in order of preference:

1. **Resolve it yourself** before delegating — the cheapest fix when you know the answer.
2. **Instruct the agent to ask** on the specific ambiguous point before proceeding (a checkpoint on the fork).
3. **State an explicit assumption** the agent should make *and surface*, so you can catch a wrong one at review.

```text
Objective is ambiguous
│
├─ Do I know the right interpretation?      ── yes ─► resolve it in the brief
├─ Is it a single, identifiable fork?       ── yes ─► tell the agent to pause and ask
└─ Many small unknowns I can't pre-answer?         ─► tell it to state its assumptions
                                                      and flag them for review
```

The trap answer is "let the agent decide and move fast". Speed bought by an unstated wrong assumption is a false economy — you pay it back in rework, and worse if the assumption drove an irreversible action.

## 2.7 A failed delegation is a brief defect first

When an agent returns the wrong thing, the disciplined response is to interrogate the brief before the model. Most failures map to a missing brief element.

| Symptom | Likely brief defect |
| --- | --- |
| Did something adjacent to what you wanted | Goal stated as a task, not an outcome |
| "Finished" but half the work is missing | No definition of done to check against |
| Used a stale or wrong number | No source of truth named |
| Took an action you didn't want | No hard constraint, or no checkpoint |
| Guessed and got it wrong | Ambiguity left for the agent to resolve silently |

This mindset — brief first, model second — is developed fully in D6. Here the point is that a strong objective prevents most failures before they happen.

## Decision framework

Use the **GDCSC brief** — Goal, Done, Constraints, Sources, Checkpoints — as the five-part template for every delegation. If any row is blank, the task isn't ready to hand off.

| Element | The question it answers | A weak answer looks like | A strong answer looks like |
| --- | --- | --- | --- |
| **Goal** | What outcome do I want? | "Help with X" | "Produce Y that does Z, ready for review" |
| **Done** | How will I know it's finished and correct? | "When it's good" | A checkable list, size-bounded, failure-aware |
| **Constraints** | What must it always/never do? | (unstated) | Hard rules marked `must`/`never`; preferences marked as such |
| **Sources** | Which inputs are authoritative? | (unstated) | Named and ranked, with conflict order |
| **Checkpoints** | Where should it pause for me? | (none, or only at the end) | Before irreversible steps and before scaling one→many |

Apply it tomorrow: take any task you'd delegate, fill all five rows, and hand off only when none is blank. A blank row is exactly where the delegation will fail.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Briefing steps instead of the outcome | You describe how *you'd* do it | State the result; let the agent plan the path |
| No definition of done | The goal felt obvious | Write the acceptance checklist before delegating |
| Mixing hard constraints with preferences | Everything feels important | Mark `must`/`never` vs "prefer"; agents need the difference |
| Not naming a source of truth | You assumed it would use the right data | Name and rank authoritative sources explicitly |
| Checkpoints only at the very end | End review feels sufficient | Put a checkpoint before irreversible steps and before one→many scaling |
| Letting the agent resolve ambiguity silently | "Move fast" bias | Resolve it, or tell the agent to pause and ask |
| Telling it to fill gaps rather than flag them | You want a complete-looking answer | Instruct: flag what you can't verify, don't fabricate |
| Blaming the model for a vague brief | The output was wrong, so the model must be | Interrogate the brief first — most failures are brief defects |

## Scenario challenge

**Scenario.** Devon, a marketing manager on ChatGPT Work, delegates: "Put together a competitive summary of our top competitors' pricing for the board deck next week." An hour later the agent returns a two-page document. It covers three competitors (Devon has four), mixes current prices with figures that turn out to be a year old, invents a price for one tier it couldn't find, and is written as flowing prose that won't fit on a slide. Devon's first instinct is that the agent "isn't good enough" and asks whether a more capable model would fix it.

**Expert reasoning trace.**

1. **Interrogate the brief before the model.** Every failure here maps to a missing GDCSC element, so a bigger model would produce the same shaped errors faster. This is a brief defect, not a capability defect.
2. **Goal defect.** "Put together a competitive summary" is a task, not an outcome — there's no statement of what the finished artefact must contain, so the agent chose its own scope and stopped at three competitors.
3. **Definition-of-done defect.** No acceptance test meant nothing checked coverage of all four competitors, one price per tier, or slide-fit. "Done" was left to the agent's judgment.
4. **Sources-of-truth defect.** No authoritative source was named, so the agent blended current and year-old figures freely — the classic missing-source symptom.
5. **Failure-handling defect.** Nothing told the agent what to do when it couldn't find a price, so it fabricated one instead of flagging the gap.
6. **Constraint/format defect.** No format constraint ("fits one slide, ≤ 6 bullets") meant it produced two pages of prose.
7. **The fix.** Rewrite the objective with all five GDCSC rows: goal = a one-slide competitor pricing summary; done = all four competitors, one sourced price per tier, a change-since-last-quarter line, ≤ 6 bullets, gaps flagged not filled; sources = the shared pricing sheet first, then the competitor's public page, flag anything unverifiable; constraints = only public pricing, never internal deal terms; checkpoint = pause after the first competitor so Devon can approve the pattern before the agent does the other three.

**The point.** A more capable model was the wrong lever. The delegation failed because the brief was missing four of its five parts; fixing the brief fixes the output, and the "one before many" checkpoint would have caught the scope and format problems after one competitor instead of after all of them.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Use a bigger model to fix a wrong result" | Capability feels like the missing ingredient | Most wrong results are brief defects; fix the objective first |
| "The goal was obvious, no need to define done" | It was obvious *to you* | The agent can't check a definition of done that doesn't exist |
| "Let the agent decide the ambiguous point to save time" | Speed bias | Silent wrong assumptions cost more in rework and risk |
| "Tell it to fill in missing data so the answer is complete" | Complete-looking answers feel finished | Instruct it to flag gaps, not fabricate them |
| "One end-of-run review is enough" | End review feels thorough | Put a checkpoint before irreversible steps and before one→many scaling |
| "List every preference as a firm requirement" | Everything seems important | Marking preferences as hard constraints wastes effort and blocks good work |

## Practice questions

Each item states how many responses to select. Attempt before revealing.

<Accordions>
  <AccordionItem title="Q1 · Which objective is MOST delegable to an agent? (Select one)">
    A. "Help with the quarterly report."
    B. "Look into the sales numbers and see what's interesting."
    C. "Produce a one-page summary of which two regions missed their Q1 target and the single largest cause for each, citing figures from the attached report."
    D. "Do something about the low sales."

    **Answer: C.** It states a concrete outcome, a checkable scope, and a named source of truth — everything an agent needs to execute and you need to verify. A, B and D are wishes, not objectives: none defines what the finished artefact contains or how you'd know it's done.
  </AccordionItem>

  <AccordionItem title="Q2 · What is a 'definition of done' for a delegated task? (Select one)">
    A. The model and temperature to use.
    B. The observable acceptance test you'll check the result against.
    C. The time limit for the run.
    D. The list of tools the agent may use.

    **Answer: B.** The definition of done is the checkable criteria that tell the agent when to stop and tell you what to verify. Model settings (A), time limits (C) and tool lists (D) are other parts of a brief but none is the acceptance test.
  </AccordionItem>

  <AccordionItem title="Q3 · An agent returns work that looks finished but silently omits part of the task. Which brief element was MOST likely missing? (Select one)">
    A. A larger model.
    B. A definition of done covering every part of the goal.
    C. A friendlier tone instruction.
    D. A longer time limit.

    **Answer: B.** Silent omission happens when there's no acceptance test to check coverage against, so the agent stops when it *feels* done. A bigger model (A) or more time (D) doesn't add a completeness check. Tone (C) is unrelated to coverage.
  </AccordionItem>

  <AccordionItem title="Q4 · The brief says 'never include a customer's full account number' and 'prefer bullet points'. How should the agent treat these? (Select one)">
    A. Both are suggestions it can ignore under time pressure.
    B. The account-number rule is a hard constraint it must not violate; the bullet-point rule is a preference.
    C. Both are hard constraints.
    D. Both are preferences.

    **Answer: B.** A safety/privacy rule stated as 'never' is a hard constraint; a stylistic 'prefer' is a preference the agent can override when needed. Treating the account-number rule as optional (A, D) is unsafe; treating the bullet preference as hard (C) can block otherwise good work.
  </AccordionItem>

  <AccordionItem title="Q5 · An agent mixes a current price with a year-old figure in its output. What is the MOST likely root cause and fix? (Select one)">
    A. Temperature too high; lower it.
    B. No source of truth named; specify and rank the authoritative sources.
    C. The model is too small; upgrade it.
    D. The prompt was too short; make it longer.

    **Answer: B.** Blending stale and current data is the classic missing-source-of-truth symptom; the fix is to name which source is authoritative and how to resolve conflicts. Temperature (A) governs variance, not which source is used. Model size (C) and prompt length (D) don't tell the agent which data to trust.
  </AccordionItem>

  <AccordionItem title="Q6 · Where does the highest-value checkpoint usually belong in a task that repeats an action over fifty items? (Select one)">
    A. Only after all fifty are complete.
    B. After the first item, to approve the pattern before the other forty-nine run.
    C. Every ten seconds regardless of progress.
    D. Never; checkpoints slow the agent down.

    **Answer: B.** The 'one before many' checkpoint converts a fifty-way mistake into a one-way one by verifying the approach on the first item before scaling. Reviewing only at the end (A) means the error already repeated fifty times. Time-based pauses (C) don't align with risk. No checkpoints (D) removes your steering.
  </AccordionItem>

  <AccordionItem title="Q7 · An objective is ambiguous on one specific, identifiable point. What is the BEST instruction to the agent? (Select one)">
    A. Decide the point yourself and keep moving.
    B. Pause and ask about that point before proceeding.
    C. Ignore the ambiguous part entirely.
    D. Produce two full versions covering both interpretations.

    **Answer: B.** A single identifiable fork is best handled by a checkpoint: have the agent pause and ask before the assumption propagates. Deciding silently (A) risks a wrong path; ignoring it (C) leaves the task incomplete; producing two full versions (D) is wasteful when a one-line question resolves it.
  </AccordionItem>

  <AccordionItem title="Q8 · An agent couldn't find a required data point, so it produced a plausible value anyway. What should the brief have instructed? (Select one)">
    A. To always fill in missing values so the output looks complete.
    B. To flag any value it cannot verify rather than fabricating one.
    C. To stop the entire task at the first missing value.
    D. To use a bigger model next time.

    **Answer: B.** A failure-aware definition of done tells the agent to surface gaps instead of inventing data, so you can see what's unverified. Filling gaps (A) hides fabrication. Halting the whole task (C) is disproportionate when one gap can be flagged. Model size (D) doesn't change fabrication behaviour.
  </AccordionItem>

  <AccordionItem title="Q9 · A delegation fails. Which sequence reflects the disciplined diagnosis? (Select one)">
    A. Upgrade the model, then the tools, then the brief.
    B. Interrogate the brief first (goal, done, constraints, sources, checkpoints), then consider the model.
    C. Re-run the same brief several times and average the results.
    D. Assume the task is impossible.

    **Answer: B.** Most failures are brief defects, so the objective and its five elements are the first thing to check before touching the model. Model/tool changes first (A) waste effort on the wrong cause. Re-running the same brief (C) reproduces the same defect. Declaring impossibility (D) skips diagnosis.
  </AccordionItem>

  <AccordionItem title="Q10 · Which pair belongs in a strong definition of done for a slide-ready summary? (Select two)">
    A. Covers all four named competitors.
    B. Uses the newest available model.
    C. Fits on one slide with no more than six bullets.
    D. Runs in under five minutes.
    E. Uses a warm, friendly tone throughout.

    **Answer: A and C.** Coverage of all four competitors (A) and a size bound of one slide / six bullets (C) are observable acceptance criteria you can tick. Model choice (B) and run time (D) are operational, not acceptance tests. Tone (E) is a preference, not a done-criterion for a data summary.
  </AccordionItem>

  <AccordionItem title="Q11 · A manager delegates 'improve our onboarding docs' with no further detail. What is the FIRST thing to fix before an agent runs? (Select one)">
    A. Choose a faster model.
    B. Turn the vague goal into a concrete, checkable outcome with a definition of done.
    C. Grant the agent access to every connector available.
    D. Set the time limit to one hour.

    **Answer: B.** 'Improve' is a wish with no outcome or acceptance test, so the first fix is to specify what 'improved' means and how you'll check it. Model speed (A) and time limits (D) don't define the outcome. Broad access (C) adds blast radius without clarifying the goal.
  </AccordionItem>

  <AccordionItem title="Q12 · When several sources give conflicting figures, what should the brief provide? (Select one)">
    A. Nothing — let the agent average them.
    B. A ranked conflict-resolution order stating which source wins and when to flag uncertainty.
    C. Instructions to always trust the public web.
    D. A larger context window.

    **Answer: B.** Ranking the sources and telling the agent when to flag uncertainty is how you keep the output grounded when data disagrees. Averaging (A) blends good and bad data. Always trusting the web (C) is often the least authoritative choice. A bigger context window (D) doesn't resolve which source to believe.
  </AccordionItem>

  <AccordionItem title="Q13 · A project lead wants an agent to migrate formatting across 120 documents. They write a clear outcome and definition of done but no checkpoints. Which TWO checkpoints add the most safety? (Select two)">
    A. Pause after the first document so the pattern can be approved before the rest run.
    B. Pause after every single document for individual approval.
    C. Pause before any step that would overwrite an original file irreversibly.
    D. Pause every thirty seconds regardless of progress.
    E. Never pause, to finish faster.

    **Answer: A and C.** Approving the pattern on the first document (A) prevents a 120-way mistake, and gating irreversible overwrites (C) protects the originals. Pausing on every document (B) defeats the point of delegating. Time-based pauses (D) don't align with risk. No pauses (E) removes all steering on an irreversible bulk action.
  </AccordionItem>

  <AccordionItem title="Q14 · Which statement best captures why 'move fast, let the agent assume' is a poor default for an ambiguous objective? (Select one)">
    A. It's always slower in practice.
    B. Speed bought by an unstated wrong assumption is repaid in rework, and worse if the assumption drove an irreversible action.
    C. Agents cannot make assumptions at all.
    D. It uses more tokens than asking.

    **Answer: B.** An unsurfaced wrong assumption compounds into rework and, on an irreversible step, into unrecoverable harm — so the apparent speed is a false economy. It isn't always literally slower up front (A). Agents can make assumptions (C). Token use (D) is not the substantive risk.
  </AccordionItem>
</Accordions>

## Key takeaways

- A delegable objective has five parts — **Goal, Done, Constraints, Sources, Checkpoints (GDCSC)**; a blank row is where the delegation will fail.
- State the **outcome**, not the keystrokes; steps constrain the agent and hide the acceptance test.
- The **definition of done** is an observable, bounded, failure-aware checklist — it tells the agent when to stop and you what to verify.
- Distinguish **hard constraints** (`must`/`never`) from **preferences**; agents need to know which they can override.
- **Name and rank sources of truth**; blended stale and current data is a missing-source symptom, not a model problem.
- Put **checkpoints** where risk is highest — before irreversible steps and before scaling one item to many.
- Handle **ambiguity** by resolving it, telling the agent to pause and ask, or having it surface assumptions — never by letting it guess silently.
- Diagnose a failed delegation as a **brief defect first**; a bigger model rarely fixes a vague objective.
