# D3 · Inputs, Outputs and Contracts

Treating every workflow step as a contract with named inputs, required fields, an output schema or template, acceptance criteria, and a defined behaviour on missing input.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain is worth **16%** of the mock — roughly **8 of 50 items**. It tests whether you can make a workflow *predictable* by defining, for every step, exactly what goes in and exactly what comes out. Decomposition ([Domain 2](/openai/applied-ai/domains/d2-decomposing-work-into-steps/)) gives you steps; this domain gives each step a **contract** so the steps fit together reliably and a colleague running the workflow gets the same result you do.

## What you need to know

A contract for a step specifies five things: the **named inputs** it consumes, which of them are **required** versus optional, the **output shape** it must produce (a schema, a template, or fixed fields), the **acceptance criteria** that define a good output, and the **behaviour on missing or malformed input**. Contracts turn "it usually works" into "it works predictably" because every step knows what it can rely on receiving and what it is obliged to return. The single most common source of workflow flakiness is an undefined output shape feeding into the next step, or an unhandled missing input that the model silently guesses around.

## Learning objectives

By the end of this page you should be able to:

1. **Specify** a step's named inputs and mark each required or optional.
2. **Define** an output as a schema, template or fixed field set that downstream steps can rely on.
3. **Write** acceptance criteria that make "good output" checkable rather than subjective.
4. **Decide** what a step does when a required input is missing or malformed — ask, default, or fail.
5. **Design** the handoff between two steps so the producer's output exactly matches the consumer's expected input.

---

## 3.1 The five parts of a step contract

```text
CONTRACT for step "Categorise invoice"
  INPUTS         invoice_text (required), known_categories (required),
                 prior_vendor_map (optional)
  OUTPUT         { vendor, amount, currency, category, confidence }  (JSON)
  ACCEPTANCE     category ∈ known_categories; amount is a number;
                 confidence in [0,1]; every field present
  ON MISSING     if invoice_text absent → FAIL with "no invoice supplied"
                 if known_categories absent → FAIL (cannot categorise safely)
```

| Part | What it pins down | Failure it prevents |
| --- | --- | --- |
| Named inputs | What the step consumes, by name | Ambiguity about what to feed it |
| Required vs optional | What must be present to run | Silent guessing when something is missing |
| Output shape | The exact structure returned | Downstream steps breaking on unexpected format |
| Acceptance criteria | What "correct" means, checkably | "Looks fine" passing when it isn't |
| On missing input | The defined fallback behaviour | The model inventing data to fill a gap |

:::tip[Assessment signal]
When a stem says the workflow "sometimes returns a different format", "breaks when a field is empty", or "a colleague got a different result", it is a contract problem. The fix is to pin the output shape and define the missing-input behaviour, not to reword the prompt.
:::

## 3.2 Named inputs and required fields

Every step should declare its inputs by name and mark each **required** or **optional**. This matters most at handoffs: if step 2 requires a `total` field, step 1's contract must promise to produce it.

| Input style | When to use | Example |
| --- | --- | --- |
| Required | The step cannot run correctly without it | The document to summarise |
| Optional with default | Nice to have; a sensible default exists | Tone = "neutral" if unspecified |
| Optional, changes behaviour | Presence switches a mode | A glossary → enforce terminology |

A frequent bug: a field that is *actually* required is treated as optional, so when it is missing the model fabricates a plausible value. Declaring it required with an explicit missing-input rule stops that.

## 3.3 Output as a schema, template or fixed fields

The output shape is the contract's spine because it is what the next step depends on. Choose the tightest shape the consumer needs.

<Tabs>
  <TabItem label="Schema (machine-consumed)">
    For structured data a downstream system or step parses.

    ```json
    {
      "vendor": "string",
      "amount": 0.0,
      "currency": "USD",
      "category": "one of the supplied categories",
      "confidence": 0.0
    }
    ```
    Specify field names, types and allowed values. Ask for *only* the JSON, no prose around it, so parsing is reliable.
  </TabItem>
  <TabItem label="Template (human-consumed)">
    For a document a person reads, fix the sections and order.

    ```text
    ## Weekly digest — {week}
    ### Highlights (max 5 bullets)
    ### Risks (each: risk — owner — mitigation)
    ### Decisions needed
    ```
    A template makes completeness checkable: a missing section is visible.
  </TabItem>
  <TabItem label="Fixed fields (lightweight)">
    When full JSON is overkill but you still need consistency.

    ```text
    Return exactly these three lines:
    CATEGORY: <one of the list>
    AMOUNT: <number>
    NEEDS_REVIEW: <yes|no>
    ```
    Cheap, human-readable, still parseable.
  </TabItem>
</Tabs>

The rule: **match the output shape to the consumer.** A human reader wants a template; a parser wants a schema; a quick human triage wants fixed fields.

## 3.4 Acceptance criteria — making "good" checkable

Acceptance criteria turn a subjective "is this good?" into a checklist a person or a later step can apply. Good criteria are **observable**.

| Weak (unobservable) | Strong (observable) |
| --- | --- |
| "The summary is high quality" | "≤ 200 words; covers all five agenda items; no claim absent from the source" |
| "The reply is on-brand" | "Second person; no jargon from the banned list; ends with a next step" |
| "The table is correct" | "Rows sum to the stated total; every category is from the allowed set" |

Acceptance criteria are also what the **draft-then-critique** step and the **human review gate** check against — so writing them here pays off in [Domain 5](/openai/applied-ai/domains/d5-review-points-and-human-oversight/).

:::tip[Assessment signal]
"How do you know the output is good enough?" is asking for acceptance criteria. The best answer names observable, checkable conditions — not "it reads well" or "the model is capable".
:::

## 3.5 Behaviour on missing or malformed input

The most dangerous default in any workflow is *silent guessing*. When a required input is absent, a step must do one of three defined things — never invent data.

```text
Required input missing?
│
├─ Can a human quickly supply it?  ─► ASK   (route back / flag for input)
│
├─ Is there a safe, explicit default? ─► DEFAULT  (and mark it as defaulted)
│
└─ Otherwise ─► FAIL LOUDLY  (stop, return a clear reason; do not fabricate)
```

| Strategy | Use when | Danger if misused |
| --- | --- | --- |
| **Ask** | A person is in the loop and can supply it cheaply | Interrupts high-volume automation |
| **Default** | A safe, documented default exists | A silent default hides that data was missing |
| **Fail** | No safe default; correctness depends on it | Over-failing blocks the workflow on trivia |

The wrong answer is a fourth option nobody writes down: **guess**. A model asked to categorise an invoice with no category list will happily invent categories — plausible, wrong, and undetectable downstream.

## 3.6 Designing the handoff between steps

A workflow is a chain of contracts, and it only holds if each producer's **output contract** matches the next consumer's **input contract**.

```text
Step 1  OUTPUT: { scope[], dates[], constraints[] }
                 │  must match  ▼
Step 2  INPUT (required): scope[], dates[]     ← consumed
        INPUT (optional): constraints[]        ← used if present
        OUTPUT: { timeline, pricing_table }
                 │  must match  ▼
Step 3  INPUT (required): timeline, pricing_table
```

When you add or change a step, re-check both edges: does it get what it needs, and does it produce what the next step requires? Most "the workflow broke after I tweaked step 2" bugs are a contract mismatch at one of these edges.

## 3.7 Contracts as living documentation

A written contract is also the thing that lets a **colleague run your workflow**. The contract *is* the interface: "give it these named inputs, expect this output shape, it is done when these criteria hold." That is why this domain feeds directly into repeatability ([Domain 6](/openai/applied-ai/domains/d6-repeatability-and-improvement/)) — an undocumented step lives only in your head and cannot be handed over, versioned, or improved.

---

## Decision framework

**The CONTRACT card** — before you build a step, fill in one card. If any row is blank, the step is not ready.

| Row | Question | Example answer |
| --- | --- | --- |
| **In** | What named inputs does it consume? | `ticket_text`, `issue_types` |
| **Req** | Which are required vs optional? | both required; `customer_tier` optional |
| **Out** | What exact shape does it return? | fixed fields: `DRAFT`, `NEEDS_REVIEW` |
| **Accept** | What observable conditions define a good output? | answers every question asked; no unverified account claim |
| **Missing** | What happens if a required input is absent? | `issue_types` missing → FAIL; `ticket_text` missing → FAIL |
| **Handoff** | Does the output match the next step's required input? | yes — next step consumes `DRAFT` |

The card's discipline is that **Missing** and **Handoff** are mandatory rows. Most flaky workflows have those two blank.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Leaving the output shape unspecified | The first result looked fine | Pin a schema/template/fixed fields the consumer can rely on |
| Treating a required input as optional | It was present in every test | Mark it required and define the missing-input behaviour |
| Letting the model guess when input is missing | Guessing looks like resilience | Ask, default (marked), or fail loudly — never fabricate |
| Writing subjective acceptance criteria | "High quality" feels sufficient | Use observable, checkable conditions |
| Mismatched handoff after editing a step | Only the edited step was tested | Re-check both edges: does it get and give what neighbours need |
| Wrapping JSON output in prose | The model is chatty by default | Ask for only the JSON so parsing is reliable |
| Silent defaults that hide missing data | A default keeps the workflow moving | Mark any defaulted field so review can see it was missing |
| No written contract | The workflow lives in the author's head | Write the contract card so a colleague can run it |

## Scenario challenge

**Scenario.** Lena automated a step that reads customer feedback emails and outputs a structured record for the CRM: sentiment, product area, and a one-line summary. It worked for weeks. Then two problems appeared. First, a colleague who took over the workflow got *different-looking* output — sometimes a paragraph, sometimes bullet points, occasionally a sentiment value like "mostly positive" instead of the expected `positive/neutral/negative`. Second, when an email arrives with no clear product mention, the step confidently outputs a product area anyway — and those guessed values have been polluting the CRM. Lena's instinct is to rewrite the prompt to be "clearer".

**Expert reasoning trace.**

1. **Diagnose problem one as an output-shape failure.** The output has no fixed schema, so the model varies format and even invents new sentiment values. A clearer prose prompt will not fix this reliably — the contract has no pinned output shape. The fix is to specify a schema: `sentiment ∈ {positive, neutral, negative}`, `product_area ∈ <allowed list>`, `summary` as one line, and to demand *only* the JSON.
2. **Diagnose problem two as a missing-input failure.** The step treats `product_area` as always-derivable, so when the email lacks a clear product it *guesses* — the forbidden fourth option. The contract must define the missing-input behaviour: if no product can be identified with confidence, output `product_area: "unknown"` and `needs_review: yes` (a marked default plus a flag), never a fabricated area.
3. **Reject the "clearer prompt" reflex.** The problem is not persuasion; it is an under-specified contract. Rewording might reduce format drift briefly but will not constrain the value set or handle the missing product, and the colleague will still get drift because the *shape* is not pinned.
4. **Fix the handoff.** The CRM (the consumer) requires an enumerated sentiment and a valid or explicitly-unknown product area. The producer's output contract must promise exactly that enum and that unknown-handling, so the CRM never receives a free-text sentiment again.
5. **Add acceptance criteria for review.** "Every record has all fields; sentiment is one of three values; product_area is from the list or 'unknown'; guessed areas are flagged." Now a sampling check ([Domain 5](/openai/applied-ai/domains/d5-review-points-and-human-oversight/)) can catch violations, and the contract is documentation the colleague can follow to get identical output.

**The decision:** replace the vague prompt with a **defined output schema (enumerated values, only-JSON)** and an **explicit missing-input rule (unknown + needs_review, never guess)**, matched to the CRM's required input, plus observable acceptance criteria — not a "clearer" prose prompt. Both symptoms — format drift across users and polluted product areas — trace to the two contract rows people most often leave blank: output shape and missing-input behaviour.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Reword the prompt to be clearer" | Prompting is the familiar lever | Format drift and guessing are contract gaps; pin the output shape and missing-input rule |
| "Let the step guess a value when input is missing" | Guessing looks robust | Guessing fabricates undetectable data; ask, default-and-mark, or fail |
| "'High quality' is a fine acceptance criterion" | It sounds like a standard | Criteria must be observable/checkable, not subjective |
| "The output looks fine, so the shape is defined" | One good sample feels like proof | An unpinned shape varies across inputs and users; specify it explicitly |
| "Wrap the JSON in a friendly explanation" | The model is chatty | Prose around JSON breaks parsers; return only the JSON |
| "A silent default keeps things flowing" | Fewer interruptions | Unmarked defaults hide missing data from review; always flag a default |

## Practice questions

Each item states how many responses to select. Commit before revealing.

<Accordions>
  <AccordionItem title="Q1 · A step 'sometimes returns a paragraph, sometimes bullets' and a colleague gets a different format than you. What is the ROOT problem? (Select one)">
    A. The model is too small
    B. The output shape is not pinned in the step's contract
    C. The temperature is too high
    D. The colleague used the wrong account

    **Answer: B.** Inconsistent format across inputs and users means the contract never specified an output shape; pinning a schema/template fixes it. Model size (A) and temperature (C) don't define structure. The account (D) is irrelevant to format.
  </AccordionItem>

  <AccordionItem title="Q2 · Which set best defines the parts of a step contract? (Select one)">
    A. Model, temperature, max tokens, stop sequence
    B. Named inputs, required vs optional, output shape, acceptance criteria, missing-input behaviour
    C. Prompt, response, cost, latency
    D. Trigger, owner, deadline, budget

    **Answer: B.** A step contract pins what goes in, what must be present, what comes out, what 'good' means, and what happens on missing input. Option A lists model settings. Option C lists observability metrics. Option D lists project-management fields, not a contract.
  </AccordionItem>

  <AccordionItem title="Q3 · A required category list is missing when a categorisation step runs. What should the step do? (Select one)">
    A. Invent plausible categories to keep going
    B. Fail loudly with a clear reason, because it cannot categorise safely without the list
    C. Return an empty string silently
    D. Pick the first category alphabetically

    **Answer: B.** With no safe default and correctness depending on the list, the step must fail loudly rather than guess. Inventing categories (A) fabricates undetectable data. A silent empty string (C) hides the failure. An arbitrary pick (D) is a disguised guess.
  </AccordionItem>

  <AccordionItem title="Q4 · Which is a strong, observable acceptance criterion for a weekly digest? (Select one)">
    A. "The digest is high quality"
    B. "The digest reads professionally"
    C. "≤ 200 words, covers all five agenda items, no claim absent from the source"
    D. "The digest is comprehensive"

    **Answer: C.** Acceptance criteria must be checkable conditions a person or step can verify. 'High quality' (A), 'reads professionally' (B) and 'comprehensive' (D) are subjective and unverifiable.
  </AccordionItem>

  <AccordionItem title="Q5 · You want a step's output parsed by a downstream system. What should you require? (Select one)">
    A. A friendly explanation followed by the data
    B. Only the JSON, with specified field names, types and allowed values, and no surrounding prose
    C. A narrative paragraph
    D. Whatever format the model prefers

    **Answer: B.** Machine consumption needs a strict schema and only the JSON so parsing is reliable. Prose around the data (A) and a narrative (C) break parsers. Letting the model choose (D) reintroduces drift.
  </AccordionItem>

  <AccordionItem title="Q6 · After you edit step 2, the workflow breaks at step 3. What is the MOST likely cause? (Select one)">
    A. Step 3 needs a bigger model now
    B. Step 2's output no longer matches step 3's required input — a handoff/contract mismatch
    C. The temperature drifted
    D. Step 1 is broken

    **Answer: B.** Editing a step commonly changes its output shape, breaking the consumer that relied on the old contract; re-check both edges. Model size (A) and temperature (C) aren't implied. Step 1 (D) wasn't touched.
  </AccordionItem>

  <AccordionItem title="Q7 · When is a marked DEFAULT the right missing-input strategy? (Select one)">
    A. When correctness fully depends on the missing value
    B. When a safe, documented default exists and you flag that the value was defaulted
    C. Whenever you want to avoid interruptions
    D. Never — always fail

    **Answer: B.** A default is appropriate only when it is safe and documented, and it must be marked so review sees the value was missing. If correctness depends on it (A), fail instead. Avoiding interruptions at any cost (C) leads to silent guessing. 'Always fail' (D) is too rigid when a safe default exists.
  </AccordionItem>

  <AccordionItem title="Q8 · Which TWO are the contract rows most often left blank in flaky workflows? (Select two)">
    A. Output shape / behaviour on missing input
    B. The model's release date
    C. The handoff match to the next step's required input
    D. The prompt's word count
    E. The color of the output

    **Answer: A and C.** Unpinned output shape (with undefined missing-input behaviour) and unchecked handoffs are the classic gaps that make workflows flaky. Release date (B), word count (D) and formatting color (E) are not contract rows.
  </AccordionItem>

  <AccordionItem title="Q9 · An extraction step outputs sentiment as free text like 'mostly positive' instead of the expected enum. What fixes it? (Select one)">
    A. Ask the model to be more careful
    B. Constrain the output to an enumerated set `{positive, neutral, negative}` in the schema and return only that structure
    C. Increase max tokens
    D. Add more examples of good emails

    **Answer: B.** An enumerated value set in the schema is what forces one of the allowed values; free text means the value set was never constrained. 'Be careful' (A) is not a contract. Max tokens (C) is unrelated. More email examples (D) don't constrain the output enum.
  </AccordionItem>

  <AccordionItem title="Q10 · Why does writing the contract matter for repeatability and handover? (Select one)">
    A. It makes the prompt shorter
    B. The contract is the interface: it tells a colleague exactly what inputs to give, what output to expect, and when the step is done
    C. It lets you skip acceptance criteria
    D. It removes the need for review

    **Answer: B.** A written contract is the runnable interface a colleague follows to get the same result; an undocumented step lives only in the author's head. It doesn't shorten prompts (A), skip criteria (C), or remove review (D).
  </AccordionItem>

  <AccordionItem title="Q11 · A step outputs a product area even when the email mentions no product, polluting the CRM. Which TWO changes fix this? (Select two)">
    A. Define a missing-input rule: if no product is identifiable, output `unknown` and flag `needs_review`
    B. Let the model keep guessing but log it
    C. Constrain product_area to the allowed list, with `unknown` as the only fallback
    D. Increase the temperature for variety
    E. Remove the product_area field entirely

    **Answer: A and C.** The fix is an explicit missing-input rule (unknown + review flag, never guess) plus constraining the field to the allowed list with a single safe fallback. Guess-and-log (B) still pollutes with fabricated values. Temperature (D) worsens variance. Removing the field (E) drops a required output.
  </AccordionItem>

  <AccordionItem title="Q12 · A downstream step requires a numeric `total`, but the producing step's contract only promises a formatted string like '$1,204.00'. What is the BEST fix? (Select one)">
    A. Parse the string downstream and hope the format never changes
    B. Change the producer's output contract to return `total` as a number, matching the consumer's required input
    C. Add a bigger model to the consumer
    D. Ask the producer to be more consistent

    **Answer: B.** The clean fix is to align the producer's output contract to the consumer's required type — a number — rather than fragile downstream parsing. Parsing-and-hoping (A) is brittle to format changes. A bigger model (C) doesn't fix a type mismatch. 'Be more consistent' (D) is not a contract change.
  </AccordionItem>
</Accordions>

## Key takeaways

- Every step is a **contract**: named inputs, required vs optional, output shape, acceptance criteria, and missing-input behaviour.
- **Pin the output shape** to the consumer — schema for machines, template for readers, fixed fields for quick triage — and return *only* that shape for parseable outputs.
- **Acceptance criteria must be observable** and checkable, not "high quality"; they are what critique steps and human gates verify against.
- On missing required input, **ask, default-and-mark, or fail loudly** — never let the model **guess**.
- A workflow is a chain of contracts: after editing a step, **re-check both handoff edges**.
- The written contract is the **interface** that lets a colleague run, version and improve the workflow — the bridge to Domains 5 and 6.
- The two rows most often left blank — **output shape** and **missing-input behaviour** — are the usual cause of flaky workflows.
