# D4 · Choosing the Right Capability

The central decision of the track – choosing between a prompt, a saved instruction, a Project, a custom GPT, a workspace agent and an API application, and when to reach for search, deep research, data analysis, Canvas or file uploads.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This is the heaviest domain on the mock — **20%**, roughly **10 of 50 items** — and it is the section the Applied AI material calls out as the single most valuable thing to master. It tests one judgment repeatedly: given a task, *what should you actually build it with?* The answer runs along a ladder from the lightest option (a well-written prompt) to the heaviest (an API application), and the skill is reaching for the **lightest capability that meets the need** rather than the most impressive one.

## What you need to know

ChatGPT offers a ladder of ways to package a workflow: a **one-off prompt**, a **saved instruction** (custom instructions / a reusable prompt), a **Project** (a workspace with its own instructions and knowledge files for recurring context), a **custom GPT** (a shareable, configured assistant for a repeatable task others can run), a **workspace agent** (delegated multi-step work with oversight), and — when you outgrow ChatGPT — an **API application** (programmatic, integrated, scalable). Orthogonal to that ladder are *in-conversation capabilities* you switch on when a task needs them: **search** for current facts, **deep research** for sourced multi-source synthesis, **data analysis** for computation over files, **Canvas** for iterating on a document or code, and **file uploads** for working over your own documents. The winning judgment is to climb the ladder only as far as repeatability, sharing and integration actually demand.

## Learning objectives

By the end of this page you should be able to:

1. **Place** a task on the prompt → saved instruction → Project → custom GPT → workspace agent → API application ladder.
2. **Justify** choosing the lightest rung that meets the need, and recognise over- and under-reaching.
3. **Select** the right in-conversation capability — search, deep research, data analysis, Canvas, file uploads — from signals in the task.
4. **Distinguish** a Project from a custom GPT, and a custom GPT from a workspace agent.
5. **Decide** when a workflow has outgrown ChatGPT and belongs in an API application.

---

## 4.1 The capability ladder

```text
LIGHTEST ─────────────────────────────────────────────────────► HEAVIEST
 Prompt      Saved         Project        Custom GPT     Workspace     API
 (one-off)   instruction   (recurring     (shareable,    agent         application
             (reusable)    context +      configured     (delegated    (programmatic,
                           knowledge)     assistant)     multi-step)    integrated)
 ▲                                                                            ▲
 use for a   use when you  use when a     use when       use when a     use when it
 single      repeat the    task needs     OTHERS run     multi-step     must run in
 task        same prompt   its own        the same       task can be    code, at scale,
             often         context/files  task without   delegated w/   or wired into
                           every time     you            oversight      other systems
```

Each rung adds capability *and* cost (setup, maintenance, governance). The default posture is: start at the left, move right only when a concrete need forces you.

| Rung | What it is | The need that justifies it |
| --- | --- | --- |
| **Prompt** | A single instruction in a chat | A one-off task |
| **Saved instruction** | Custom instructions or a reusable prompt you keep | You repeat the same prompt often |
| **Project** | A ChatGPT workspace with its own instructions + knowledge files | A recurring task needs the same context/files every time |
| **Custom GPT** | A configured, shareable assistant | *Other people* need to run the same task without you |
| **Workspace agent** | Delegated multi-step execution with oversight | A multi-step task can run semi-autonomously under review |
| **API application** | Code calling the API (Responses API) | It must run programmatically, at scale, or integrate with other systems |

:::tip[Assessment signal]
"I keep re-pasting the same background" → **Project**. "My teammates need to run this too" → **custom GPT**. "It has to run inside our app / on a schedule at scale / wired to our database" → **API application**. Match the *need* in the stem to the rung; the trap is a stem describing a Project need and a "build a custom GPT / API app" answer.
:::

## 4.2 Prompt vs saved instruction vs Project

The first three rungs are all inside your own ChatGPT use; the difference is *what persists*.

| | Persists? | Shareable? | Carries files? | Choose when |
| --- | --- | --- | --- | --- |
| One-off prompt | No | No | Attach ad hoc | You do this once |
| Saved instruction | The instruction | Copy/paste | No | You repeat the same ask often |
| Project | Instructions + knowledge files + chats | Shared projects with your workspace | Yes (knowledge files) | A task needs the same standing context every time |

**Worked example.** You write a weekly digest that always needs the team's style guide and last quarter's OKRs.

- *Prompt each time:* you re-paste the style guide and OKRs weekly — wasteful and error-prone.
- *Saved instruction:* the ask is saved, but you still attach the files each time.
- *Project:* the style guide and OKRs live as knowledge files; the Project instruction says "produce the Friday digest in the house template." You just drop in the week's updates. This is the right rung — recurring context is exactly what a Project is for.

## 4.3 Project vs custom GPT

Both hold standing instructions and knowledge. The dividing question is **who runs it**.

```text
Do OTHER people need to run this task, on their own, without you?
│
├─ No  ─► PROJECT      (your recurring workspace; shareable within your team
│                       as a shared project, but centred on your work)
│
└─ Yes ─► CUSTOM GPT   (a packaged assistant others invoke; configured once,
                        used by many; can be shared across the workspace)
```

A Project is *your* recurring context. A custom GPT is a *product you hand to others*: it wraps the instructions, knowledge and behaviour so a colleague gets consistent results without knowing how you set it up. If the answer to "who runs it" is "just me, repeatedly", a Project is lighter and sufficient; building a custom GPT is over-reaching.

## 4.4 Custom GPT vs workspace agent

A custom GPT still *responds to a person turn by turn*. A **workspace agent** *executes a multi-step task* on your behalf, using tools and connectors, under oversight.

| | Custom GPT | Workspace agent |
| --- | --- | --- |
| Interaction | Conversational, one turn at a time | Delegated: give an objective, it works multiple steps |
| Best for | A repeatable *assistant* others chat with | A repeatable *task* to be executed with checkpoints |
| Oversight | You read each reply | You set boundaries and review at checkpoints |
| Track link | This domain | Deepened in [Agents and Workflows](/openai/agents/) |

Reach for a workspace agent only when the task genuinely has *steps to execute*, not just *questions to answer* — and always with the oversight design from [Domain 5](/openai/applied-ai/domains/d5-review-points-and-human-oversight/).

## 4.5 When to leave ChatGPT for an API application

ChatGPT stops being the right container when the workflow must be **embedded, scheduled at scale, integrated, or governed programmatically**.

| Signal | Why ChatGPT isn't enough | Where it belongs |
| --- | --- | --- |
| Must run inside another product's UI | ChatGPT is a separate surface | API application |
| Thousands of runs on a schedule | Manual/agent use doesn't scale to that volume | API (with Batch/Flex for cost) |
| Wired to a database or internal service | No programmatic data path in a chat | API application (+ tools/MCP) |
| Deterministic contracts and logging required | You need code around the model calls | API application |
| A one-time or occasional human task | Code is overkill | Stay in ChatGPT |

The trap is jumping to "build an API app" for a task a Project or custom GPT handles. Building software has real cost; it is justified by scale, integration and programmatic control — not by ambition.

## 4.6 In-conversation capabilities — the second axis

Independent of which rung you are on, a task may need a specific capability switched on. Match the signal to the capability.

| Capability | Reach for it when the task needs… | Signal words |
| --- | --- | --- |
| **Search** | Current facts, recent events, live info beyond the model's knowledge | "latest", "current", "today", "recent news" |
| **Deep research** | A sourced synthesis across many sources, with citations | "compare across the market", "cite sources", "thorough report" |
| **Data analysis** | Computation, charts or stats over an uploaded file | "analyse this spreadsheet", "compute", "plot", "trend" |
| **Canvas** | Iterating on a document or code side-by-side | "draft and revise", "edit this together", "work on the doc" |
| **File uploads** | Working over *your own* documents | "summarise this PDF", "answer from these files" |

<Tabs>
  <TabItem label="Search vs deep research">
    **Search** answers a current-facts question quickly ("what's the latest on X?"). **Deep research** does a longer, multi-source, cited investigation ("produce a sourced comparison of the five leading vendors"). Reaching for deep research on a quick factual lookup wastes time; using plain search for a report that must be *sourced and comprehensive* under-delivers.
  </TabItem>
  <TabItem label="Data analysis vs file uploads">
    **File uploads** let the model *read* your documents. **Data analysis** lets it *compute* over them — totals, trends, charts, statistical checks. If the task is "summarise this report", uploads suffice; if it is "find the outliers and plot the trend", you need data analysis.
  </TabItem>
  <TabItem label="Canvas">
    **Canvas** is for iterative deliverables — a document or code you revise across several turns in a dedicated editing surface. If you will refine the same artefact repeatedly, Canvas beats re-pasting; for a one-shot answer it is unnecessary.
  </TabItem>
</Tabs>

:::tip[Assessment signal]
Two capability decisions hide in most stems: *which rung* (packaging) and *which in-conversation capability* (search/research/analysis/Canvas/uploads). "Recent" or "current" → search; "sourced comparison" → deep research; "compute/plot over a file" → data analysis; "iterate on the doc" → Canvas. Don't confuse search (quick, current) with deep research (long, sourced).
:::

## 4.7 Context and the cost of over-reaching

Heavier rungs cost more than money. A Project must be maintained (stale knowledge files mislead), a custom GPT shared across a workspace needs governance, an agent needs oversight design, an API application needs engineering and monitoring. Over-reaching front-loads cost you may never recoup; under-reaching (a one-off prompt for a task ten people run weekly) wastes time and produces inconsistent results. The right rung minimises *total* cost — build + maintain + govern — for the need actually present.

---

## Decision framework

**The LADDER question set** — ask in order; the first "yes" that adds a genuine need moves you up one rung.

| # | Question | If yes → this rung (at least) |
| --- | --- | --- |
| **L** — Later? | Will you do this again? | Saved instruction |
| **A** — Attached context? | Does it need the same files/context every time? | Project |
| **D** — Delegated to others? | Do other people need to run it themselves? | Custom GPT |
| **D** — Do steps run? | Is it a multi-step task to *execute*, not just answer? | Workspace agent |
| **E** — Embedded/at scale? | Must it run in code, at scale, or wired to systems? | API application |
| **R** — Right capability? | Does it need current facts, sourcing, computation, or doc iteration? | Turn on search / deep research / data analysis / Canvas |

Stop at the **lowest** rung whose need is truly present. If only **L** is yes, a saved instruction is the answer — not a custom GPT because it "might be useful to others someday".

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Building a custom GPT when only *you* run the task | Custom GPTs feel more "real" | Use a Project; it is the lighter, sufficient rung |
| Jumping to an API application for a Project-sized task | Engineering feels like the serious option | Reserve the API for scale, integration, or programmatic control |
| Re-pasting the same context into fresh prompts weekly | It works and needs no setup | Move recurring context into a Project's knowledge files |
| Using deep research for a quick factual lookup | "Research" sounds thorough | Use search for current facts; reserve deep research for sourced reports |
| Using plain chat for a sourced market comparison | It answered, sort of | Deep research when the output must be multi-source and cited |
| Uploading a spreadsheet then asking for computed trends without data analysis | Uploads seem to cover files | Turn on data analysis for computation, stats and charts |
| Reaching for a workspace agent for a single question | Agents are the exciting rung | Agents are for multi-step *execution*; a question is a prompt or GPT |
| Leaving stale knowledge files in a Project | Set-and-forget | Maintain Project knowledge; stale files silently mislead |

## Scenario challenge

**Scenario.** Marcus supports a 30-person sales team. Every Monday he compiles a "deal-risk digest": he pulls each rep's notes, checks the latest news on the top five accounts, computes which deals slipped versus last week from the CRM export, and writes a one-page summary in the team's template with the style guide applied. Today he does it all as fresh prompts, re-pasting the style guide and template each time, and it takes two hours. Leadership now wants *every rep* to self-serve a personal version, and wants the whole thing to eventually run automatically each Monday morning and post to Slack. Marcus's manager says "just build an API application for the whole thing".

**Expert reasoning trace.**

1. **Separate the two decisions.** There is a *packaging* question (which rung) and a *capability* question (what to switch on) — and there are actually two different consumers now: Marcus's own recurring work, and the reps' self-serve version, and a future automated version. They may sit at different rungs.
2. **Fix Marcus's own recurring work first — it's a Project.** Re-pasting the style guide and template weekly is the classic Project signal (**A** — same context every time). Put the style guide, template and last-week's-baseline as knowledge files in a Project; Marcus drops in this week's notes. That alone kills most of the two hours. Jumping straight to an API app for *his* task is over-reaching.
3. **Assign the in-conversation capabilities.** "Latest news on top accounts" → **search** (current facts), not deep research, since it's a quick current-facts pull per account. "Which deals slipped vs last week from the CRM export" → **data analysis** (computation over an uploaded file), not just file upload. The write-up in the template is plain generation; iterating it could use **Canvas**.
4. **Handle the reps' self-serve version — that's a custom GPT.** "Every rep runs a personal version themselves, without Marcus" is the defining custom-GPT signal (**D** — others run it). Package the instructions and shared knowledge into a workspace-shared custom GPT so each rep gets consistent output. It is lighter than an API app and exactly fits "others run the same task."
5. **Only the automated Monday-morning-to-Slack version justifies the API.** "Run automatically on a schedule and post to Slack" is embedding + scheduling + integration (**E**) — that is the one piece that genuinely outgrows ChatGPT and belongs in an API application (with the Responses API, tools/connectors, and Batch/Flex if volume grows). But that is a *future* piece, not "the whole thing today."
6. **Reject the manager's one-shot API answer.** Building an API app for all of it front-loads engineering and governance cost for parts (Marcus's own digest, the reps' self-serve) that a Project and a custom GPT handle far more cheaply. Right-sizing means: Project now, custom GPT for reps, API only for the scheduled integration when it's actually needed.

**The decision:** a **Project** (with search + data analysis + Canvas) for Marcus's recurring digest, a **shared custom GPT** for reps to self-serve, and an **API application** *only* for the future scheduled Slack automation — not a single API application for everything. Each need maps to the lightest rung that satisfies it.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Build a custom GPT" for a task only you run | Custom GPTs feel more capable | Others-run-it is the custom-GPT trigger; solo recurring work is a Project |
| "Build an API application" for the whole thing | Engineering sounds serious/scalable | Only scale, integration or programmatic scheduling justifies the API |
| "Use deep research" for a quick current fact | 'Research' implies thoroughness | Current facts → search; deep research is for sourced multi-source reports |
| "File upload is enough" to compute trends | Uploads do read the file | Computation/plots/stats need data analysis switched on |
| "Use a workspace agent" for a single question | Agents are the exciting rung | Agents execute multi-step tasks; a question is a prompt or GPT |
| "One prompt each week is fine" for recurring context | It works with no setup | Recurring standing context is exactly what a Project is for |

## Practice questions

Each item states how many responses to select. Commit before revealing.

<Accordions>
  <AccordionItem title="Q1 · You repeatedly run the same task and it needs the same style guide and reference files every time. Which rung fits BEST? (Select one)">
    A. A one-off prompt with the files re-pasted each time
    B. A Project with the files as knowledge and standing instructions
    C. An API application
    D. A workspace agent

    **Answer: B.** Recurring context that must be present every time is the defining signal for a Project. Re-pasting (A) is the waste a Project removes. An API app (C) and agent (D) over-reach for a solo recurring task.
  </AccordionItem>

  <AccordionItem title="Q2 · The decisive question separating a Project from a custom GPT is: (Select one)">
    A. Which model it uses
    B. Whether other people need to run the task themselves without you
    C. How many knowledge files it has
    D. Its temperature setting

    **Answer: B.** A Project is your recurring workspace; a custom GPT packages the task so others run it independently. Model (A), file count (C) and temperature (D) don't determine which of the two you need.
  </AccordionItem>

  <AccordionItem title="Q3 · A task must run inside your company's web app, on demand, wired to your database. Which capability is required? (Select one)">
    A. A saved instruction
    B. A Project
    C. A custom GPT
    D. An API application

    **Answer: D.** Embedding in another product, on demand, integrated with a database is exactly what an API application is for. Saved instructions (A), Projects (B) and custom GPTs (C) all live inside ChatGPT and cannot embed programmatically into your app.
  </AccordionItem>

  <AccordionItem title="Q4 · A user needs the latest news about a named company from this week. Which in-conversation capability fits BEST? (Select one)">
    A. Deep research
    B. Search
    C. Data analysis
    D. Canvas

    **Answer: B.** A quick current-facts lookup is what search is for. Deep research (A) is heavier — for sourced, multi-source reports. Data analysis (C) is for computation over files. Canvas (D) is for iterating on documents.
  </AccordionItem>

  <AccordionItem title="Q5 · The task is 'produce a thoroughly sourced comparison of the five leading vendors, with citations'. Which capability fits BEST? (Select one)">
    A. Plain chat with no tools
    B. Search for one headline
    C. Deep research
    D. Canvas only

    **Answer: C.** A multi-source, cited, comprehensive comparison is the deep-research use case. Plain chat (A) can't source it reliably. A single search headline (B) under-delivers. Canvas (D) is an editing surface, not a research tool.
  </AccordionItem>

  <AccordionItem title="Q6 · You uploaded a sales spreadsheet and need the outliers found and a trend plotted. What must be switched on? (Select one)">
    A. File uploads alone
    B. Data analysis
    C. Search
    D. A custom GPT

    **Answer: B.** Computation, outlier detection and plotting over a file require data analysis; uploads alone only let the model read the file (A). Search (C) is for current facts. A custom GPT (D) is a packaging rung, not a compute capability.
  </AccordionItem>

  <AccordionItem title="Q7 · Ten colleagues need to run the same configured assistant themselves, getting consistent results without your involvement. Which rung? (Select one)">
    A. A Project you keep to yourself
    B. A custom GPT shared across the workspace
    C. A one-off prompt you send them
    D. An API application

    **Answer: B.** Others running the same task independently, with consistent behaviour, is the custom-GPT signal. A solo Project (A) doesn't hand the task to others. A one-off prompt (C) yields inconsistency. An API app (D) over-reaches for in-ChatGPT self-serve.
  </AccordionItem>

  <AccordionItem title="Q8 · Why prefer the lightest rung that meets the need? (Select one)">
    A. Heavier rungs are always slower to run
    B. Each heavier rung adds build, maintenance and governance cost that is only justified by a real need
    C. Lighter rungs are always more accurate
    D. It is required by OpenAI policy

    **Answer: B.** Climbing the ladder adds total cost of ownership; you pay it only when a concrete need (recurring context, sharing, execution, scale/integration) justifies it. Heavier rungs aren't inherently slower (A) or less accurate (C). It's a design principle, not a policy rule (D).
  </AccordionItem>

  <AccordionItem title="Q9 · A workflow currently done as weekly prompts must eventually run automatically every Monday and post to Slack. Which rung does THAT specific requirement justify? (Select one)">
    A. A saved instruction
    B. A Project
    C. A custom GPT
    D. An API application

    **Answer: D.** Scheduled, automated, integrated-with-Slack execution is embedding at scale — the justification for an API application. Saved instructions (A), Projects (B) and custom GPTs (C) don't run on a schedule wired into Slack programmatically.
  </AccordionItem>

  <AccordionItem title="Q10 · Which TWO signals indicate a workspace agent rather than a custom GPT? (Select two)">
    A. The work is a multi-step task to be executed, not a single question answered
    B. Other people simply need to chat with a configured assistant
    C. The task can run semi-autonomously with review at checkpoints
    D. You need the same reference files present every time
    E. The output is a one-line answer

    **Answer: A and C.** A workspace agent executes multi-step tasks semi-autonomously under oversight; a custom GPT answers turn by turn. Others chatting with an assistant (B) is a custom GPT. Standing files (D) point to a Project. A one-line answer (E) is a prompt.
  </AccordionItem>

  <AccordionItem title="Q11 · A manager says 'just build an API application' for a task that is: (a) your own recurring digest, (b) a version reps run themselves, (c) a future scheduled Slack post. What is the BEST right-sizing? (Select one)">
    A. One API application for all three
    B. A Project for (a), a shared custom GPT for (b), and an API application only for (c)
    C. A custom GPT for all three
    D. Keep everything as weekly prompts

    **Answer: B.** Each need maps to a different rung: recurring solo context → Project; others self-serve → custom GPT; scheduled integrated automation → API. One API app for all (A) over-builds (a) and (b). A custom GPT for all (C) can't do scheduled Slack posting. Weekly prompts (D) don't scale or self-serve.
  </AccordionItem>

  <AccordionItem title="Q12 · A recurring deliverable is a document you refine over several turns in the same session. Which in-conversation capability best supports the iteration? (Select one)">
    A. Search
    B. Deep research
    C. Canvas
    D. Data analysis

    **Answer: C.** Canvas is the dedicated surface for iterating on a document or code across turns. Search (A) and deep research (B) gather information; data analysis (D) computes over files — none is an editing/iteration surface.
  </AccordionItem>

  <AccordionItem title="Q13 · You keep re-pasting the same 4-page background into new chats for a monthly task and results vary. Which TWO fixes are appropriate, cheapest first? (Select two)">
    A. Move the background into a Project's knowledge files with standing instructions
    B. Immediately build a full API application
    C. If colleagues also run it, package it as a shared custom GPT
    D. Increase the model temperature for consistency
    E. Keep re-pasting but use a bigger model

    **Answer: A and C.** The recurring context belongs in a Project (cheapest fix), and if others run it too, a shared custom GPT packages it for them. An API app (B) over-reaches for this. Temperature (D) worsens consistency. A bigger model (E) doesn't fix the re-pasting or the variance from missing standing context.
  </AccordionItem>

  <AccordionItem title="Q14 · A stem says the output 'must be a sourced, cited market report' AND 'must then be refined into a polished brief over several edits'. Which TWO capabilities fit? (Select two)">
    A. Deep research for the sourced report
    B. Search for a single current fact
    C. Canvas for the iterative refinement of the brief
    D. Data analysis for computing over a file
    E. A workspace agent to answer one question

    **Answer: A and C.** A sourced, cited report calls for deep research, and refining it across several edits calls for Canvas. A single search fact (B) under-delivers the report. Data analysis (D) isn't needed with no file to compute over. A one-question agent (E) doesn't match a multi-edit deliverable.
  </AccordionItem>
</Accordions>

## Key takeaways

- The capability ladder runs **prompt → saved instruction → Project → custom GPT → workspace agent → API application**; reach for the **lightest rung that meets the need**.
- **Project vs custom GPT** turns on *who runs it*: your recurring context is a Project; a task others run themselves is a custom GPT.
- **Custom GPT vs workspace agent** turns on *answer vs execute*: a GPT responds turn by turn; an agent executes multi-step work under oversight.
- Leave ChatGPT for an **API application** only when the workflow must be **embedded, scheduled at scale, integrated, or programmatically governed**.
- The second axis is **in-conversation capability**: search (current facts), deep research (sourced synthesis), data analysis (computation over files), Canvas (iterating on a doc/code), file uploads (your documents).
- Don't confuse **search** (quick, current) with **deep research** (long, sourced), or **file uploads** (read) with **data analysis** (compute).
- Over-reaching front-loads build/maintain/govern cost; under-reaching wastes time and yields inconsistent results — right-size to the need actually present.
