Agents and Workflows
D1 · What an Agent Is and When to Use One
Distinguishing a prompt from a workflow from an agent on autonomy, tool access, duration and reversibility, and judging honestly when an agent is the wrong tool.
This domain is worth 16% of the material — roughly 8 of 50 items on our mock. It tests one judgment above all others: given a piece of work, is the right tool a single prompt, a repeatable workflow, or a delegated agent — and when is an agent actively the wrong answer? The Academy course opens here because most delegation failures are chosen before a brief is ever written: someone reached for an agent when a workflow would have been safer, cheaper and easier to verify.
What you need to know
A prompt is a single request you supervise turn by turn. A workflow is a repeatable sequence of prompting steps with defined inputs, outputs and review points, still driven by you. An agent is given an objective and the autonomy to plan, use tools and take multiple steps toward it with limited supervision. The four axes that separate them are autonomy (how much it decides for itself), tool access (what it can touch), duration (how long it runs unsupervised) and reversibility (how hard its actions are to undo). An agent is the right choice when a task is genuinely multi-step, benefits from tool use, and is worth the cost of setting boundaries and verifying unwatched work. It is the wrong choice when the task is trivial, when you cannot specify a clear definition of done, or when the actions are irreversible and cannot be gated.
Learning objectives
By the end of this page you should be able to:
- Distinguish a prompt, a workflow and an agent on autonomy, tool access, duration and reversibility.
- Decide which of the three fits a given task, and justify the choice against those axes.
- Recognise the situations where an agent is the wrong tool and a workflow or a single prompt is safer.
- Describe what “human oversight” means for an agent versus for a workflow.
- Estimate the setup and verification cost an agent adds, and weigh it against the work it saves.
1.1 The three modes of working with AI
The whole domain rests on one distinction. Read the axes, not the labels: the same task can be a prompt today and an agent tomorrow depending on how much you hand off.
| Mode | Autonomy | Tool access | Typical duration | Reversibility need | You are… |
|---|---|---|---|---|---|
| Prompt | You decide every step | Usually none, or one you invoke | Seconds to a minute, watched | Low — you see output before acting | Doing the work with help |
| Workflow | You sequence the steps; each is a prompt | Whatever each step needs, invoked by you | Minutes, with review points | Medium — review points catch errors | Running a repeatable process |
| Agent | It plans and chooses steps toward a goal | A set of tools it can use on its own | Minutes to a long run, mostly unwatched | High — it may act before you see it | Delegating an outcome |
autonomy ▲ │ ┌──────────┐ high │ │ AGENT │ goal in, plans + acts, │ └──────────┘ uses tools, runs unwatched │ ┌────────────┐ medium │ │ WORKFLOW │ you sequence steps + review points │ └────────────┘ │ ┌────────┐ low │ │ PROMPT │ one request, you supervise the turn │ └────────┘ └──────────────────────────────────────────► tool access + duration + blast radiusAssessment signal
When a stem says the work is “one clear question”, “a single rewrite” or “you will read the answer before doing anything with it”, it is a prompt — even if it is complex. Reach for agent only when the stem says “multiple steps”, “use these tools”, “run on its own” or “come back when it is done”.
1.2 Autonomy: who chooses the next step
Autonomy is the axis people misjudge most. A long, detailed prompt is still low-autonomy if you decide what happens next. Autonomy is about who chooses the path, not how big the request is.
| Signal in the task | Autonomy level | Right mode |
|---|---|---|
| You will read the output and decide what to do next | Low | Prompt |
| The steps are fixed and you will run each one in order | Medium | Workflow |
| The steps depend on what earlier steps find, and you want it to figure that out | High | Agent |
An agent earns its keep exactly when the path cannot be scripted in advance — when step three depends on what step two discovered. If you already know every step, a workflow is cheaper to build and far easier to verify.
1.3 Tool access: capability and blast radius are the same thing
An agent is only as powerful as the tools it can reach, and only as dangerous. Every connector, search tool or file action you grant increases both what it can accomplish and what it can break.
| Tool access granted | What it enables | What it risks |
|---|---|---|
| Read-only search over supplied documents | Grounded answers, synthesis | Very little — it cannot change anything |
| A connector to shared files (read) | Company-knowledge answers | Reading data you did not intend to expose |
| A connector that can edit or send | Real work completed end to end | Irreversible external actions |
| The open web plus the ability to act on it | Broad research and execution | Largest blast radius; hardest to verify |
The rule this domain teaches, and D3 develops: grant the least access that lets the work succeed. A summariser needs read, not write. A research agent needs search, not send.
1.4 Duration and unwatched work
Duration changes the problem qualitatively, not just quantitatively. A ten-second prompt you watch is self-correcting — you see a wrong turn and stop it. A twenty-minute agent run is not: by the time you look, it has already taken the steps.
watched work unwatched work────────────── ────────────────you see each step │ you see a plan, then a resultyou can stop instantly │ you can only stop at checkpointserrors caught live │ errors caught by verification aftertrust = observation │ trust = evidence trail + spot-checksThis is why longer autonomy demands the boundaries of D4 and the verification of D5. The moment work runs unwatched, “I’ll notice if it goes wrong” stops being a control.
1.5 Reversibility: the axis that decides how careful to be
Reversibility is the single most important axis for judging risk. It does not decide whether to use an agent — it decides how much of D4 (boundaries) and D5 (review) you must add before you do.
| Action the work performs | Reversible? | What it forces |
|---|---|---|
| Draft a document, summarise, analyse | Yes — nothing left the workspace | Light review is enough |
| Save a file inside the workspace | Mostly — you can delete it | Spot-check before relying on it |
| Send an email, post a message, publish | No — the recipient has it | Approval gate before the action |
| Pay, delete, change a system of record | No — often costly to undo | System-enforced gate plus sign-off |
Assessment signal
“Irreversible”, “external”, “sends”, “publishes”, “pays”, “cannot be undone” in a stem push the answer toward add an approval gate or this needs human sign-off — regardless of how capable the agent is. Capability never removes the need for a gate on an irreversible step.
1.6 When an agent is the wrong answer
The course is refreshingly honest that agents are frequently the wrong tool. Reaching for one when a workflow would do is the classic beginner mistake, because agents cost more to set up, cost more to run, and are harder to verify.
An agent is the wrong tool when:
- The task is trivial or one-shot — a single rewrite, one lookup, one calculation. A prompt is faster.
- The steps are fully known and fixed — a documented workflow is cheaper to build and trivial to verify.
- You cannot state a clear definition of done — an agent given a fuzzy goal will produce fuzzy work and you will not be able to tell if it succeeded.
- The actions are irreversible and cannot be gated — if you cannot put an approval checkpoint before the dangerous step, do not automate it.
- Verification would cost more than the work — if checking the unwatched output takes longer than doing the task yourself, delegation is a net loss.
- The task needs judgment you cannot specify — nuanced tone, sensitive stakeholder calls, anything where “you’ll know it when you see it” is the real acceptance test.
Should I use an agent?│├─ Is it one step I'll watch? ── yes ─► PROMPT├─ Are the steps fixed and known? ── yes ─► WORKFLOW├─ Can I state a clear "done"? ── no ─► not yet: sharpen the objective (D2)├─ Are irreversible actions gate-able? ── no ─► don't automate the dangerous step└─ otherwise ─► AGENT, with boundaries + verification1.7 What “oversight” means at each level
Human oversight is not one thing. Its shape changes with autonomy, and matching the wrong kind of oversight to the mode is a common trap.
| Mode | Oversight looks like |
|---|---|
| Prompt | You read the output before you use it. That is the whole control. |
| Workflow | You review at defined points between steps; a bad step is caught before the next runs. |
| Agent | You set boundaries up front, approve gated actions during the run, and verify artefacts after. Oversight is designed in, not applied by watching. |
The mistake is assuming an agent can be supervised the way a prompt is — by watching. You cannot watch a long unwatched run; you can only bound it and check it.
Decision framework
Use the PADR test to place any task before you build anything. Score each axis, and the highest-demanding axis usually decides the mode.
| Axis | Ask | Prompt | Workflow | Agent |
|---|---|---|---|---|
| P — Path | Do I know every step in advance? | I’ll decide live | Yes, fixed | No, it must plan |
| A — Access | What must it touch? | Nothing / one tool I invoke | Tools I invoke per step | A tool set it uses itself |
| D — Duration | How long does it run unwatched? | Seconds, watched | Minutes, with reviews | Minutes to long, unwatched |
| R — Reversibility | How hard to undo its actions? | Trivial | Reviewable | Needs gates on irreversible steps |
How to read it: if every axis lands in the Prompt column, do not build a workflow. If Path is “fixed” but you are tempted by an agent, build the workflow — it is cheaper and more verifiable. Only when Path is “must plan” and Access is “a tool set it uses itself” does an agent clearly win, and then R tells you how many boundaries to add before you let it run.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| Using an agent for a one-shot task | Agents feel more powerful and impressive | If you’ll read the answer and act yourself, it’s a prompt |
| Treating a long prompt as an agent | Length is mistaken for autonomy | Autonomy is who chooses the next step, not prompt size |
| Building an agent for fully known steps | “Automation” sounds better than “workflow” | A fixed sequence is a workflow — cheaper, more verifiable |
| Granting broad tool access “to be safe” | Fear of the agent getting stuck for lack of access | Grant least access; widen only when a run actually needs it |
| Delegating without a definition of done | The goal felt obvious in your head | If you can’t state “done”, you can’t delegate it yet |
| Assuming you’ll supervise by watching | Habit from single-prompt use | Long runs are unwatched; design boundaries and verification |
| Ignoring reversibility when choosing | The output looked harmless | Irreversible actions need gates before the mode is even chosen |
| Judging fit by task difficulty | Hard tasks feel like “agent work” | Fit is decided by path, access, duration and reversibility, not difficulty |
Scenario challenge
Scenario. Priya leads operations at a mid-size firm on ChatGPT Work. Every Monday she compiles a supplier-status report: she pulls the latest delivery figures from a shared spreadsheet, checks three supplier portals for open issues, drafts a one-page summary, and emails it to the leadership list. It takes her about ninety minutes and the steps are the same every week. A colleague suggests she “build an agent to just do the whole thing and send it”. Priya is tempted — the sending is the part she most wants off her plate.
Expert reasoning trace.
- Run the PADR test, not the enthusiasm. Path: the steps are identical every week — that is a fixed path, which points to a workflow, not an agent. Access: it needs to read a spreadsheet and read three portals, and to send an email. Duration: a few minutes, and she could watch it. Reversibility: the email to leadership is irreversible and external — the highest-risk axis in the whole task.
- Separate the reversible bulk from the irreversible tail. Pulling figures, checking portals and drafting the summary are all reversible read-and-draft work — safe to automate. Sending to leadership is the one irreversible step, and it is exactly the step she wanted to hand off.
- Resist “one agent that sends it”. The temptation is to automate end to end, but the send is where blast radius lives. A wrong figure emailed to leadership cannot be recalled.
- Choose the mode per axis. Because the path is fixed, the right build is a workflow with review points — cheaper and far easier to verify than an agent. If she later wants more autonomy (e.g., it decides which portals to check based on which suppliers are active), it becomes an agent — but even then the send stays behind an approval gate.
- Design the gate. The workflow drafts the email and pauses; Priya reviews the one-page summary and the figures, then approves the send. She has moved the ninety minutes of assembly off her plate while keeping the two-minute judgment — the review and the send — where it belongs.
The point. The task did not need an agent; it needed a workflow with a human gate on the one irreversible step. Reaching for an agent “because it’s more powerful” would have added setup and verification cost for no benefit, and automating the send outright would have removed the only control that mattered.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| “Complex task, so use an agent” | Difficulty feels like agent territory | Fit is decided by path/access/duration/reversibility, not difficulty |
| “Long detailed prompt = agent” | It’s big, so it must be autonomous | Autonomy is who picks the next step, not prompt length |
| “Fixed weekly steps, build an agent” | Automation sounds more advanced | Known fixed steps are a workflow — cheaper and more verifiable |
| “Automate it end to end including the send” | Removes the most tedious step | Irreversible external steps need a gate; don’t automate them away |
| “Give it every tool so it never gets stuck” | Feels safer than under-provisioning | Broad access is broad blast radius; grant least access |
| “Supervise the long run by watching it” | Works fine for single prompts | Unwatched runs need designed boundaries and after-the-fact verification |
Practice questions
Each item states how many responses to select. Attempt before revealing.
Q1 · A user needs a single paragraph of marketing copy rewritten in a friendlier tone, which they will read and paste themselves. Which mode fits BEST? (Select one)
A. An agent with a web-search tool. B. A prompt — one request the user reads and uses. C. A workflow with three review points. D. An agent that emails the copy to the team.
Answer: B. One request the user reads and acts on is a prompt; autonomy, tool access and duration are all minimal. An agent (A, D) adds setup, cost and blast radius for no benefit. A three-step workflow (C) over-engineers a single rewrite.
Q2 · Which four axes best distinguish a prompt from a workflow from an agent? (Select one)
A. Model size, temperature, token count and price. B. Autonomy, tool access, duration and reversibility. C. Prompt length, formatting, language and tone. D. Speed, cost, popularity and vendor.
Answer: B. The course frames the distinction on autonomy (who chooses the next step), tool access (what it can touch), duration (how long it runs unwatched) and reversibility (how hard actions are to undo). Model settings (A), stylistic properties (C) and commercial factors (D) do not separate the three modes.
Q3 · A task has the same fixed five steps every week and you will run each in order with a quick check between them. Which mode is MOST appropriate? (Select one)
A. A prompt. B. A workflow. C. An agent that plans its own steps. D. An agent with broad tool access.
Answer: B. Fixed, known steps with review points between them are the definition of a workflow — cheaper to build and easier to verify than an agent. A single prompt (A) can’t carry five sequenced steps. Agents (C, D) add autonomy you don’t need when the path is already fixed.
Q4 · A colleague argues 'this task is really hard, so it needs an agent'. What is the BEST response? (Select one)
A. Agree — hard tasks are what agents are for. B. Difficulty alone doesn’t decide the mode; judge path, access, duration and reversibility instead. C. Use the most expensive model for hard tasks. D. Split the hard task into ten prompts regardless.
Answer: B. Fit is decided by the four axes, not by how hard the task feels; a hard task with a fixed path is still a workflow. Difficulty as the deciding factor (A) is the trap. Model choice (C) is unrelated to the mode. Blindly splitting (D) ignores whether the steps are even fixed.
Q5 · An agent will read three supplier portals and draft a report, then send it to leadership. Which part MOST demands a control before you automate? (Select one)
A. Reading the first portal. B. Drafting the report. C. Sending the report to leadership. D. Formatting the report as one page.
Answer: C. Sending to an external audience is irreversible — the highest-blast-radius step — so it needs an approval gate before automation. Reading portals (A) and drafting (B) are reversible and safe to automate. Formatting (D) carries no risk.
Q6 · You cannot clearly state what 'done' looks like for a task. What does this tell you about delegating it to an agent? (Select one)
A. Delegate anyway; the agent will infer the goal. B. It is not yet ready to delegate — sharpen the objective and definition of done first. C. Use a bigger model to compensate. D. Grant more tools so it has more options.
Answer: B. Without a definition of done you cannot brief the agent or verify its work, so the task isn’t delegable yet — the fix is in D2, not in more model or tools. Delegating a fuzzy goal (A) produces fuzzy work. A bigger model (C) or more tools (D) can’t substitute for a missing acceptance test.
Q7 · What does 'autonomy' mean when distinguishing an agent from a long prompt? (Select one)
A. The number of tokens in the request. B. Who chooses the next step — you, or the agent. C. How polite the model is. D. Whether the output is formatted as a table.
Answer: B. Autonomy is about who decides the path: in a prompt you decide the next move, in an agent it plans and chooses. Token count (A) measures length, not autonomy — a long prompt is still low autonomy. Tone (C) and formatting (D) are irrelevant.
Q8 · An operations analyst wants to hand off a task whose steps depend on what earlier steps discover, involving several tools, running unwatched for about fifteen minutes, producing only a draft. Which mode fits and why? (Select two)
A. An agent, because the path can’t be scripted in advance. B. A prompt, because it produces only a draft. C. An agent, because it needs to use several tools on its own over an unwatched run. D. A workflow, because the steps are fixed. E. A single prompt with a very long instruction.
Answer: A and C. The path is discovered as it goes (A) and it uses several tools autonomously over an unwatched run (C) — both are the defining signals of an agent, and the reversible draft output keeps the risk manageable. It is not a prompt (B, E) because it isn’t one watched step, and not a workflow (D) because the steps are not fixed in advance.
Q9 · Why is a twenty-minute unwatched agent run qualitatively different from a ten-second prompt you watch? (Select one)
A. It uses more tokens. B. You can’t catch a wrong turn live, so trust must come from boundaries and after-the-fact verification, not observation. C. It always costs more money. D. It is always less accurate.
Answer: B. Watched work is self-correcting because you can stop a wrong step; unwatched work has already acted by the time you look, so control shifts to designed boundaries and verification. Token use (A) and cost (C) are consequences, not the qualitative difference. Unwatched runs aren’t inherently less accurate (D).
Q10 · When is granting an agent broad tool access the WRONG default? (Select one)
A. Always — broad access maximises capability with no downside. B. Almost always at the start — broad access is broad blast radius; grant the least access the work needs and widen only if a run requires it. C. Only when the model is small. D. Only on weekends.
Answer: B. Least-privilege thinking treats every added tool as added blast radius, so you start narrow and widen on evidence of need. Broad-by-default (A) maximises risk, not just capability. Model size (C) and timing (D) are irrelevant to the principle.
Q11 · A manager insists on 'one agent that does the whole weekly report and sends it automatically' for a task with fixed steps. Which TWO points should you raise? (Select two)
A. Fixed steps point to a workflow, which is cheaper and easier to verify than an agent. B. The automatic send is irreversible and external, so it needs an approval gate rather than full automation. C. Agents are always better than workflows, so proceed. D. The report should be sent without review to save time. E. A bigger model removes the need for any gate.
Answer: A and B. With fixed steps a workflow is the right build (A), and the one irreversible step — the send — should stay behind a human gate rather than being fully automated (B). Claiming agents always win (C) ignores cost and verifiability. Sending without review (D) removes the only control that matters. No model (E) removes the need to gate an irreversible action.
Q12 · A task would take you five minutes to do, but building and verifying an agent for it would take an hour every time it runs. What does the reversibility-and-cost view say? (Select one)
A. Automate it — automation is always worth it. B. Don’t delegate: if verification costs more than the work, delegation is a net loss. C. Delegate but skip verification to save time. D. Use two agents to be safe.
Answer: B. When checking the unwatched output costs more than simply doing the task, an agent is a net loss — a core ‘wrong tool’ signal. Automating regardless (A) ignores the cost. Skipping verification (C) trades a small time saving for unbounded risk. Adding a second agent (D) increases cost and verification burden.
Key takeaways
- Prompt, workflow and agent differ on four axes: autonomy, tool access, duration and reversibility — not on difficulty or prompt length.
- Autonomy is who chooses the next step. A long prompt is still low-autonomy if you decide what happens next.
- Use an agent when the path can’t be scripted in advance and it needs to use tools over an unwatched run; use a workflow when the steps are fixed.
- Reversibility decides how many boundaries and how much verification you must add, not whether you may use an agent.
- An agent is the wrong tool for trivial tasks, fully-known steps, fuzzy goals, un-gateable irreversible actions, and cases where verification costs more than the work.
- Grant the least tool access the work needs; every added tool is added blast radius.
- Oversight for an agent is designed in (boundaries plus verification), not applied by watching a long unwatched run.
Last updated Sep 18, 2026