AI Cert Prep
Type to search documentation.

API Developer Path

D1 · Scoping AI Solutions

Framing a problem for AI, eliciting requirements, defining success metrics and a cost envelope, deciding build versus buy, judging feasibility and writing a solution plan.

This domain is about 12% of the OAI-API mock – roughly 7 of 60 items – and it mirrors the Academy Scope AI Solutions course (30 min). It tests the work you do before a single API call: turning a vague business ask into a defined job, a measurable success criterion and a cost envelope, and deciding whether AI is even the right tool. Most items describe a stakeholder request and ask what you should establish first, or whether the problem is a good fit for a language model at all.

What you need to know

Scoping is requirement elicitation for probabilistic software. You clarify the job to be done, the inputs and outputs, the quality bar and how it will be measured, the volume and latency the system must sustain, and the cost and risk envelope the business will accept. You decide build-vs-buy (a Responses API integration, a hosted product, or no AI at all), and you judge feasibility honestly – some tasks a model does well, some it does badly, and some it cannot verify. The output of scoping is a short solution plan that names the metric you will move and the smallest thing you can ship to test it.

Learning objectives

By the end of this page you should be able to:

  1. Reframe a business request as a well-formed AI task with defined inputs, outputs and constraints.
  2. Elicit requirements – volume, latency, quality bar, data sensitivity, integration points.
  3. Define a success metric that is measurable before you build.
  4. Estimate a cost envelope from token volume and model price, and check it against the business value.
  5. Judge feasibility and choose build-vs-buy, including “do not use AI here”.
  6. Write a solution plan that scopes the smallest useful first version.

1.1 Reframing a request as a task

Stakeholders describe outcomes (“make support faster”); you need a task (“draft a first-response reply from the ticket text and the last three tickets from the same customer”). A well-formed task names the inputs the model will see, the output shape it must produce, and the boundary of what it must not do.

Vague requestReframed taskInputsOutput
“Summarise our meetings”Produce a 5-bullet action list with owners from a transcriptTranscript textMarkdown list, one owner per action
“Answer policy questions”Answer HR policy questions grounded only in the policy PDFs, with citationsQuestion + retrieved policy chunksAnswer + cited sources
“Triage tickets”Classify a ticket into one of 8 queues and a priorityTicket subject + bodyJSON {queue, priority}
“Help engineers”Suggest a code fix for a failing test, for human reviewRepo + failing test outputPatch proposal

Assessment signal

Stems that say “a stakeholder asks you to…”, “the business wants…”, or “before you start building” are scoping items. The correct answer usually clarifies or measures something first rather than jumping to a model or a prompt.

1.2 Eliciting requirements

Six questions decide almost every design choice downstream. Ask them before you pick a model.

RequirementQuestionWhy it drives design
VolumeHow many requests per day, and peak per minute?Sets rate limits, batch vs realtime, cost
LatencyInteractive (user waiting) or background?Streaming, background mode, fast vs flex
Quality barWhat is “good enough”, and who decides?Model choice, reasoning effort, review gate
Data sensitivityPII, regulated data, residency constraints?Retention, store, residency, network controls
InputsText, images, files, structured data?Model capability, file inputs, RAG
IntegrationWhere does the output go, in what format?Structured outputs, function calling

1.3 Defining a success metric

If you cannot measure it, you cannot improve it and you cannot prove it works. A scoping metric must be defined before build and be evaluable on a held-out dataset (this is where D1 hands off to D3, evals).

text
Task type Good success metric Bad metric
───────────────── ─────────────────────────────── ─────────────────
Classification accuracy / F1 on a labelled set "feels accurate"
Extraction field-level exact match "looks complete"
Grounded Q&A % answers supported by a citation "sounds right"
Summarisation rubric score by an LLM grader "reads well"
Support drafting % accepted by agent without edit "faster" (unmeasured)

Assessment signal

When an option says “we will know it is working when customers are happier” and another says “we will measure the percentage of drafts an agent accepts unedited on a 200-ticket sample”, the measurable, pre-defined metric is correct. Vague outcomes are distractors.

1.4 The cost envelope

Scope the cost before you build; a solution that works but costs more than the value it creates is a failed scope. Estimate from expected tokens and the model price. The September 2026 prices per million tokens:

ModelInput $/MTokOutput $/MTok
gpt-6-astra$10$50
gpt-5.6-sol$4$20
gpt-5.6-terra$2$12
gpt-5.6-luna$0.20$1.20

Worked example: a triage classifier on gpt-5.6-luna, 800 input tokens and 40 output tokens per ticket, 50,000 tickets/day.

text
Input: 50,000 × 800 = 40,000,000 tok/day = 40 MTok × $0.20 = $8.00/day
Output: 50,000 × 40 = 2,000,000 tok/day = 2 MTok × $1.20 = $2.40/day
Daily total ≈ $10.40 → ≈ $312/month

The same task on gpt-5.6-sol would be roughly 40 × $4 + 2 × $20 = $200/day (~$6,000/month) for no accuracy gain on a task Luna handles – the cost envelope is what rejects that choice.

1.5 Build versus buy versus don’t

OptionChoose whenWatch out for
Build on the APIYou need control, custom data, or the task is core to your productYou own evals, ops, cost and safety
Buy a hosted productA commodity capability exists (transcription, generic chat) and speed-to-value winsData flow, lock-in, per-seat cost at scale
Use ChatGPT / no codeA knowledge worker can do it in the UI with Projects and filesNot repeatable or auditable at volume
Don’t use AIThe task needs a guaranteed-correct, deterministic answerForcing AI onto arithmetic, lookups, or exact rules

The “AI can’t verify itself” test

If the task requires an answer that must be provably correct and there is no cheap external check (regulatory lookups, financial reconciliation with a ground truth, safety-critical instructions), a probabilistic model is the wrong core – use it to assist a deterministic system, not to replace it.

1.6 Judging feasibility

text
Does a ground-truth or a cheap check exist?
│
┌──────────────┴───────────────┐
│ Yes │ No
▼ ▼
Measurable → good AI fit. Can a human review each output
Build with an eval loop. at the required volume?
│
┌──────────┴──────────┐
│ Yes │ No
▼ ▼
Human-in-the-loop Reconsider scope: narrow
assist is feasible. the task, add a check, or
don't ship AI here.

1.7 The solution plan template

The deliverable of scoping is a one-page plan. Every item on the mock that asks “what should the plan include” expects these fields.

text
Problem: the business pain, in one sentence
Job to be done: the specific task, with inputs and outputs
Users & volume: who, how many requests/day, peak/min
Success metric: measurable, on a held-out dataset, target value
Constraints: latency, data sensitivity, residency, budget
Cost envelope: $/request estimate × volume vs. value created
Approach: build/buy/no-AI, model, retrieval? tools?
Risks & gates: failure modes, human review points
First version: smallest shippable slice + how you'll evaluate it

Decision framework

Use the SCOPE framework to turn any request into a plan you could act on tomorrow.

LetterStepQuestion to answerOutput
SState the jobWhat exact task, with what inputs and outputs?One-sentence task definition
CCriteriaHow will we measure “good enough”?A metric and a target on a dataset
OOperating limitsVolume, latency, data sensitivity, budget?Constraint list + cost envelope
PPathBuild, buy, or no AI? Which model and tools?Approach and model choice
EExperimentWhat is the smallest version that tests the metric?First-version scope + eval plan

The most valuable habit the course teaches is finishing C before starting P: teams that pick a model before defining the metric almost always over-buy.

Common mistakes

MistakeWhy it happensWhat to do instead
Picking the model firstIt feels like progress; models are excitingDefine the metric and cost envelope, then pick the cheapest model that meets them
No measurable success metric“It looks good” is easy; a dataset is workWrite a metric evaluable on a held-out set before building
Scoping the ideal system, not the first versionAmbition and stakeholder pressureScope the smallest slice that tests the metric
Ignoring volume and peakDemos run once; production runs constantlyEstimate requests/day and peak/min; size limits and cost from them
Forcing AI onto deterministic workAI is the mandate of the quarterUse deterministic code for exact rules; let AI assist, not decide
Skipping data-sensitivity questionsIt surfaces late, as a launch blockerAsk about PII, regulated data and residency during elicitation
Treating cost as an afterthoughtToken math is tediousCompute $/request × volume and compare to value in the plan
Confusing “faster” with a metricSpeed is intuitive but unmeasuredDefine what faster means numerically (e.g., median handle time)

Scenario challenge

Scenario. You are a developer at a logistics company. The VP of Operations says: “Our dispatchers waste hours reading driver incident reports. Build an AI that handles them.” There is no dataset, no defined output, and the VP expects a demo in a week. Incident reports are free text, sometimes contain injury details (regulated), arrive at roughly 3,000/day with spikes after storms, and currently a dispatcher reads each one and either files it, escalates to safety, or requests more detail.

Expert reasoning trace.

  1. Reframe the request into a job. “Handle them” is an outcome, not a task. The real job is: given a report, classify it into file / escalate / need-more-detail and extract structured fields (date, location, severity, injury flag). That is a classification-plus-extraction task, not open-ended chat.
  2. Define the metric before building. Ask the safety team for 200 historical reports with the decision a dispatcher actually made. That labelled set gives an accuracy target and, critically, a way to check the escalation decision – the high-stakes one where a false negative is dangerous.
  3. Surface data sensitivity early. Injury details are regulated. That drives retention (store: false or short retention), residency questions, and a mandatory human review gate on any escalate decision – you never auto-close a safety-relevant report.
  4. Estimate the cost envelope. 3,000/day at ~1,200 input + 60 output tokens on gpt-5.6-luna is ~3.6 MTok × $0.20 + 0.18 MTok × $1.20 ≈ $0.72 + $0.22 = $0.94/day. Trivial versus dispatcher hours saved – the envelope is not the constraint here; correctness on escalation is.
  5. Scope the first version, not the ideal one. Ship the classifier with human confirmation on every escalate, measured against the 200-report set. Defer auto-filing until the eval proves the escalation recall is high enough.
  6. Reject the demo trap. A week-one demo that classifies with no eval and no review gate would look impressive and be dangerous. The plan says: build the eval set first, then the classifier, then measure.

Exam-correct decision: establish the labelled dataset and the escalation metric first, scope a human-in-the-loop first version, and size cost from real volume. Not “start prompting gpt-6-astra and demo it”, not “auto-close low-severity reports before measuring recall”.

Assessment traps

TrapWhy it is temptingThe discriminator
“Pick gpt-6-astra so quality is never the problem”Best model feels safestScope the metric and cost first; the cheapest model that passes the eval wins
“Ship the demo this week to show progress”Stakeholder pressureA demo without a metric or a review gate is scope failure, not progress
“The metric is that users are happier”Sounds customer-focusedA metric must be measurable on a held-out dataset before build
“Automate the whole workflow end to end”AmbitionScope the smallest slice that tests the metric; keep humans on high-stakes steps
“AI can do anything, so it fits”HypeIf the task needs provable correctness with no cheap check, AI is the wrong core
“Cost doesn’t matter, tokens are cheap”Individually trueAt volume, $/request × requests/day is the number that kills or saves a project

Practice questions

Each item states how many responses to select. Commit before revealing.

Q1 · A product manager asks you to 'use AI to improve onboarding'. What should you establish FIRST? (Select one)

A. Which model has the largest context window. B. The specific task, its inputs and outputs, and a measurable success metric. C. Whether to use the Agents SDK or the Agents API. D. The prompt wording for the first version.

Answer: B. Scoping starts by reframing an outcome into a defined, measurable task. Model context (A), runtime (C) and prompt wording (D) are all downstream P-step choices that depend on the job and metric you have not defined yet.

Q2 · Which of the following is a well-formed success metric for a ticket-triage classifier? (Select one)

A. Customers report being happier with support. B. The model feels accurate in testing. C. Classification accuracy of at least 92% on a held-out set of 500 labelled tickets. D. Responses are generated in under two seconds.

Answer: C. A metric must be measurable on a held-out dataset with a target. Happiness (A) is an unmeasured outcome, “feels accurate” (B) is subjective, and latency (D) is a constraint, not a quality metric for classification.

Q3 · A task requires returning a customer's exact contractual renewal date, which exists in a structured billing database. What is the BEST approach? (Select one)

A. Ask gpt-6-astra at max reasoning effort to recall the date. B. Have the model call a function that queries the billing database and return the exact value. C. Fine-tune a model on all past renewal dates. D. Put the entire billing table in the prompt every request.

Answer: B. Exact, provable facts belong in a deterministic lookup the model calls as a tool, not in model recall. Recall (A) can hallucinate, fine-tuning (C) bakes in stale data, and dumping the whole table (D) is expensive and still not authoritative.

Q4 · You are scoping a summarisation feature for 20,000 documents per day, ~1,000 input and ~150 output tokens each, quality bar met by `gpt-5.6-luna`. Roughly what is the daily cost? (Select one)

A. About $0.76. B. About $7.60. C. About $76. D. About $760.

Answer: B. Input: 20,000 × 1,000 = 20 MTok × $0.20 = $4.00. Output: 20,000 × 150 = 3 MTok × $1.20 = $3.60. Total ≈ $7.60/day. The discriminator is doing the tokens ÷ 1,000,000 × price/MTok arithmetic correctly: option A is 10× too low, and C and D are 10× and 100× too high.

Q5 · A stakeholder wants an AI that computes exact monthly financial reconciliations that must always match the ledger. Which TWO statements should shape your scope? (Select two)

A. Exact reconciliation is deterministic work; the model should orchestrate deterministic calculations, not perform the arithmetic itself. B. Any output touching financial figures needs a verification step against ground truth before use. C. gpt-6-astra at max effort guarantees correct arithmetic. D. Fine-tuning on past reconciliations makes the arithmetic exact. E. Because it is internal, no verification is needed.

Answer: A and B. Exact arithmetic is deterministic and must be computed by code (or a tool) with a verification step against the ledger; the model coordinates. No reasoning effort (C) or fine-tuning (D) makes a probabilistic model provably exact, and “internal” (E) does not remove the need to verify financial figures.

Q6 · Your VP wants a full autonomous incident-handling system in one week, with no labelled data available. What is the MOST appropriate first step? (Select one)

A. Build the full autonomous system and demo it. B. Assemble a small labelled dataset of past incidents and their correct dispositions, then scope a human-in-the-loop first version measured against it. C. Pick the largest model and enable every tool. D. Skip the metric and iterate on the prompt until the demo looks good.

Answer: B. With no data and a high-stakes decision, you must build the evaluation dataset and scope a reviewed first version. Building everything (A) or maximising tools (C) skips measurement, and prompt-tuning to a good-looking demo (D) has no metric behind it.

Q7 · A commodity need – transcribing support calls to text – has a mature hosted product available. Your team is small and speed-to-value matters. Which factor MOST justifies buying over building? (Select one)

A. Building would let you claim you built it. B. Transcription is a solved commodity; buying delivers value faster and lets your team focus on the differentiated part. C. Hosted products are always cheaper at every scale. D. Building avoids all data-flow considerations.

Answer: B. Buy commodity capability, build the differentiator. Bragging rights (A) are not a business reason, hosted is not always cheaper at scale (C), and buying introduces, not removes, data-flow considerations (D).

Q8 · Which requirements are ESSENTIAL to elicit during scoping because they change the architecture? (Select two)

A. The daily request volume and peak per minute. B. Whether inputs contain PII or regulated data. C. The founder’s favourite model. D. The colour of the dashboard. E. Whether the office uses Mac or Windows.

Answer: A and B. Volume/peak sizes rate limits, batch-vs-realtime and cost; data sensitivity drives retention, residency and review gates – both change the architecture. Model preference (C), UI colour (D) and OS (E) do not.

Q9 · A feature must respond while a user waits in a chat UI, and another must process a nightly backlog of 2M documents. What does correct scoping conclude? (Select one)

A. Both should use realtime streaming for consistency. B. The interactive feature needs low latency (streaming, possibly fast mode); the backlog is a good fit for Batch, which trades latency for lower cost. C. Both should use the Batch API. D. Both should use background mode with webhooks.

Answer: B. Latency requirements differ: interactive work needs streaming/low latency, while a large non-urgent backlog fits Batch’s cheaper, higher-latency processing. Forcing one mode on both (A, C) ignores the latency requirement; background mode (D) suits the backlog but not the waiting user.

Q10 · Which item belongs in the 'Success metric' line of a solution plan? (Select one)

A. The model ID you will use. B. A target value on a held-out dataset, e.g. ‘field-level exact match ≥ 95% on 300 labelled invoices’. C. The list of engineers on the project. D. The prompt template.

Answer: B. The success-metric line names a measurable target on a dataset. Model ID (A) is the approach line, staffing (C) is not a metric, and the prompt (D) is an implementation detail, not a success criterion.

Q11 · A director insists on `gpt-6-astra` for a high-volume extraction task that `gpt-5.6-luna` handles at 96% on your eval. What is the BEST response? (Select one)

A. Comply; the best model is always safest. B. Show the eval and the cost envelope: Luna meets the metric at roughly a fiftieth of the input price, so Astra adds cost without measurable benefit. C. Use Astra but at low effort to save money. D. Skip the eval and split traffic between both.

Answer: B. The scoping discipline is ‘the cheapest model that passes the eval’, backed by the cost math. Complying by default (A) over-buys, Astra-at-low (C) still costs far more per token than Luna, and splitting traffic without an eval (D) has no basis for the decision.

Q12 · During scoping, you discover the task's outputs feed an irreversible action (auto-cancelling shipments). How should the plan reflect this? (Select one)

A. Nothing changes; the model is accurate enough. B. Add a mandatory human review gate before the irreversible action and define a metric specifically for the decision that triggers it. C. Increase reasoning effort to max and skip review. D. Log the action after the fact for auditing only.

Answer: B. Irreversible actions require a review gate and a targeted metric on the triggering decision. Trusting accuracy (A) or raising effort (C) does not make an irreversible auto-action safe, and after-the-fact logging (D) does not prevent the harm.

Key takeaways

  • Scoping turns an outcome into a defined task with inputs, outputs and a measurable metric.
  • Define the success metric on a held-out dataset before choosing a model.
  • Elicit volume, latency, quality bar, data sensitivity, inputs and integration – they drive the architecture.
  • Estimate the cost envelope from tokens × price/MTok × volume and compare it to business value.
  • Build the differentiator, buy the commodity, and don’t force AI onto provable-exact deterministic work.
  • Use the SCOPE framework; finish Criteria before choosing the Path.
  • Ship the smallest first version that tests the metric, with review gates on high-stakes or irreversible steps.

Last updated Sep 18, 2026