API Developer Path
D7 · Production Safety and Operations
Safety classifiers and moderation, red teaming, misalignment monitoring, content provenance, rate and spend limits, error codes and retries, RBAC, Private Link, IP allowlist, mTLS and workload identity federation.
This domain is about 8% of the OAI-API mock – roughly 5 of 60 items – and draws on the operations and safety material threaded through the Academy API pathway. It tests whether you can take a working application to production safely: moderating content, red-teaming and monitoring for misalignment, marking provenance, setting rate and spend limits, handling errors with retries, and locking down access with RBAC and network controls.
What you need to know
Production is where an application meets untrusted inputs, real money and real users. You screen content with moderation and safety classifiers, probe weaknesses with red teaming, and watch for misalignment in live behaviour. You mark AI-generated media with content provenance. You cap blast radius with rate limits and spend limits, and you handle transient failures with retries and exponential backoff keyed off the right error codes. You restrict who and what can reach the API with RBAC, Private Link, IP allowlists, mutual TLS and workload identity federation – so credentials are scoped, short-lived and network-bounded. A deployment checklist ties these together as launch gates.
Learning objectives
By the end of this page you should be able to:
- Apply moderation and safety classifiers to inputs and outputs.
- Explain red teaming and misalignment monitoring.
- Set rate limits and spend limits to bound blast radius.
- Handle error codes with correct retry and backoff behaviour.
- Choose access controls: RBAC, Private Link, IP allowlist, mTLS, workload identity federation.
- Run a deployment checklist before launch.
7.1 Safety: moderation, classifiers, red teaming, misalignment
| Control | What it does | When |
|---|---|---|
| Moderation | Flags disallowed content in inputs/outputs | Screen user input and model output at runtime |
| Safety classifiers | Detect specific risk categories (incl. cyber-safety checks) | Gate high-risk flows |
| Red teaming | Deliberate adversarial testing before and after launch | Find jailbreaks, injection, data-exfiltration paths |
| Misalignment monitoring | Watch live behaviour for drift from intended goals | Ongoing, especially for agents with tools |
Red teaming is proactive (you attack your own system); misalignment monitoring is continuous (you watch what it actually does). Content provenance (marking AI-generated media) supports transparency obligations.
Assessment signal
“Untrusted user input”, “jailbreak”, “prompt injection”, “before launch we should test adversarially” point to red teaming and guardrails/moderation. “Watch the deployed agent for drift” points to misalignment monitoring.
7.2 Rate limits and spend limits
Two different blast-radius controls:
| Limit | Bounds | Failure it prevents |
|---|---|---|
| Rate limit | Requests/tokens per minute | A bug or spike overwhelming the service; noisy-neighbour |
| Spend limit | Dollars per period | A runaway loop or abuse draining the budget |
Runaway agent loops calling the API 1000×/min │ rate limit → 429 caps the request rate spend limit → hard cap on dollars, alerting before the ceilingSet both before launch. A spend limit is the difference between a bug costing $50 and costing $50,000.
7.3 Error codes and retries
| Code | Meaning | Correct handling |
|---|---|---|
429 | Rate limit / quota exceeded | Retry with exponential backoff (respect Retry-After) |
500 / 503 | Server / service error | Retry with backoff; it is transient |
400 | Bad request (malformed) | Do not retry; fix the request |
401 / 403 | Auth / permission | Do not retry; fix credentials/scope |
import timefor attempt in range(5): try: return client.responses.create(model="gpt-5.6-terra", input=q) except RateLimitError: # 429 time.sleep(2 ** attempt) # 1, 2, 4, 8, 16s except BadRequestError: # 400 — do not retry raiseThe discriminator on many items: retry transient errors (429/5xx) with backoff; never blindly retry a 400/401/403 – retrying a malformed or unauthorised request just wastes quota and can worsen rate limiting.
7.4 Access control and network security
| Control | What it does | Use when |
|---|---|---|
| RBAC / permissions | Scope who can do what (keys, roles, projects) | Always; least privilege on keys and roles |
| Private Link | Private network path to the API, off the public internet | Sensitive workloads needing network isolation |
| IP allowlist | Only listed source IPs may call the API | Fixed egress infrastructure |
| mutual TLS (mTLS) | Both client and server authenticate with certs | Strong mutual authentication requirements |
| Workload identity federation | Short-lived federated credentials (K8s, AWS, Azure, GCP, OCI, GitHub Actions, SPIFFE, X.509) instead of long-lived keys | Cloud workloads that should not hold static secrets |
Prefer federated, short-lived credentials
Long-lived API keys in environment variables are a standing liability. Workload identity federation lets a cloud workload exchange its own identity for a short-lived token, so there is no static key to leak. Prefer it wherever your platform supports it.
7.5 The deployment checklist
Before launch, confirm:[ ] Moderation on untrusted inputs and outputs[ ] Red-team pass for injection / jailbreak / exfiltration[ ] Guardrails + human approval on irreversible actions[ ] Rate limit set; spend limit set with alerting below the ceiling[ ] Retry with exponential backoff on 429/5xx; no retry on 4xx auth/validation[ ] RBAC least privilege on keys and roles[ ] Network controls as required (Private Link / IP allowlist / mTLS)[ ] Federated, short-lived credentials instead of static keys where possible[ ] Data residency / retention verified against requirements (see D4)[ ] Eval gate green (see D3); tracing/observability on (see D4)[ ] Misalignment monitoring for deployed agentsDecision framework
Use GUARD to decide the controls a production launch needs.
| Letter | Step | Question |
|---|---|---|
| G | Gate content | Is moderation on inputs and outputs, and have you red-teamed? |
| U | Usage caps | Are rate and spend limits set with alerting? |
| A | Auth & access | Is RBAC least-privilege, with the right network controls? |
| R | Resilience | Do you retry transient errors with backoff and not retry auth/validation errors? |
| D | Detect drift | Is misalignment monitoring and tracing in place for agents? |
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| No spend limit | “It won’t loop” | Set a spend limit with alerting; it caps runaway cost |
Retrying a 400/401 | Generic retry-everything wrapper | Retry only transient 429/5xx with backoff; fix 4xx |
| Long-lived keys in env vars | Easiest to set up | Use workload identity federation for short-lived tokens |
| Moderating input but not output | Focus on user text | Screen model outputs too, especially agent actions |
| Skipping red teaming | “We tested the happy path” | Adversarially test injection/jailbreak before launch |
| Over-broad API key scope | Convenience | RBAC least privilege per key/role/project |
| No misalignment monitoring for agents | Launch-and-forget | Continuously watch deployed agent behaviour |
| Trusting the network by default | Public internet is fine | Add Private Link / IP allowlist / mTLS as risk requires |
Scenario challenge
Scenario. You are launching an autonomous support agent that reads customer messages (untrusted input), calls internal tools to issue refunds, and runs on Kubernetes. Security asks: how do we stop it draining budget, stop a malicious message hijacking it, and avoid static API keys in the cluster? The team’s current plan wraps every API call in a “retry up to 5 times on any error” loop and stores a long-lived API key in a Kubernetes secret.
Expert reasoning trace.
- Budget blast radius → spend limit + rate limit. An agent that loops or is abused could call the API thousands of times. Set a spend limit with alerting below the ceiling and a rate limit, so a runaway loop hits a hard dollar cap and a
429rather than an unbounded bill. - Malicious message → red team + moderation + guardrails + approval. Customer messages are untrusted, so a prompt-injection payload could try to make the agent issue an unauthorised refund. Red-team for injection before launch, moderate inputs, add output/tool guardrails, and gate the refund action behind human approval (D4) since it is irreversible.
- Retry loop is wrong. “Retry any error 5×” will happily retry a
400(malformed) and a401/403(auth/permission), wasting quota and masking real bugs. Correct: retry only429/5xxwith exponential backoff, and fail fast on4xxauth/validation. - Static key is a liability. A long-lived key in a K8s secret can leak. Use workload identity federation so the pod exchanges its Kubernetes identity for a short-lived token – no static key to steal.
- Network controls. If the workload’s egress is fixed, add an IP allowlist, and use Private Link if the traffic must stay off the public internet; RBAC scopes the credential to only the refund/support tools.
- Monitor after launch. Turn on misalignment monitoring and tracing so a drifting or hijacked agent is caught in production.
Exam-correct decision: spend + rate limits, red teaming and moderation with an approval gate on refunds, selective retry (429/5xx with backoff only), workload identity federation instead of a static key, RBAC and network controls, and misalignment monitoring. Not retry-everything, not a long-lived key in a secret, not an auto-refund with no gate.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| “Retry every error a few times” | One wrapper for all failures | Retry only transient 429/5xx with backoff; never auth/validation 4xx |
| “A rate limit is enough to control cost” | It limits requests | A rate limit bounds throughput; a spend limit bounds dollars – set both |
| “Store the API key in an env var / secret” | Simple | Use workload identity federation for short-lived credentials |
| “Moderate the user input” | User text feels like the risk | Screen outputs and tool actions too, not just inputs |
| “We tested it works, we’re ready” | Happy path passed | Red-team injection/jailbreak/exfiltration before launch |
| “The agent is smart, no approval needed for refunds” | Autonomy bias | Irreversible actions need a human approval gate |
| “Public internet is fine for sensitive traffic” | Default setup | Add Private Link / IP allowlist / mTLS to the risk level |
Practice questions
Each item states how many responses to select. Commit before revealing.
Q1 · An API call returns a `429`. What is the correct handling? (Select one)
A. Retry immediately in a tight loop.
B. Retry with exponential backoff, respecting any Retry-After header.
C. Do not retry; fix the request.
D. Escalate to a larger model.
Answer: B. A 429 is a transient rate-limit error, handled by backing off exponentially and honouring Retry-After. A tight retry loop (A) worsens the limiting, a 429 is not a malformed request to fix (C), and model size (D) is unrelated.
Q2 · Which error should you NOT automatically retry? (Select one)
A. 429 rate limit.
B. 503 service unavailable.
C. 400 bad request (malformed input).
D. 500 server error.
Answer: C. A 400 means the request itself is malformed; retrying it will fail again and waste quota – fix the request instead. 429 (A), 503 (B) and 500 (D) are transient and appropriate to retry with backoff.
Q3 · A team fears a runaway agent loop could generate a huge bill. Which control MOST directly caps the dollar cost? (Select one)
A. A rate limit only. B. A spend limit with alerting below the ceiling. C. A larger context window. D. Streaming responses.
Answer: B. A spend limit caps dollars directly and can alert before the ceiling. A rate limit (A) bounds throughput, not total spend, and context size (C) and streaming (D) do not control cost.
Q4 · A Kubernetes workload should call the API without holding a long-lived key. What is the BEST approach? (Select one)
A. Store a long-lived key in a Kubernetes secret. B. Use workload identity federation to exchange the pod’s identity for a short-lived token. C. Hard-code the key in the container image. D. Pass the key as a command-line argument.
Answer: B. Workload identity federation issues short-lived, federated credentials so there is no static key to leak. A secret (A), a baked-in key (C) and a CLI argument (D) are all long-lived-key anti-patterns.
Q5 · An agent processes untrusted customer messages and can take actions. Which TWO controls address prompt-injection risk BEST? (Select two)
A. Input/output guardrails and moderation to detect and block injected instructions. B. Red teaming before launch to find injection and exfiltration paths. C. A larger context window. D. Higher reasoning effort. E. Removing all logging.
Answer: A and B. Guardrails/moderation block injected instructions at runtime and red teaming finds the paths before launch. Context size (C) and effort (D) do not stop injection, and removing logging (E) reduces the visibility you need.
Q6 · What is the difference between red teaming and misalignment monitoring? (Select one)
A. They are the same activity. B. Red teaming is proactive adversarial testing; misalignment monitoring is continuous observation of deployed behaviour for drift. C. Red teaming only applies to images. D. Misalignment monitoring happens only before launch.
Answer: B. Red teaming attacks the system to find weaknesses; misalignment monitoring watches live behaviour over time. They are distinct (A), red teaming is not image-only (C), and monitoring is ongoing after launch, not pre-launch only (D).
Q7 · A sensitive workload must reach the API without traversing the public internet. Which control fits? (Select one)
A. A spend limit. B. Private Link, providing a private network path to the API. C. A larger model. D. Predicted outputs.
Answer: B. Private Link gives a private network path off the public internet for sensitive traffic. A spend limit (A) is a cost control, and model size (C) and predicted outputs (D) are unrelated to network isolation.
Q8 · An API key used by a read-only reporting job also has permission to delete resources. What principle is violated and what is the fix? (Select one)
A. None; broad keys are convenient. B. Least privilege; scope the key via RBAC to only the read permissions the job needs. C. Statelessness; store nothing. D. Caching; enable it.
Answer: B. An over-scoped key violates least privilege and enlarges blast radius; RBAC should scope it to read-only. Convenience (A) is not a justification, and statelessness (C) and caching (D) are unrelated.
Q9 · Which items belong on a pre-launch deployment checklist for an agent that takes irreversible actions? (Select two)
A. A human approval gate on the irreversible actions. B. Rate and spend limits with alerting. C. Retrying every error type indefinitely. D. Disabling tracing to reduce noise. E. Storing static keys in the repo.
Answer: A and B. Irreversible actions need an approval gate, and rate/spend limits cap blast radius – both are launch gates. Retrying everything indefinitely (C) is wrong, disabling tracing (D) removes needed observability, and static keys in the repo (E) is a serious anti-pattern.
Q10 · A deployed support agent gradually starts taking actions outside its intended scope. Which capability is designed to catch this? (Select one)
A. Prompt caching. B. Misalignment monitoring of live agent behaviour, with tracing to investigate. C. Batch processing. D. A larger context window.
Answer: B. Misalignment monitoring watches deployed behaviour for drift from intended goals, and tracing lets you investigate. Caching (A) and Batch (C) are performance features, and context size (D) does not detect behavioural drift.
Key takeaways
- Screen untrusted inputs and model outputs with moderation and safety classifiers; red-team before launch.
- Set both a rate limit (throughput) and a spend limit (dollars) with alerting to bound blast radius.
- Retry only transient errors (
429,5xx) with exponential backoff; never blindly retry400/401/403. - Scope credentials with RBAC least privilege and add Private Link, IP allowlist or mTLS to the risk level.
- Prefer workload identity federation and short-lived tokens over long-lived static keys.
- Gate irreversible agent actions behind human approval and monitor deployed agents for misalignment.
- Treat the deployment checklist as launch gates: safety, limits, resilience, access, residency, evals and observability.
Last updated Sep 18, 2026