# OpenAI API Cheat Sheet

Responses API request and response shapes in Python and TypeScript, conversation state, streaming, background mode, structured outputs, tools, reasoning effort, caching, batch, errors and limits.

import { Tabs, TabItem, Steps } from '@prosefly/astro-components';

The **Responses API** is the primary interface for new work. **Chat Completions is legacy** — it still runs, but the Responses API is where conversation state, background mode, reasoning items and the built-in tool surface live, so build against it. Everything here reflects the September 2026 surface described in the [OpenAI tracks](/openai/); re-verify shapes against [developers.openai.com/api/docs](https://developers.openai.com/api/docs).

## Minimal request

<Tabs>
<TabItem label="Python">
```python
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY

resp = client.responses.create(
    model="gpt-5.6-terra",
    input="Give three risks of vendor lock-in.",
    reasoning={"effort": "medium"},
)
print(resp.output_text)
```
</TabItem>
<TabItem label="TypeScript">
```typescript
import OpenAI from 'openai';

const client = new OpenAI(); // reads OPENAI_API_KEY

const resp = await client.responses.create({
  model: 'gpt-5.6-terra',
  input: 'Give three risks of vendor lock-in.',
  reasoning: { effort: 'medium' },
});
console.log(resp.output_text);
```
</TabItem>
</Tabs>

`input` accepts a string or an array of typed items (messages, tool outputs, files). `output_text` is a convenience aggregate; the authoritative content is in the `output` array of items.

## Request fields

| Field | Notes |
| --- | --- |
| `model` | Pinned ID: `gpt-6-astra`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna` |
| `input` | String or array of typed input items |
| `instructions` | System-level guidance for this response |
| `reasoning.effort` | `none`…`max` (5.6), `low`…`max` (Astra); lowest that works |
| `max_output_tokens` | Output cap; watch the `incomplete` status if hit |
| `tools` / `tool_choice` | Built-in tools and your functions |
| `text.format` | Structured output (JSON Schema) |
| `previous_response_id` | Server-side conversation state |
| `store` | Persist the response for later retrieval / state |
| `stream` | `true` for SSE |
| `background` | `true` to run asynchronously |
| `metadata` | Your key/value tags (e.g. hashed user id) |

## Response shape

```json
{
  "id": "resp_01",
  "object": "response",
  "model": "gpt-5.6-terra",
  "status": "completed",
  "output": [
    { "type": "reasoning", "id": "rs_01", "summary": [] },
    { "type": "message", "role": "assistant",
      "content": [{ "type": "output_text", "text": "The notice period is 60 days." }] }
  ],
  "usage": {
    "input_tokens": 1200,
    "input_tokens_details": { "cached_tokens": 1024 },
    "output_tokens": 42,
    "output_tokens_details": { "reasoning_tokens": 18 },
    "total_tokens": 1242
  }
}
```

### status — branch on it

| `status` | Meaning | Action |
| --- | --- | --- |
| `completed` | Finished normally | Read `output` |
| `incomplete` | Stopped early (e.g. `max_output_tokens`) | Inspect `incomplete_details`; continue or raise the cap |
| `in_progress` | Background/streamed, not done | Poll or keep streaming |
| `failed` | Errored | Inspect `error`; retry only if transient |

`output_tokens_details.reasoning_tokens` are billed as output — high effort spends real money here.

## Conversation state

Two ways to carry state; do not mix them for the same thread.

| Approach | How | Use when |
| --- | --- | --- |
| Server-side | Set `store: true`, then pass `previous_response_id` on the next call | You want OpenAI to hold the thread; less to send each turn |
| Client-side | Resend the full `input` array yourself | You need full control / your own store |

```python
first = client.responses.create(model="gpt-5.6-terra", input="My name is Dana.", store=True)
second = client.responses.create(
    model="gpt-5.6-terra",
    input="What is my name?",
    previous_response_id=first.id,
)
```

## Streaming

Server-sent events emit **semantic** events, not raw token deltas — branch on the event `type`.

<Tabs>
<TabItem label="Python">
```python
stream = client.responses.create(
    model="gpt-5.6-terra", input="Write a haiku about latency.", stream=True,
)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print("\n", event.response.usage)
```
</TabItem>
<TabItem label="TypeScript">
```typescript
const stream = await client.responses.create({
  model: 'gpt-5.6-terra', input: 'Write a haiku about latency.', stream: true,
});
for await (const event of stream) {
  if (event.type === 'response.output_text.delta') process.stdout.write(event.delta);
  else if (event.type === 'response.completed') console.log('\n', event.response.usage);
}
```
</TabItem>
</Tabs>

Common event types: `response.created`, `response.output_item.added`, `response.output_text.delta`, `response.function_call_arguments.delta`, `response.output_item.done`, `response.completed`, `response.error`. Tool-call arguments stream as fragments — buffer and parse only on `done`.

## Background mode

For long jobs, start with `background: true`, get an id back immediately, then poll or subscribe to a webhook.

```python
job = client.responses.create(model="gpt-6-astra", input=big_task, background=True)
# later
resp = client.responses.retrieve(job.id)
while resp.status in ("queued", "in_progress"):
    time.sleep(2)
    resp = client.responses.retrieve(job.id)
```

Background mode pairs with webhooks so you are not holding an open connection for minutes. It is the API-level analogue of the Agents API's durable sessions for one-shot long work.

## Structured outputs

```python
schema = {
  "type": "object",
  "properties": {
    "vendor": {"type": "string"},
    "total": {"type": "number"},
    "currency": {"type": "string", "enum": ["USD", "EUR", "GBP"]},
  },
  "required": ["vendor", "total", "currency"],
  "additionalProperties": False,
}
resp = client.responses.create(
    model="gpt-5.6-terra",
    input="Extract the invoice: Acme, 1240.50 USD.",
    text={"format": {"type": "json_schema", "name": "invoice", "schema": schema, "strict": True}},
)
```

`strict: true` constrains generation to the schema. Still validate downstream and retry with the specific error fed back — structured output guarantees shape, not business correctness.

## Function calling

```python
tools = [{
  "type": "function",
  "name": "get_order",
  "description": "Look up one order by ID. Returns status and ETA.",
  "parameters": {
    "type": "object",
    "properties": {"order_id": {"type": "string"}},
    "required": ["order_id"], "additionalProperties": False,
  },
  "strict": True,
}]

resp = client.responses.create(model="gpt-5.6-terra", input="Where is ORD-12345?", tools=tools)
# resp.output contains a function_call item; execute it, then send the output back:
followup = client.responses.create(
    model="gpt-5.6-terra",
    previous_response_id=resp.id,
    input=[{"type": "function_call_output", "call_id": call_id,
            "output": '{"status":"shipped","eta":"2026-09-17"}'}],
)
```

The loop: model emits a `function_call` item → you run it → you send a `function_call_output` item back (referencing `previous_response_id` or resending state) → repeat until a plain message.

## Reasoning effort

```python
client.responses.create(model="gpt-5.6-luna", input=simple_transform, reasoning={"effort": "none"})
client.responses.create(model="gpt-6-astra", input=hard_problem, reasoning={"effort": "xhigh"})
```

Use the **lowest effort that gets the result**. Reasoning tokens are billed as output. There is no exact GPT-5.5 → 5.6 effort mapping — re-tune per model. See the [model lineup](/appendix/openai/model-lineup/).

## File inputs

```python
f = client.files.create(file=open("contract.pdf", "rb"), purpose="user_data")
resp = client.responses.create(
    model="gpt-5.6-terra",
    input=[{"role": "user", "content": [
        {"type": "input_file", "file_id": f.id},
        {"type": "input_text", "text": "Summarise the termination clause."},
    ]}],
)
```

Images use `input_image` with a `file_id` or URL. Upload once and reference by id across many calls rather than re-uploading.

## Compaction and token counting

- **Compaction** summarises older turns server-side so a long session stays inside the context window while preserving the narrative. It is the API-side lever against unbounded context growth; the Agents API applies context summarisation automatically inside a session.
- **Token counting** — inspect `usage.input_tokens`, `usage.output_tokens`, `input_tokens_details.cached_tokens` (cache hits) and `output_tokens_details.reasoning_tokens` (billed reasoning) on every response to keep cost honest.

## Built-in tools

| Tool | One-line purpose |
| --- | --- |
| `web_search` | Answer from the live web with citations |
| `file_search` | Retrieve over your uploaded/indexed files |
| retrieval | Grounded answers over a managed store |
| MCP / connectors | Reach external systems via MCP servers |
| secure MCP tunnel | Reach private MCP servers without exposing them |
| `code_interpreter` | Run code in a sandbox for data/analysis |
| `image_generation` | Generate images inline |
| `computer_use` | Drive a computer/browser UI |
| shell / local shell | Execute shell commands (sandboxed / local) |
| apply patch | Apply code edits |
| tool search | Discover tools from a large catalogue |
| programmatic tool calling | Invoke tools from generated code |
| async tool calling | Long-running tools without blocking the turn |

## Quality and cost features

| Feature | What it does | Use it when |
| --- | --- | --- |
| **Prompt caching** | Reuses a stable prefix; cached input is discounted; cache diagnostics report hits | The same long system prompt/context repeats across calls |
| **Batch** | Offline processing at a discount, results within a window | Latency-tolerant, high-volume jobs |
| **Flex processing** | Lower-priced, best-effort latency tier | Non-urgent traffic that tolerates variable latency |
| **Fast mode** | Latency-optimised path | Interactive, latency-critical calls |
| **Predicted outputs** | Supply expected text to speed up edits | Regenerating a document with small changes |

Confirm cache hits via `usage.input_tokens_details.cached_tokens`; caching lowers cost more than it lowers rate-limit pressure.

## Error codes and retry policy

| HTTP | Type | Retry? |
| --- | --- | --- |
| 400 | `invalid_request_error` | No — fix the request |
| 401 | `authentication_error` | No — key/credentials |
| 403 | `permission_error` | No — entitlement/region |
| 404 | `not_found_error` | No — model/resource id |
| 409 | `conflict` | Sometimes — resolve state then retry |
| 422 | `unprocessable` | No — fix the payload |
| 429 | `rate_limit_error` | Yes — backoff, honour `retry-after` |
| 500 | `server_error` | Yes — backoff |
| 503 | `service_unavailable` | Yes — backoff, consider a fallback model |

Use exponential backoff with jitter, honour `retry-after`, log the request id from response headers, and pass an idempotency key on side-effecting requests so a retry does not double-act.

```python
import time, random
from openai import OpenAI, RateLimitError, APIStatusError

client = OpenAI()
RETRYABLE = {429, 500, 503}

def call_with_retry(**params):
    for attempt in range(6):
        try:
            return client.responses.create(**params)
        except RateLimitError as e:
            wait = float(e.response.headers.get("retry-after", 0)) or min(60, 2 ** attempt)
            time.sleep(wait + random.uniform(0, 0.5))
        except APIStatusError as e:
            if e.status_code in RETRYABLE:
                time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
            else:
                raise
    raise RuntimeError("exhausted retries")
```

## Rate and spend limits

- **Rate limits** bind on requests and tokens per minute; you hit whichever binds first. Watch the rate-limit response headers and throttle before a 429 rather than after.
- **Spend limits** cap cost per period at the org/project level; hitting them returns an error, not a silent stop.
- Manage both from the dashboard; enforce project-level budgets so one runaway job cannot exhaust the org.

:::tip[Assessment signal]
"Chat Completions" in a stem about *new* work is usually the distractor — the correct surface is the **Responses API**. "Long-running", "don't hold the connection", "come back later" points at **background mode**; "same system prompt every call" points at **prompt caching**; "overnight, cheap" points at **Batch**.
:::

## Key facts to memorise

- Responses API is primary; **Chat Completions is legacy** for new work.
- Conversation state: `store: true` + `previous_response_id` (server-side) or resend `input` (client-side) — not both.
- Streaming emits semantic events; branch on event `type`, buffer tool-arg fragments.
- Structured output guarantees shape (`strict: true`), not business correctness — still validate and retry.
- Retry only 429/5xx with backoff and `retry-after`; fix 4xx. Use idempotency keys on side-effecting calls.
- Cached input, Batch, Flex, Fast mode and predicted outputs are the cost/latency levers.
