OpenAI Batch API
OpenAI's async bulk-processing endpoint. Hand GPT thousands of requests at once as a single file, let them run within a 24-hour window, and collect the results — at roughly half the standard token cost. Learn the real SDK calls, how to match results back by custom_id, per-item error handling, and when to batch vs stream.
Learning objectives
- Explain what the Batch API is and the cost/latency trade it makes.
- Build a JSONL input file and submit it with
client.batches.create. - Poll a batch to completion and retrieve its output file.
- Match each result back to its request by
custom_id, handling per-item errors. - Decide between batch, realtime, and streaming for a workload.
OPENAI_API_KEY) from platform.openai.com + pip install openai. Keep the key in a .env file and let OpenAI() read it from the environment.1 · What the Batch API is essential
Not every request needs an answer in two seconds. Overnight evals, bulk classification of a backlog, embedding a whole corpus, generating thousands of product descriptions — this is work measured in throughput, not latency. Paying the premium realtime price (and babysitting rate limits) for work nobody is waiting on is a waste. The Batch API is OpenAI's answer: submit a big pile of requests as one job, let it process asynchronously within a window, and pay less for the privilege.
The deal is explicit. You give up speed — a batch completes within a 24-hour window, not immediately — and in exchange you get two things: roughly 50% off the per-token price, and a much higher throughput ceiling because the work doesn't count against your realtime rate limits the same way. For work that isn't time-sensitive, that's close to free money.
The common mistake is reaching for batch when a user is actually waiting (a chat reply, a live search) — the 24-hour window makes that a terrible experience — or, conversely, hammering the realtime endpoint in a for loop for a million-row offline job and getting rate-limited into oblivion. Match the tool to the deadline: a human waiting → realtime; a cron job → batch.
- The blue box is the realtime tier — fast, full price, for anything a user is waiting on.
- The green box is batch — up to 24 hours, about half the cost, for offline/bulk work.
- The trade is one-directional: you spend latency to buy cost and throughput. There's no tier that's both instant and cheap.
In short: batch is realtime's offline sibling — same model, cheaper, slower, for work nobody is watching.
2 · The batch lifecycle essential
A batch is a file-in, file-out job, and the lifecycle has four moves. First you build a JSONL file — one line per request, each line carrying a custom_id you choose plus the request body. You upload that file (client.files.create with purpose="batch"), then create the batch pointing at the uploaded file id. OpenAI processes it asynchronously; you poll its status until it's completed, then download the output file — another JSONL, one result line per request, each echoing the custom_id so you can line results back up with inputs.
That custom_id is the thread that ties the whole thing together. Results don't necessarily come back in order, so the id is how you say "this answer belongs to row 4,217 of my dataset." Choose it meaningfully (a database row id, a document hash) and matching results back is trivial.
custom_id), upload it, create the batch, poll to completion, then download the output JSONL and match results back by custom_id.
3 · Recipe 1 — submit a batch essential
Build the JSONL, upload it, and create the batch. Each line's body is exactly what you'd pass to responses.create, and url names the endpoint the batch runs against.
submit.pyimport json
from openai import OpenAI
client = OpenAI()
# 1. build a JSONL file — one request per line, each with a custom_id
rows = ["Summarize photosynthesis.", "Explain TCP in one line."]
with open("requests.jsonl", "w") as f:
for i, prompt in enumerate(rows):
f.write(json.dumps({
"custom_id": f"req-{i}", # YOUR id — how results map back
"method": "POST",
"url": "/v1/responses", # the endpoint this batch runs
"body": {"model": "gpt-5.5", "input": prompt},
}) + "\n")
# 2. upload the file
upload = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")
# 3. create the batch pointing at the uploaded file
batch = client.batches.create(
input_file_id=upload.id,
endpoint="/v1/responses",
completion_window="24h", # the only window today
)
print(batch.id, batch.status)
- Each JSONL line is a self-contained request: a
custom_idyou own, themethod/urlnaming the endpoint, and abodyidentical to a normalresponses.createcall. client.files.create(..., purpose="batch")uploads the JSONL and returns a file object; its.idis what the batch references.client.batches.createties it together — the input file, theendpoint(must match theurlin your lines), and thecompletion_window(currently"24h"). You get back a batch with anidand astatus.
Try this: point url and endpoint at /v1/embeddings with an embedding model to batch-embed a whole corpus overnight — same lifecycle, different endpoint.
body — instructions, input, max_output_tokens, tools, structured output. A batch is just a thousand of those requests in one file.4 · Recipe 2 — poll until it ends intermediate
A batch isn't instant, so you poll its status until it reaches a terminal state. The lifecycle runs validating → in_progress → finalizing → completed (or failed/expired/cancelled). Poll gently — this is a job that takes minutes to hours, so check every so often, not in a tight loop.
poll.pyimport time
from openai import OpenAI
client = OpenAI()
def wait_for(batch_id, every=30):
while True:
b = client.batches.retrieve(batch_id)
print(b.status, b.request_counts) # completed/total progress
if b.status in {"completed", "failed", "expired", "cancelled"}:
return b
time.sleep(every) # poll gently — this takes a while
5 · Recipe 3 — retrieve & match by custom_id intermediate
When the batch is completed it has an output_file_id. Download that file, parse it as JSONL, and use each line's custom_id to line results back up with your inputs.
collect.pyimport json
from openai import OpenAI
client = OpenAI()
def collect(batch):
content = client.files.content(batch.output_file_id).text
results = {}
for line in content.splitlines():
row = json.loads(line)
cid = row["custom_id"] # your id, echoed back
if row.get("error"):
results[cid] = ("ERROR", row["error"])
else:
body = row["response"]["body"] # the full Responses object
results[cid] = ("OK", body["output_text"])
return results
client.files.content(output_file_id).textdownloads the output JSONL as a string.- Each line has your
custom_idand either anerroror aresponse; keying a dict bycustom_idmaps every answer back to its input regardless of order. - The successful
response.bodyis a full Responses object — readoutput_textfrom it just like a live call.
Try this: there's also an error_file_id for batch-level failures — download it the same way to see which rows the API rejected before processing.
6 · Error handling per item advanced
A batch of 10,000 rarely comes back 10,000 clean. Some rows will error — a malformed body, a content-policy refusal, a token overflow — and the critical property of batch is that one bad row doesn't sink the job. Each line in the output carries its own success-or-error, so you handle failure per item, not per batch. The pattern: collect the OK results, bucket the errors by custom_id, and re-submit just the failures as a new, smaller batch.
len(results) == len(inputs). A silent gap — a row that never came back — is the batch bug that corrupts a dataset weeks later. The custom_id map makes the missing ones obvious.7 · Batch vs realtime vs streaming advanced
Three delivery modes, one decision rule — who, if anyone, is waiting? Streaming (ox2 §5) is for a human watching tokens appear right now. Realtime (a normal responses.create) is for a request a program needs answered in-line to continue. Batch is for work with no live consumer — the cheaper, higher-throughput tier for anything you can collect later.
| Mode | Latency | Use for |
|---|---|---|
| Streaming | tokens as they generate | chat UIs, live assistants |
| Realtime | seconds | in-line program steps, user-facing calls |
| Batch | <24 hours, ~50% cost | evals, bulk classification, corpus embedding |
8 · Cost math (runs offline) professional
The batch discount is only worth claiming if you can quantify it. This estimator is pure Python — no key needed — and makes the saving a number you can put in a design doc.
batch_savings.py# illustrative rates $/1M tokens — check current pricing
RT_IN, RT_OUT = 5.0, 15.0
BATCH_DISCOUNT = 0.5 # ~50% off
def cost(reqs, in_tok, out_tok, batch=False):
c = reqs * (in_tok*RT_IN + out_tok*RT_OUT) / 1_000_000
return c * (BATCH_DISCOUNT if batch else 1)
n, i, o = 100_000, 600, 300
print(f"realtime ${cost(n,i,o):,.0f} batch ${cost(n,i,o,batch=True):,.0f}")
9 · Tech-lead — production batch pipelines tech-lead
A production batch pipeline is a data pipeline that happens to call an LLM. The LLM call is the easy part; the engineering is everything around it: chunk a large dataset into batches under the file-size and request-count limits, persist each batch id and its custom_id→row mapping to a database, poll from a scheduled job (not a blocking process), reconcile counts on completion, and auto-resubmit the error rows. Treat the whole thing as idempotent and resumable — a batch that fails at hour 20 shouldn't mean re-running the first 19.
The capstone for this track (the batch document pipeline project) builds exactly this end to end.
🪜 Practice ladder beginner → industry
- Beginner: write a 2-line JSONL by hand and submit it with Recipe 1.
- Easy: poll your batch with Recipe 2 and watch
request_countsclimb. - Core: collect results with Recipe 3 and print each answer next to its
custom_id. - Stretch: deliberately include one malformed row and handle its per-item error without crashing.
- Hard: re-submit only the failed rows as a second batch and merge the results.
- Industry: persist batch ids + the id→row map to SQLite and resume collection in a fresh process.
✓ Checkpoint — you can move on when you can…
- Explain the batch cost/latency trade and when to use it.
- Build a JSONL, upload it, and create a batch.
- Poll to completion and collect results matched by
custom_id. - Handle per-item errors and reconcile counts.
Knowledge check check yourself
What does the Batch API trade away, what do you get in return, and what single field ties inputs to outputs?
Show answer
custom_id you set on each JSONL input line is echoed on each output line, so you map results back to inputs regardless of order.Walk the batch lifecycle from a list of prompts to collected results, naming the SDK calls.
Show answer
custom_id, method, url, and a body); upload it with client.files.create(..., purpose="batch"); submit with client.batches.create(input_file_id=…, endpoint="/v1/responses", completion_window="24h"); poll client.batches.retrieve(id) until status == "completed"; then download client.files.content(batch.output_file_id) and match each line by custom_id.