AI EngineeringZero to ProductionHome·About·What’s new·Contact
OpenAI API in Practice · Part 1

OpenAI Batch API

OpenAI's async bulk-processing endpoint. Hand GPT thousands of requests at once as a single file, let them run within a 24-hour window, and collect the results — at roughly half the standard token cost. Learn the real SDK calls, how to match results back by custom_id, per-item error handling, and when to batch vs stream.

⏱️ ~1.5 hours🧪 5 recipes🎯 Beginner→Tech-lead

Learning objectives

  • Explain what the Batch API is and the cost/latency trade it makes.
  • Build a JSONL input file and submit it with client.batches.create.
  • Poll a batch to completion and retrieve its output file.
  • Match each result back to its request by custom_id, handling per-item errors.
  • Decide between batch, realtime, and streaming for a workload.
⚙️ To run this for realThe code here is complete and correct as written. To execute it you'll need an OpenAI API key (OPENAI_API_KEY) from platform.openai.com + pip install openai. Keep the key in a .env file and let OpenAI() read it from the environment.

1 · What the Batch API is essential

Not every request needs an answer in two seconds. Overnight evals, bulk classification of a backlog, embedding a whole corpus, generating thousands of product descriptions — this is work measured in throughput, not latency. Paying the premium realtime price (and babysitting rate limits) for work nobody is waiting on is a waste. The Batch API is OpenAI's answer: submit a big pile of requests as one job, let it process asynchronously within a window, and pay less for the privilege.

The deal is explicit. You give up speed — a batch completes within a 24-hour window, not immediately — and in exchange you get two things: roughly 50% off the per-token price, and a much higher throughput ceiling because the work doesn't count against your realtime rate limits the same way. For work that isn't time-sensitive, that's close to free money.

The common mistake is reaching for batch when a user is actually waiting (a chat reply, a live search) — the 24-hour window makes that a terrible experience — or, conversely, hammering the realtime endpoint in a for loop for a million-row offline job and getting rate-limited into oblivion. Match the tool to the deadline: a human waiting → realtime; a cron job → batch.

same model, two speed/price tiers — pick by the deadline realtime seconds · full price a human is waiting batch <24h · ~50% off offline / bulk work give up latency → get cost + throughput Latency for cost. Realtime answers in seconds at full price; batch answers within 24 hours at about half price and far higher throughput. The only question is whether anyone is waiting on the result.
🗺️ How to read this diagram
  • The blue box is the realtime tier — fast, full price, for anything a user is waiting on.
  • The green box is batch — up to 24 hours, about half the cost, for offline/bulk work.
  • The trade is one-directional: you spend latency to buy cost and throughput. There's no tier that's both instant and cheap.

In short: batch is realtime's offline sibling — same model, cheaper, slower, for work nobody is watching.

2 · The batch lifecycle essential

A batch is a file-in, file-out job, and the lifecycle has four moves. First you build a JSONL file — one line per request, each line carrying a custom_id you choose plus the request body. You upload that file (client.files.create with purpose="batch"), then create the batch pointing at the uploaded file id. OpenAI processes it asynchronously; you poll its status until it's completed, then download the output file — another JSONL, one result line per request, each echoing the custom_id so you can line results back up with inputs.

That custom_id is the thread that ties the whole thing together. Results don't necessarily come back in order, so the id is how you say "this answer belongs to row 4,217 of my dataset." Choose it meaningfully (a database row id, a document hash) and matching results back is trivial.

file in → async processing → file out, matched by custom_id build JSONL1 line/request files.createpurpose=batch batches.create+ poll status output JSONLmatch by custom_id statuses: validating → in_progress → finalizing → completed File in, file out. Build a JSONL of requests (each with a custom_id), upload it, create the batch, poll to completion, then download the output JSONL and match results back by custom_id.

3 · Recipe 1 — submit a batch essential

Build the JSONL, upload it, and create the batch. Each line's body is exactly what you'd pass to responses.create, and url names the endpoint the batch runs against.

Recipe 1
submit.pyimport json
from openai import OpenAI
client = OpenAI()

# 1. build a JSONL file — one request per line, each with a custom_id
rows = ["Summarize photosynthesis.", "Explain TCP in one line."]
with open("requests.jsonl", "w") as f:
    for i, prompt in enumerate(rows):
        f.write(json.dumps({
            "custom_id": f"req-{i}",          # YOUR id — how results map back
            "method": "POST",
            "url": "/v1/responses",          # the endpoint this batch runs
            "body": {"model": "gpt-5.5", "input": prompt},
        }) + "\n")

# 2. upload the file
upload = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")

# 3. create the batch pointing at the uploaded file
batch = client.batches.create(
    input_file_id=upload.id,
    endpoint="/v1/responses",
    completion_window="24h",            # the only window today
)
print(batch.id, batch.status)
▶ How this works
  1. Each JSONL line is a self-contained request: a custom_id you own, the method/url naming the endpoint, and a body identical to a normal responses.create call.
  2. client.files.create(..., purpose="batch") uploads the JSONL and returns a file object; its .id is what the batch references.
  3. client.batches.create ties it together — the input file, the endpoint (must match the url in your lines), and the completion_window (currently "24h"). You get back a batch with an id and a status.

Try this: point url and endpoint at /v1/embeddings with an embedding model to batch-embed a whole corpus overnight — same lifecycle, different endpoint.

The body is just a normal requestEverything you know from the OpenAI API chapter applies inside body — instructions, input, max_output_tokens, tools, structured output. A batch is just a thousand of those requests in one file.

4 · Recipe 2 — poll until it ends intermediate

A batch isn't instant, so you poll its status until it reaches a terminal state. The lifecycle runs validating → in_progress → finalizing → completed (or failed/expired/cancelled). Poll gently — this is a job that takes minutes to hours, so check every so often, not in a tight loop.

Recipe 2
poll.pyimport time
from openai import OpenAI
client = OpenAI()

def wait_for(batch_id, every=30):
    while True:
        b = client.batches.retrieve(batch_id)
        print(b.status, b.request_counts)        # completed/total progress
        if b.status in {"completed", "failed", "expired", "cancelled"}:
            return b
        time.sleep(every)                         # poll gently — this takes a while
Don't busy-waitA batch can take hours. Polling every few seconds just wastes calls and may hit rate limits. Poll on the order of tens of seconds to minutes, or trigger the retrieval from a scheduled job rather than a blocking loop.

5 · Recipe 3 — retrieve & match by custom_id intermediate

When the batch is completed it has an output_file_id. Download that file, parse it as JSONL, and use each line's custom_id to line results back up with your inputs.

Recipe 3
collect.pyimport json
from openai import OpenAI
client = OpenAI()

def collect(batch):
    content = client.files.content(batch.output_file_id).text
    results = {}
    for line in content.splitlines():
        row = json.loads(line)
        cid = row["custom_id"]                 # your id, echoed back
        if row.get("error"):
            results[cid] = ("ERROR", row["error"])
        else:
            body = row["response"]["body"]       # the full Responses object
            results[cid] = ("OK", body["output_text"])
    return results
▶ How this works
  1. client.files.content(output_file_id).text downloads the output JSONL as a string.
  2. Each line has your custom_id and either an error or a response; keying a dict by custom_id maps every answer back to its input regardless of order.
  3. The successful response.body is a full Responses object — read output_text from it just like a live call.

Try this: there's also an error_file_id for batch-level failures — download it the same way to see which rows the API rejected before processing.

6 · Error handling per item advanced

A batch of 10,000 rarely comes back 10,000 clean. Some rows will error — a malformed body, a content-policy refusal, a token overflow — and the critical property of batch is that one bad row doesn't sink the job. Each line in the output carries its own success-or-error, so you handle failure per item, not per batch. The pattern: collect the OK results, bucket the errors by custom_id, and re-submit just the failures as a new, smaller batch.

Always reconcile countsAfter collection, assert that len(results) == len(inputs). A silent gap — a row that never came back — is the batch bug that corrupts a dataset weeks later. The custom_id map makes the missing ones obvious.

7 · Batch vs realtime vs streaming advanced

Three delivery modes, one decision rule — who, if anyone, is waiting? Streaming (ox2 §5) is for a human watching tokens appear right now. Realtime (a normal responses.create) is for a request a program needs answered in-line to continue. Batch is for work with no live consumer — the cheaper, higher-throughput tier for anything you can collect later.

ModeLatencyUse for
Streamingtokens as they generatechat UIs, live assistants
Realtimesecondsin-line program steps, user-facing calls
Batch<24 hours, ~50% costevals, bulk classification, corpus embedding

8 · Cost math (runs offline) professional

The batch discount is only worth claiming if you can quantify it. This estimator is pure Python — no key needed — and makes the saving a number you can put in a design doc.

Offline
batch_savings.py# illustrative rates $/1M tokens — check current pricing
RT_IN, RT_OUT = 5.0, 15.0
BATCH_DISCOUNT = 0.5                       # ~50% off

def cost(reqs, in_tok, out_tok, batch=False):
    c = reqs * (in_tok*RT_IN + out_tok*RT_OUT) / 1_000_000
    return c * (BATCH_DISCOUNT if batch else 1)

n, i, o = 100_000, 600, 300
print(f"realtime ${cost(n,i,o):,.0f}  batch ${cost(n,i,o,batch=True):,.0f}")

9 · Tech-lead — production batch pipelines tech-lead

A production batch pipeline is a data pipeline that happens to call an LLM. The LLM call is the easy part; the engineering is everything around it: chunk a large dataset into batches under the file-size and request-count limits, persist each batch id and its custom_id→row mapping to a database, poll from a scheduled job (not a blocking process), reconcile counts on completion, and auto-resubmit the error rows. Treat the whole thing as idempotent and resumable — a batch that fails at hour 20 shouldn't mean re-running the first 19.

The capstone for this track (the batch document pipeline project) builds exactly this end to end.

🪜 Practice ladder beginner → industry

  1. Beginner: write a 2-line JSONL by hand and submit it with Recipe 1.
  2. Easy: poll your batch with Recipe 2 and watch request_counts climb.
  3. Core: collect results with Recipe 3 and print each answer next to its custom_id.
  4. Stretch: deliberately include one malformed row and handle its per-item error without crashing.
  5. Hard: re-submit only the failed rows as a second batch and merge the results.
  6. Industry: persist batch ids + the id→row map to SQLite and resume collection in a fresh process.

✓ Checkpoint — you can move on when you can…

  • Explain the batch cost/latency trade and when to use it.
  • Build a JSONL, upload it, and create a batch.
  • Poll to completion and collect results matched by custom_id.
  • Handle per-item errors and reconcile counts.

Knowledge check check yourself

✓ Knowledge check

What does the Batch API trade away, what do you get in return, and what single field ties inputs to outputs?

Show answer
You trade latency — a batch completes within a 24-hour window rather than instantly — for roughly 50% lower token cost and higher throughput. The custom_id you set on each JSONL input line is echoed on each output line, so you map results back to inputs regardless of order.
✓ Knowledge check

Walk the batch lifecycle from a list of prompts to collected results, naming the SDK calls.

Show answer
Build a JSONL file (one line per request with a custom_id, method, url, and a body); upload it with client.files.create(..., purpose="batch"); submit with client.batches.create(input_file_id=…, endpoint="/v1/responses", completion_window="24h"); poll client.batches.retrieve(id) until status == "completed"; then download client.files.content(batch.output_file_id) and match each line by custom_id.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in