AI EngineeringZero to ProductionHome·About·What’s new·Contact
Codex & OpenAI · Chapter O2

The OpenAI API — Responses, Structured Output & Tools

One API call is the atom every larger system is built from. This chapter is the OpenAI counterpart to the Anthropic API lesson: the single responses.create call, how to get machine-parseable output, how tools (function calling) work, and how to stream — all with the real openai SDK.

⏱️ ~2 hours🧪 5 labs🎯 Beginner→Advanced

Learning objectives

  • Make a single request with the Responses API and read the reply and token usage.
  • Separate the system role (instructions) from the user turn (input).
  • Get schema-valid structured output with responses.parse and Pydantic.
  • Wire tools (function calling) and run the function-call loop.
  • Stream tokens, and handle errors, retries, and cost like a service.
⚙️ To run this for realThe code here is complete and correct as written. To execute it you'll need an OpenAI API key (OPENAI_API_KEY) from platform.openai.com + pip install openai. Keep the key in a .env file (never in code) and let the zero-argument OpenAI() client pick it up from the environment. Reading and learning works without any of this — run when you're ready.

1 · The single call essential

Everything OpenAI does for you is one function call, repeated. Agents, RAG, chatbots — strip away the orchestration and each is a loop around the same primitive: send some text, get some text back. Learn that primitive cold and the rest of the course stops feeling like magic and starts feeling like plumbing you control.

OpenAI gives you two ways to make that call, and picking the right one matters. The Responses API (client.responses.create) is the primary, current interface — it's what you should learn and build on. The older Chat Completions API (client.chat.completions.create, which returns resp.choices[0].message.content) still works and is supported indefinitely, so you'll see it in older tutorials — but it's the previous standard, and this course teaches Responses. The two are shaped differently enough that mixing them up is a common early bug.

The Responses call takes three things you care about: model (which GPT), instructions (the system role — who the assistant is and its rules), and input (the user's content). It hands back a response object whose output_text property is the reply as a plain string. That's the whole atom.

OpenAI() reads the key responses.create() model · instructions · input resp .output_text three top boxes are the whole program — larger systems just repeat this Build client · call · read. That is the entire API — responses.create with model/instructions/input, then resp.output_text. Everything else in the course is this primitive in a loop.
🗺️ How to read this diagram
  • The first box builds the client — OpenAI() reads your key from the environment.
  • The middle box is the one call: responses.create() with the three arguments you care about.
  • The green box is the reply; resp.output_text is the convenience property that gives you the text as a string, no block-list to unpack.

In short: client → call → output_text. Memorize this shape; it never changes.

Lab O2.1
first_call.pyfrom dotenv import load_dotenv
from openai import OpenAI
load_dotenv()
client = OpenAI()                      # reads OPENAI_API_KEY

resp = client.responses.create(
    model="gpt-5.5",
    instructions="You are a concise assistant.",
    input="What is an API token? Answer in two sentences.",
)
print(resp.output_text)
print(resp.usage.input_tokens, resp.usage.output_tokens)
Same shape as Claude, different namesOn Claude this is client.messages.create(system=…, messages=[…]) and you pull text out of a resp.content block list. On OpenAI it's client.responses.create(instructions=…, input=…) and resp.output_text. Same idea, different field names — which is exactly what the Claude | OpenAI tabs across this course let you compare.

2 · System vs user — instructions and input essential

The biggest early leverage isn't a cleverer prompt — it's putting the right text in the right slot. The Responses API separates two channels: instructions is the system/developer role — stable rules about who the assistant is, what format to use, what never to do — and input is the user turn, the actual question or task. Mixing them (stuffing your rules into the user message) works until it doesn't: the model starts treating your rules as negotiable content rather than standing orders.

Why the split matters, concretely: the instructions are where you encode the parts that don't change between requests — tone, output format, refusal rules, the persona. The input is the volatile part that changes every call. Keeping the stable rules in instructions makes behavior consistent and is also what lets the platform cache the stable prefix for you (cheaper repeated calls).

The common mistake: re-sending the same long rules inside every user message. Put them in instructions once, keep input lean, and both your consistency and your bill improve.

Multi-turn: resend the historyThe model is stateless — it remembers nothing between calls. To hold a conversation you pass input as a list of role/content turns and append each reply before the next call. "Memory" is just you resending the transcript; there's no hidden server-side state in a basic call.

3 · Structured output with Pydantic expert

Picture the 2am page. Your classifier has hummed along for weeks, its output scraped by a regex that plucks a priority out of the model's reply. Then one night the model, entirely reasonably, answers "Sure! The priority here is high 😊" instead of the exact phrasing your regex expected — the parse returns nothing, the ticket routes nowhere, and your phone lights up. Nothing broke on your side; the model just rephrased. That fragility is exactly what structured output eliminates.

The idea is a shift in who guarantees the shape of the data. Instead of hoping the model phrases things the way your parser expects, you declare a Pydantic schema up front and the shape becomes a constraint the model must satisfy while generating. With OpenAI you use client.responses.parse(..., text_format=YourModel) and read a validated instance from resp.output_parsed — a real typed object, no parsing step to break.

Lab O2.2
extract.pyfrom openai import OpenAI
from pydantic import BaseModel
from typing import Literal
client = OpenAI()

class Ticket(BaseModel):
    category: Literal["bug","feature","billing","other"]
    priority: Literal["low","medium","high"]
    summary: str
    needs_human: bool

resp = client.responses.parse(          # .parse() validates for you
    model="gpt-5.5", max_output_tokens=512,
    input="I was charged twice and support hasn't replied in 3 days!",
    text_format=Ticket,
)

t = resp.output_parsed                  # a real Ticket instance
print(t.category, t.priority, t.needs_human)
▶ How this works
  1. class Ticket(BaseModel) declares the exact shape you want back — a Literal field becomes an enum the model cannot step outside.
  2. client.responses.parse(..., text_format=Ticket) makes the model return data matching that schema. (On Claude the same thing is messages.parse(output_format=…).)
  3. resp.output_parsed hands you a validated Ticket object — t.category, t.priority — not a string to regex. (On Claude it's resp.parsed_output.)

Try this: ask for something ambiguous and watch the fields still come back valid — the schema makes "the model phrased it differently" bugs impossible.

The one difference to rememberOpenAI: responses.parse(text_format=Model) → resp.output_parsed. Claude: messages.parse(output_format=Model) → resp.parsed_output. The Pydantic model itself is byte-for-byte identical across both.

4 · Tools — function calling advanced

A model that can only emit text can answer questions; a model that can call your functions can do things. Tools are how you let GPT reach out to the world — look up weather, query a database, send an email — while you keep control of what actually runs. The model never executes anything; it asks you to, and your code decides.

In the Responses API a tool is a flat dict: {"type":"function","name":…,"description":…,"parameters":{…}}, where parameters is a JSON schema of the arguments. You pass a list of these as tools=[…]. When the model wants a tool, it returns an item in resp.output with type == "function_call", carrying .name, .arguments (a JSON string you json.loads), and a .call_id. You run the real function and feed the result back as a function_call_output item keyed by that same call_id. There is no stop_reason — you loop while any function_call items remain, and when none do, the answer is in resp.output_text.

the model asks; your code disposes — loop until no function_calls remain responses.create tools=[…] function_call? name · arguments · call_id run fn → function_call_output none → output_text (done) Ask · run · feed · repeat. The model emits function_call items; you run the real function and append a function_call_output with the matching call_id; loop until the output has no more function calls, then read output_text.
🗺️ How to read this diagram
  • The first box is the call with tools=[…] passed in.
  • The middle box is the decision: does resp.output contain a function_call item? Unlike Claude there's no stop_reason flag — you check the output items themselves.
  • The amber box is the tool branch: run the function, append a function_call_output with the matching call_id, loop back.
  • The green box is the exit: no function calls left → the answer is in output_text.

In short: flat tool defs, function_call items in, function_call_output items back, loop on the output — not on a stop reason.

Lab O2.3
tool_loop.pyimport json
from openai import OpenAI
client = OpenAI()

WEATHER_TOOL = {
    "type": "function",
    "name": "get_weather",
    "description": "Current weather for a city. Call when the user asks about weather.",
    "parameters": {"type":"object",
        "properties":{"city":{"type":"string"}}, "required":["city"]},
}
def get_weather(city): return f"18°C and cloudy in {city}"
DISPATCH = {"get_weather": get_weather}
MAX_STEPS = 6

def run_agent(user_input):
    input_list = [{"role":"user", "content": user_input}]
    for _ in range(MAX_STEPS):          # hard cap — always terminates
        resp = client.responses.create(
            model="gpt-5.5", tools=[WEATHER_TOOL], input=input_list)
        calls = [o for o in resp.output if o.type == "function_call"]
        if not calls:                   # no tool calls → done
            return resp.output_text
        input_list += resp.output         # append output items verbatim
        for call in calls:
            out = DISPATCH[call.name](**json.loads(call.arguments))
            input_list.append({"type":"function_call_output",
                               "call_id": call.call_id, "output": str(out)})
    return "Stopped: hit step limit."

print(run_agent("What's the weather in Tokyo?"))
Always cap the loopAn agent without a hard MAX_STEPS cap can spin forever — ask for a tool, get a result, ask again — quietly burning money. The cap is a seatbelt, not a nicety. (This is the same lesson as the Claude agent loop in Ch 4; only the protocol names differ.)

5 · Streaming advanced

A ten-second wait feels broken; the same ten seconds with words appearing feels fast. Streaming sends tokens as they're generated instead of making the user wait for the whole reply. In the Responses API you open a stream as a context manager and iterate events — filtering for the text-delta event type and reading each chunk off event.delta.

Lab O2.4
stream.pyfrom openai import OpenAI
client = OpenAI()

with client.responses.stream(
    model="gpt-5.5",
    input="Write a one-sentence bedtime story about a robot.",
) as stream:
    for event in stream:
        if event.type == "response.output_text.delta":
            print(event.delta, end="", flush=True)
▶ How this works
  1. with client.responses.stream(...) as stream: opens a streaming connection and closes it cleanly when done.
  2. You iterate the event stream and filter for event.type == "response.output_text.delta" — the text-chunk events.
  3. Each chunk's text is on event.delta; print it immediately so the reply appears live.

Try this: there's no .text_stream helper like Claude's — OpenAI streams typed events, so you filter by event.type. That's the one gotcha when porting a streaming loop between the two SDKs.

6 · Errors, retries & cost professional

Three things separate a script from a service: it catches the specific errors it can handle, it retries only the transient ones, and it reads its own cost on every call. The OpenAI SDK raises typed exceptions you can branch on — openai.RateLimitError (a 429, worth retrying with backoff), openai.APIStatusError (check the status code — a 4xx is your bug, don't retry; a 5xx is transient), and the base openai.APIError. Catch most-specific first, and never wrap everything in a bare except that treats a fatal 400 like a retryable 429.

Lab O2.5
robust.pyimport time, openai
from openai import OpenAI
client = OpenAI()

def ask(prompt, tries=3):
    for i in range(tries):
        try:
            r = client.responses.create(model="gpt-5.5", input=prompt)
            return r.output_text
        except openai.RateLimitError:
            time.sleep(2 ** i)            # 429 → back off and retry
        except openai.APIStatusError as e:
            if e.status_code < 500: raise   # 4xx is your bug — fail fast
            time.sleep(2 ** i)            # 5xx is transient
    raise RuntimeError("exhausted retries")
Cost lives on resp.usageEvery response carries resp.usage.input_tokens and resp.usage.output_tokens; cached-prefix reuse shows up as resp.usage.input_tokens_details.cached_tokens. Log them per request and your cost dashboard builds itself.

🪜 Practice ladder beginner → industry

  1. Beginner: run Lab O2.1; print output_text and both token counts.
  2. Easy: move your rules from input into instructions and confirm behavior gets more consistent.
  3. Core: run Lab O2.2 and change a Literal field — watch the model stay inside the enum.
  4. Stretch: add a second tool to Lab O2.3 and confirm the loop needs no changes to dispatch it.
  5. Hard: make Lab O2.4 stream into a running string and return the full text at the end.
  6. Industry: combine Lab O2.5's retry wrapper with per-request usage logging and a cost-per-feature ledger.

✓ Checkpoint — you can move on when you can…

  • Make a Responses API call and read output_text and usage.
  • Explain instructions vs input and why the split matters.
  • Get a validated Pydantic object from responses.parse.
  • Run the function-call loop and stream token deltas.

Knowledge check check yourself

✓ Knowledge check

What is the difference between the OpenAI Responses API and Chat Completions, and which should you build on today?

Show answer
The Responses API (client.responses.create → resp.output_text) is OpenAI's primary, current interface and the one to build on. Chat Completions (client.chat.completions.create → resp.choices[0].message.content) is the previous standard — still supported indefinitely, so you'll see it in older code, but new work should use Responses.
✓ Knowledge check

How do you get schema-valid structured output from GPT, and how does the tool-calling loop know when it's done?

Show answer
Use client.responses.parse(..., text_format=PydanticModel) and read the validated instance from resp.output_parsed. For tools, the model returns function_call items in resp.output (with name, a JSON-string arguments, and call_id); you run the function and append a function_call_output with the matching call_id. There is no stop_reason — you loop while function-call items remain, and when none do, the answer is in resp.output_text.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in