The OpenAI API — Responses, Structured Output & Tools
One API call is the atom every larger system is built from. This chapter is the OpenAI counterpart to the Anthropic API lesson: the single responses.create call, how to get machine-parseable output, how tools (function calling) work, and how to stream — all with the real openai SDK.
Learning objectives
- Make a single request with the Responses API and read the reply and token usage.
- Separate the system role (
instructions) from the user turn (input). - Get schema-valid structured output with
responses.parseand Pydantic. - Wire tools (function calling) and run the function-call loop.
- Stream tokens, and handle errors, retries, and cost like a service.
OPENAI_API_KEY) from platform.openai.com + pip install openai. Keep the key in a .env file (never in code) and let the zero-argument OpenAI() client pick it up from the environment. Reading and learning works without any of this — run when you're ready.1 · The single call essential
Everything OpenAI does for you is one function call, repeated. Agents, RAG, chatbots — strip away the orchestration and each is a loop around the same primitive: send some text, get some text back. Learn that primitive cold and the rest of the course stops feeling like magic and starts feeling like plumbing you control.
OpenAI gives you two ways to make that call, and picking the right one matters. The Responses API (client.responses.create) is the primary, current interface — it's what you should learn and build on. The older Chat Completions API (client.chat.completions.create, which returns resp.choices[0].message.content) still works and is supported indefinitely, so you'll see it in older tutorials — but it's the previous standard, and this course teaches Responses. The two are shaped differently enough that mixing them up is a common early bug.
The Responses call takes three things you care about: model (which GPT), instructions (the system role — who the assistant is and its rules), and input (the user's content). It hands back a response object whose output_text property is the reply as a plain string. That's the whole atom.
responses.create with model/instructions/input, then resp.output_text. Everything else in the course is this primitive in a loop.
- The first box builds the client —
OpenAI()reads your key from the environment. - The middle box is the one call:
responses.create()with the three arguments you care about. - The green box is the reply;
resp.output_textis the convenience property that gives you the text as a string, no block-list to unpack.
In short: client → call → output_text. Memorize this shape; it never changes.
first_call.pyfrom dotenv import load_dotenv
from openai import OpenAI
load_dotenv()
client = OpenAI() # reads OPENAI_API_KEY
resp = client.responses.create(
model="gpt-5.5",
instructions="You are a concise assistant.",
input="What is an API token? Answer in two sentences.",
)
print(resp.output_text)
print(resp.usage.input_tokens, resp.usage.output_tokens)
client.messages.create(system=…, messages=[…]) and you pull text out of a resp.content block list. On OpenAI it's client.responses.create(instructions=…, input=…) and resp.output_text. Same idea, different field names — which is exactly what the Claude | OpenAI tabs across this course let you compare.2 · System vs user — instructions and input essential
The biggest early leverage isn't a cleverer prompt — it's putting the right text in the right slot. The Responses API separates two channels: instructions is the system/developer role — stable rules about who the assistant is, what format to use, what never to do — and input is the user turn, the actual question or task. Mixing them (stuffing your rules into the user message) works until it doesn't: the model starts treating your rules as negotiable content rather than standing orders.
Why the split matters, concretely: the instructions are where you encode the parts that don't change between requests — tone, output format, refusal rules, the persona. The input is the volatile part that changes every call. Keeping the stable rules in instructions makes behavior consistent and is also what lets the platform cache the stable prefix for you (cheaper repeated calls).
The common mistake: re-sending the same long rules inside every user message. Put them in instructions once, keep input lean, and both your consistency and your bill improve.
input as a list of role/content turns and append each reply before the next call. "Memory" is just you resending the transcript; there's no hidden server-side state in a basic call.3 · Structured output with Pydantic expert
Picture the 2am page. Your classifier has hummed along for weeks, its output scraped by a regex that plucks a priority out of the model's reply. Then one night the model, entirely reasonably, answers "Sure! The priority here is high 😊" instead of the exact phrasing your regex expected — the parse returns nothing, the ticket routes nowhere, and your phone lights up. Nothing broke on your side; the model just rephrased. That fragility is exactly what structured output eliminates.
The idea is a shift in who guarantees the shape of the data. Instead of hoping the model phrases things the way your parser expects, you declare a Pydantic schema up front and the shape becomes a constraint the model must satisfy while generating. With OpenAI you use client.responses.parse(..., text_format=YourModel) and read a validated instance from resp.output_parsed — a real typed object, no parsing step to break.
extract.pyfrom openai import OpenAI
from pydantic import BaseModel
from typing import Literal
client = OpenAI()
class Ticket(BaseModel):
category: Literal["bug","feature","billing","other"]
priority: Literal["low","medium","high"]
summary: str
needs_human: bool
resp = client.responses.parse( # .parse() validates for you
model="gpt-5.5", max_output_tokens=512,
input="I was charged twice and support hasn't replied in 3 days!",
text_format=Ticket,
)
t = resp.output_parsed # a real Ticket instance
print(t.category, t.priority, t.needs_human)
class Ticket(BaseModel)declares the exact shape you want back — aLiteralfield becomes an enum the model cannot step outside.client.responses.parse(..., text_format=Ticket)makes the model return data matching that schema. (On Claude the same thing ismessages.parse(output_format=…).)resp.output_parsedhands you a validatedTicketobject —t.category,t.priority— not a string to regex. (On Claude it'sresp.parsed_output.)
Try this: ask for something ambiguous and watch the fields still come back valid — the schema makes "the model phrased it differently" bugs impossible.
responses.parse(text_format=Model) → resp.output_parsed. Claude: messages.parse(output_format=Model) → resp.parsed_output. The Pydantic model itself is byte-for-byte identical across both.4 · Tools — function calling advanced
A model that can only emit text can answer questions; a model that can call your functions can do things. Tools are how you let GPT reach out to the world — look up weather, query a database, send an email — while you keep control of what actually runs. The model never executes anything; it asks you to, and your code decides.
In the Responses API a tool is a flat dict: {"type":"function","name":…,"description":…,"parameters":{…}}, where parameters is a JSON schema of the arguments. You pass a list of these as tools=[…]. When the model wants a tool, it returns an item in resp.output with type == "function_call", carrying .name, .arguments (a JSON string you json.loads), and a .call_id. You run the real function and feed the result back as a function_call_output item keyed by that same call_id. There is no stop_reason — you loop while any function_call items remain, and when none do, the answer is in resp.output_text.
function_call items; you run the real function and append a function_call_output with the matching call_id; loop until the output has no more function calls, then read output_text.
- The first box is the call with
tools=[…]passed in. - The middle box is the decision: does
resp.outputcontain afunction_callitem? Unlike Claude there's nostop_reasonflag — you check the output items themselves. - The amber box is the tool branch: run the function, append a
function_call_outputwith the matchingcall_id, loop back. - The green box is the exit: no function calls left → the answer is in
output_text.
In short: flat tool defs, function_call items in, function_call_output items back, loop on the output — not on a stop reason.
tool_loop.pyimport json
from openai import OpenAI
client = OpenAI()
WEATHER_TOOL = {
"type": "function",
"name": "get_weather",
"description": "Current weather for a city. Call when the user asks about weather.",
"parameters": {"type":"object",
"properties":{"city":{"type":"string"}}, "required":["city"]},
}
def get_weather(city): return f"18°C and cloudy in {city}"
DISPATCH = {"get_weather": get_weather}
MAX_STEPS = 6
def run_agent(user_input):
input_list = [{"role":"user", "content": user_input}]
for _ in range(MAX_STEPS): # hard cap — always terminates
resp = client.responses.create(
model="gpt-5.5", tools=[WEATHER_TOOL], input=input_list)
calls = [o for o in resp.output if o.type == "function_call"]
if not calls: # no tool calls → done
return resp.output_text
input_list += resp.output # append output items verbatim
for call in calls:
out = DISPATCH[call.name](**json.loads(call.arguments))
input_list.append({"type":"function_call_output",
"call_id": call.call_id, "output": str(out)})
return "Stopped: hit step limit."
print(run_agent("What's the weather in Tokyo?"))
MAX_STEPS cap can spin forever — ask for a tool, get a result, ask again — quietly burning money. The cap is a seatbelt, not a nicety. (This is the same lesson as the Claude agent loop in Ch 4; only the protocol names differ.)5 · Streaming advanced
A ten-second wait feels broken; the same ten seconds with words appearing feels fast. Streaming sends tokens as they're generated instead of making the user wait for the whole reply. In the Responses API you open a stream as a context manager and iterate events — filtering for the text-delta event type and reading each chunk off event.delta.
stream.pyfrom openai import OpenAI
client = OpenAI()
with client.responses.stream(
model="gpt-5.5",
input="Write a one-sentence bedtime story about a robot.",
) as stream:
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
with client.responses.stream(...) as stream:opens a streaming connection and closes it cleanly when done.- You iterate the event stream and filter for
event.type == "response.output_text.delta"— the text-chunk events. - Each chunk's text is on
event.delta; print it immediately so the reply appears live.
Try this: there's no .text_stream helper like Claude's — OpenAI streams typed events, so you filter by event.type. That's the one gotcha when porting a streaming loop between the two SDKs.
6 · Errors, retries & cost professional
Three things separate a script from a service: it catches the specific errors it can handle, it retries only the transient ones, and it reads its own cost on every call. The OpenAI SDK raises typed exceptions you can branch on — openai.RateLimitError (a 429, worth retrying with backoff), openai.APIStatusError (check the status code — a 4xx is your bug, don't retry; a 5xx is transient), and the base openai.APIError. Catch most-specific first, and never wrap everything in a bare except that treats a fatal 400 like a retryable 429.
robust.pyimport time, openai
from openai import OpenAI
client = OpenAI()
def ask(prompt, tries=3):
for i in range(tries):
try:
r = client.responses.create(model="gpt-5.5", input=prompt)
return r.output_text
except openai.RateLimitError:
time.sleep(2 ** i) # 429 → back off and retry
except openai.APIStatusError as e:
if e.status_code < 500: raise # 4xx is your bug — fail fast
time.sleep(2 ** i) # 5xx is transient
raise RuntimeError("exhausted retries")
resp.usage.input_tokens and resp.usage.output_tokens; cached-prefix reuse shows up as resp.usage.input_tokens_details.cached_tokens. Log them per request and your cost dashboard builds itself.🪜 Practice ladder beginner → industry
- Beginner: run Lab O2.1; print
output_textand both token counts. - Easy: move your rules from
inputintoinstructionsand confirm behavior gets more consistent. - Core: run Lab O2.2 and change a
Literalfield — watch the model stay inside the enum. - Stretch: add a second tool to Lab O2.3 and confirm the loop needs no changes to dispatch it.
- Hard: make Lab O2.4 stream into a running string and return the full text at the end.
- Industry: combine Lab O2.5's retry wrapper with per-request
usagelogging and a cost-per-feature ledger.
✓ Checkpoint — you can move on when you can…
- Make a Responses API call and read
output_textandusage. - Explain
instructionsvsinputand why the split matters. - Get a validated Pydantic object from
responses.parse. - Run the function-call loop and stream token deltas.
Knowledge check check yourself
What is the difference between the OpenAI Responses API and Chat Completions, and which should you build on today?
Show answer
client.responses.create → resp.output_text) is OpenAI's primary, current interface and the one to build on. Chat Completions (client.chat.completions.create → resp.choices[0].message.content) is the previous standard — still supported indefinitely, so you'll see it in older code, but new work should use Responses.How do you get schema-valid structured output from GPT, and how does the tool-calling loop know when it's done?
Show answer
client.responses.parse(..., text_format=PydanticModel) and read the validated instance from resp.output_parsed. For tools, the model returns function_call items in resp.output (with name, a JSON-string arguments, and call_id); you run the function and append a function_call_output with the matching call_id. There is no stop_reason — you loop while function-call items remain, and when none do, the answer is in resp.output_text.