Request Parameters, in Full
You've used model, instructions, and input. The Responses API has more dials — reasoning effort, output caps, sampling, structured output, tools, state, and metadata — and knowing what each does (and when to leave it alone) is the difference between tuning on purpose and cargo-culting. A reference you'll come back to.
Learning objectives
- Name the core Responses parameters and what each controls.
- Use the reasoning-effort and output-length dials deliberately.
- Understand sampling (
temperature/top_p) and when to touch it. - Know the structured-output, tools, state, and metadata parameters.
pip install openai. Parameter values here are illustrative — check current docs for the exact allowed ranges per model.1 · The core three essential
Three parameters do the heavy lifting on every call, and you already know them. model picks which GPT (keep it in config — ox1). instructions is the system/developer role — the stable rules. input is the user content — a string for a one-shot, or a list of role/content turns for a conversation or multimodal request. Master these and you can make any call; the rest of this chapter is the dials you reach for on top of them.
core.pyresp = client.responses.create(
model="gpt-5.5", # which model (from config)
instructions="You are terse.", # system/developer role
input="Define idempotency.", # user content (str or list)
)
2 · Reasoning effort & output length essential
Two dials control how hard the model thinks and how much it says. reasoning={"effort": …} (low/medium/high/max) sets how much hidden reasoning a reasoning-capable model does — more for hard problems, less for easy ones. max_output_tokens caps the reply length — a hard ceiling that both controls cost (output bills high) and prevents runaway generations.
effort.pyresp = client.responses.create(
model="gpt-5.5",
reasoning={"effort": "high"}, # low | medium | high | max
max_output_tokens=800, # hard cap on the reply
input="Plan a zero-downtime database migration.",
)
effort: "low", small max_output_tokens. Hard reasoning/agents → effort: "high" or "max". Paying for high effort or an uncapped output on a trivial task is money burned — these are the two dials that most directly move your bill.3 · Sampling — temperature & top_p intermediate
These control randomness, and the honest advice is: usually leave them alone. temperature (roughly 0–2) scales how adventurous the next-token choice is — low is focused and near-deterministic, high is creative and varied. top_p (nucleus sampling, 0–1) instead limits the choice to the most-probable mass of tokens. The key discipline: tune one, not both, and only when you have a reason — low temperature for extraction/classification where you want stable, repeatable output; higher for brainstorming or creative writing where variety is the point.
The common mistake is reflexively setting temperature=0 "for determinism" and expecting byte-identical output every time. Low temperature makes output more consistent, not provably identical — and on reasoning models the effort dial often matters more than temperature anyway. Default to leaving sampling untouched unless a measured problem tells you to change it.
temperature and top_p together — they interact confusingly. Pick one, change it only when evals show a reason (too random → lower; too repetitive → raise), and remember low temperature ≠ guaranteed-identical output.4 · Structured output & tools intermediate
Two parameters turn free text into something a program can trust. For structured output, use client.responses.parse(..., text_format=PydanticModel) and read resp.output_parsed (covered in ox2) — the response is constrained to your schema. For tools, pass tools=[…] — a mix of your function tools ({"type":"function",…}) and hosted tools ({"type":"web_search"} etc., oap5); tool_choice can force or forbid tool use when you need to.
structured_and_tools.py# structured output
r = client.responses.parse(model="gpt-5.5", input="...", text_format=MyModel)
obj = r.output_parsed
# tools (function + hosted), optionally forced
r = client.responses.create(
model="gpt-5.5", input="...",
tools=[MY_FUNCTION_TOOL, {"type": "web_search"}],
tool_choice="auto", # auto | required | none
)
5 · State & metadata advanced
A few parameters manage conversation state and bookkeeping. store (bool) controls whether OpenAI retains the response server-side so it can be referenced later; previous_response_id can chain a new call onto a stored prior response, letting the platform carry context instead of you re-sending the whole transcript. metadata attaches your own key/value tags to a response — handy for correlating calls with a trace_id or feature label in your logs (oap4). These are optional conveniences; a basic stateless app ignores them and just re-sends its input list each turn.
stop_reason. You check resp.status ("completed", etc.) and, for tools, whether resp.output still contains function_call items (ox2) — not a single stop field.6 · Quick reference professional
| Parameter | Controls | Default stance |
|---|---|---|
model | which model | from one config value |
instructions | system/developer role | stable rules, cache-friendly |
input | user content | str, or role/content list |
reasoning.effort | how hard it thinks | medium; raise with evidence |
max_output_tokens | reply length cap | always set a sane cap |
temperature / top_p | randomness | leave default; tune one if needed |
text_format | structured output (via .parse) | use for machine-read output |
tools / tool_choice | function & hosted tools | as needed |
store / previous_response_id | server-side state | optional; off for stateless |
metadata | your tags | trace_id / feature label |
7 · Professional — a sane default profile professional
The professional move is to not re-decide every parameter on every call. Wrap responses.create in a small helper that encodes your defaults — model from config, a sensible max_output_tokens cap, effort chosen per task-type, sampling left untouched, and metadata auto-stamped with a trace_id and feature label — and let call sites override only what's genuinely different. That turns "which parameters?" from a per-call question into a one-time policy, and makes a parameter change (a model bump, a new cap) a one-line edit instead of a find-and-replace.
🪜 Practice ladder beginner → industry
- Beginner: make a call with just the core three parameters.
- Easy: add
max_output_tokensand watch the reply get capped. - Core: compare
effort: "low"vs"high"on a hard prompt (depth vs tokens). - Stretch: lower
temperatureon an extraction task and check consistency across runs. - Hard: add
metadatawith a trace_id and confirm it round-trips on the response. - Industry: write the default-profile helper and have three call sites override only one parameter each.
✓ Checkpoint — you can move on when you can…
- Name the core three parameters and what each controls.
- Set
reasoning.effortandmax_output_tokensdeliberately. - Explain when (and when not) to touch sampling.
- Describe the structured-output, tools, state, and metadata params.
Knowledge check check yourself
Which two parameters most directly control cost and quality, and how should you set them?
Show answer
reasoning={"effort": …} (low/medium/high/max) controls how much the model reasons, and max_output_tokens caps reply length (output bills at the higher rate). Set effort low for easy tasks and high/max only for hard reasoning, and always set a sane output cap. Paying for high effort or an uncapped output on a trivial task is money burned.What's the right discipline for temperature/top_p, and what replaces stop_reason in the Responses API?
Show answer
stop_reason: check resp.status and, for tools, whether resp.output still contains function_call items.