AI EngineeringZero to ProductionHome·About·What’s new·Contact
OpenAI API in Practice · Part 8

Request Parameters, in Full

You've used model, instructions, and input. The Responses API has more dials — reasoning effort, output caps, sampling, structured output, tools, state, and metadata — and knowing what each does (and when to leave it alone) is the difference between tuning on purpose and cargo-culting. A reference you'll come back to.

⏱️ ~1 hour📖 Reference🎯 Beginner→Professional

Learning objectives

  • Name the core Responses parameters and what each controls.
  • Use the reasoning-effort and output-length dials deliberately.
  • Understand sampling (temperature/top_p) and when to touch it.
  • Know the structured-output, tools, state, and metadata parameters.
⚙️ To run this for realNeeds an OpenAI API key + pip install openai. Parameter values here are illustrative — check current docs for the exact allowed ranges per model.

1 · The core three essential

Three parameters do the heavy lifting on every call, and you already know them. model picks which GPT (keep it in config — ox1). instructions is the system/developer role — the stable rules. input is the user content — a string for a one-shot, or a list of role/content turns for a conversation or multimodal request. Master these and you can make any call; the rest of this chapter is the dials you reach for on top of them.

core.pyresp = client.responses.create(
    model="gpt-5.5",                 # which model (from config)
    instructions="You are terse.",     # system/developer role
    input="Define idempotency.",       # user content (str or list)
)

2 · Reasoning effort & output length essential

Two dials control how hard the model thinks and how much it says. reasoning={"effort": …} (low/medium/high/max) sets how much hidden reasoning a reasoning-capable model does — more for hard problems, less for easy ones. max_output_tokens caps the reply length — a hard ceiling that both controls cost (output bills high) and prevents runaway generations.

Lab OP8.1
effort.pyresp = client.responses.create(
    model="gpt-5.5",
    reasoning={"effort": "high"},     # low | medium | high | max
    max_output_tokens=800,            # hard cap on the reply
    input="Plan a zero-downtime database migration.",
)
Set these on purposeClassification/routing → effort: "low", small max_output_tokens. Hard reasoning/agents → effort: "high" or "max". Paying for high effort or an uncapped output on a trivial task is money burned — these are the two dials that most directly move your bill.

3 · Sampling — temperature & top_p intermediate

These control randomness, and the honest advice is: usually leave them alone. temperature (roughly 0–2) scales how adventurous the next-token choice is — low is focused and near-deterministic, high is creative and varied. top_p (nucleus sampling, 0–1) instead limits the choice to the most-probable mass of tokens. The key discipline: tune one, not both, and only when you have a reason — low temperature for extraction/classification where you want stable, repeatable output; higher for brainstorming or creative writing where variety is the point.

The common mistake is reflexively setting temperature=0 "for determinism" and expecting byte-identical output every time. Low temperature makes output more consistent, not provably identical — and on reasoning models the effort dial often matters more than temperature anyway. Default to leaving sampling untouched unless a measured problem tells you to change it.

One knob, with a reasonDon't set temperature and top_p together — they interact confusingly. Pick one, change it only when evals show a reason (too random → lower; too repetitive → raise), and remember low temperature ≠ guaranteed-identical output.

4 · Structured output & tools intermediate

Two parameters turn free text into something a program can trust. For structured output, use client.responses.parse(..., text_format=PydanticModel) and read resp.output_parsed (covered in ox2) — the response is constrained to your schema. For tools, pass tools=[…] — a mix of your function tools ({"type":"function",…}) and hosted tools ({"type":"web_search"} etc., oap5); tool_choice can force or forbid tool use when you need to.

structured_and_tools.py# structured output
r = client.responses.parse(model="gpt-5.5", input="...", text_format=MyModel)
obj = r.output_parsed

# tools (function + hosted), optionally forced
r = client.responses.create(
    model="gpt-5.5", input="...",
    tools=[MY_FUNCTION_TOOL, {"type": "web_search"}],
    tool_choice="auto",              # auto | required | none
)

5 · State & metadata advanced

A few parameters manage conversation state and bookkeeping. store (bool) controls whether OpenAI retains the response server-side so it can be referenced later; previous_response_id can chain a new call onto a stored prior response, letting the platform carry context instead of you re-sending the whole transcript. metadata attaches your own key/value tags to a response — handy for correlating calls with a trace_id or feature label in your logs (oap4). These are optional conveniences; a basic stateless app ignores them and just re-sends its input list each turn.

No stop_reason hereIf you're coming from the Chat Completions API or from Claude, note the Responses object has no stop_reason. You check resp.status ("completed", etc.) and, for tools, whether resp.output still contains function_call items (ox2) — not a single stop field.

6 · Quick reference professional

ParameterControlsDefault stance
modelwhich modelfrom one config value
instructionssystem/developer rolestable rules, cache-friendly
inputuser contentstr, or role/content list
reasoning.efforthow hard it thinksmedium; raise with evidence
max_output_tokensreply length capalways set a sane cap
temperature / top_prandomnessleave default; tune one if needed
text_formatstructured output (via .parse)use for machine-read output
tools / tool_choicefunction & hosted toolsas needed
store / previous_response_idserver-side stateoptional; off for stateless
metadatayour tagstrace_id / feature label

7 · Professional — a sane default profile professional

The professional move is to not re-decide every parameter on every call. Wrap responses.create in a small helper that encodes your defaults — model from config, a sensible max_output_tokens cap, effort chosen per task-type, sampling left untouched, and metadata auto-stamped with a trace_id and feature label — and let call sites override only what's genuinely different. That turns "which parameters?" from a per-call question into a one-time policy, and makes a parameter change (a model bump, a new cap) a one-line edit instead of a find-and-replace.

🪜 Practice ladder beginner → industry

  1. Beginner: make a call with just the core three parameters.
  2. Easy: add max_output_tokens and watch the reply get capped.
  3. Core: compare effort: "low" vs "high" on a hard prompt (depth vs tokens).
  4. Stretch: lower temperature on an extraction task and check consistency across runs.
  5. Hard: add metadata with a trace_id and confirm it round-trips on the response.
  6. Industry: write the default-profile helper and have three call sites override only one parameter each.

✓ Checkpoint — you can move on when you can…

  • Name the core three parameters and what each controls.
  • Set reasoning.effort and max_output_tokens deliberately.
  • Explain when (and when not) to touch sampling.
  • Describe the structured-output, tools, state, and metadata params.

Knowledge check check yourself

✓ Knowledge check

Which two parameters most directly control cost and quality, and how should you set them?

Show answer
reasoning={"effort": …} (low/medium/high/max) controls how much the model reasons, and max_output_tokens caps reply length (output bills at the higher rate). Set effort low for easy tasks and high/max only for hard reasoning, and always set a sane output cap. Paying for high effort or an uncapped output on a trivial task is money burned.
✓ Knowledge check

What's the right discipline for temperature/top_p, and what replaces stop_reason in the Responses API?

Show answer
Leave sampling at default unless a measured problem calls for it; then tune one of temperature or top_p (not both) — lower for stable extraction, higher for creative variety — and remember low temperature means more consistent, not provably identical. The Responses API has no stop_reason: check resp.status and, for tools, whether resp.output still contains function_call items.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in