AI EngineeringZero to ProductionHome·About·What’s new·Contact
Codex & OpenAI · Chapter O1

Introduction to GPT & the OpenAI Model Family

This section targets OpenAI's models the same way the Claude section targets Claude. Before you write code against GPT, get the mental model straight: what the models are, how they differ, and how to pick the right one for a task — so your choices are deliberate, not cargo-culted.

⏱️ ~1.5 hours🧪 3 labs🎯 Beginner→Tech-lead

Learning objectives

  • Explain what a GPT model is and how it's trained to be helpful and safe.
  • Reason about the OpenAI model family by role — flagship, mini, realtime — without leaning on fixed tier names.
  • Reason about tokens, the context window, and cost.
  • Pick the right model per task, and tune the reasoning-effort dial, like a lead sizing a fleet.
▶ Where this fitsThis is the OpenAI counterpart to the Claude & Anthropic section. Everything you learned about prompting, RAG, agents, and evaluation applies to both vendors — only the SDK call changes. Many code pages in this course show Claude | OpenAI tabs side by side; this section is the deep dive on the OpenAI side.

1 · What GPT is essential

Here's a mistake almost everyone makes on day one. They open the docs, copy the first responses.create snippet, swap in their prompt, and treat the model as a magic box that "understands English." Then it does something baffling — invents a fact, refuses a reasonable request, or trails off mid-sentence — and they have no mental model to explain why. The fix isn't more prompt-tweaking; it's knowing, at a gut level, what the thing actually is. Ten minutes here saves you a week of confused debugging later.

So here's the honest one-sentence version: GPT is a very, very good autocomplete. Think of the predictive text on your phone that suggests the next word — now imagine that same trick scaled up on a supercomputer, trained on a huge slice of the written world, until "predict the next word" starts producing working code, coherent essays, and step-by-step reasoning. That's not a metaphor that breaks down later; it's literally the machinery. GPT stands for Generative Pre-trained Transformer — "generative" because it produces text, "pre-trained" because it learns from vast data before you ever call it, and "transformer" after the neural-network architecture underneath. Everything it does is choosing the next token (roughly, the next chunk of text) given everything it has seen so far, one token at a time, left to right.

But raw autocomplete would be useless — and often dangerous. A model that only predicts "what text usually comes next" will happily continue a harmful rant, or make up a citation because made-up citations look like real ones. So OpenAI does a second stage of training called alignment — chiefly RLHF (reinforcement learning from human feedback) — that reshapes which continuations the model prefers, steering it toward answers that are helpful, honest, and safe. The everyday analogy: the base model is a brilliant new hire who has read everything but has no judgment; alignment is the onboarding that teaches them what's actually appropriate to say at work.

The common mistake to name and avoid: believing GPT "knows" things the way a database knows things. It doesn't look up facts — it predicts plausible text, and plausible is not the same as true. That single insight explains hallucinations (confident wrong answers), why grounding it in retrieved documents (RAG, Ch 3) matters so much, and why you verify anything load-bearing. Keep the autocomplete picture in your head and GPT stops being mysterious.

two training stages turn raw autocomplete into a helpful assistant vast text the written world base model predict next token, one at a time aligned GPT helpful · honest · safe pre-training alignment (RLHF) plausible ≠ true — it predicts text, it does not look facts up Autocomplete, then aligned. A base model learns to predict the next token from vast text; a second alignment stage (RLHF) reshapes which continuations it prefers toward helpful, honest, and safe ones. Because it predicts plausible text rather than looking facts up, plausible is not the same as true.
🗺️ How to read this diagram

Read the three boxes left to right — they are the pipeline that produces the GPT you call.

  • The blue "vast text" box is the training data — a huge slice of the written world the model learns from.
  • The middle "base model" box is the raw autocomplete: pre-training teaches it only to predict the next token, one at a time, left to right. Brilliant, but with no sense of what is appropriate to say.
  • The green "aligned GPT" box is the payoff: a second stage (RLHF) reshapes which continuations it prefers, steering it toward helpful, honest, and safe answers.
  • The arrows show the one-way flow — each stage feeds the next; the bottom line is the warning that survives all of it: it predicts plausible text, so plausible is not the same as true (hallucinations).

In short: base model = predict the next token; alignment = prefer the good continuations; and it never looks facts up, which is why grounding it in real documents matters.

Next-token prediction, alignedUnder the hood GPT predicts the most likely next token given everything so far. Alignment training (RLHF) shapes which continuations it prefers — toward helpful, safe, truthful ones. That's why it follows instructions and refuses harmful ones.

2 · The model family essential

Picture the tool wall in a decent workshop. There's a cordless screwdriver you grab a hundred times a day, a heavy drill for the jobs the screwdriver can't manage, and a tiny precision driver for fiddly, repetitive work. A good tradesperson doesn't reach for the biggest drill every time "to be safe" — they match the tool to the task. OpenAI ships GPT the same way: not one model, but a family, and your job is to grab the right one.

Here's where OpenAI differs from Anthropic in a way worth naming up front. Anthropic publishes three tidy, stable tier names — Opus, Sonnet, Haiku. OpenAI's lineup is less tidy and changes faster: there's a general-purpose flagship model (this course's examples use the id gpt-5.5), smaller and cheaper "mini" variants for high-volume work, and specialized realtime/audio models (for example gpt-realtime-2 for voice and gpt-4o-family models for multimodal). The important discipline: don't memorize or hard-code a tier name as if it were permanent. Think in roles — flagship for the hard stuff, mini for the cheap high-volume stuff, realtime for voice — and pin the exact model id in one config value you can swap when OpenAI ships the next generation.

Why have a family at all, rather than one model that does everything? Because the three axes you care about — capability, speed, and cost — genuinely trade against each other. A bigger model reasons over more steps and makes fewer mistakes on hard problems, but every token it produces costs more and arrives a little slower. There is no single model that is simultaneously the smartest, the fastest, and the cheapest, any more than there's one vehicle that's the roomiest, quickest, and most fuel-efficient. The family exists precisely so you can pick your point on that triangle per task.

The common mistake — the expensive one this chapter is built to cure — is defaulting to the flagship for everything because it's "the best." At the volume of a real production system that habit can multiply your bill many times over for tasks a mini model handles identically. The professional instinct is the opposite: pick by the task, not by habit, start from a mid/mini model, and justify every move up to the flagship with evidence.

one scale, opposite ends — you cannot have all three at once ← cheaper · faster more capable · pricier → mini models fastest · cheapest volume · routing realtime/audio voice · streaming specialized flagship deepest reasoning hard tasks · agents think in roles, not fixed names — pin the exact model id in config Roles on one trade-off line. Mini models sit at the cheap-fast end, the flagship at the capable-pricey end, and realtime/audio models are a specialized branch. Pick a point per task — and keep the exact id in config, because OpenAI's names change faster than Anthropic's.
🗺️ How to read this diagram

Read the boxes along the single grey scale at the bottom: left is cheaper and faster, right is more capable and pricier.

  • The green "mini models" box is the cheap-fast end — the default for high-volume, simple work like classification and routing.
  • The purple "flagship" box is the capable-pricey end — reach for it only when a task's evals show the smaller model failing.
  • The blue "realtime/audio" box sits apart because it's a specialized branch, not just a bigger/smaller general model — you pick it for voice and streaming, not for raw reasoning.

In short: pick a point on this line per task, keep the id in config, and never assume today's model name is permanent.

RoleStrengthUse for
Flagship (e.g. gpt-5.5)most capable, deepest reasoninghard reasoning, agents, complex code
Mini variantsfast, cheaphigh-volume classification, routing, simple extraction
Realtime / audio (e.g. gpt-realtime-2, gpt-4o)low-latency, multimodalvoice agents, streaming, image/audio input
Don't hard-code the model name everywhereModel ids like gpt-5.5 are placeholders that OpenAI will supersede. Put the id in one config value and reference that everywhere, so upgrading the whole app is a one-line change — not a find-and-replace across your codebase.

3 · Tokens & the context window essential

Here's the unit that quietly runs your whole bill. The model doesn't read characters or words — it reads tokens, chunks of text roughly ¾ of a word on average ("tokenization" splits "unbelievable" into something like un, bel, iev, able). Every request is priced in tokens: what you send (input tokens) and what comes back (output tokens), and output usually costs several times more per token than input. If you only remember one thing for cost control, make it this: fewer tokens in and out means a smaller bill and a faster reply.

The second concept is the context window — the maximum number of tokens the model can "see" in a single request, input plus output combined. Think of it as the model's working desk: everything it can reason about has to fit on the desk at once. Modern GPT models have large windows, but they are not infinite, and two things go wrong when you forget this. First, a long conversation or a big pile of retrieved documents can simply overflow the window and the call fails. Second — subtler — stuffing the window full costs real money and can actually dilute the model's attention on what matters. More context is not automatically better.

The common mistake: treating the context window as free scratch space and dumping entire documents, whole chat histories, and verbose system prompts into every call "just in case." At scale that's the difference between a sustainable product and a runaway bill. The professional move is to send the least context that still answers the question — which is exactly why retrieval (RAG) exists: fetch only the relevant chunks instead of pasting the whole corpus.

Read your usage on every callThe Responses API returns token counts on resp.usage — resp.usage.input_tokens and resp.usage.output_tokens. Logging these per request is how you turn "the bill looks high" into "this feature is sending 8k tokens of context it doesn't need." Beginners ignore usage; professionals read it on every call.

4 · How the models differ in practice intermediate

On paper the family is a tidy ladder; in practice the differences show up as behavior under pressure. A mini model is perfectly capable of classifying a support ticket or extracting a date — the easy, well-specified work — and it will do it faster and cheaper than the flagship. Where the gap opens up is on tasks that need chains of reasoning: multi-step math, debugging a subtle logic error, planning an agent's next five moves, or holding many constraints in mind at once. There, the flagship's extra capacity turns into measurably fewer mistakes.

OpenAI also exposes a reasoning-effort dial on its reasoning-capable models: reasoning={"effort": "low"} through "high" (and "max"). This is the knob that lets one model span a range — spend little effort on an easy question (fast, cheap) and more on a hard one (slower, pricier, more careful). It's the OpenAI analog of asking the model to "think longer before answering," and like model choice itself, it's a cost/quality trade you should set deliberately per task, not leave at a reflexive maximum.

Effort is a dial, not a defaultClassification and routing → low. Most application work → medium. Hard reasoning, agents, tricky code → high or max. Paying for high effort on a trivial task is the same mistake as calling the flagship for everything — correctness that beats cost, nowhere else.

5 · Pick the right model per task advanced

This is the skill that separates a demo from a system. The decision has two knobs, not one: which model and how much reasoning effort. Treat them as a single question — "what's the cheapest (model, effort) pair whose evals clear this task's quality bar?" — and you'll make deliberate choices instead of defaulting to the biggest, slowest, most expensive option out of anxiety.

The method is the same one the Claude section teaches, because it's vendor-neutral: start from the middle, move with evidence. Begin with a mid or mini model at medium effort. Run your evals (Ch 5). If it's failing the hard cases, escalate — a bigger model, or more effort, whichever your evals say helps. If it's acing everything, downgrade — a smaller model or less effort — and bank the savings. The only thing that justifies a move up or down is a measured result, never a hunch.

Lab O1.1

Your first OpenAI call. The Responses API is the primary way to talk to GPT: give it instructions (the system role) and input (the user's message), and read the reply from resp.output_text.

first_call.pyfrom dotenv import load_dotenv
from openai import OpenAI
load_dotenv()
client = OpenAI()                      # reads OPENAI_API_KEY from the env

resp = client.responses.create(
    model="gpt-5.5",                 # one config value you can swap later
    instructions="You are a concise assistant.",  # the system role
    input="Explain what an API token is in two sentences.",
)
print(resp.output_text)                # the generated text
print(resp.usage.input_tokens, resp.usage.output_tokens)  # cost, every call
▶ How this works
  1. client = OpenAI() builds the client; it reads your key from OPENAI_API_KEY in the environment, so the key never appears in code.
  2. client.responses.create(...) is the single API call. model picks which GPT; instructions sets the assistant's role (the system message); input is the user's content.
  3. resp.output_text is the convenience property that gives you the reply as a plain string — no digging through a list of content blocks.
  4. resp.usage carries input_tokens and output_tokens — log them so cost is visible from day one.

Try this: add reasoning={"effort": "low"} then "high" on a genuinely hard prompt and compare the depth of the answer against the token counts in resp.usage. That quality-vs-cost trade is a theme of the whole course.

Same code, both vendorsCompare this with the Anthropic API lesson: the retrieval, prompting, and agent logic you build is identical — only the client and the call line differ (client.responses.create(...) → resp.output_text here, client.messages.create(...) → content blocks on Claude). That's exactly why so many pages in this course show the two SDKs on a toggle.

6 · Cost at scale professional

At one request a day, model choice is a rounding error. At a million requests a month, it's a budget line with your name on it. This is where the "pick by task" discipline pays for itself. The single most powerful pattern is tiered routing: send the easy majority of traffic to a cheap mini model and escalate only the genuinely hard minority to the flagship. If 90% of requests are easy and you route them to a model that costs a fraction of the flagship, you save most of your bill versus sending everything to the big model — and the savings are largest exactly when the easy majority is large.

The second lever is output length. Output tokens bill at the higher rate, so a verbose model that pads every answer with preamble costs more than a terse one that says only what's needed — for identical quality. Cap max_output_tokens, ask for concise answers in the instructions, and don't pay for politeness filler.

Lab O1.2

A tiny cost estimator makes the trade-off concrete. Rates are illustrative — plug in current OpenAI prices.

estimate.py# $ per 1M tokens (illustrative — check current pricing)
RATES = {
    "flagship": {"in": 5.0, "out": 15.0},
    "mini":     {"in": 0.4, "out": 1.6},
}

def monthly(model, reqs, in_tok, out_tok):
    r = RATES[model]
    return reqs * (in_tok*r["in"] + out_tok*r["out"]) / 1_000_000

# 1M requests: all-flagship vs route 90% to mini
all_big = monthly("flagship", 1_000_000, 800, 400)
routed  = monthly("mini", 900_000, 800, 400) \
        + monthly("flagship", 100_000, 800, 400)
print(f"all-flagship: ${all_big:,.0f}  routed: ${routed:,.0f}")
▶ How this works

The estimator prices a workload from its request volume and average tokens. The two scenarios — everything on the flagship vs routing 90% of traffic to a mini model — make the saving visible as a number, which is what gets a model-routing design approved. The exact figures depend on current pricing; the shape of the saving does not.

Try this: raise the hard-traffic fraction from 10% toward 50% and watch the routed cost climb toward all-flagship. Routing pays off most when the easy majority is large — which is the condition to check before you build it.

7 · A model policy for the team tech-lead

Model choice is an architecture decision, not a per-engineer whim. Left unmanaged, every developer picks the flagship "to be safe," and your bill is the sum of everyone's anxiety. A one-page team policy turns it into a deliberate, reviewable choice — the same way you'd review a database or a capacity decision.

Lab O1.3

Your task: Write a 5-line "Model Policy" your team would actually follow, covering: the default (model + effort), when to escalate, when to downgrade, how routing works at scale, and the cost guardrail.

💡 Hint: "with evidence" is the whole policy — every move off the default should point at an eval result, not an opinion.

Show solution

Team Model Policy

  • Default: a mid/mini model at reasoning={"effort":"medium"} for all new work, with the exact id in one shared config value.
  • Escalate to the flagship (or higher effort) only with evidence: move up when evals show the default failing the hard cases — not "to be safe."
  • Downgrade with evidence: drop to a smaller model or lower effort when evals show it's good enough for the volume.
  • Route by difficulty at scale: send the easy majority to a mini model, escalate only hard cases; a task fine on mini but run on the flagship can cost many times more for no quality gain.
  • Cost guardrails: every workload declares expected tokens/volume and an estimated monthly cost at design time, logs resp.usage in production, and alerts when actual spend deviates. Trim output length — it bills at the higher rate.

This operationalizes the chapter: default small, escalate or downgrade only on eval evidence, route by difficulty, and treat which-model-where as a deliberate, budgeted design choice.

🪜 Practice ladder beginner → industry

  1. Beginner: run Lab O1.1 and read off resp.output_text plus the two token counts from resp.usage.
  2. Easy: change the model config value and re-run — confirm nothing else in your code had to change.
  3. Core: add reasoning={"effort": …} and compare answer depth vs output tokens on an easy and a hard prompt.
  4. Stretch: extend the Lab O1.2 estimator to a three-way split (mini / flagship / realtime) and chart the cross-over point.
  5. Hard: write a route(prompt) function that sends short/simple prompts to a mini model and long/complex ones to the flagship, and log which tier each request hit.
  6. Industry: draft the one-page Model Policy from Lab O1.3 for a real workload you know, with concrete numbers and an alerting threshold.

✓ Checkpoint — you can move on when you can…

  • Explain what GPT is and how RLHF alignment shapes it.
  • Describe the OpenAI family by role and reason about tokens/context/cost.
  • Pick a (model, effort) pair per task and justify it with evals.
  • Set a team model policy and route by difficulty at scale.

Knowledge check check yourself

✓ Knowledge check

OpenAI does not publish fixed named tiers like small/medium/large. How should an engineer choose which GPT model to call?

Show answer
Choose by capability, speed, and cost for the specific task rather than by a tier name. OpenAI ships a flagship general model, faster and cheaper mini variants, and realtime/audio models; pick the smallest model whose evals clear the task's quality bar, and tune the reasoning-effort dial (low/medium/high/max) to buy more reasoning only when a task needs it. Keep the exact id in one config value so upgrades are a one-line change.
✓ Knowledge check

In the OpenAI Python SDK, which call is primary for a single request and how do you read the text and token usage?

Show answer
Use the Responses API as primary: client.responses.create(model="gpt-5.5", instructions=..., input=...) and read the generated text from resp.output_text. Token usage is on resp.usage.input_tokens and resp.usage.output_tokens. The older chat.completions endpoint (with resp.choices[0].message) still works but Responses is the current standard.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in