Prompting GPT & Migrating from Claude
Two practical skills in one chapter: how to prompt GPT models well (and where the reasoning models change the old advice), and — if you already have a Claude codebase — exactly what to change to run it on OpenAI. A line-by-line migration map plus the prompting habits that survive a vendor switch.
Learning objectives
- Apply prompting habits that work across models (clear role, examples, structure).
- Adjust for reasoning models — where "think step by step" is now redundant.
- Translate a Claude API call to the OpenAI Responses API, field by field.
- Plan a migration: what changes, what stays, and how to verify with evals.
1 · Prompting GPT well essential
Most good prompting is vendor-neutral, and that's the reassuring news. The habits that make Claude reliable make GPT reliable: give it a clear role and task, put stable rules in the system channel (instructions), show a few examples of the input→output you want, and ask for structure when a program will read the result. None of that changes when you switch the logo on the API. If your prompts are well-built, they port.
The GPT-specific adjustments are mostly about where text goes and which model you're on. Put rules in instructions, not stuffed into the user turn (ox2); keep the stable prefix first so caching fires (oap2); and prefer the schema-constrained responses.parse over asking for JSON in prose. These are the same instincts the course teaches throughout, applied to the OpenAI surface.
The common mistake is porting prompts that lean on Anthropic-specific scaffolding — heavy XML-tag structuring that Claude loves, or prompt tricks tuned to one model's quirks — and assuming they're optimal on GPT. They usually still work, but the right move after a migration is to re-run your evals and let measurement, not folklore, tell you what to adjust.
2 · Reasoning models change the old advice essential
One classic prompting trick is now often redundant: "let's think step by step." On older models, that chain-of-thought nudge measurably improved hard-problem accuracy. On modern reasoning models — which do hidden reasoning internally, dialed by reasoning={"effort": …} — telling them to think step by step is usually unnecessary and sometimes counterproductive; the reasoning is happening whether you ask or not. The modern lever is the effort dial, not the magic phrase.
So the updated habit: for a reasoning model, raise effort for hard tasks rather than appending reasoning incantations, and keep prompts clean and direct. Save explicit "show your work" instructions for when you actually want the reasoning surfaced in the output, not as a performance crutch.
reasoning={"effort":"high"} replaces "think step by step" as the way to buy more careful reasoning. The phrase isn't harmful, but the dial is the real control — and it's measurable, which the phrase never was.3 · The Claude → OpenAI migration map advanced
If you have a working Claude codebase, migration is mostly mechanical — a handful of renames, not a rewrite. The concepts line up one-to-one; the SDK surface differs. Here's the field-by-field map.
| Concept | Anthropic (Claude) | OpenAI (Responses) |
|---|---|---|
| Client | Anthropic() | OpenAI() |
| The call | client.messages.create(...) | client.responses.create(...) |
| System prompt | system=… | instructions=… |
| User content | messages=[{role,content}] | input=… (str or list) |
| Output cap | max_tokens | max_output_tokens |
| Read the text | loop resp.content text blocks | resp.output_text |
| Structured output | messages.parse(output_format=…) → .parsed_output | responses.parse(text_format=…) → .output_parsed |
| Tool definition | {name, input_schema} | {"type":"function", name, parameters} |
| Tool request | tool_use block, tool_use_id | function_call item, call_id |
| Tool result back | tool_result in a user message | function_call_output item |
| Loop control | stop_reason == "tool_use" | any function_call in resp.output |
| Reasoning | thinking + output_config.effort | reasoning={"effort":…} |
| Usage | usage.input_tokens / cache_read_input_tokens | usage.input_tokens / input_tokens_details.cached_tokens |
| Exceptions | anthropic.APIError | openai.APIError |
before_claude.pyfrom anthropic import Anthropic
client = Anthropic()
resp = client.messages.create(
model="claude-opus-4-8", max_tokens=512,
system="You are terse.",
messages=[{"role":"user","content":"Define idempotency."}],
)
text = next(b.text for b in resp.content if b.type=="text")
after_openai.pyfrom openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-5.5", max_output_tokens=512,
instructions="You are terse.",
input="Define idempotency.",
)
text = resp.output_text
- The two programs do the identical thing; every change is a rename from the map above — client, method,
system→instructions,messages→input,max_tokens→max_output_tokens. - The biggest readability win is reading the reply: Claude's content-block loop collapses to
resp.output_text. - Tool-using code is the one place with real structural change — the
tool_use/stop_reasonloop becomes thefunction_call/resp.outputloop (Ch 4 teaches both).
Try this: take one real Claude call from your code and port it using only the table — you'll find text calls are a 5-minute change; tool loops take a bit more care.
4 · Planning a migration professional
A migration isn't done when the code compiles — it's done when the evals pass. The plan: (1) swap the client + call surface using the map; (2) move system→instructions and restructure tool loops; (3) re-point the model id to a config value; and critically (4) re-run your eval suite (Ch 5) against the OpenAI version. Different model, different behavior — a prompt tuned on Claude may need adjustment on GPT, and only your evals will tell you where. Migrate behind a flag so you can run both and compare before cutting over.
5 · Tech-lead — vendor flexibility as architecture tech-lead
The strategic lesson of this whole track: treat the vendor as a boundary, not an assumption. Keep the model id and client construction in config (oap6), keep tool functions plain, keep prompts in a registry, and maintain an eval suite that runs against whichever vendor you point at. Do that and switching from Claude to OpenAI (or running both, routing by task) becomes a configuration and eval exercise — not a rewrite. In a market where the best model changes every few months, that flexibility is itself a design goal.
🪜 Practice ladder beginner → industry
- Beginner: rewrite the
before_claude.pytext call as OpenAI using only the map. - Easy: remove a "think step by step" line and raise
effortinstead; compare. - Core: port a
messages.parsestructured-output call toresponses.parse. - Stretch: port a tool-using loop —
tool_use/stop_reason→function_call/resp.output. - Hard: put a migration behind an env flag and run the same evals against both vendors.
- Industry: design the config + eval boundary that makes vendor choice a one-flag decision for a real app.
✓ Checkpoint — you can move on when you can…
- Apply vendor-neutral prompting habits to GPT.
- Explain why the effort dial replaces "think step by step" on reasoning models.
- Translate a Claude call to OpenAI field by field.
- Plan a migration that ends with passing evals, not just compiling code.
Knowledge check check yourself
Translate a Claude messages.create text call to OpenAI. Which fields change?
Show answer
Anthropic()→OpenAI(); client.messages.create→client.responses.create; system=→instructions=; messages=[{role,content}]→input= (string or list); max_tokens→max_output_tokens; and reading the reply changes from looping resp.content text blocks to just resp.output_text. Tool loops change more: tool_use/tool_use_id/stop_reason become function_call/call_id/checking resp.output.Why is "let's think step by step" now often redundant, and when is a migration actually complete?
Show answer
reasoning={"effort":…} dial, so the chain-of-thought phrase that helped older models is usually unnecessary — raise effort instead. A migration is complete not when the code compiles but when your eval suite passes on the new vendor: a different model behaves differently, so re-run evals (ideally behind a flag running both) and let measurement, not folklore, tell you what to adjust.