AI EngineeringZero to ProductionHome·About·What’s new·Contact
OpenAI API in Practice · Part 2

Prompt Caching with OpenAI

OpenAI caches long, repeated prompt prefixes automatically — no breakpoints to place, no flags to set. This chapter shows how it works, how to prove you're getting hits via the usage fields, what the discount is worth, and — crucially — how to structure prompts so the cache actually fires.

⏱️ ~1 hour🧪 2 labs🎯 Beginner→Tech-lead

Learning objectives

  • Explain what prompt caching is and why OpenAI's is automatic.
  • Read cached_tokens from usage to prove a cache hit.
  • Quantify the cost and latency saving.
  • Structure prompts (stable prefix first) so the cache fires — and avoid silent cache-killers.
⚙️ To run this for realNeeds an OpenAI API key (OPENAI_API_KEY) + pip install openai. The cost-math lab runs offline with no key.

1 · What prompt caching is essential

Most of your prompt is the same every single call. A big system instruction, a few-shot block, a tool catalog, a long document the user keeps asking about — that prefix is identical request after request, yet naively you pay to process it from scratch every time. Prompt caching is the provider recognizing "I've seen this exact prefix before" and reusing the computation, so you're billed far less for the repeated part.

Here's the headline difference from Anthropic, and it's the whole reason this page is shorter than its Claude counterpart: OpenAI caching is automatic. On Claude you explicitly mark a cache breakpoint with cache_control on a content block. On OpenAI you do nothing — the platform automatically caches long prompt prefixes (above an internal length threshold) and reuses them on subsequent calls that start with the same bytes. There's no API surface to learn; there's a discipline to learn (keep the prefix stable), and a metric to watch (cached_tokens).

The common mistake flows directly from "automatic": people assume automatic means "nothing I do matters," then scatter volatile content (a timestamp, a per-request id, a shuffled tool order) near the front of the prompt and silently get zero cache hits. Automatic caching still only works on an unchanged prefix — so where you put your volatile content is entirely in your hands.

keep the stable prefix first — only the volatile tail is re-billed stable prefix (cached) system · few-shot · tools · long doc volatile tail this request's question cached_tokens reports how much of the prefix was reused Prefix cached, tail fresh. Automatic caching reuses the unchanged leading bytes of your prompt; only the volatile tail (the per-request question) is processed anew. Put stable content first and the cache does the rest.
🗺️ How to read this diagram
  • The green block is the stable prefix — the bytes identical on every call, which the platform caches.
  • The amber block is the volatile tail — this request's unique question — which is always freshly processed.
  • The lesson is positional: anything volatile must sit after everything stable, or it poisons the prefix and kills the hit.

In short: stable first, volatile last — that ordering is the entire skill, since the caching itself is automatic.

2 · Reading the usage fields — proof of a hit essential

"I added caching" and "I have caching" are different claims — the usage fields settle it. Every response reports how many input tokens were served from cache. Run the same large-prefix request twice; on the second call cached_tokens should be large. If it stays zero, something in your prefix is changing between calls.

Lab OP2.1
prove_cache.pyfrom openai import OpenAI
client = OpenAI()

BIG_STABLE_INSTRUCTIONS = "...a long, identical system prompt..."  # same bytes every call

def ask(question):
    resp = client.responses.create(
        model="gpt-5.5",
        instructions=BIG_STABLE_INSTRUCTIONS,   # stable prefix
        input=question,                          # volatile tail
    )
    u = resp.usage
    print("input:", u.input_tokens,
          "cached:", u.input_tokens_details.cached_tokens)  # >0 on 2nd+ call = win
    return resp.output_text

ask("First question")      # cached: 0  (nothing cached yet)
ask("Second question")     # cached: large (prefix reused)
▶ How this works
  1. The instructions are a big, byte-for-byte identical prefix; only input changes per call.
  2. resp.usage.input_tokens_details.cached_tokens is the money metric — how many input tokens were served from cache. (Contrast Claude's cache_read_input_tokens.)
  3. First call caches nothing (cached_tokens == 0); the second identical-prefix call reuses it, so cached_tokens jumps.

Try this: inject a datetime.now() string into BIG_STABLE_INSTRUCTIONS and re-run — cached_tokens drops to 0. That's a silent cache-killer caught in the act.

3 · The cost math (runs offline) intermediate

Cached input tokens bill at a steep discount. This offline estimator shows the saving on a workload that re-sends a big stable prompt many times.

Offline
cache_savings.pyIN_RATE = 5.0          # $/1M input tokens (illustrative)
CACHED_RATE = 0.5      # cached input is much cheaper

def monthly(reqs, prefix_tok, tail_tok, cached_frac):
    cached = prefix_tok * cached_frac
    fresh = prefix_tok * (1 - cached_frac) + tail_tok
    per = (cached*CACHED_RATE + fresh*IN_RATE) / 1_000_000
    return reqs * per

# 100k calls, 4k stable prefix, 200-token tail
cold = monthly(100_000, 4000, 200, cached_frac=0.0)
warm = monthly(100_000, 4000, 200, cached_frac=0.9)
print(f"no cache ${cold:,.0f}  90% cached ${warm:,.0f}")

4 · What invalidates the cache advanced

Because caching keys on the unchanged leading bytes of the prompt, anything that perturbs the prefix drops your hit rate to zero — with no error to warn you. The usual culprits:

Silent cache-killers
  • A datetime.now(), UUID, or request id placed in the system/instruction prefix
  • Unsorted json.dumps() of config/tools (add sort_keys=True)
  • A per-user or reordered tool catalog near the front
  • Conditional prompt sections that vary request to request
Rule: stable content first; volatile content (timestamps, ids, the user's question) last.

This is the same engineering rule as Claude's prompt caching — the difference is only that OpenAI decides when to cache for you, so your sole job is to not break the prefix.

5 · Tech-lead — a caching strategy that gets hits tech-lead

A caching win isn't something you turn on; it's something you architect your prompt layout for. The lead-level move is to standardize prompt construction across the team so the stable prefix really is stable: a shared builder that always emits [fixed system] + [fixed few-shot] + [fixed tool catalog] + [volatile user turn] in that order, with the volatile content strictly last. Add a lightweight guard that logs cached_tokens / input_tokens as a cache-hit-rate metric and alerts when it drops — a sudden fall to zero usually means someone slipped a timestamp into the prefix.

Pair this with the Batch API for offline work and the usage logging from token counting & usage, and cost control stops being guesswork.

🪜 Practice ladder beginner → industry

  1. Beginner: run Lab OP2.1 twice and watch cached_tokens go 0 → large.
  2. Easy: shrink the prefix below the caching threshold and confirm hits stop.
  3. Core: inject a timestamp into the prefix and prove it kills the cache.
  4. Stretch: run the offline estimator at 0%, 50%, 90% cached and chart the curve.
  5. Hard: write a prompt builder that guarantees stable-first ordering and returns the assembled prompt.
  6. Industry: add a cache-hit-rate metric (cached_tokens/input_tokens) with an alert threshold to a real call path.

✓ Checkpoint — you can move on when you can…

  • Explain why OpenAI caching is automatic and what you still control.
  • Prove a cache hit with cached_tokens.
  • Name the common silent cache-killers and the stable-first rule.
  • Design a prompt layout and metric that keep hit rates high.

Knowledge check check yourself

✓ Knowledge check

How does OpenAI prompt caching differ from Anthropic's, and what do you still have to get right?

Show answer
OpenAI caches long prompt prefixes automatically — there's no cache_control breakpoint to place like on Claude. What you still control is prefix stability: caching only reuses unchanged leading bytes, so you must put all stable content (system, few-shot, tools, long docs) first and all volatile content (timestamps, ids, the user's question) last, or the hit rate collapses.
✓ Knowledge check

Which usage field proves a cache hit, and what's a common silent cache-killer?

Show answer
resp.usage.input_tokens_details.cached_tokens reports how many input tokens were served from cache — large on a repeated-prefix call, zero if nothing matched. A classic silent killer is putting a datetime.now() or a per-request UUID in the system/instruction prefix, which changes the leading bytes and drops the hit rate to zero with no error.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in