Prompt Caching with OpenAI
OpenAI caches long, repeated prompt prefixes automatically — no breakpoints to place, no flags to set. This chapter shows how it works, how to prove you're getting hits via the usage fields, what the discount is worth, and — crucially — how to structure prompts so the cache actually fires.
Learning objectives
- Explain what prompt caching is and why OpenAI's is automatic.
- Read
cached_tokensfrom usage to prove a cache hit. - Quantify the cost and latency saving.
- Structure prompts (stable prefix first) so the cache fires — and avoid silent cache-killers.
OPENAI_API_KEY) + pip install openai. The cost-math lab runs offline with no key.1 · What prompt caching is essential
Most of your prompt is the same every single call. A big system instruction, a few-shot block, a tool catalog, a long document the user keeps asking about — that prefix is identical request after request, yet naively you pay to process it from scratch every time. Prompt caching is the provider recognizing "I've seen this exact prefix before" and reusing the computation, so you're billed far less for the repeated part.
Here's the headline difference from Anthropic, and it's the whole reason this page is shorter than its Claude counterpart: OpenAI caching is automatic. On Claude you explicitly mark a cache breakpoint with cache_control on a content block. On OpenAI you do nothing — the platform automatically caches long prompt prefixes (above an internal length threshold) and reuses them on subsequent calls that start with the same bytes. There's no API surface to learn; there's a discipline to learn (keep the prefix stable), and a metric to watch (cached_tokens).
The common mistake flows directly from "automatic": people assume automatic means "nothing I do matters," then scatter volatile content (a timestamp, a per-request id, a shuffled tool order) near the front of the prompt and silently get zero cache hits. Automatic caching still only works on an unchanged prefix — so where you put your volatile content is entirely in your hands.
- The green block is the stable prefix — the bytes identical on every call, which the platform caches.
- The amber block is the volatile tail — this request's unique question — which is always freshly processed.
- The lesson is positional: anything volatile must sit after everything stable, or it poisons the prefix and kills the hit.
In short: stable first, volatile last — that ordering is the entire skill, since the caching itself is automatic.
2 · Reading the usage fields — proof of a hit essential
"I added caching" and "I have caching" are different claims — the usage fields settle it. Every response reports how many input tokens were served from cache. Run the same large-prefix request twice; on the second call cached_tokens should be large. If it stays zero, something in your prefix is changing between calls.
prove_cache.pyfrom openai import OpenAI
client = OpenAI()
BIG_STABLE_INSTRUCTIONS = "...a long, identical system prompt..." # same bytes every call
def ask(question):
resp = client.responses.create(
model="gpt-5.5",
instructions=BIG_STABLE_INSTRUCTIONS, # stable prefix
input=question, # volatile tail
)
u = resp.usage
print("input:", u.input_tokens,
"cached:", u.input_tokens_details.cached_tokens) # >0 on 2nd+ call = win
return resp.output_text
ask("First question") # cached: 0 (nothing cached yet)
ask("Second question") # cached: large (prefix reused)
- The
instructionsare a big, byte-for-byte identical prefix; onlyinputchanges per call. resp.usage.input_tokens_details.cached_tokensis the money metric — how many input tokens were served from cache. (Contrast Claude'scache_read_input_tokens.)- First call caches nothing (
cached_tokens == 0); the second identical-prefix call reuses it, socached_tokensjumps.
Try this: inject a datetime.now() string into BIG_STABLE_INSTRUCTIONS and re-run — cached_tokens drops to 0. That's a silent cache-killer caught in the act.
3 · The cost math (runs offline) intermediate
Cached input tokens bill at a steep discount. This offline estimator shows the saving on a workload that re-sends a big stable prompt many times.
cache_savings.pyIN_RATE = 5.0 # $/1M input tokens (illustrative)
CACHED_RATE = 0.5 # cached input is much cheaper
def monthly(reqs, prefix_tok, tail_tok, cached_frac):
cached = prefix_tok * cached_frac
fresh = prefix_tok * (1 - cached_frac) + tail_tok
per = (cached*CACHED_RATE + fresh*IN_RATE) / 1_000_000
return reqs * per
# 100k calls, 4k stable prefix, 200-token tail
cold = monthly(100_000, 4000, 200, cached_frac=0.0)
warm = monthly(100_000, 4000, 200, cached_frac=0.9)
print(f"no cache ${cold:,.0f} 90% cached ${warm:,.0f}")
4 · What invalidates the cache advanced
Because caching keys on the unchanged leading bytes of the prompt, anything that perturbs the prefix drops your hit rate to zero — with no error to warn you. The usual culprits:
- A
datetime.now(), UUID, or request id placed in the system/instruction prefix - Unsorted
json.dumps()of config/tools (addsort_keys=True) - A per-user or reordered tool catalog near the front
- Conditional prompt sections that vary request to request
This is the same engineering rule as Claude's prompt caching — the difference is only that OpenAI decides when to cache for you, so your sole job is to not break the prefix.
5 · Tech-lead — a caching strategy that gets hits tech-lead
A caching win isn't something you turn on; it's something you architect your prompt layout for. The lead-level move is to standardize prompt construction across the team so the stable prefix really is stable: a shared builder that always emits [fixed system] + [fixed few-shot] + [fixed tool catalog] + [volatile user turn] in that order, with the volatile content strictly last. Add a lightweight guard that logs cached_tokens / input_tokens as a cache-hit-rate metric and alerts when it drops — a sudden fall to zero usually means someone slipped a timestamp into the prefix.
Pair this with the Batch API for offline work and the usage logging from token counting & usage, and cost control stops being guesswork.
🪜 Practice ladder beginner → industry
- Beginner: run Lab OP2.1 twice and watch
cached_tokensgo 0 → large. - Easy: shrink the prefix below the caching threshold and confirm hits stop.
- Core: inject a timestamp into the prefix and prove it kills the cache.
- Stretch: run the offline estimator at 0%, 50%, 90% cached and chart the curve.
- Hard: write a prompt builder that guarantees stable-first ordering and returns the assembled prompt.
- Industry: add a cache-hit-rate metric (
cached_tokens/input_tokens) with an alert threshold to a real call path.
✓ Checkpoint — you can move on when you can…
- Explain why OpenAI caching is automatic and what you still control.
- Prove a cache hit with
cached_tokens. - Name the common silent cache-killers and the stable-first rule.
- Design a prompt layout and metric that keep hit rates high.
Knowledge check check yourself
How does OpenAI prompt caching differ from Anthropic's, and what do you still have to get right?
Show answer
cache_control breakpoint to place like on Claude. What you still control is prefix stability: caching only reuses unchanged leading bytes, so you must put all stable content (system, few-shot, tools, long docs) first and all volatile content (timestamps, ids, the user's question) last, or the hit rate collapses.Which usage field proves a cache hit, and what's a common silent cache-killer?
Show answer
resp.usage.input_tokens_details.cached_tokens reports how many input tokens were served from cache — large on a repeated-prefix call, zero if nothing matched. A classic silent killer is putting a datetime.now() or a per-request UUID in the system/instruction prefix, which changes the leading bytes and drops the hit rate to zero with no error.