AI EngineeringZero to ProductionHome·About·What’s new·Contact
OpenAI API in Practice · Part 4

Token Counting & Usage with OpenAI

Tokens are the unit your whole bill is denominated in. This chapter shows how to count them before you call (with tiktoken, offline and free) and how to read exactly what you spent after (resp.usage) — so cost stops being a monthly surprise and becomes a number you design around.

⏱️ ~1 hour🧪 3 labs🎯 Beginner→Tech-lead

Learning objectives

  • Count tokens locally with tiktoken before sending a request.
  • Read input_tokens, output_tokens, and cached_tokens from a response.
  • Turn token counts into dollar estimates.
  • Log usage per request to build a cost-per-feature view.
⚙️ To run this for realThe tiktoken labs run fully offline (pip install tiktoken). The live-usage lab needs an OpenAI API key + pip install openai.

1 · Why tokens, not words essential

The model doesn't see words — it sees tokens, and so does your invoice. A token is a chunk of text, roughly ¾ of a word on average; "tokenization" splits text into these sub-word pieces using a fixed vocabulary. Every request is priced by tokens in two buckets — input (what you send) and output (what comes back) — and output almost always costs several times more per token than input. Internalize that and two instincts follow: shorten prompts, and cap output length.

There are two moments you care about token counts. Before the call, you want to estimate — will this prompt fit the context window, roughly what will it cost? That's a local calculation with tiktoken, no network, no key. After the call, you want the truth — exactly what was billed — which the API reports on resp.usage. Estimate up front, reconcile after.

estimate before · reconcile after before · tiktoken local · free · estimate after · resp.usage exact · billed truth Two moments, two tools. tiktoken estimates token counts locally before you spend anything; resp.usage reports the exact billed counts after the call.

2 · Counting locally with tiktoken essential

tiktoken is OpenAI's tokenizer library. You load the encoding for your model and count — instantly, offline, for free. This is how you check a prompt fits before paying to find out it doesn't.

Lab OP4.1
count.pyimport tiktoken

enc = tiktoken.encoding_for_model("gpt-4o")   # encoding for the model family

def n_tokens(text: str) -> int:
    return len(enc.encode(text))

prompt = "Summarize the quarterly report in three bullets."
print(n_tokens(prompt))                        # e.g. 9
print(enc.encode("unbelievable"))             # one word → several tokens
▶ How this works
  1. tiktoken.encoding_for_model(...) returns the tokenizer that model uses — the same splitting the API applies.
  2. enc.encode(text) turns text into the list of integer token ids; its length is the token count.
  3. Encoding "unbelievable" shows the sub-word reality — one word becomes several tokens, which is why word-count is a bad proxy for cost.

Try this: count a long document before sending it — if it's near the context window, you know to chunk or summarize before spending a cent.

Estimate ≠ exact, but closetiktoken counts the text tokens precisely, but the billed total also includes a little message/role overhead and (for multimodal) image tokens. Use it to estimate and to guard the context window; use resp.usage for the exact bill.

3 · Reading resp.usage — the billed truth intermediate

Every response reports what it actually cost. These are the numbers to log on every call.

Lab OP4.2
usage.pyfrom openai import OpenAI
client = OpenAI()

resp = client.responses.create(model="gpt-5.5", input="Hello!")
u = resp.usage
print("input:  ", u.input_tokens)
print("output: ", u.output_tokens)
print("total:  ", u.total_tokens)
print("cached: ", u.input_tokens_details.cached_tokens)   # reused prefix
▶ How this works
  1. resp.usage.input_tokens / output_tokens are the billed counts for the two buckets; total_tokens is their sum.
  2. resp.usage.input_tokens_details.cached_tokens shows how many input tokens were served from cache — your discount, made visible.
  3. These field names are the OpenAI shape — note it's input_tokens/output_tokens, not the legacy prompt_tokens/completion_tokens from chat.completions.

Try this: make the same call twice behind a big stable prefix and watch cached_tokens climb on the second — the usage object is where caching proves itself.

4 · Tokens → dollars intermediate

A tiny pricing function turns counts into money. Keep rates as inputs (they change), and remember output bills higher than input.

Lab OP4.3 · offline
price.pydef dollars(usage, in_rate, out_rate):
    # rates are $ per 1M tokens
    return (usage.input_tokens*in_rate + usage.output_tokens*out_rate) / 1_000_000

# works with a tiny stand-in too (offline):
class U: input_tokens, output_tokens = 1200, 350
print(f"${dollars(U, 5.0, 15.0):.4f}")

5 · A per-feature cost ledger professional

"The OpenAI bill went up" is useless; "the summarizer feature's cost tripled on Tuesday" is actionable. The move is to tag every call with a feature label and accumulate usage per label. Then a monthly report tells you which feature spends what — and a sudden jump points at the exact code path.

Pattern · offline-runnable core
ledger.pyfrom collections import defaultdict
ledger = defaultdict(lambda: {"in": 0, "out": 0})

def record(feature, usage):
    ledger[feature]["in"]  += usage.input_tokens
    ledger[feature]["out"] += usage.output_tokens

# after each real call: record("summarizer", resp.usage)
# monthly: price each feature's totals and sort — biggest spender first
Log usage from day oneWrap your responses.create in a helper that logs trace_id, feature label, model, and the three usage numbers as one JSON line per call. That one habit turns cost debugging from archaeology into a database query — the same observability pattern as the production chapter.

6 · Tech-lead — budgets & guardrails tech-lead

At the lead level, token discipline becomes policy: every workload declares an expected token profile and a monthly cost estimate at design time (via tiktoken), production logs actual usage per feature, and an alert fires when actual spend diverges from the estimate. Cap max_output_tokens globally so a runaway generation can't produce a surprise bill, and treat a feature's cost-per-request as a reviewable number like latency or error rate.

🪜 Practice ladder beginner → industry

  1. Beginner: count tokens in three strings with Lab OP4.1.
  2. Easy: read resp.usage from a live call and compare to your tiktoken estimate.
  3. Core: price a call with Lab OP4.3 at input and output rates.
  4. Stretch: wrap responses.create to log usage as one JSON line per call.
  5. Hard: build the per-feature ledger and print a sorted monthly cost report.
  6. Industry: add a design-time token estimate + a production spend-vs-estimate alert for one real workload.

✓ Checkpoint — you can move on when you can…

  • Count tokens locally with tiktoken before a call.
  • Read input/output/cached tokens from resp.usage.
  • Convert counts to dollars with model rates.
  • Log usage per feature and reason about budgets.

Knowledge check check yourself

✓ Knowledge check

What's the difference between counting with tiktoken and reading resp.usage, and when do you use each?

Show answer
tiktoken counts text tokens locally and for free — use it before a call to estimate cost and check the prompt fits the context window. resp.usage reports the exact billed input_tokens/output_tokens/total_tokens (plus input_tokens_details.cached_tokens) after the call. The billed total can exceed tiktoken's text count due to message overhead and image tokens, so estimate with tiktoken and reconcile with usage.
✓ Knowledge check

Why tag calls with a feature label, and which usage field names does the OpenAI Responses API use?

Show answer
Tagging each call with a feature label and accumulating usage per label lets a monthly report show which feature spends what, so a cost jump points at a specific code path instead of "the bill went up." The Responses API reports resp.usage.input_tokens and resp.usage.output_tokens (not the legacy prompt_tokens/completion_tokens), with cached reuse under input_tokens_details.cached_tokens.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in