AI EngineeringZero to ProductionHome·About·What’s new·Contact
OpenAI API in Practice · Part 3

Files API, Vision & PDF with OpenAI

Text in, text out is only half the API. GPT can also see — read images and PDFs and answer questions about them. This chapter covers the two ways to get a file in front of the model (inline or via the Files API), how multimodal input is structured, and the practical limits.

⏱️ ~1.5 hours🧪 3 labs🎯 Beginner→Advanced

Learning objectives

  • Send an image to GPT inline (URL or base64) and ask about it.
  • Upload a file with the Files API and reference it by id.
  • Structure multimodal input — mixing text and image parts.
  • Reason about PDFs, size limits, and cost of visual tokens.
⚙️ To run this for realNeeds an OpenAI API key (OPENAI_API_KEY) + pip install openai, and a multimodal-capable model (e.g. gpt-5.5 / a gpt-4o-family model).

1 · Two ways to give the model a file essential

There are two doors, and picking the right one saves you grief. For a one-off image you pass it inline — a URL the model fetches, or base64 bytes embedded in the request. For anything you'll reuse, or anything large (a multi-page PDF, a file referenced across several calls), you upload it once with the Files API and then reference it by its file_id. Inline is simplest; the Files API avoids re-sending the same big payload on every call.

The mental model mirrors email: inline is pasting a screenshot straight into the message body; the Files API is uploading an attachment once and linking it. For a quick question about one picture, paste it. For a document your app will interrogate repeatedly, upload it and pass the handle.

The common mistake is base64-embedding a large PDF into every request — you pay to upload those bytes on each call and can blow the request-size limit. Upload once, reference by id.

one-off & small → inline · reused or large → Files API inline image URL or base64 in the request Files API files.create once → reference by file_id Paste vs attach. Inline (URL/base64) is simplest for a one-off small image; the Files API uploads once and references by file_id for large or reused files — avoiding re-sending bytes every call.
🗺️ How to read this diagram
  • The blue box is inline delivery — the image rides along in the request as a URL or base64.
  • The green box is the Files API — upload once, then every call references the lightweight file_id.
  • Choose by reuse and size: tiny one-off → inline; big or repeated → upload.

In short: inline for a quick look, Files API for anything you'll send more than once or that's large.

2 · Recipe 1 — an image inline essential

Multimodal input is just input structured as a list of content parts — some text, some image. You pass a message whose content mixes an input_text part and an input_image part.

Recipe 1
see_image.pyfrom openai import OpenAI
client = OpenAI()

resp = client.responses.create(
    model="gpt-5.5",
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": "What is in this image?"},
            {"type": "input_image",
             "image_url": "https://ai.studybydoing.in/chart.png"},
        ],
    }],
)
print(resp.output_text)
▶ How this works
  1. Instead of a plain string, input is a list of messages; the user message's content is a list of typed parts.
  2. An input_text part carries your question; an input_image part carries the picture (here an image_url the model fetches; you can also pass base64 bytes).
  3. The reply comes back on resp.output_text exactly like a text-only call — the multimodality is all in how you build input.

Try this: add a second input_image part and ask the model to compare the two images — content parts compose, so multi-image questions just work.

3 · Recipe 2 — upload with the Files API intermediate

For large or reused files, upload once and reference the returned id. The Files API is also how you supply a PDF for the model to read.

Recipe 2
upload_pdf.pyfrom openai import OpenAI
client = OpenAI()

# upload once — returns a file object with an id
f = client.files.create(file=open("report.pdf", "rb"), purpose="user_data")

resp = client.responses.create(
    model="gpt-5.5",
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": "Summarize this report."},
            {"type": "input_file", "file_id": f.id},  # reference by id
        ],
    }],
)
print(resp.output_text)
▶ How this works
  1. client.files.create(file=…, purpose=…) uploads the bytes once and returns a file object; f.id is the handle.
  2. An input_file content part references that file_id — no re-uploading on subsequent calls that reuse the same file.
  3. The Files API also backs other features: purpose="batch" for batch jobs, and vector stores for file_search.

Try this: list your uploads with client.files.list() and clean one up with client.files.delete(f.id) — uploaded files persist until you remove them.

The Files API is a shared building blockThe same client.files.create uploads JSONL for the Batch API (purpose="batch"), documents for vision/reading (purpose="user_data"), and files for retrieval. Learn it once; it recurs across the whole platform.

4 · PDFs, limits & the cost of pixels advanced

Vision isn't free, and images aren't tiny. An image is converted into a block of tokens before the model reasons over it, and a high-resolution image can be a surprising number of tokens — so a "quick look at a screenshot" can cost more than a paragraph of text. PDFs are read page by page; a long document is many pages of visual/text tokens. Two practical consequences: watch resp.usage on multimodal calls (they run higher than you'd guess), and downscale images to the smallest resolution that still answers the question.

There are also hard limits — maximum request size, maximum file size, supported formats — so a giant file must go through the Files API rather than inline, and extremely large documents may need chunking (process a few pages at a time) just like a text corpus.

Downscale before you sendA full-resolution phone photo can cost many times more tokens than a sensibly downscaled copy — for identical answer quality on most questions. Resize to the smallest resolution that preserves the detail you're asking about; your bill is counting pixels.

🪜 Practice ladder beginner → industry

  1. Beginner: run Recipe 1 against a public image URL and read the description.
  2. Easy: swap the URL for base64-encoded local bytes.
  3. Core: upload a PDF with Recipe 2 and summarize it; check resp.usage.
  4. Stretch: send two images in one request and ask the model to compare them.
  5. Hard: downscale an image to three resolutions and chart token cost vs answer quality.
  6. Industry: build a reusable uploader that caches file_ids so the same document is never uploaded twice.

✓ Checkpoint — you can move on when you can…

  • Send an image inline via URL and base64.
  • Upload a file and reference it by file_id.
  • Build multimodal input from text and image parts.
  • Reason about PDF handling, limits, and visual-token cost.

Knowledge check check yourself

✓ Knowledge check

When should you pass an image inline versus uploading it via the Files API?

Show answer
Pass inline (an input_image part with an image_url or base64) for a one-off, small image. Upload via client.files.create and reference by file_id (an input_file part) for large files, PDFs, or anything reused across calls — so you don't re-send the bytes every request or hit the request-size limit.
✓ Knowledge check

How is multimodal input structured, and why watch usage closely on vision calls?

Show answer
Multimodal input is a list of messages whose content is a list of typed parts — input_text for the question and input_image/input_file for the media. Watch resp.usage because images are converted into a block of tokens before reasoning, and high-resolution images can cost far more than you'd expect — downscaling to the smallest sufficient resolution saves real money.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in