Files API, Vision & PDF with OpenAI
Text in, text out is only half the API. GPT can also see — read images and PDFs and answer questions about them. This chapter covers the two ways to get a file in front of the model (inline or via the Files API), how multimodal input is structured, and the practical limits.
Learning objectives
- Send an image to GPT inline (URL or base64) and ask about it.
- Upload a file with the Files API and reference it by id.
- Structure multimodal
input— mixing text and image parts. - Reason about PDFs, size limits, and cost of visual tokens.
OPENAI_API_KEY) + pip install openai, and a multimodal-capable model (e.g. gpt-5.5 / a gpt-4o-family model).1 · Two ways to give the model a file essential
There are two doors, and picking the right one saves you grief. For a one-off image you pass it inline — a URL the model fetches, or base64 bytes embedded in the request. For anything you'll reuse, or anything large (a multi-page PDF, a file referenced across several calls), you upload it once with the Files API and then reference it by its file_id. Inline is simplest; the Files API avoids re-sending the same big payload on every call.
The mental model mirrors email: inline is pasting a screenshot straight into the message body; the Files API is uploading an attachment once and linking it. For a quick question about one picture, paste it. For a document your app will interrogate repeatedly, upload it and pass the handle.
The common mistake is base64-embedding a large PDF into every request — you pay to upload those bytes on each call and can blow the request-size limit. Upload once, reference by id.
file_id for large or reused files — avoiding re-sending bytes every call.
- The blue box is inline delivery — the image rides along in the request as a URL or base64.
- The green box is the Files API — upload once, then every call references the lightweight
file_id. - Choose by reuse and size: tiny one-off → inline; big or repeated → upload.
In short: inline for a quick look, Files API for anything you'll send more than once or that's large.
2 · Recipe 1 — an image inline essential
Multimodal input is just input structured as a list of content parts — some text, some image. You pass a message whose content mixes an input_text part and an input_image part.
see_image.pyfrom openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-5.5",
input=[{
"role": "user",
"content": [
{"type": "input_text", "text": "What is in this image?"},
{"type": "input_image",
"image_url": "https://ai.studybydoing.in/chart.png"},
],
}],
)
print(resp.output_text)
- Instead of a plain string,
inputis a list of messages; the user message'scontentis a list of typed parts. - An
input_textpart carries your question; aninput_imagepart carries the picture (here animage_urlthe model fetches; you can also pass base64 bytes). - The reply comes back on
resp.output_textexactly like a text-only call — the multimodality is all in how you buildinput.
Try this: add a second input_image part and ask the model to compare the two images — content parts compose, so multi-image questions just work.
3 · Recipe 2 — upload with the Files API intermediate
For large or reused files, upload once and reference the returned id. The Files API is also how you supply a PDF for the model to read.
upload_pdf.pyfrom openai import OpenAI
client = OpenAI()
# upload once — returns a file object with an id
f = client.files.create(file=open("report.pdf", "rb"), purpose="user_data")
resp = client.responses.create(
model="gpt-5.5",
input=[{
"role": "user",
"content": [
{"type": "input_text", "text": "Summarize this report."},
{"type": "input_file", "file_id": f.id}, # reference by id
],
}],
)
print(resp.output_text)
client.files.create(file=…, purpose=…)uploads the bytes once and returns a file object;f.idis the handle.- An
input_filecontent part references thatfile_id— no re-uploading on subsequent calls that reuse the same file. - The Files API also backs other features:
purpose="batch"for batch jobs, and vector stores for file_search.
Try this: list your uploads with client.files.list() and clean one up with client.files.delete(f.id) — uploaded files persist until you remove them.
client.files.create uploads JSONL for the Batch API (purpose="batch"), documents for vision/reading (purpose="user_data"), and files for retrieval. Learn it once; it recurs across the whole platform.4 · PDFs, limits & the cost of pixels advanced
Vision isn't free, and images aren't tiny. An image is converted into a block of tokens before the model reasons over it, and a high-resolution image can be a surprising number of tokens — so a "quick look at a screenshot" can cost more than a paragraph of text. PDFs are read page by page; a long document is many pages of visual/text tokens. Two practical consequences: watch resp.usage on multimodal calls (they run higher than you'd guess), and downscale images to the smallest resolution that still answers the question.
There are also hard limits — maximum request size, maximum file size, supported formats — so a giant file must go through the Files API rather than inline, and extremely large documents may need chunking (process a few pages at a time) just like a text corpus.
🪜 Practice ladder beginner → industry
- Beginner: run Recipe 1 against a public image URL and read the description.
- Easy: swap the URL for base64-encoded local bytes.
- Core: upload a PDF with Recipe 2 and summarize it; check
resp.usage. - Stretch: send two images in one request and ask the model to compare them.
- Hard: downscale an image to three resolutions and chart token cost vs answer quality.
- Industry: build a reusable uploader that caches
file_ids so the same document is never uploaded twice.
✓ Checkpoint — you can move on when you can…
- Send an image inline via URL and base64.
- Upload a file and reference it by
file_id. - Build multimodal
inputfrom text and image parts. - Reason about PDF handling, limits, and visual-token cost.
Knowledge check check yourself
When should you pass an image inline versus uploading it via the Files API?
Show answer
input_image part with an image_url or base64) for a one-off, small image. Upload via client.files.create and reference by file_id (an input_file part) for large files, PDFs, or anything reused across calls — so you don't re-send the bytes every request or hit the request-size limit.How is multimodal input structured, and why watch usage closely on vision calls?
Show answer
input is a list of messages whose content is a list of typed parts — input_text for the question and input_image/input_file for the media. Watch resp.usage because images are converted into a block of tokens before reasoning, and high-resolution images can cost far more than you'd expect — downscaling to the smallest sufficient resolution saves real money.