← Blog

Claude Vision API: Images, PDFs, and Multimodal Inputs

September 20, 2026

How to send images, PDFs, and documents to the Claude API: accepted formats, base64-vs-URL, document blocks, token costs, and CCA-F exam concepts.

Claude accepts images, PDFs, and plain-text documents alongside text in the messages array. Supported image formats are JPEG, PNG, GIF, and WebP — passed as base64 or a publicly accessible URL. PDFs are passed as document blocks. Token cost scales with pixel dimensions; a 1,000 × 1,000 px image costs roughly 1,500–2,000 input tokens.

What Can Claude See? Multimodal Capabilities at a Glance

Every current Claude model — Opus 4.8, Sonnet 4.6, and Haiku 4.5 — accepts mixed content in the messages array. A single user turn can combine images, PDFs, plain-text documents, and instructions. You can send a chart, a specification PDF, and a question in one request and receive a single coherent response.

Multimodal inputs arrive as content blocks inside the content array of a message. The type field on each block determines how the model processes it: "image" for visual content, "document" for PDFs and plain text, and "text" for the prompt itself.

Which Image Formats Does Claude Accept?

Claude accepts four image formats:

  • JPEG (image/jpeg) — compressed photographs and general imagery
  • PNG (image/png) — lossless images, screenshots, and diagrams
  • GIF (image/gif) — static or animated images; for animated GIFs, only the first frame is analyzed
  • WebP (image/webp) — modern compressed format with broad browser support

No other formats are natively supported. If your pipeline produces TIFF, BMP, or SVG files, convert them to one of the four before sending. The media_type field in the source block must match the actual image format — a mismatch returns an error.

How Do You Pass an Image to the API: base64 vs. URL?

Two source types are available for images. Choosing between them depends on where your image lives and whether it is publicly reachable.

Base64 — for local or private images

Encode the raw image bytes in base64 and include them directly in the request payload. This works for any image the application can read, regardless of access control. The tradeoff is payload size: a 1 MB image adds roughly 1.3 MB to the request body after encoding.

import anthropic, base64

client = anthropic.Anthropic()

with open("chart.png", "rb") as f:
    image_data = base64.standard_b64encode(f.read()).decode("utf-8")

response = client.messages.create(
    model="claude-opus-4-8",
    max_tokens=1024,
    messages=[{
        "role": "user",
        "content": [
            {
                "type": "image",
                "source": {
                    "type": "base64",
                    "media_type": "image/png",
                    "data": image_data
                }
            },
            {"type": "text", "text": "Describe the trend shown in this chart."}
        ]
    }]
)
print(response.content[0].text)

URL — for publicly accessible images

Pass the image URL directly. Anthropic fetches the image server-side. The URL must be publicly accessible — authenticated URLs and private bucket links will fail.

response = client.messages.create(
    model="claude-opus-4-8",
    max_tokens=1024,
    messages=[{
        "role": "user",
        "content": [
            {
                "type": "image",
                "source": {
                    "type": "url",
                    "url": "https://example.com/architecture-diagram.png"
                }
            },
            {"type": "text", "text": "Identify the components in this diagram."}
        ]
    }]
)

Use base64 for private or generated images; use URL for images already served from a public CDN or storage bucket. URL requests carry lighter payloads but add a network dependency on the remote host.

Can Claude Read PDFs and Other Documents?

PDFs use a "document" content block rather than an "image" block. The source structure is the same as for images — type, media_type, and data for base64, or type and url for a hosted file — but the outer block type changes to "document" and the media type becomes "application/pdf".

import base64

with open("spec.pdf", "rb") as f:
    pdf_data = base64.standard_b64encode(f.read()).decode("utf-8")

response = client.messages.create(
    model="claude-opus-4-8",
    max_tokens=2048,
    messages=[{
        "role": "user",
        "content": [
            {
                "type": "document",
                "source": {
                    "type": "base64",
                    "media_type": "application/pdf",
                    "data": pdf_data
                }
            },
            {"type": "text", "text": "Summarize the key requirements."}
        ]
    }]
)

The document block also accepts plain-text content via "type": "text" in the source — useful for passing large text files without wrapping them as inline text blocks. For documents you query repeatedly across many sessions, the Files API lets you upload once and reference by file_id, avoiding repeated base64 encoding on every call.

How Much Do Images and Documents Cost in Tokens?

Image token cost is driven by pixel dimensions, not file size. A 1,000 × 1,000 pixel image costs roughly 1,500–2,000 input tokens on standard models. Larger images cost proportionally more.

On Opus 4.7 and Opus 4.8, which support high-resolution vision up to 2,576 pixels on the long edge, a full-resolution image can reach approximately 4,784 tokens — roughly three times the cost on earlier models. For cost-sensitive pipelines, sending images at 1,080p is a practical tradeoff that preserves most of the accuracy gain while reducing token spend.

PDF and plain-text document blocks are billed by the token count of their extracted text content, similar to long text blocks. Use the token-counting endpoint (client.messages.count_tokens()) to estimate costs before sending large documents.

Three levers for controlling multimodal token spend:

  • Resize before encoding — halving linear dimensions reduces token cost by roughly 75%
  • Use the Files API for documents queried across many requests
  • Use prompt caching on a large static document — pay the cache-write cost once and read cheaply on subsequent requests

When Should You Use Vision vs. Text-Only Inputs?

Use vision when the information is inherently visual: charts, screenshots, scanned documents, photographs, and UI mockups. Use text-only when the content can be extracted before the API call — if you can run OCR, parse a CSV, or extract text from a programmatically generated PDF, sending plain text is cheaper and more predictable.

The deciding factor is whether extracting the text is practical in your pipeline. Multimodal adds token cost; it makes sense when the visual channel carries information that text extraction would lose or distort.

What Does the Claude Certified Architect Exam Test About Multimodal?

The Claude Certified Architect Foundations exam covers multimodal inputs as part of its API and developer usage domains. Expect questions on:

  • Accepted image formats — JPEG, PNG, GIF, WebP; no others are natively supported
  • Image source types — the distinction between "base64" (any accessible image) and "url" (publicly reachable only)
  • Document vs. image blocks — PDFs use "document" type, not "image"
  • Token cost model — costs scale with pixel dimensions, not file size; high-resolution vision on Opus 4.7 and 4.8 increases per-image cost significantly
  • When to use multimodal — scenarios where text extraction is impractical or lossy

The CCA-F tends to test structural precision: which fields are required, what values are valid, and what happens when you use the wrong block type. Knowing that media_type lives inside source, and that document and image blocks have distinct structural requirements, is the kind of detail the exam rewards.

This post explains the mechanics of multimodal inputs. What it cannot give you is practice applying them under exam conditions. PlinthPrep is an independent study resource — not affiliated with Anthropic — with practice questions aligned to the multimodal section of the Claude Certified Architect Foundations exam. Visit plinthprep.com to find out whether you can answer these questions correctly, not just recognise the concepts.

Frequently asked questions

What image formats does Claude accept?
Claude accepts JPEG, PNG, GIF, and WebP image files. Images can be passed to the API either as base64-encoded data with the matching media type (image/jpeg, image/png, etc.) or as a publicly accessible URL. All four formats work with both the base64 and URL source types.
How do you pass an image to the Claude API?
Include an image block in the content array of a user message. Set type to "image" and provide a source object with type "base64" (including media_type and data) or type "url" (including the URL string). Add a text block in the same content array with your prompt or question about the image.
How much does an image cost in Claude API tokens?
Image token cost scales with pixel dimensions. A 1,000 × 1,000 pixel image costs roughly 1,500–2,000 input tokens on standard models. High-resolution images on Opus 4.7 and later can reach approximately 4,784 tokens at full resolution. Resizing images before sending is the most effective way to control vision costs.
How do you pass a PDF to the Claude API?
PDFs are sent as document blocks rather than image blocks. Set the content block type to "document" and provide a source with type "base64" and media_type "application/pdf", or use a URL source pointing to a publicly accessible PDF. You can include both a document block and a text prompt in the same user message.
What does the CCA-F exam test about multimodal inputs?
The Claude Certified Architect Foundations exam tests knowledge of supported input types (images, PDFs, plain text), the structure of image and document content blocks, the difference between base64 and URL image sources, how token costs scale with image resolution, and when to choose multimodal over text-only approaches.

Share this post

Plinth Prep is an independent study resource and is not affiliated with, endorsed by, or sponsored by Anthropic. Practice material is written by Plinth Prep and does not reproduce real exam content.