Claude Message Batches API: Async Bulk Processing for CCA-F
September 25, 2026
The Claude Message Batches API runs up to 100,000 requests asynchronously at 50% of standard pricing — the right surface when throughput beats latency for CCA-F and production workloads.
The Claude Message Batches API processes up to 100,000 requests in a single asynchronous job at half the standard per-token price, with results available within 24 hours. Use it when throughput matters more than latency — bulk evaluations, dataset annotation, and overnight report generation are the canonical fits.
What Is the Claude Message Batches API?
The Claude Message Batches API is an endpoint (POST /v1/messages/batches) that accepts a list of Messages API requests, processes them asynchronously, and stores results for retrieval once the job finishes. It is designed for workloads where you have many independent requests to run and can afford to wait minutes or hours rather than seconds.
Every request in the batch is a full Messages API call — same model selection, same system prompts, same tool definitions — submitted as part of a larger job rather than fired individually. The Batches API handles queuing, execution, and result storage on your behalf.
How Does Batch Processing Differ from a Standard Messages Call?
A standard client.messages.create() call blocks until a response arrives, typically within seconds. That synchronous contract is ideal for interactive features: a chatbot reply, a real-time classification, a streaming answer appearing in a UI.
Batch processing inverts the contract. You submit a job and receive a batch ID. The API processes requests in the background and you check back by polling. When the job's processing_status reaches "ended", you retrieve results as a stream of per-request outcomes.
The practical consequence is that batch processing is not appropriate anywhere a user is waiting. It is appropriate for any workflow that can be structured as "submit now, collect results later" — nightly pipelines, offline scoring jobs, large-scale annotation tasks, or evaluation harnesses running against hundreds of test cases.
What Are the Cost and Throughput Trade-offs?
The Batches API applies a 50 percent discount on all token usage — both input and output tokens, across every model available through the standard Messages API. The discount is flat, not tiered. As a reference using pricing cached at the time of writing: Claude Sonnet 4.6 at $3.00 per million input tokens and $15.00 per million output tokens becomes $1.50 and $7.50 per million under batch pricing. Verify current figures at Anthropic's pricing page before building cost models.
The throughput ceiling is high. A single batch can contain up to 100,000 requests or 256 MB of data, whichever limit is reached first. Most batches complete within one hour; the hard maximum is 24 hours. Results are stored for retrieval for 29 days after creation.
The cost you give up is latency. If a workload is time-sensitive, the Batches API is the wrong tool. If latency is irrelevant and throughput or cost matters, it is the right one.
How Do You Submit, Poll, and Retrieve a Batch Job?
The three-step lifecycle is: create the batch, poll for completion, stream the results.
Submit
Pass a list of request objects to client.messages.batches.create(). Each object requires a custom_id — your identifier returned with the result — and a params object that mirrors a standard messages.create() call:
import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request
client = anthropic.Anthropic()
batch = client.messages.batches.create(
requests=[
Request(
custom_id="classify-001",
params=MessageCreateParamsNonStreaming(
model="claude-haiku-4-5",
max_tokens=50,
messages=[{
"role": "user",
"content": "Classify as positive/negative/neutral: 'Great product!'"
}]
)
),
Request(
custom_id="classify-002",
params=MessageCreateParamsNonStreaming(
model="claude-haiku-4-5",
max_tokens=50,
messages=[{
"role": "user",
"content": "Classify as positive/negative/neutral: 'Terrible experience.'"
}]
)
),
]
)
print(batch.id) # msgbatch_...
Poll
The processing_status field on a batch object reflects current state. Poll until it equals "ended":
import time
while True:
batch = client.messages.batches.retrieve(batch.id)
if batch.processing_status == "ended":
break
print(f"Processing: {batch.request_counts.processing} remaining")
time.sleep(60)
Polling every 60 seconds is reasonable for most workloads. The request_counts object gives you a live breakdown of how many requests are still processing, succeeded, errored, or canceled.
Retrieve
Call client.messages.batches.results() to iterate over per-request outcomes. Each result carries the custom_id you assigned and a result object discriminated by type:
for result in client.messages.batches.results(batch.id):
match result.result.type:
case "succeeded":
msg = result.result.message
text = next((b.text for b in msg.content if b.type == "text"), "")
print(f"[{result.custom_id}] {text}")
case "errored":
print(f"[{result.custom_id}] Error: {result.result.error.type}")
case "expired":
print(f"[{result.custom_id}] Expired — resubmit")
What Limits and Edge Cases Should You Know?
Several constraints matter for production use of the Claude Message Batches API:
- Result expiration. Batch results are stored for 29 days. If you do not retrieve them within that window, the data is gone. Build result retrieval into your pipeline's completion logic rather than treating it as optional cleanup.
- Per-request errors are isolated. Individual requests within a batch can fail independently. An
"errored"result with type"invalid_request"indicates a problem with that specific request's parameters — fix it and resubmit. A server-side error type is safe to retry wholesale. - Cancellation is partial. Calling
client.messages.batches.cancel(batch_id)sends a cancellation signal, but requests already in processing may complete before it takes effect. Results for canceled requests appear with type"canceled". - 256 MB body limit. For batches with long system prompts or document inputs, the total payload size can hit the 256 MB ceiling before the 100,000-request limit. Monitor payload size when each request is large.
- All Messages API features are supported. Vision inputs, tool definitions, prompt caching, and structured outputs all work inside batch requests. Prompt caching is especially valuable in batches where many requests share a large system prompt — each request that hits the cache pays the reduced read cost on top of the batch discount.
How Does the Batches API Appear on the CCA-F Exam?
The Claude Certified Architect Foundations exam tests architectural judgment, not API syntax. For the Batches API, that means scenario questions framed as trade-off decisions rather than recall of method names.
Typically where candidates lose points is misidentifying the batch pattern as appropriate for latency-sensitive paths. An item might describe a workload — "nightly scoring of 50,000 product descriptions" versus "real-time customer support reply" — and ask which API surface to use. The correct answer for the overnight pipeline is batch; the correct answer for the live reply is standard or streaming. Understanding the synchronous versus asynchronous contract is the core concept being tested.
Secondary exam signals include: knowing that the discount is 50 percent and applies uniformly across all models; knowing that the maximum processing window is 24 hours; and knowing that the Batches API supports all Messages API features rather than being a reduced-capability surface. You may also encounter questions probing the poll-and-retrieve pattern — the API does not push results and a client must check status actively rather than awaiting a callback.
This post covers the Batches API conceptually and walks through the submit/poll/retrieve mechanics. If you are preparing for the Claude Certified Architect Foundations certification, PlinthPrep's practice questions let you test whether you have genuinely absorbed these architectural distinctions — including edge cases around result types, expiration windows, and when to choose batch over streaming. PlinthPrep is an independent study resource, not affiliated with Anthropic. Visit plinthprep.com to see the exam-prep questions.
Frequently asked questions
- What is the Claude Message Batches API?
- The Claude Message Batches API processes large sets of Messages API requests asynchronously in a single job. You submit up to 100,000 requests at once, and results become available within 24 hours at 50 percent of standard per-token pricing. Use it for bulk classification, dataset annotation, offline evaluation, and report generation.
- How does batch processing differ from a standard Messages API call?
- A standard Messages API call returns a response synchronously within seconds, suitable for interactive features. Batch processing submits requests as an asynchronous job and returns no immediate result. You poll for status and retrieve results when the job finishes, typically within one hour. The trade-off is latency for cost and throughput.
- How much does the Claude Message Batches API cost compared to standard API pricing?
- The Batches API charges 50 percent of standard per-token rates for both input and output tokens across all Claude models. There is no tiered discount structure — the reduction is flat and applies to every model available through the standard Messages API. Verify current model pricing at Anthropic's pricing documentation.
- How do you submit, poll, and retrieve a Claude batch job?
- Submit a batch by posting a list of requests, each with a unique custom_id and a params object matching the Messages API. Poll with batches.retrieve until processing_status equals ended, which typically takes under one hour. Then stream results with batches.results, filtering by result type to handle successes, errors, and expirations.
- How does the Claude Message Batches API appear on the CCA-F exam?
- The CCA-F exam tests architectural trade-offs, not syntax. Expect scenario questions asking which API surface — standard, streaming, or batch — fits a given workload. The exam also tests awareness of the 50 percent cost discount, the 24-hour maximum processing window, and the asynchronous poll-and-retrieve pattern rather than real-time response delivery.