← Blog

Claude Streaming API: SSE Events, When to Use It, CCA-F

September 23, 2026

How the Claude streaming API delivers tokens over SSE, the six event types in sequence, when to stream vs. batch, and what the CCA-F exam tests.

Claude's streaming API delivers response tokens over server-sent events (SSE) as they are generated, rather than waiting for the full completion. This cuts perceived latency for chat interfaces, is required for showing extended thinking in real time, and carries no extra token cost. Skip it only for background batch jobs where the user never waits.

What Is Claude Streaming and Why Does It Reduce Latency?

Claude API streaming mode is enabled by adding "stream": true to a standard POST /v1/messages request. Instead of holding the HTTP connection open until generation completes, the server writes tokens to the response body as they are produced. Your application can begin rendering immediately rather than waiting for the full output.

Perceived latency drops substantially because the first token typically arrives within a second, even for responses that take many seconds to complete. The total generation time is the same whether you stream or not — streaming reshapes when the user sees output, not how fast the model runs.

Streaming does not add a cost surcharge. Token billing is identical whether you stream or not: you pay the same input and output rates for every token consumed. No streaming-specific fee exists in Anthropic's current pricing model. Streaming reshapes when your application receives output, not how many tokens are consumed.

One practical consequence: the Anthropic SDK raises an error if you attempt a non-streaming request with max_tokens above roughly 16,000, since a response that large risks timing out before the HTTP connection completes. For large-output requests, client.messages.stream() combined with .get_final_message() is the standard pattern.

How Does the SSE Event Sequence Work?

The Claude streaming API sends six SSE event types: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, and message_stop. The content_block_delta events carry incremental text characters to render in real time. The message_delta event delivers the final stop_reason and usage totals. The sequence is consistent across all models and response sizes.

Each event arrives as a line-delimited pair in the response body — an event: name and a data: field containing JSON:

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

Walking through the complete sequence for a plain text response:

  • message_start — fires once at the beginning. Contains the message ID, model, and an initial usage object (output tokens are zero at this point).
  • content_block_start — opens a new content block. For plain text, content_block.type is "text" with an empty initial text field.
  • content_block_delta — fires per token. Contains delta.type of "text_delta" and the new characters in delta.text. This is the event your application renders in real time.
  • content_block_stop — closes the current block.
  • message_delta — fires once near the end. Contains the final stop_reason — typically "end_turn" — and the completed usage object with output token totals.
  • message_stop — the stream is complete.

In Python with the SDK, iterating stream.text_stream yields only the text delta strings. Iterate the full event iterator when you need to inspect non-text blocks such as tool use or thinking.

When Should You Use Streaming vs. Waiting for the Full Response?

Use Claude API streaming when a user is waiting in real time — chat interfaces, voice applications, or any interactive product where a blank screen feels broken. Also use it when max_tokens exceeds roughly 16,000, since the Anthropic SDK refuses non-streaming calls at that scale to prevent HTTP timeouts. Skip streaming only for background batch processing.

The Anthropic SDK's client.messages.stream() helper accumulates the response behind the scenes so you can call .get_final_message() after iteration to obtain the complete response object. Streaming everywhere and calling the helper when you need the assembled result is a safe default: you get timeout protection on large outputs without manually buffering deltas.

For background jobs where no user is waiting — batch document processing, nightly evaluation pipelines, bulk classification — the Batches API is a better fit. It processes requests asynchronously and is designed for throughput rather than latency.

How Does Streaming Interact with Tool Use and Extended Thinking?

When Claude decides to call a tool during a streaming request, a content_block_start event opens a tool_use block, followed by content_block_delta events carrying input_json_delta fragments. Check the stop_reason in the message_delta event: if it is "tool_use", execute your tools, then send a new streaming request with the results appended to the message history.

Extended thinking adds two additional delta types. When adaptive thinking is enabled, the API opens a block with content_block.type of "thinking" and streams its content via delta.type of "thinking_delta". On Opus 4.8 and Opus 4.7, the thinking text is omitted from the stream by default — the block still appears but its text field is empty. Setting thinking.display to "summarized" restores visible reasoning content. This detail is typically where candidates lose points on exam questions about the thinking-and-streaming interaction: the block exists regardless, but the visible content depends on the display parameter.

How Do You Handle Streaming Errors and Interrupted Connections?

Wrap your streaming loop in a try/except block and handle the SDK's typed exceptions: APIConnectionError for dropped connections, RateLimitError for rate limiting, and APIStatusError for HTTP-level errors. Partial responses from interrupted streams can be discarded for simple chat; for long agentic runs, buffer the accumulated text and pass it as context on retry.

try:
    with client.messages.stream(
        model="claude-opus-4-8",
        max_tokens=4096,
        messages=[{"role": "user", "content": prompt}]
    ) as stream:
        for text in stream.text_stream:
            print(text, end="", flush=True)
except anthropic.APIConnectionError:
    print("\nConnection lost.")
except anthropic.RateLimitError:
    print("\nRate limited — retry after a delay.")
except anthropic.APIStatusError as e:
    print(f"\nAPI error {e.status_code}.")

The message_delta event carries a usage field with output tokens consumed at that point. Logging this value is useful for monitoring even on interrupted streams, since partial output still counts toward your token usage.

What Does the CCA-F Exam Expect You to Know About Streaming?

The CCA-F exam tests streaming in the API domain across three areas.

Event sequence. Know the six event types by name and identify which one carries the final stop_reason and usage totals. The answer is message_delta, not message_stop — a common distractor in scenario-based questions.

When to stream. Exam scenarios present an application type and ask you to justify the approach. Real-time chat and large-output requests point toward streaming. Background batch jobs point toward the Batches API. The SDK's behavior at high max_tokens values — refusing non-streaming calls — is a tested edge case.

Streaming and extended thinking. Questions may describe a scenario where reasoning text does not appear in the streamed output and ask you to identify the cause. The expected answer on Opus 4.8 and Opus 4.7: thinking content is omitted by default, and setting thinking.display to "summarized" restores it.

This post explains how Claude API streaming works conceptually. If you are working toward the CCA-F certification, PlinthPrep's practice questions on streaming let you test whether you have genuinely understood the event sequence, decision scenarios, and thinking interaction — not just read about them. PlinthPrep is an independent study resource, not affiliated with Anthropic. Visit plinthprep.com to see the exam-prep questions.

Frequently asked questions

What are the SSE event types in the Claude streaming API?
The Claude streaming API sends six SSE event types: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, and message_stop. The content_block_delta events carry incremental text characters to render in real time. The message_delta event delivers the final stop_reason and usage totals. The sequence is consistent across all models and response sizes.
When should I use the Claude API streaming mode instead of waiting for the full response?
Use Claude API streaming when a user is waiting in real time — chat interfaces, voice applications, or any interactive product where a blank screen feels broken. Also use it when max_tokens exceeds roughly 16,000, since the Anthropic SDK refuses non-streaming calls at that scale to prevent HTTP timeouts. Skip streaming only for background batch processing.
Does the Claude streaming API cost more tokens than a regular request?
Streaming does not add a cost surcharge. Token billing is identical whether you stream or not: you pay the same input and output rates for every token consumed. No streaming-specific fee exists in Anthropic's current pricing model. Streaming reshapes when your application receives output, not how many tokens are consumed.
How does the Claude streaming API handle tool use?
When Claude decides to call a tool during a streaming request, a content_block_start event opens a tool_use block, followed by content_block_delta events carrying input_json_delta fragments. Check the stop_reason in the message_delta event: if it is tool_use, execute your tools, then send a new streaming request with the results appended to the message history.
How should I handle errors and dropped connections in Claude API streaming?
Wrap your streaming loop in a try/except block and handle the SDK's typed exceptions: APIConnectionError for dropped connections, RateLimitError for rate limiting, and APIStatusError for HTTP-level errors. Partial responses from interrupted streams can be discarded for simple chat; for long agentic runs, buffer the accumulated text and pass it as context on retry.

Share this post

Plinth Prep is an independent study resource and is not affiliated with, endorsed by, or sponsored by Anthropic. Practice material is written by Plinth Prep and does not reproduce real exam content.