Claude API Response: stop_reason, usage, and Content Blocks
September 30, 2026
Decode every Claude API response field: what each stop_reason value means, how to read usage token counters, and which content block types your code must handle.
A Claude API response always returns three top-level fields: stop_reason, usage, and content. The stop_reason tells you exactly why generation stopped — end_turn (natural completion), max_tokens (truncated), stop_sequence (custom trigger matched), or tool_use (Claude is requesting a tool call). Every production integration and every Claude Certified Architect Foundations question about agentic loops starts with reading this field correctly.
What Does a Full Claude API Response Object Look Like?
When you call POST /v1/messages, the response body contains several fields, but three drive all your application logic:
stop_reason— why generation stoppedusage— token accounting for the requestcontent— an array of content blocks (text, tool calls, or reasoning)
A minimal response looks like this:
{
"id": "msg_01XFDUDYJgAACzvnptvVoYEL",
"type": "message",
"role": "assistant",
"model": "claude-opus-4-8",
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 25,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0,
"output_tokens": 305
},
"content": [
{
"type": "text",
"text": "Hello! How can I help you today?"
}
]
}
The id, type, role, and model fields are metadata. stop_sequence is null unless you defined custom stop strings and one matched. Your code branches entirely on stop_reason, reads costs from usage, and extracts output from content.
What Are the Four stop_reason Values — and What Should Your Code Do for Each?
The API defines four primary stop reasons. Each demands a different response from your code.
end_turn
Generation completed naturally. Claude finished its response. This is the success case — extract text from content, return it to the user, and close the loop.
max_tokens
Output was cut off because it reached your max_tokens value. The response is incomplete. Your code should increase max_tokens and retry, or switch to streaming for large outputs. Never treat max_tokens as equivalent to end_turn — doing so silently serves truncated responses to users.
stop_sequence
You defined one or more custom stop strings via the stop_sequences parameter, and the model generated one of them. The matching string is returned in stop_sequence on the response object. Use this when you want generation to halt at a specific delimiter — for example, extracting a single JSON object and stopping before Claude writes anything after the closing brace.
tool_use
Claude has decided to call one of your defined tools. This is not an error — it is a signal to your agentic loop. The content array will contain one or more tool_use blocks. Your code must execute the function, collect the result, and send it back before generation can continue. Misreading this as a failure state is one of the most common mistakes in production agentic integrations, and a reliable source of CCA-F exam questions.
Two additional values appear in specific contexts: pause_turn (Claude paused mid-task in a server-side tool loop and can be resumed by resending the conversation) and refusal (Claude declined to respond for safety reasons, with a structured stop_details object attached).
How Do You Read the usage Object (Including Cache Token Fields)?
The usage object always contains four fields:
input_tokens— Prompt tokens processed at full price. This is the portion of your prompt that was not served from cache.output_tokens— Tokens generated in the response.cache_creation_input_tokens— Tokens written to the prompt cache on this request. These cost approximately 1.25× the standard input price because they are being stored for future reads.cache_read_input_tokens— Tokens served from an existing cache entry. These cost approximately 0.1× the standard input price — the primary savings from prompt caching.
The total prompt size for a request is the sum of all three input fields:
total prompt tokens = input_tokens + cache_creation_input_tokens + cache_read_input_tokens
If you track only input_tokens in a production integration that uses prompt caching, both your cost estimates and your context-window calculations will be wrong. A request where input_tokens shows 4,000 but cache_read_input_tokens shows 90,000 has processed 94,000 total prompt tokens — and your context-window math needs to account for all of them.
What Content Block Types Can Claude Return?
The content field is always an array. Each element is a block with a type field that must be checked before accessing type-specific properties. Accessing block.text on a tool_use block will throw an error in typed languages and silently produce undefined in others.
text
The most common block type. Access block.text for Claude's prose response. In a simple request with no tools and no extended thinking, the content array typically contains exactly one text block.
tool_use
Present when stop_reason is tool_use. Each block contains three key fields:
id— a unique identifier you must echo back when sending the resultname— the name of the tool Claude has calledinput— a parsed JSON object with the tool's arguments
A single response can contain multiple tool_use blocks if Claude wants to call several tools in parallel. Collect all of them, execute each, and return all results before continuing.
thinking
Present when you have enabled adaptive thinking (thinking: {"type": "adaptive"}). These blocks contain Claude's internal reasoning, accessible via block.thinking. In Claude Opus 4.7 and 4.8, thinking content is omitted from the response by default — set "display": "summarized" on the thinking parameter to receive it. If your UI streams reasoning to users and you've migrated to Opus 4.7 or 4.8, the absence of this flag causes a long silent pause before output begins.
How Does stop_reason: tool_use Drive an Agentic Tool Loop?
The loop follows a precise cycle:
- Send the messages array to the API.
- Check
stop_reason. Ifend_turn, exit. Iftool_use, continue. - Extract all
tool_useblocks fromcontent. - Execute each tool.
- Append the assistant's full
contentarray as an assistant message. - Append a user message containing
tool_resultblocks — one per tool, each referencing the originaltool_useblock'sid. - Send the updated message history back to the API and return to step 2.
while True:
response = client.messages.create(
model="claude-opus-4-8",
max_tokens=16000,
tools=tools,
messages=messages
)
if response.stop_reason == "end_turn":
break
tool_results = []
for block in response.content:
if block.type == "tool_use":
result = execute_tool(block.name, block.input)
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": result
})
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": tool_results})
The critical constraint: every tool_use block must receive a matching tool_result. The API rejects the follow-up request if any tool use is unmatched. Note also that you must append the assistant's full response.content array — not just the text — so that tool_use blocks are preserved in history.
CCA-F Exam Quick Reference: Response Anatomy at a Glance
| Field | What it signals | Your code should… |
|---|---|---|
stop_reason: "end_turn" |
Natural completion | Return output to the user |
stop_reason: "max_tokens" |
Output truncated | Increase max_tokens or stream |
stop_reason: "stop_sequence" |
Custom delimiter matched | Extract output up to that point |
stop_reason: "tool_use" |
Tool call requested | Execute tools, send results, loop |
usage.input_tokens |
Uncached prompt tokens | Include in total prompt size |
usage.cache_read_input_tokens |
Tokens read from cache | Include in total prompt size |
usage.cache_creation_input_tokens |
Tokens written to cache | Expect higher cost this request |
usage.output_tokens |
Generated tokens | Track for output cost |
content[].type: "text" |
Prose response | Access via block.text |
content[].type: "tool_use" |
Tool call request | Access id, name, input |
content[].type: "thinking" |
Reasoning trace | Access via block.thinking |
Understanding the response schema at this level is necessary but not sufficient for the CCA-F exam. The exam tests application — whether you can identify the correct branching logic for a given stop_reason, diagnose where a tool loop breaks down, or calculate what the usage fields mean for context-window management. Plinth Prep is an independent study resource, not affiliated with Anthropic, and its practice questions on response anatomy, agentic loops, and tool use are designed to close exactly that gap. Visit plinthprep.com to see the exam-aligned practice questions.
Frequently asked questions
- What are the four stop_reason values in a Claude API response?
- The four stop_reason values are: end_turn (Claude finished naturally), max_tokens (output was truncated at your requested limit), stop_sequence (a custom string you defined triggered a halt), and tool_use (Claude is requesting a tool call and needs a result before continuing). Each requires a different code branch.
- What does stop_reason tool_use mean in a Claude API response?
- stop_reason tool_use means Claude has decided to invoke one of your defined tools and is pausing to receive the result. It is not an error. Your code should extract the tool_use block from content[], execute the function, and send back a tool_result in the next API call.
- What fields does the usage object contain in a Claude API response?
- The usage object contains four fields: input_tokens (uncached prompt tokens billed at full price), output_tokens (generated tokens), cache_creation_input_tokens (tokens written to the prompt cache at about 1.25x cost), and cache_read_input_tokens (tokens served from cache at about 0.1x cost). All four are always present.
- What content block types can Claude return in the content array?
- Claude can return three content block types: text (prose output, accessed via block.text), tool_use (a function call request with a name, id, and input), and thinking (internal reasoning from adaptive thinking, accessed via block.thinking). Always check block.type before accessing type-specific fields.
- What should you do when stop_reason is max_tokens in a Claude response?
- When stop_reason is max_tokens, Claude's output was cut off at your requested limit — the response is incomplete. Increase max_tokens and retry, or switch to streaming for large outputs. Never treat max_tokens as natural completion; always distinguish it from end_turn in your application logic.