Claude Context Window: Token Limits and Management Strategies
August 31, 2026
Learn Claude's per-model context window limits, what counts toward the token budget, how to count tokens via the API, and which management strategy fits your use case.
Claude's context window is the total token budget for one API request — system prompt, conversation history, images, and tool outputs combined. Opus and Sonnet models accept up to 1 million input tokens; Haiku 4.5 accepts 200,000. Output tokens are capped separately. Exceeding the limit throws an error rather than silently truncating, making proactive management essential.
How Many Tokens Does Each Claude Model Support?
Anthropic's current lineup splits into two tiers by context window size. Claude Opus 4.8, Opus 4.7, Opus 4.6, and Claude Sonnet 4.6 all support 1 million input tokens — the same ceiling regardless of tier. Claude Haiku 4.5 is the one exception: its context window is 200,000 tokens, one-fifth the size of its siblings.
Output tokens are accounted for separately and carry their own per-model ceiling. Opus models and Sonnet 4.6 support up to 128,000 output tokens; Haiku 4.5 tops out at 64,000. Verify the current per-model figures in Anthropic's model documentation before relying on large output budgets in production — these numbers can be updated between releases.
A common exam trap is assuming Haiku shares the 1M window of the premium models. It does not. An architecture tested with Sonnet may silently hit Haiku's smaller window when you swap models for cost reasons. Always verify the model you ship is the model you tested with.
What Counts Toward the Context Window?
Every token in your request draws from the same pool:
- System prompt — the operator-defined instructions sent at the start of every request
- Conversation history — all prior user and assistant turns; the API is stateless, so you send the full history on each call
- Tool definitions — the JSON schemas you declare in the
toolsarray; a rich schema can run to several hundred tokens - Tool results — the content returned in
tool_resultmessages, including code execution outputs and web search results, which can be substantial - Images — base64-encoded images or URL references; a single high-resolution image sent to a current Opus model can consume thousands of tokens depending on resolution
- Thinking blocks — when extended thinking is enabled, reasoning tokens appear in the response and count toward input on the next turn if you include them in the conversation history
Images and tool outputs are the most common sources of unexpected context exhaustion. A multi-turn agent loop that passes screenshots or large tool results on every turn can exhaust even a 1M-token window faster than the text alone suggests. Auditing token consumption by component — not just total tokens — is the first step toward a stable budget.
How Do You Count Tokens Before Sending a Request?
Anthropic exposes a dedicated token counting endpoint at POST /v1/messages/count_tokens. Pass the same model, system prompt, tools, and messages you intend to send; the API returns an input_tokens count without running inference or charging for output tokens.
In Python:
count = client.messages.count_tokens(
model="claude-opus-4-8",
system="You are a helpful assistant.",
messages=[{"role": "user", "content": user_input}]
)
print(count.input_tokens)
One critical warning: do not use tiktoken or any OpenAI-compatible tokenizer to estimate Claude token counts. Those libraries are calibrated for GPT-family models and consistently undercount Claude's usage — often by 15–20% on typical text, and by significantly more on code or non-English input. An undercount means you may send a request that appears within budget when it is actually over the limit.
For ongoing monitoring in a production agent, track the usage object in every response. The fields input_tokens, cache_creation_input_tokens, and cache_read_input_tokens together give the true size of each request, and their sum tells you how close you are to the window ceiling.
What Happens When You Exceed the Context Limit?
Claude does not silently truncate. If your assembled request exceeds the model's input limit, the API returns an error — either a 400 invalid_request_error before generation starts, or a completed response with stop_reason set to "model_context_window_exceeded" if the model ran out of context mid-generation.
This design is intentional. Silent truncation would discard the beginning of your context without warning, potentially removing your system prompt or the earliest turns of a long conversation — producing incoherent responses that are difficult to debug. The error-on-exceed behavior forces explicit context management and surfaces the problem at the boundary rather than burying it in output quality.
Practically, any application that accumulates conversation history must implement a strategy for keeping the assembled request under budget. Relying on the model to handle overflow is not an option.
Which Strategy Should You Use to Manage Long Contexts?
Four patterns address different needs. The right choice depends on your latency budget, accuracy requirements, and whether you control a retrieval layer.
Truncation
Drop the oldest messages from conversation history when you approach the limit. This is the simplest approach and adds no latency, but it is lossy — important context from early in a conversation can disappear without warning. Truncation works well for stateless Q&A flows where each question is independent, and poorly for long agentic tasks where early instructions constrain later actions.
Summarization and Server-Side Compaction
Replace older turns with a condensed summary. The Anthropic API provides server-side compaction as a beta feature: opt in with the compact-2026-01-12 beta header and set the context_management parameter on your request. When triggered, the API automatically summarizes earlier context and returns a compaction block in the response. You must preserve that block and pass it back verbatim on subsequent turns — the API uses it to reconstruct the summarized history. Appending only the text and discarding the compaction block breaks this mechanism silently.
Manual summarization — asking Claude to summarize the conversation so far before truncating it — achieves a similar result without the beta dependency, at the cost of an extra inference call. Either approach trades a small amount of fidelity for dramatically longer effective conversations.
Retrieval-Augmented Generation (RAG)
Rather than keeping the full document corpus in context, index content externally and retrieve only the chunks relevant to the current query. RAG is the right choice when the context problem is document scale — a large codebase or a library of policy documents that will never fit in any context window — and when accurate retrieval is feasible. It shifts the engineering complexity from token budgeting to retrieval quality, and it removes the window constraint at the source rather than managing around it.
Prompt Caching
Prompt caching does not reduce your token count — a cached request still occupies the full context window. What it does is make large, stable contexts dramatically cheaper to reuse across many requests. If your system prompt and document context are the same for a batch of requests, marking that stable prefix with cache_control cuts the input token cost for those tokens by up to 90% on cache hits. This is a cost optimization, not a window expansion. PlinthPrep's post on Claude prompt caching economics covers when the math favors caching and when it does not.
What Does the CCA-F Exam Expect You to Know About Context Windows?
The CCA-F exam tests context window knowledge at the architecture level, not just the specification level. Expect scenario-based questions that require you to reason about trade-offs, not recall a single number.
Key areas the exam covers:
- Input versus output budgets — the context window governs input tokens; output tokens are a separate, lower ceiling that must also be planned for
- What contributes to token count — system prompts, all conversation turns, images, tool definitions, and tool results all draw from the same pool
- Error behavior on overflow — the API errors rather than truncating silently;
model_context_window_exceededas a stop reason is distinct frommax_tokens - Model differences — Haiku 4.5's 200,000-token window is meaningfully smaller than the 1 million-token window on Opus and Sonnet models; this affects architecture decisions when mixing models
- Strategy selection — matching truncation, summarization, RAG, or caching to the requirements of a given scenario, including their respective trade-offs
- Token counting — using
/v1/messages/count_tokensrather than third-party tokenizers calibrated for other model families
A scenario question might describe an agent that processes growing conversation histories on a cost-sensitive workload and ask which strategy minimizes data loss while keeping costs predictable. Knowing that compaction preserves meaning while truncation discards it — and that RAG removes the window constraint at the source — is the kind of reasoning the exam is looking for.
Reading about context window mechanics is different from being tested on them under exam conditions. PlinthPrep's practice questions on model capabilities and context management are designed to surface exactly that gap — the difference between recognizing a concept and applying it correctly in a scenario you haven't seen before. PlinthPrep is an independent study resource, not affiliated with Anthropic. Visit plinthprep.com to see the exam-aligned questions on this domain.
Frequently asked questions
- What is Claude's context window?
- Claude's context window is the total token budget for a single API request, covering system prompt, messages, images, and tool outputs. Opus and Sonnet models support up to 1 million input tokens; Haiku 4.5 supports 200,000. Output tokens are counted separately and capped per model. Check Anthropic's documentation for current per-model ceilings.
- What counts toward Claude's context window?
- Every input in the request draws from the same pool: the system prompt, all conversation turns, tool definitions, tool results, and any images. High-resolution images can consume thousands of tokens each. Thinking blocks also count toward input on subsequent turns if included in history. Complex tool schemas alone can be several hundred tokens.
- How do I count tokens before sending a Claude API request?
- Use Anthropic's dedicated endpoint at POST /v1/messages/count_tokens, passing the same model, messages, and system prompt you intend to use. The SDK exposes this as client.messages.count_tokens() and returns an input_tokens integer. Avoid tiktoken — it is calibrated for OpenAI models and substantially undercounts Claude's token usage, leading to budget miscalculations.
- What happens when a Claude request exceeds the context window?
- Claude does not silently truncate. Exceeding the input limit produces an API error — a 400 response before generation starts, or a completed response with stop_reason set to model_context_window_exceeded. This forces explicit context management in your application and prevents the subtle, hard-to-diagnose bugs that silent truncation causes in long-running agentic loops.
- Which context management strategy should I use for Claude?
- Choose based on your use case: truncation is fast but lossy; server-side compaction preserves meaning but requires a beta opt-in; RAG avoids window limits by retrieving only relevant chunks from external storage; prompt caching reduces cost on stable prefixes without reducing token count. Most production agents combine two or more of these strategies.