LIMITS

Claude token limits

Distinguish Claude context windows, output caps, API rate limits and subscription allowances. Check a manual context budget.

REFERENCESources checked
On this page 6 sections

Which Claude token limit do you mean?

A Claude token limit can refer to a single request's context, a maximum response, an API rate, or a product usage allowance. Identify the kind first. A larger context window does not give you a fixed number of messages per day.

Four different questions
LimitQuestionWhere to check
Context windowWill this request fit?Model specifications and complete request count
Maximum outputHow long may one answer be?Selected model and max_tokens setting
API rateHow much traffic may this account send?Account API limits and response headers
Product allowanceHow much plan usage remains?Account usage screen and reset message

API context and maximum output

API specifications · checked 3 October 2026
ModelContext windowMaximum output
Claude Sonnet 5.51,000,000128,000
Claude Opus 5.51,000,000128,000
Claude Haiku 4.5200,00064,000

These are model specifications for the direct Claude API, checked against the model overview. Available product settings and account access can differ. Check the chosen model ID instead of treating “Claude” as one unchanging specification.

Context includes system instructions, history, tools and other request material, plus generated output. Cached prefixes still occupy it. Use the context documentation when accounting for a complete request.

A quick context budget

This check accepts quantities you provide. It compares input plus reserved output with the chosen model's context and checks the output cap separately. It is not a model-specific tokenizer or a guarantee that an API request will be accepted.

Free, Pro and Max usage

Subscription allowances belong to the product and account. Do not convert a model context size into a daily token quota for a plan. Check your account's usage screen for the current allowance and reset information, and use the product's displayed message to identify the actual restriction.

When estimating whether a plan fits your work, record your own workload: task length, model selection, attachments and coding activity. A simple table of observed tasks and displayed usage is more useful than a universal “messages per day” claim.

API rate limits

Rate limits govern how much request traffic an account can send over time. They are different from the amount of text one request holds. An API request may fit its context and still encounter rate limiting.

Check the account's current API limits and response headers. The provider's rate-limit reference documents the applicable measurements and retry behavior. Your account tier and model can affect the result.

What to do next

  • If a request is too large, count the complete material and reserve room for the desired output.
  • If the requested output is too large, lower it or split the job into smaller responses.
  • If an account allowance is exhausted, follow the account's displayed reset or billing options.
  • If API traffic is rate limited, use the provider's retry guidance and account limits.

For terminal workflows, the Claude Code limits guide helps connect a symptom to the right screen. For an API budget, use the Claude cost calculator.

Sources & review

Checked 3 October 2026. Product rules can change; consult the linked official references.