Which Claude token limit do you mean?
A Claude token limit can refer to a single request's context, a maximum response, an API rate, or a product usage allowance. Identify the kind first. A larger context window does not give you a fixed number of messages per day.
| Limit | Question | Where to check |
|---|---|---|
| Context window | Will this request fit? | Model specifications and complete request count |
| Maximum output | How long may one answer be? | Selected model and max_tokens setting |
| API rate | How much traffic may this account send? | Account API limits and response headers |
| Product allowance | How much plan usage remains? | Account usage screen and reset message |
API context and maximum output
| Model | Context window | Maximum output |
|---|---|---|
| Claude Sonnet 5.5 | 1,000,000 | 128,000 |
| Claude Opus 5.5 | 1,000,000 | 128,000 |
| Claude Haiku 4.5 | 200,000 | 64,000 |
These are model specifications for the direct Claude API, checked against the model overview. Available product settings and account access can differ. Check the chosen model ID instead of treating “Claude” as one unchanging specification.
Context includes system instructions, history, tools and other request material, plus generated output. Cached prefixes still occupy it. Use the context documentation when accounting for a complete request.
A quick context budget
This check accepts quantities you provide. It compares input plus reserved output with the chosen model's context and checks the output cap separately. It is not a model-specific tokenizer or a guarantee that an API request will be accepted.
Free, Pro and Max usage
Subscription allowances belong to the product and account. Do not convert a model context size into a daily token quota for a plan. Check your account's usage screen for the current allowance and reset information, and use the product's displayed message to identify the actual restriction.
When estimating whether a plan fits your work, record your own workload: task length, model selection, attachments and coding activity. A simple table of observed tasks and displayed usage is more useful than a universal “messages per day” claim.
API rate limits
Rate limits govern how much request traffic an account can send over time. They are different from the amount of text one request holds. An API request may fit its context and still encounter rate limiting.
Check the account's current API limits and response headers. The provider's rate-limit reference documents the applicable measurements and retry behavior. Your account tier and model can affect the result.
What to do next
- If a request is too large, count the complete material and reserve room for the desired output.
- If the requested output is too large, lower it or split the job into smaller responses.
- If an account allowance is exhausted, follow the account's displayed reset or billing options.
- If API traffic is rate limited, use the provider's retry guidance and account limits.
For terminal workflows, the Claude Code limits guide helps connect a symptom to the right screen. For an API budget, use the Claude cost calculator.
Sources & review
Checked 3 October 2026. Product rules can change; consult the linked official references.