What is an AI token?
An AI token is a unit of text that a language model reads or generates. A tokenizer converts text into a sequence of numbered units. A token might contain a whole word, part of a word, punctuation, whitespace, or only part of a character.
Token counting measures that sequence. It matters when you estimate API costs, prepare a long prompt, or work within a model's context window. The units depend on the tokenizer: one piece of text can produce different counts for different models.
Try English, Chinese and code
Raw-text example · o200k_base · edit up to 2,000 characters
This live example uses the o200k_base raw-text tokenizer from the existing counter. The colored pieces show token boundaries, including punctuation and spaces. Switch the example or edit the text to see how the count changes. These counts describe this tokenizer, rather than every AI model.
Notice that the English sentence has a different relationship between characters and tokens than the Chinese sentence. Code introduces operators, quotes and repeated identifiers. Treat each language and workload as something to measure.
Tokens, words and characters
Characters describe written symbols. Words describe linguistic units, and their boundaries can depend on language. Tokens describe the model's encoded input. There is no fixed conversion that makes these three counts interchangeable.
For example, a short compound word may split into multiple tokens while a common word fits into one. A space or punctuation mark can also affect how the next unit is formed. A document with many words is not necessarily the document with the most tokens.
When you need a word-to-token estimate, use a sample that resembles your real text. Count its words and tokens, compute a sample ratio, then apply that ratio cautiously to similar material. Measure the complete text before making a final budget.
Input tokens and output tokens
Input is the material sent to a model. Output is what the model generates. An API can price them differently, so an input count alone cannot predict the complete request cost.
Consider a hypothetical request with 10,000 input tokens and 1,000 output tokens. To estimate the cost, multiply each quantity by its own price per million tokens and add them. A repeated workflow also needs a request-count assumption. Our Claude cost calculator lets you enter these quantities separately.
A prompt is only part of a request
The text you paste into a counter might exclude system instructions, previous messages, tools and attachments. Counting that text is useful, but a complete API request can contain more material. Context limits also apply to the model's generation, not just your visible prompt.
For a reproducible comparison, keep the same text and select the exact tokenizer you intend to compare. For billing, check the provider's request usage. Read the counting methodology for the scope of each local engine and estimate.
How to count tokens in practice
- Paste a representative prompt or upload a plain UTF-8 text file in the token counter.
- Choose the appropriate tokenizer and read the accuracy label beside the result.
- Check whether conversation history or tools add material outside the pasted text.
- Enter expected output separately in a cost calculation.
If you are investigating a limit, first determine whether it concerns context length, output size, account usage or request rate. The Claude limits reference explains those different questions.