Learn
GuideAI & agentsCompanion to Context and memory
AI tokens: what a bot is counting
An AI token is a unit an AI model processes. For text, a token can represent a word, part of a word, a character, or another piece of text. Splitting text into these units is called tokenization. One word does not reliably equal one token. [1]
You do not need to count tokens while writing every request. Understanding them helps explain why a short question can still involve a large amount of processing.
Your visible message is only one ingredient
Suppose you ask a bot: “Write a reminder for Saturday.”
The application might also supply your writing preferences, the event note, earlier conversation, and descriptions of available tools. If the bot searches files, retrieved passages may become input for another step. The visible request is only part of the information involved. [2][5]
The reminder it generates is output. An agent performing several steps can make several model requests, each with its own inputs and outputs.
Swipe sideways to see the whole diagram, or open it full size.
Figure explanation: Five possible inputs are grouped together: current request, relevant conversation, instructions and preferences, a retrieved event note, and available tool descriptions. This group supplies the model request. The generated reminder is its output. The diagram is a simplified text example and does not show all possible usage categories. What is included and counted depends on the model and application.
Three limits to keep separate
| Term | What it describes | A useful question |
|---|---|---|
| Context capacity | How much a model can work with in a request, subject to its input and output rules | Will the relevant material fit? |
| Output limit | How much the model can generate in a response | Can it produce the requested deliverable in one response? |
| Account usage allowance | What your product or plan lets you use over a period | What does this service count, and when does it reset? |
The first two are related technical limits. [2] The third is a product rule. A subscription might count messages, credits, compute, or a combination. Check its current documentation instead of assuming a token total predicts its allowance.
Some services also report categories such as cached input or reasoning usage. Images, audio, and other inputs can have their own accounting rules. Use the provider’s supported counter and usage report for an exact measurement. [1]
A better way to trim a request
Imagine your bot needs a public reminder with the time and entrance. You attach the approved event note and a 60-page planning archive.
Before removing essential instructions, ask whether the archive is needed. Supplying the current note may reduce irrelevant material while preserving the facts that matter.
Do keep “registration is optional” if readers need to know it. A shorter request that causes a wrong answer creates more work. Clear, sufficient information is the goal.
Other meanings of “token”
An access token is a credential used to authorize access. [3] A design token is a named design value, such as a color or spacing size. [4] Neither is an AI token. If a bot uses the word without explanation, ask which meaning applies.
Also, a long response is not evidence of better work. For the reminder task, two accurate sentences may satisfy the goal better than a page of polished speculation.
What to ask your bot
Try:
Use the approved event note and my tone preference. Keep the reminder under 80 words. Identify any missing facts.That controls the useful output in ordinary language. For exact token usage, consult the product’s own reporting. Continue with context and memory to understand what information reaches the model.
Sources
- Google, Understand and count tokens, checked September 25, 2026. Supports tokenization, input/output counting, and provider-specific counting of different input types. No Gemini limits, prices, or text-to-token ratios are applied to Grok Bot.
- OpenAI, Conversation state, checked September 25, 2026. Supports context/output limits and the inclusion of earlier conversation in model requests. The event example and advice are original.
- IETF, RFC 6749, section 1.4: Access Token, published October 2012; checked September 25, 2026. Used for the meaning of access token, not implementation advice.
- Design Tokens Community Group, Design Tokens Format Module, checked September 25, 2026. Supports the distinct design meaning of token.
- Anthropic, Token counting, checked September 25, 2026. Supports including messages, system information, and tools when counting a request.
