What is a token?
A token is the basic unit of text a language model works with: a word, part of a word or a punctuation mark. A model’s context limit, speed and price are all measured in tokens.
How text becomes tokens
Each model family has its own tokenizer, so the same text can come out as a different number of tokens on different models. In English, a token is roughly four characters on average. Languages with long, inflected words tend to need more tokens for the same content, and code counts differently from prose because of indentation, symbols and long identifiers.
Input and output tokens
Pricing and limits treat token types differently.
- Input: everything sent to the model, including instructions, history, files and tool results.
- Output: the model’s reply, usually priced higher per token than input.
- Cached input: a repeated prefix read from the cache costs much less.
- Thinking tokens: with reasoning models, the thinking produced before the answer counts as output.
Why agents use so many tokens
A coding agent resends the whole context to the model on every turn. Over a long session, files read, test output and history pile up, and each new step adds to that load. That’s why most of the spend comes not from the code an agent writes but from the files and command output it reads.
Cutting usage
The biggest savings come from text that never needed to reach the model.
- Shorten long command output or pass only the error lines.
- Give small jobs to a cheaper model.
- Start a new conversation when the topic changes.
- Turn off tools and MCP servers you don’t use.
Tokens in AgentVera
AgentVera has tools both to cut token use and to measure it.
- RTK shortens agents’ test, build and git output before it reaches the model, typically 60–90% fewer tokens.
- Smart tools cut Claude Code’s repeat reads and needless full test runs.
- Automatic model gives small jobs to a cheaper model.
- The AI ROI tab shows tokens per successful task, a model comparison and tokens saved; cost is shown in tokens, not money.
- Token savingsContext profiles, shorter command output with RTK, live context size and automatic /compact.
- Smart toolsCuts repeat reads, full test runs and repeated commands before they reach the model.
- AI ROI analyticsAgent success rate, tokens per task, follow-ups and commits that needed fixing later.
FAQ
How many words is one token?
There’s no fixed ratio. In English a token is roughly four characters, or about three quarters of a word; other languages and code often need more.
Why does the same text cost different amounts on different models?
Each model family uses its own tokenizer and its own price per token, so both the count and the rate can differ.
How can I see how many tokens I use?
Providers offer token counting tools, and API responses include usage fields. Coding agents also usually show a session’s context size or usage.
Do tokens matter on a subscription?
Yes. Even if you aren’t billed per token, usage limits fill up based on how much you use, so fewer tokens means hitting the limit later.