What is a token?

In short

A token is the basic unit of text a language model works with: a word, part of a word or a punctuation mark. A model’s context limit, speed and price are all measured in tokens.

How text becomes tokens

Each model family has its own tokenizer, so the same text can come out as a different number of tokens on different models. In English, a token is roughly four characters on average. Languages with long, inflected words tend to need more tokens for the same content, and code counts differently from prose because of indentation, symbols and long identifiers.

Input and output tokens

Pricing and limits treat token types differently.

  • Input: everything sent to the model, including instructions, history, files and tool results.
  • Output: the model’s reply, usually priced higher per token than input.
  • Cached input: a repeated prefix read from the cache costs much less.
  • Thinking tokens: with reasoning models, the thinking produced before the answer counts as output.

Why agents use so many tokens

A coding agent resends the whole context to the model on every turn. Over a long session, files read, test output and history pile up, and each new step adds to that load. That’s why most of the spend comes not from the code an agent writes but from the files and command output it reads.

Cutting usage

The biggest savings come from text that never needed to reach the model.

  • Shorten long command output or pass only the error lines.
  • Give small jobs to a cheaper model.
  • Start a new conversation when the topic changes.
  • Turn off tools and MCP servers you don’t use.

Tokens in AgentVera

AgentVera has tools both to cut token use and to measure it.

  • RTK shortens agents’ test, build and git output before it reaches the model, typically 60–90% fewer tokens.
  • Smart tools cut Claude Code’s repeat reads and needless full test runs.
  • Automatic model gives small jobs to a cheaper model.
  • The AI ROI tab shows tokens per successful task, a model comparison and tokens saved; cost is shown in tokens, not money.

FAQ

How many words is one token?

There’s no fixed ratio. In English a token is roughly four characters, or about three quarters of a word; other languages and code often need more.

Why does the same text cost different amounts on different models?

Each model family uses its own tokenizer and its own price per token, so both the count and the rate can differ.

How can I see how many tokens I use?

Providers offer token counting tools, and API responses include usage fields. Coding agents also usually show a session’s context size or usage.

Do tokens matter on a subscription?

Yes. Even if you aren’t billed per token, usage limits fill up based on how much you use, so fewer tokens means hitting the limit later.

Related terms

All terms

See these ideas at work.

Download AgentVera for free, add a project folder and run Claude Code, Codex and other agents side by side.