If you want to reduce Claude Code token usage, the prompt is rarely the place to start. The real weight is the baseline context added to every request before you type a single word. In long conversations it drives both cost and how fast you hit usage limits. This post explains where that baseline comes from, how we measured it, and which settings cut it by up to 61%.
Where do the tokens go?
When you message a coding agent, the model sees more than what you wrote. The request also contains:
- The system prompt
- Tool definitions (Bash, Read, Edit, Write, Glob, Grep and others)
- User-level plugins, hooks and skills
- Tool definitions from connected MCP servers
- The project’s instruction files (CLAUDE.md, AGENTS.md)
- The conversation so far
The first five are re-read on every turn, and the last one grows as you go. A small baseline adds up quickly over a hundred-turn session.
Measured: how much context is in the first request?
We measured the context sent to the model on the first request of a new session with Claude Code 2.1 and Codex 0.155:
| Profile | Claude Code | Codex |
|---|---|---|
| Standard | ≈34.4k | ≈13.3k |
| Balanced | ≈23.7k (−31%) | — |
| Lean | ≈13.3k (−61%) | ≈9.7k (−27%) |
The balanced profile limits the tool set to Bash, Read, Edit, Write, Glob and Grep. The lean profile also skips user-level plugins, hooks, skills and external MCP servers. On Codex, the lean profile turns off apps and web search.
Settings that didn’t help
While measuring, we found that some flags did not shrink the first request on their own:
--strict-mcp-config--exclude-dynamic-system-prompt-sections--disable-slash-commands
Claude Code already loads those parts only when needed. The real savings come from narrowing the tool set and not loading user-level plugins, hooks and skills.
We deliberately left two options out: --safe-mode also turns off the project’s CLAUDE.md, and --bare only works with an API key. Both are too restrictive for daily use.
How to reduce Claude Code token usage, step by step
1. Give the task only the tools it needs
An agent updating docs or refactoring one file doesn’t need a browser, image tools or a dozen MCP servers. Start these jobs with the balanced or lean profile.
2. Attach MCP servers selectively
Every MCP server adds its tool definitions to the context. Give a server only to the agents that will use it, not to all of them. AgentVera’s MCP manager shows the approximate token cost of each tool when you test a connection, so you spot a heavy server before it bloats every turn.
3. Watch the context size
You can’t save what you can’t see. AgentVera shows each agent’s live context size in its pane, so you look at a number instead of guessing when a conversation got expensive.
4. Compact long conversations
/compact summarizes the conversation so far and shrinks every turn after it. Run it by hand, or set a threshold in AgentVera to compact automatically once context passes it.
5. Limit what moves between agents
When a flow passes one agent’s output to another, it doesn’t have to send all of it. The Transform node in flows can truncate a reply, keep the first or last N lines, or keep only the last code block, so the target agent’s context doesn’t fill with noise.
6. Polish prompts with a cheaper model
If you want to tidy a draft before sending it, you don’t need the expensive main model. AgentVera’s Revise runs claude -p --model sonnet with tools off and no session log. File contents are not sent; only the draft, the agent’s name and role, and the project name.
Choosing a profile
| Work | Suggested profile | Why |
|---|---|---|
| Browser testing, MCP-heavy integration | Standard | Needs the full tool set |
| Everyday feature work | Balanced | Core file and shell tools are enough |
| Docs, small refactors | Lean | Smallest baseline, cheapest turns |
| Long research session | Balanced + /compact at a threshold | Keeps context growth in check |
Check the numbers on your own setup
Your numbers depend on your plugins, skills and MCP servers. To see your own baseline, open a new session, send one short message and look at the context size of the first request. Then try the same with a narrower profile. The difference is what you pay on every turn of the session.
Working out the difference over a session
Percentages can feel abstract, so think in terms of a session. Say you run 40 turns with one agent. Leaving conversation history aside, the baseline is re-read on every turn.
| Profile | Baseline on the first request | Baseline alone over 40 turns |
|---|---|---|
| Standard | ≈34.4k | ≈1.38M |
| Balanced | ≈23.7k | ≈948k |
| Lean | ≈13.3k | ≈532k |
This table only shows the baseline; real usage is higher once conversation history, tool output and the model’s replies are added. Provider mechanisms such as caching can also affect billing. Still, the table makes one thing clear: a smaller baseline makes every turn cheaper, and the gap grows as the session gets longer.
Which profile for which job?
The profile isn’t a set-and-forget setting. It depends on the agent’s role:
- Research agent: reads many files and may need a browser or docs MCPs. Start with the standard profile and automatic compaction at a threshold.
- Implementation agent: mostly reads and writes files and runs shell commands. Balanced is enough.
- Test or docs agent: works in a narrow area. Lean saves the most here.
With parallel agents these differences add up. Giving each agent a profile that fits its role, instead of standard for all three, makes the same limit last longer.
Setting automatic /compact at a threshold
Manual compaction is easy to forget. Automatic compaction is simple: once an agent’s context passes your threshold, /compact runs. When picking the threshold:
- Too low: the agent is summarized often and may lose earlier details.
- Too high: compaction comes late, and every turn until then is expensive.
In practice an earlier threshold works for long research sessions and a later one for short implementation sessions. Since AgentVera shows the live context size in the pane, you can tune the threshold by watching it.
Common mistakes
- Every MCP for every agent. Unused tool definitions are read on every turn.
- Never compacting. A session with hundreds of turns keeps carrying old details it no longer needs.
- Passing output as is. Sending a long reply in full to another agent bloats that agent’s context too.
- Only trimming the prompt. The prompt is usually a small part of the total.
Wrapping up
Most Claude Code token savings come from the baseline that’s re-read on every turn. Narrowing tools to the task, attaching MCP servers selectively and compacting at the right time lets you do the same work with less context. On the setup we measured, that meant 61% less on the first request.
AgentVera lets you set context profiles, the live context indicator and automatic /compact per agent. See token savings and the Claude Code page for details, or start on the free plan.