What is prompt caching?
Prompt caching means the model provider stores the unchanged opening part of a request and reuses it on later requests instead of processing it again. Input read from the cache is much cheaper and faster than regular input.
How it works
The cache works on the prefix: if a request matches the previous one exactly from the start, that part is read from the cache. The first request writes the cache and later ones read it. A single change in the prefix means everything after it is processed again.
Why it matters for coding agents
An agent resends its system instructions, tool definitions and conversation history on every turn. Since that part barely changes, a warm cache means most of a long session’s input comes from cheap cache reads. Agents like Claude Code and Codex use caching on their own; you don’t need to configure anything.
What breaks the cache
A cache miss means the whole conversation gets written again.
- Long pauses: with Anthropic the cache lives 5 minutes by default; wait longer and it’s rewritten.
- Switching models: the cache is specific to a model.
- Changing the prefix: editing the system prompt, tool list or MCP servers mid-session.
- Compaction: /compact rewrites the conversation, so the cache starts over.
Pricing
Providers price cached input well below regular input. With Anthropic, writing to the cache costs a bit more than regular input, while OpenAI applies caching automatically to long enough prompts. Check your provider’s pricing page for current rates.
Prompt caching in AgentVera
The prompt cache card in the AI ROI tab shows how well your agents’ conversations use the cache. It is all worked out on your computer.
- The share of input read from cache and the tokens the cache saved.
- Why the cache was rewritten: waits over 5 minutes, or model and context changes.
- A list of large conversations that used the cache least.
FAQ
Do I need to turn prompt caching on?
Not in agents like Claude Code and Codex; they use it themselves. If you’re building on the API, check your provider’s docs: some need you to mark cache points in the request, others cache automatically.
How long does the cache last?
It depends on the provider. With Anthropic the default is 5 minutes, refreshed each time the cache is used, and a longer option is available.
Does caching change the answers?
No. It only avoids reprocessing identical input; the model sees the same context.
Why does a long pause make the next turn expensive?
Once the pause outlasts the cache, the next request writes the whole conversation to the cache again. In a long conversation that’s a large input cost all at once.