Prompt caching in practice: cutting repeat-context costs on the Claude API

•By Blacdisk Team

Prompt caching in practice: cutting repeat-context costs on the Claude API

If your application sends the same 20,000 tokens of instructions, tool definitions, or reference documents on every request, you are paying full input price for work the API has already done. Prompt caching fixes that by letting the API reuse a previously processed prefix at a fraction of the cost. In the worked example below, a workload that costs $4.10 uncached drops to about $0.55.

This guide covers what gets cached, what it costs, how to set it up, and how to confirm it's working. Pricing and thresholds are taken from Anthropic's documentation as retrieved on 2026-09-28; both change, so check the linked pages before you budget.

Key Takeaways

  • Cache reads cost 0.1x the base input price on most models (0.05x on Claude Opus 5.5, 0.025x on Claude Fable 5.1), while 5-minute cache writes cost 1.25x. A cached prefix is cheaper than an uncached one from the second request onward.
  • Put cache_control on the last block that is identical across requests. A breakpoint on a block that changes every time (a timestamp, the user's message) never produces a hit.
  • Use the 5-minute TTL when traffic is steady, since every hit refreshes it for free. Use the 1-hour TTL when gaps between requests are longer than five minutes.
  • Track cache_read_input_tokens and cache_creation_input_tokens on every response. If both are zero, nothing was cached.

What does prompt caching actually cache?

Prompt caching stores a prefix of your request, not a response. The API caches everything from the start of the prompt up to and including the block you mark, in a fixed order: tools, then system, then messages (Claude Platform Docs: Prompt caching, retrieved 2026-09-28). The model still runs on every request and the output is identical to what you'd get without caching. Only the input side gets cheaper.

Three properties shape everything else in this article:

  1. Exact matching. A cache hit requires the prefix to be identical, including all text and images up to the marked block. Change one character early in the prompt and everything after it is a miss.
  2. Default lifetime of five minutes, refreshed on use. Each time the cached content is read, the timer resets at no extra charge. The lifetime is measured from the start of the request, so a slow streaming response eats into it.
  3. Writes happen only at your breakpoints. The docs describe the write as a hash of the prefix ending at the marked block. Reads look backward for entries that earlier requests wrote, over a window of 20 blocks.

How much does prompt caching save?

Cache pricing is a multiplier on the base input rate. Per the same documentation, a 5-minute write costs 1.25x, a 1-hour write costs 2x, and a read costs 0.1x, with lower read multipliers on Claude Opus 5.5 (0.05x) and Claude Fable 5.1 and Mythos 5.1 (0.025x).

Here is a worked example on Claude Sonnet 5.5, which lists at $2 per million input tokens, $2.50 for 5-minute writes, and $0.20 for reads. Assume a 20,000-token static prefix (system prompt, tools, reference material), a 500-token unique message on each request, and 100 requests inside the cache window.

Tokens Rate (per MTok) Cost
No caching: 100 × 20,500 tokens 2,050,000 $2.00 $4.10
Caching: one 5-minute write 20,000 $2.50 $0.05
Caching: 99 reads 1,980,000 $0.20 $0.40
Caching: 100 fresh user messages 50,000 $2.00 $0.10
Caching total ≈ $0.55

That is roughly an 87% reduction on input cost for this workload. Output tokens are billed the same either way, so your real saving depends on how input-heavy your traffic is.

Input cost for 100 requests, 20K-token shared prefix $4.10 ≈ $0.55 No caching With caching
Source: author's calculation from Claude Sonnet 5.5 rates in the Claude Platform Docs, retrieved 2026-09-28. Input tokens only.

When does caching start paying off?

Take a request that would cost 1.0 unit uncached. With a 5-minute cache, the first request costs 1.25 units (a write) and every later request costs 0.1 units (a read). Two requests cost 1.35 units cached versus 2.0 uncached, so a single reuse is enough to come out ahead.

The 1-hour TTL has a higher write cost of 2.0 units. Two requests cost 2.1 units versus 2.0 uncached, which is slightly worse. Three requests cost 2.2 versus 3.0, which is better. Use the 1-hour option only when you expect at least a few reads and the gaps are too long for the 5-minute cache to survive.

How do you turn prompt caching on?

There are two ways, and both use the same cache_control field.

Option 1: Automatic caching

Add one top-level cache_control field. The API places the breakpoint on the last cacheable block and moves it forward as a conversation grows. This is the simplest path for multi-turn chat.

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    cache_control={"type": "ephemeral"},
    system="You are a support assistant for Acme. ...long instructions...",
    messages=[{"role": "user", "content": "How do I reset my password?"}],
)
print(response.usage)

Option 2: Explicit breakpoints

Place cache_control on individual blocks when different parts of your prompt change at different rates. You can use up to four breakpoints per request, and adding breakpoints doesn't add cost; you pay only for what is actually written and read.

A common layout for a retrieval or agent workload:

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    tools=TOOLS,  # cache_control on the last tool caches all tool definitions
    system=[
        {"type": "text", "text": STATIC_INSTRUCTIONS,
         "cache_control": {"type": "ephemeral"}},   # rarely changes
        {"type": "text", "text": KNOWLEDGE_BASE_DOCS,
         "cache_control": {"type": "ephemeral"}},   # changes daily
    ],
    messages=conversation,  # growing history; final block can carry a breakpoint too
)

With this structure, updating the knowledge base invalidates only the documents segment and the conversation after it. The tools and instructions segments stay cached.

Where should the cache breakpoint go?

Place it on the last block whose content is identical across the requests you want to share a cache. This is the most common source of "caching doesn't work" reports, so it's worth spelling out.

Consider a prompt with a large static context followed by a per-request block holding a timestamp and the user's question. If you mark that final block, the hash includes the timestamp, so it changes every time. The docs describe this outcome directly: you pay for a fresh cache write on every request and never get a read, because the lookback only finds entries that earlier requests wrote at their own breakpoints.

The fix is to mark the end of the static prefix and keep the varying content after it:

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    system=[
        {"type": "text", "text": STATIC_CONTEXT,
         "cache_control": {"type": "ephemeral"}},  # breakpoint on stable content
    ],
    messages=[
        {"role": "user",
         "content": f"Current time: {now}\n\n{question}"},  # varies, uncached
    ],
)

Automatic caching has the same trap, because it places the breakpoint on the last block. For a prompt shaped like this one, use an explicit breakpoint on the static block instead.

For long conversations, remember the 20-block lookback window. If a single turn appends 20 or more blocks past your last write, the lookback misses it. A second breakpoint closer to the recent content prevents that.

Which TTL should you use?

Pick by how often the same prefix is requested:

system=[{
    "type": "text",
    "text": STATIC_CONTEXT,
    "cache_control": {"type": "ephemeral", "ttl": "1h"},
}]

You can mix TTLs in one request, but longer-TTL entries must appear before shorter ones. The documentation also notes that cache hits aren't deducted from your rate limit, which can matter as much as cost for high-volume workloads (see the 1-hour cache section of the docs).

How do you check that caching is working?

Read three fields from usage on every response:

Total input is the sum of all three. A small helper makes this easy to log:

def cache_hit_rate(usage) -> float:
    total = (usage.cache_read_input_tokens
             + usage.cache_creation_input_tokens
             + usage.input_tokens)
    return usage.cache_read_input_tokens / total if total else 0.0

On the first request, expect a large cache_creation_input_tokens and zero reads. On later requests inside the TTL, that should flip: reads large, writes small, input_tokens near the size of the new user message. If both cache fields are zero, the prompt was not cached, and the API returns no error in that case.

Why are you getting cache misses?

Work through this checklist in order:

  1. Prompt below the minimum length. Minimums vary by model: 512 tokens for Claude Sonnet 5.5 and Opus 5.5, 1,024 for Sonnet 5 and Sonnet 4.6, 4,096 for Claude Haiku 4.5. Shorter prompts are processed without caching and without an error.
  2. Breakpoint on a changing block. See the timestamp example above.
  3. Something upstream changed. Editing tool definitions invalidates tools, system, and messages. Toggling web search or citations changes the system prompt. Changing tool_choice, adding or removing images, or altering thinking or effort settings invalidates the message cache. The docs have a full table.
  4. Parallel requests fired before the first finished. A cache entry becomes available only after the first response begins. If you fan out 50 requests at once, send one first and wait for it to start responding.
  5. Unstable serialization. Some languages, such as Go and Swift, randomize key order when converting to JSON. That changes tool_use block contents between requests and breaks the match.
  6. Different workspaces. On the Claude API, caches are isolated per workspace, so two workspaces never share an entry.

Anthropic also offers cache diagnostics, which compare consecutive requests and report where the prefix diverged.

Can you warm the cache before users arrive?

Yes. Send a request with max_tokens: 0 and a breakpoint on your shared prefix. The API processes the prompt, writes the cache, and returns immediately with no output tokens billed. You still pay the write charge if the prefix wasn't already cached.

client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=0,
    system=[{"type": "text", "text": STATIC_CONTEXT,
             "cache_control": {"type": "ephemeral"}}],
    messages=[{"role": "user", "content": "warmup"}],
)

Two cautions from the docs: put the breakpoint on the shared system content, not the placeholder message, and keep thinking and effort settings identical to your real requests so the warmed entry is the one your traffic actually hits. Pre-warming doesn't work with streaming, extended thinking, structured outputs, or inside batch requests.

Frequently Asked Questions

Does prompt caching change Claude's responses?

No. Caching affects only how the input prefix is processed and billed. The output is identical to what you would receive without it.

Is prompt caching safe for sensitive data?

Anthropic states that prompt caching is eligible for zero data retention. Cache entries are isolated between organizations, and on the Claude API also between workspaces. See the data retention section of the docs for details on your compliance situation.

Should I use automatic caching or explicit breakpoints?

Start with automatic caching for chat-style, multi-turn apps. Move to explicit breakpoints when your prompt has a large static prefix plus a varying suffix, or when different sections update at different frequencies.

Conclusion

Prompt caching is one of the few optimizations that lowers cost and latency without touching output quality. The recipe is short:

If your bill has a large, repeated prefix in it, you can likely test this in an afternoon: add one cache_control field, run a hundred requests, and compare cache_read_input_tokens to the total.


Sources: Claude Platform Docs: Prompt caching (retrieved 2026-09-28). Prices, model lineups, and minimum token thresholds change; verify against the pricing page before making budget decisions.