Why Does Your 40th Claude Code Message Cost More Than Your 4th?

•By Blacdisk Team

Why Does Your 40th Claude Code Message Cost More Than Your 4th?

Send the same one-line request at message 4 and again at message 40 of a Claude Code session, and the second one carries far more weight. In the simple model used in this post, it sends about 4.4 times as much input to the model, even though your prompt is identical. The difference is everything riding along with it.

That hidden cargo is the core of Claude Code token usage. The model doesn't remember your session between messages, so every message brings the whole conversation so far along with it. Prompt caching softens the bill without ending the growth, and a single idle coffee break can undo the savings. Below you'll find the mechanics, a worked model with numbers, and six habits that keep long sessions affordable.

Key Takeaways

  • Claude Code re-sends the whole conversation with every message, so input grows as a session lengthens. In this post's model, message 40 sends about 4.4 times the input of message 4.
  • Prompt caching bills reads at a tenth of the normal input price, which cuts the model's message-40 cost from 93,000 to 11,600 token-equivalents, but only while the cache stays warm.
  • The default cache lifetime is five minutes of inactivity. After an idle gap, the same message costs about ten times as much as it would have warm.
  • The fixes are cheap: /clear between unrelated tasks, compact on purpose with instructions, check your context between tasks, and keep each exchange lean.

Why does each Claude Code message cost more than the last?

Claude's API is stateless: nothing carries over between requests unless the client sends it again. So Claude Code sends the system prompt, the tool definitions, your CLAUDE.md, every earlier prompt, every file it read, every command's output, and every reply, again, on every message. One course on compaction puts it plainly: each new message includes the full previous context, which is why costs climb as a conversation gets longer (Steve Kinney, retrieved 2026-09-21). Anthropic's own cost guide makes the same point in practical terms: stale context wastes tokens on every message that follows (Claude Code Docs, retrieved 2026-09-21).

To see the shape of the growth, here's a deliberately simple model. Assume a session starts with 15,000 tokens of fixed context (system prompt, tool definitions, CLAUDE.md) and that each exchange adds 2,000 tokens: your prompt, any tool results, and Claude's reply. Real exchanges vary a lot, and a single large file read can add more than that in one go, so treat these as round numbers rather than measurements.

Input sent per message keeps climbing Line chart: input tokens sent per message rise from 15,000 at message 1 to 93,000 at message 40, with no caching. Input sent per message keeps climbing Illustrative model, no caching 0 25k 50k 75k 100k 1 10 20 30 40 Message 4: 21,000 Message 40: 93,000 Message number Input tokens sent
Illustrative model: 15,000-token starting context, +2,000 tokens per exchange, no caching. Source: author's calculation.

At message 4 the model sends 21,000 tokens. At message 40 it sends 93,000, about 4.4 times as many. Nothing about the request changed; the history did the growing.

The fastest way to inflate that history is tool output. A test log, a large file read, or a directory listing enters the conversation once and then gets re-sent with every later message until the session ends or is compacted.

How fast does a long session add up?

Because each message re-sends a longer history than the one before it, total input over a session grows faster than the message count. In the model, the first four messages process about 72,000 input tokens in total. Forty messages process about 2.16 million, which is 30 times as much for 10 times the messages. A straight-line pace would have landed at 720,000.

Total input grows faster than the message count Bar chart: cumulative input tokens processed after 4, 10, 20, 30 and 40 messages, compared with a straight-line pace. Total input grows faster than the message count Illustrative model, cumulative input tokens, no caching 0 500k 1M 1.5M 2M 72k 4 240k 10 680k 20 1.32M 30 2.16M 40 Messages in the session Dashed: straight-line pace of 18,000 tokens per message
Illustrative model: cumulative input tokens across a session, no caching. Source: author's calculation.

That gap is why a session can feel cheap for a while and then suddenly doesn't. Anthropic's docs put average enterprise spend at roughly $13 per developer per active day, with 90% of users staying under $30 (Claude Code Docs, retrieved 2026-09-21). The same page says that when spend runs unexpectedly high, the usual causes are sessions that were never cleared or Opus left as the default model. The first of those is the one this post is about, and it's the one you control most directly.

What does prompt caching change, and what doesn't it?

Prompt caching lets Anthropic reuse the already-processed front of a prompt instead of reprocessing it. Reads from the cache cost 10% of the normal input price. Writing to the cache carries a 25% premium for the default five-minute lifetime, and a one-hour option costs more to write (Anthropic prompt caching docs, retrieved 2026-09-21; Technspire, retrieved 2026-09-21). The break-even is quick. A five-minute write plus one read costs 1.35 times the base price, versus 2 times for sending the same prefix twice uncached (Technspire). Claude Code turns this on for you, along with auto-compaction, which summarizes history as you approach the context limit (Claude Code Docs, retrieved 2026-09-21).

Apply that to the model. Each message reads the earlier conversation from the cache at 10% and writes only the new 2,000 tokens at the 25% premium. Measured in base-input-token equivalents, message 4 costs 4,400 and message 40 costs 11,600. The ratio shrinks from 4.4x to about 2.6x. Across all 40 messages, the total drops from 2.16 million to roughly 323,000, about 6.7 times lower.

Caching flattens the slope, not the line. Every message still re-reads an ever-larger prefix, just at a tenth of the price. And the rates differ by model: Anthropic's docs list cache hits on Fable 5.1 and Mythos 5.1 at 0.025 times the base input price, while other models use the standard 0.1 times (Anthropic prompt caching docs, retrieved 2026-09-21). The exact ratios shift with the model; the shape doesn't.

What happens to the bill when the cache goes cold?

The cache only stays warm while you keep working. The default lifetime is five minutes of inactivity, and each cache read resets the clock at no extra charge (Developers Digest, retrieved 2026-09-21). Step away for lunch, and the next message finds nothing to read. The entire prefix gets rewritten at the 1.25x premium.

In the model, message 40 after a cold cache costs 116,250 token-equivalents, about ten times the warm cost of 11,600. It even lands 25% above having no caching at all (93,000), because a cache write with no read behind it is pure surcharge (Developers Digest).

Warm cache vs no cache vs cold cache Grouped bar chart: cost of message 4 and message 40 with no caching, a warm cache, and a cold cache, in base-input-token equivalents. Warm cache vs no cache vs cold cache Illustrative model, cost in base-input-token equivalents No caching Warm cache Cold cache 0 25k 50k 75k 100k 125k 21,000 4,400 26,250 Message 4 93,000 11,600 116,250 Message 40 Cost (base-input-token equivalents)
Illustrative model: cost of one message in base-input-token equivalents. Warm = cache read at 0.1x plus new tokens written at 1.25x; cold = full prefix written at 1.25x. Source: author's calculation using Anthropic's published multipliers.

One caution: five minutes is the documented default for the API, and the lifetime a specific tool uses is an implementation detail that can change. Treat it as a working assumption, not a guarantee. The habit that follows is the same either way: a long idle gap makes a big context expensive.

It also changes the compact-or-clear decision. One guide argues that /compact works best while the cache is still warm, because the summarizing pass reads your context at the discount. After a long break that pass runs at full price, and a fresh start with /clear can be cheaper (systemprompt.io, retrieved 2026-09-21). It's third-party guidance, but the logic follows directly from the pricing above.

/clear vs /compact explained

Do long sessions cost more than money?

They do, in quality. A third-party explainer on Claude Code's context window notes that a window that isn't technically full can still produce worse output, because models don't weight every position in a long context equally and material in the middle gets less reliable attention (GeoToolbox, retrieved 2026-09-21). The same piece describes compaction as happening in two stages: older tool outputs are cleared first, and the conversation is summarized only if that isn't enough. Your requests and key code snippets survive, while detailed instructions from early in the session may not.

That makes compaction a lossy step, and timing matters. Compact at a clean boundary, such as after finishing a phase of work, and the summary is built from a good state. Compact a session that has already gone off track and you bake the confusion into the summary (GeoToolbox).

[INTERNAL-LINK: subagents as a token strategy → post on subagents]

Six habits that keep long sessions affordable

1. Clear between unrelated tasks

Anthropic's docs recommend /clear when you switch to unrelated work, and suggest running /rename first so you can find the old session later with /resume (Claude Code Docs, retrieved 2026-09-21). It costs nothing and removes the whole growing prefix in one move. One task per session is the simplest rule to follow.

2. Compact on purpose, with instructions

Don't wait for auto-compaction to pick the moment. Run /compact yourself at a clean boundary and tell it what to keep, for example /compact Keep the failing test output and the schema decisions. If you type the same steer every time, the docs say you can add a "Compact instructions" section to your CLAUDE.md so every compaction in that project inherits it (Claude Code Docs, retrieved 2026-09-21).

3. Make your context visible

The docs point to /usage for token usage and a configurable status line that shows it continuously (Claude Code Docs, retrieved 2026-09-21). The /context command shows how full the window is (GeoToolbox). A ten-second check between tasks tells you whether a clear is overdue.

4. Watch the clock

If you're active and the session is getting long, compact while the cache is warm. If you're coming back from a long break, weigh a fresh start instead (systemprompt.io). Writing decisions and a task list to a markdown file before you leave makes that fresh start cheap, because Claude can re-read a short file instead of a long history.

5. Keep each exchange lean

Name the file instead of asking Claude to explore. Trim logs before they enter the conversation, for instance by sending the last 50 lines rather than the whole output. Delegate wide searches to subagents, which work in their own context window and hand back a result. What they return still joins your main session, though, and one guide warns that several subagents returning large outputs at once can fill the main context quickly (Build to Launch, retrieved 2026-09-21). Ask for short summaries.

6. Match the model to the job

Anthropic's docs say Sonnet handles most coding tasks well and costs less than Opus, which they suggest reserving for complex architectural decisions or multi-step reasoning; you can switch mid-session with /model (Claude Code Docs, retrieved 2026-09-21). Model choice changes the price per token. It doesn't change how fast your history grows, so pair it with the habits above rather than substituting for them.

Frequently Asked Questions

Does /compact itself cost tokens?

Yes. Compaction is a model request that has to read the context it's summarizing, so it costs roughly one more pass over that context. It pays for itself when the summary is much shorter than what it replaces and you'll keep working on the same thread. Running it while the cache is warm keeps that pass cheap.

Do subscription users see this too?

The mechanics are the same: bigger contexts mean more tokens processed. What differs is how they show up. Anthropic's docs frame Claude Code cost as API token consumption and send subscribers to the pricing page for plan details (Claude Code Docs, retrieved 2026-09-21), so check how your plan meters usage before drawing conclusions about dollars.

Doesn't auto-compact solve this?

It's a safety net, not a strategy. Auto-compaction runs when the window nears its limit, which means you've already carried a large context through many messages, and it fires at a moment you didn't choose. Detailed early instructions may not survive the summary (GeoToolbox). Compacting deliberately, at a clean boundary and with instructions, gives you a better result.

When is a long session worth it?

When the work is one coherent thread that keeps using what came before, like a multi-step refactor. Clearing has a cost too: wipe context you still needed and you pay to re-explain it (Build to Launch, retrieved 2026-09-21). A workable rule: unrelated task, clear; same task with a long history, compact; back from a long break, consider a fresh start with a saved notes file.

Conclusion

Your 40th message costs more than your 4th because the conversation is the cargo, and it only gets heavier. Caching cuts what you pay to carry it, but only while the cache is warm.

Try it today: run /context in your current session and note the percentage, then clear or compact at your next natural break and compare.

if hate burning tokens and claude usage quickly check out the Blacdisk tool and how it can save you up to 80 percent on input tokens.

Sources and method

Model assumptions. 15,000-token starting context; each exchange adds 2,000 tokens; message n sends 15,000 + 2,000 × (n − 1) input tokens. Warm-cache cost = 0.1 × the previous message's context + 1.25 × 2,000 new tokens (the first message writes its full context at 1.25x). Cold-cache cost = 1.25 × the full context. Output and extended-thinking tokens are excluded, and Anthropic's per-model rate differences are ignored. These are illustrative numbers, not measurements of a real session.

Sources (all retrieved 2026-09-21). Claude Code Docs: Manage costs · Anthropic prompt caching docs · Technspire: Anthropic prompt caching pricing mechanics · Developers Digest: Prompt caching economics · systemprompt.io: Claude Code cost optimisation · GeoToolbox: Claude Code context window · Build to Launch: Claude Code token optimization · Steve Kinney: Claude Code compaction

Pricing multipliers, model rates, and cache lifetimes change. Check Anthropic's current documentation before relying on any specific figure.