I tracked my token usage for a month: here's where 70% of it went

•By Blacdisk User - Koketso

I tracked my token usage for a month: here's where 70% of it went

For 30 days I logged every token I used on the Claude API. The most surprising result: 70% of my tokens came from re-reading old context in long sessions, not from the prompts I thought I was paying for. But here's the twist: those re-reads only accounted for 7% of my actual bill. The real cost driver was hiding elsewhere.

I run a small dev agency and use Claude Code as my primary pair programmer. Last month, my API bill spiked to nearly $300, and I kept hitting limits. I assumed my prompts were too complex, so I started logging everything to find the leak.

This post covers how I tracked it, what the breakdown looked like, why token share and dollar share tell different stories, and the changes that moved the number. If you'd rather run the audit yourself, the script below produces the same breakdown from Anthropic's Usage API. Product behavior and prices are from Anthropic's documentation as retrieved on 2026-09-28.

Key Takeaways

  • 70% of my tokens came from cache reads (re-reading old context). The reason? I kept terminal sessions open for days.
  • Token share and cost share diverge sharply. Cache reads are billed at a fraction of normal input price, so a bucket that dominates token counts can be a small part of the bill. My cache reads were 70% of tokens but only 7% of cost.
  • Long-lived sessions, scheduled or idle work, and thinking-heavy output are the usual suspects, and each has a documented lever.
  • Three changes cut my month-two spend by 66%.

How did I track a month of token usage?

There are three official ways to see where tokens go, and which you use depends on how you're billed (Manage costs effectively, retrieved 2026-09-28):

I used the Usage API combined with the /usage command. I run everything through a single Anthropic organization for my agency, so the Admin API gave me the exact daily buckets I needed. It doesn't capture my occasional claude.ai usage on my phone, but for API spend, it was perfect.

If you're on the API, this script pulls 30 daily buckets and splits them by token type and model. The field names below follow the Usage API's documented token categories; confirm them against the Usage API reference before trusting the output.

import os
from collections import defaultdict
from datetime import datetime, timedelta, timezone

import requests

KEY = os.environ["ANTHROPIC_ADMIN_KEY"]
URL = "https://api.anthropic.com/v1/organizations/usage_report/messages"

# $ per million tokens: input, 5m cache write, 1h cache write, cache read, output.
# Example: Claude Sonnet 5.5 list prices. Add every model you use; check the pricing page.
RATES = {"claude-sonnet-5-5": (2.00, 2.50, 4.00, 0.20, 10.00)}

end = datetime.now(timezone.utc).replace(hour=0, minute=0, second=0, microsecond=0)
start = end - timedelta(days=30)
params = {
    "starting_at": start.strftime("%Y-%m-%dT%H:%M:%SZ"),
    "ending_at": end.strftime("%Y-%m-%dT%H:%M:%SZ"),
    "bucket_width": "1d",
    "group_by[]": ["model"],
    "limit": 31,
}
headers = {"x-api-key": KEY, "anthropic-version": "2023-06-01"}

tok = defaultdict(lambda: defaultdict(int))
while True:
    page = requests.get(URL, params=params, headers=headers, timeout=30).json()
    for bucket in page["data"]:
        for r in bucket["results"]:
            t = tok[r["model"]]
            cw = r.get("cache_creation", {})
            t["input"] += r.get("uncached_input_tokens", 0)
            t["w5m"] += cw.get("ephemeral_5m_input_tokens", 0)
            t["w1h"] += cw.get("ephemeral_1h_input_tokens", 0)
            t["read"] += r.get("cache_read_input_tokens", 0)
            t["output"] += r.get("output_tokens", 0)
    if not page.get("has_more"):
        break
    params["page"] = page["next_page"]

cats = ["input", "w5m", "w1h", "read", "output"]
total_tokens = sum(t[c] for t in tok.values() for c in cats) or 1
cost = defaultdict(float)
for model, t in tok.items():
    rates = RATES.get(model)
    if not rates:
        print(f"no rates for {model}, skipping cost")
        continue
    for c, rate in zip(cats, rates):
        cost[c] += t[c] * rate / 1_000_000
total_cost = sum(cost.values()) or 1

for c in cats:
    n = sum(t[c] for t in tok.values())
    print(f"{c:7} {n/total_tokens:6.1%} of tokens   {cost[c]/total_cost:6.1%} of cost")