Setting spend limits and monitoring usage per project on the Claude API

•By Blacdisk

Setting spend limits and monitoring usage per project on the Claude API

One shared API key across three projects works fine until the month-end bill arrives and nobody can say which project caused the spike. Worse, a runaway loop in one project can burn the budget for all the others. The fix is structural: give each project its own boundary, cap what that boundary can spend, and report on it separately.

This guide covers the Claude Console (Claude Platform) setup: workspaces, the three layers of spend limits, what your code sees when a limit is hit, and how to pull cost per project programmatically. Figures and behaviors come from Anthropic's documentation as retrieved on 2026-09-28. Limits and tiers change, so confirm them on the linked pages.

Key Takeaways

  • A workspace is the unit for per-project limits. Create one per project or environment, and give each its own API keys, spend limit, and rate limits.
  • Limits stack: your tier's monthly cap ($500 Start, $1,000 Build, $200,000 Scale), an optional lower organization limit you set, and optional lower workspace limits. Workspace limits can't exceed the organization's, and they can't be set on the Default Workspace.
  • Hitting a limit you set returns HTTP 400. Hitting your tier's cap returns HTTP 429 with no retry-after header, so retries fail until the month resets.
  • For reporting, the Cost API groups spend by workspace_id and the Usage API groups tokens by workspace, model, or API key. Both need an Admin API credential.

Why should each project get its own workspace?

Workspaces exist to separate projects, environments, or teams while keeping billing and administration central (Claude Platform Docs: Workspaces, retrieved 2026-09-28). Each one holds its own members, service accounts, API keys, and limits. An organization can have up to 100 by default, not counting archived ones.

Three practical consequences follow:

Pick a naming scheme that says what the workspace is for. The docs suggest patterns like "Production - Customer Chatbot" or "Dev - Internal Tools". Splitting environments matters as much as splitting projects: development traffic with a small cap keeps experiments from eating production budget.

One trap deserves attention. Every organization has a Default Workspace, and you can't set limits on it. If production traffic runs there, it has no workspace-level cap. Create named workspaces and move real workloads out of Default.

How do spend limits work?

There are three layers, and the lowest one that applies is the one that stops you.

Layer Who sets it Where Notes
Tier cap Anthropic, by usage tier Billing page (view) Start $500, Build $1,000, Scale $200,000 per calendar month. Custom tier has no cap.
Organization limit You Settings > Billing Any value at or below your tier cap.
Workspace limit You Settings > Workspaces > workspace > Spend limits At or below the organization's. Not available on the Default Workspace.

Source: Rate limits and Workspaces, retrieved 2026-09-28.

If you don't set a workspace limit, it matches the organization's. The workspace Spend limits tab also lets you configure alerts at spending thresholds, so you can get a warning before the cap rather than a 400 after it.

Each workspace also has a Rate limits tab for requests per minute, input tokens per minute, and output tokens per minute. Spend limits control total cost per month, while rate limits control burst throughput and keep one project from starving the others of capacity.

How should you divide the budget?

The docs state that organization-wide limits always apply, even if workspace limits add up to more. That means a set of workspace caps that sums to more than the organization limit gives you soft isolation only: two projects can jointly exhaust the org cap before either reaches its own.

For hard isolation, keep the workspace limits' sum at or below the organization limit. Here is an illustrative split for an organization that sets its own limit at $800 on the Build tier (whose cap is $1,000):

Workspace Monthly limit
Prod - Customer Chatbot $400
Prod - Batch Pipeline $250
Staging $50
Dev - Internal Tools $50
Unallocated buffer $50
Example: $800 organization limit split by workspace Prod - Chatbot Prod - Batch Staging Dev - Internal Buffer $400 $250 $50 $50 $50
Illustrative allocation by the author. Tier caps from the Claude Platform Docs, retrieved 2026-09-28.

The docs don't promise to-the-cent enforcement, so treat limits as guardrails and keep some headroom rather than budgeting to the last dollar.

What does your code see when a limit is hit?

The response depends on which limit you reached, and the difference matters for retry logic.

A limit you set returns HTTP 400 with error type invalid_request_error. The message begins "You have reached your specified API usage limits", or "...specified workspace API usage limits" for a workspace cap, and says when access resumes. Raising or removing the limit restores access sooner.

Your tier's cap returns HTTP 429 with error type rate_limit_error. There is no retry-after header, so retrying (including the SDKs' automatic retries) fails until access resumes at 00:00 UTC on the first day of the next month, unless you get a higher limit. On the Messages API you can tell it apart from an ordinary rate limit by error.details.error_code, which is enforced_spend_limit_reached.

The Claude Code workspace is checked separately, and requests over its limit can instead get a 429 with a retry-after header.

A handler that separates these cases stops your app from hammering an endpoint that can't recover:

import anthropic

client = anthropic.Anthropic()

class BudgetExhausted(Exception):
    pass

def ask(prompt: str):
    try:
        return client.messages.create(
            model="claude-sonnet-5-5",
            max_tokens=1024,
            messages=[{"role": "user", "content": prompt}],
        )
    except anthropic.BadRequestError as e:
        # 400: a limit *you* set (organization or workspace)
        if "reached your specified" in str(e):
            raise BudgetExhausted(str(e)) from e
        raise
    except anthropic.RateLimitError as e:
        # 429: either a normal rate limit (retry-after present)
        # or the tier spend cap (no retry-after)
        try:
            code = e.response.json().get("error", {}).get("details", {}).get("error_code")
        except ValueError:
            code = None
        if code == "enforced_spend_limit_reached":
            raise BudgetExhausted(str(e)) from e
        raise

Catch BudgetExhausted at the edge of your app, switch to a degraded mode or a clear "temporarily unavailable" message, and page a human. Don't retry.

How do you monitor usage per project?

There are three tools, from quickest to most flexible.

The Console

The Usage and Cost pages can be viewed for a single workspace or for all workspaces. Organization Billing role holders can see cost, usage, and limit values for every workspace, so finance can check without developer access (Claude Help Center, retrieved 2026-09-28).

The anthropic-workspace-id header

Every Claude API response includes the ID of the workspace the request's credential resolved to. Logging it next to your own request metadata lets you confirm that a service is running in the workspace you think it is, and reconcile against reports later.

raw = client.messages.with_raw_response.create(
    model="claude-sonnet-5-5",
    max_tokens=256,
    messages=[{"role": "user", "content": "ping"}],
)
print(raw.headers.get("anthropic-workspace-id"))
message = raw.parse()

The Usage and Cost APIs

Both are Admin API endpoints. They require an Admin API key, an org:admin OAuth token, or a personal or service account key that isn't scoped to a workspace. Workspace keys don't work. They also aren't available on individual accounts or on Claude Platform on AWS (Usage and Cost API, retrieved 2026-09-28).

Usage API Cost API
Endpoint /v1/organizations/usage_report/messages /v1/organizations/cost_report
Measures Tokens (uncached, cached, cache creation, output), server tools Spend in USD, reported as decimal strings in cents
Granularity 1m, 1h, or 1d buckets Daily only
Group or filter by Workspace, API key, model, service tier, context window, data residency Workspace, description
Gaps No dollar amounts Priority Tier costs aren't included

Data typically appears within about five minutes of a request completing, and polling once per minute is supported for sustained use.

Here is a month-to-date report by workspace, with a check against each workspace's budget. It assumes each result carries an amount (in cents) and a workspace_id; confirm the field names against the Cost API reference before relying on it.

import os
from collections import defaultdict
from datetime import datetime, timezone

import requests

ADMIN_KEY = os.environ["ANTHROPIC_ADMIN_KEY"]
URL = "https://api.anthropic.com/v1/organizations/cost_report"
BUDGETS_USD = {
    "wrkspc_01...chatbot": 400.0,   # replace with your workspace IDs
    "wrkspc_01...batch": 250.0,
}

def month_to_date_usd() -> dict:
    now = datetime.now(timezone.utc)
    start = now.replace(day=1, hour=0, minute=0, second=0, microsecond=0)
    params = {
        "starting_at": start.strftime("%Y-%m-%dT%H:%M:%SZ"),
        "ending_at": now.strftime("%Y-%m-%dT%H:%M:%SZ"),
        "group_by[]": ["workspace_id"],
        "limit": 31,
    }
    headers = {"x-api-key": ADMIN_KEY, "anthropic-version": "2023-06-01"}
    totals = defaultdict(float)
    while True:
        page = requests.get(URL, params=params, headers=headers, timeout=30).json()
        for bucket in page["data"]:
            for row in bucket["results"]:
                totals[row["workspace_id"]] += float(row["amount"]) / 100  # cents -> USD
        if not page.get("has_more"):
            return totals
        params["page"] = page["next_page"]

for ws, spent in month_to_date_usd().items():
    budget = BUDGETS_USD.get(ws)
    if budget:
        pct = spent / budget * 100
        flag = "ALERT" if pct >= 90 else "warn" if pct >= 75 else "ok"
        print(f"{ws}: ${spent:,.2f} of ${budget:,.0f} ({pct:.0f}%) [{flag}]")

Run it on a schedule and post the output to Slack or your pager. Because reports lag by minutes, this complements the built-in alerts and caps rather than replacing them.

If you'd rather not build dashboards, the docs list ready-made integrations for CloudZero, Datadog, Grafana Cloud, Harness, Honeycomb, and Vantage.

What if two projects share one workspace?

You can still split them by API key. Give each project its own key and use group_by[]=api_key_id on the Usage API. The Cost API groups by workspace and description, not by key, so you'd estimate dollars from token counts and your model's rates. Separate workspaces are cleaner if you need real per-project dollar figures and hard caps.

What gotchas should you plan for?

Frequently Asked Questions

Can I set a spend limit on the Default Workspace?

No. The docs state that limits can't be set on the Default Workspace. Create a named workspace and issue new keys there.

Are workspace limits allowed to exceed my organization's limit?

No. Workspace limits must be at or below the organization's limit, and the organization-wide limit always applies on top.

How current is usage data?

Usage and cost data typically shows up within about five minutes of a request finishing, though longer delays can occur. Design alerts with that lag in mind.

Conclusion

Per-project cost control comes down to four steps:

Set aside twenty minutes to do the first two today, since they cost nothing and cap your worst case immediately. Then add the reporting script once you have a week of data to compare against.


Sources: Workspaces, Rate limits, and Usage and Cost API, all Claude Platform Docs, retrieved 2026-09-28. Tiers, caps, and endpoint schemas change; verify against the current docs before making budget decisions.