Token Budgets: Setting Limits Before a Big Task Instead of After
Token Budgets: Setting Limits Before a Big Task Instead of After
Most people find out they've overspent on a Claude Code task the same way: they check /usage partway through, or notice the session has been running for an hour, and only then start thinking about cost. That's the wrong order. By the time you're checking, the tokens are already spent, the budgeting decisions that actually would have controlled cost all happen before the task starts, not during it.
Here's what setting a budget up front actually looks like, and why it beats watching the meter after the fact.
Why after-the-fact budgeting doesn't really work
Checking cost partway through a task feels like control, but it mostly isn't. By the point you notice a session is expensive, the expensive part has usually already happened, a long exploration phase, a big file read, a wrong path that got explored in depth before being abandoned. Noticing late lets you stop digging, but it doesn't undo the hole.
There's also a structural reason costs climb the longer a session runs, independent of how much new work is happening: Claude Code sends the full conversation with every request, so a one-line question late in a long session still carries the cost of the whole conversation behind it. A session that's been open for hours can burn real tokens on idle background activity too, conversation summarization for /resume, scheduled tasks firing on their interval, cache misses after a break longer than the cache lifetime reprocessing your full context from scratch. None of that shows up as "work you asked for," which is exactly why checking usage only when something feels expensive misses a lot of what's actually driving the number.
What "setting a budget before" actually means
It's not just picking a dollar figure and hoping. It's a small set of decisions made at the start of a task that shape how many tokens it's going to need, before any of them get spent.
Set an actual cap, not just a mental one. The --max-budget-usd CLI flag gives you a hard limit Claude Code checks the running session cost against, rather than a number you're supposed to remember to look at. For teams, the same principle applies at the org level, workspace spend limits in the Claude Console, or spend limits on usage credits for Team and Enterprise plans, so a cap exists structurally rather than depending on someone noticing.
Choose the model for the task, not by default. Sonnet handles most coding work well and costs meaningfully less than Opus; Opus is worth reserving for genuinely complex architectural decisions or multi-step reasoning. Deciding this before you start, rather than defaulting to whatever's already selected, is one of the single highest-leverage budget decisions available, it's set once and applies to every token the task uses.
Decide how much thinking budget the task actually needs. Extended thinking is on by default and can add tens of thousands of tokens per request as output tokens. For a task that doesn't need deep reasoning, lowering the effort level with /effort (or turning thinking off in /config where supported) up front avoids paying for reasoning depth the task was never going to use.
Plan before you let it implement. Plan mode (Shift+Tab to cycle into it) has Claude explore the codebase and propose an approach before writing any code. This isn't just a correctness safeguard, it's a budget safeguard. The expensive failure mode in a big task is usually not "the code was wrong," it's "a wrong approach got implemented at length before anyone caught it." A plan reviewed up front is a cheap way to avoid paying for that at scale.
Scope what loads into context before the task starts. /context shows you what's already consuming your context window before you've typed a single task-specific message, system prompt, tool definitions, MCP servers, memory files, skills. A bloated CLAUDE.md, several unused MCP servers left enabled, or heavy tool definitions all quietly tax every single request in the task ahead, not just the ones that need them. Trimming this before a big task starts is a one-time cost that pays off on every subsequent message; trimming it after is the same fix, just applied too late to save what's already been spent.
Decide upfront what gets delegated to subagents. Verbose operations, running a large test suite, fetching documentation, processing log files, can be delegated to subagents so the noisy output stays in the subagent's own context and only a summary comes back to the main conversation. Deciding this as part of planning the task, rather than noticing afterward that a giant log dump bloated the whole session, keeps the main conversation's cost from absorbing work that didn't need to live there.
What to check before, not just during
A short pre-task checklist, all of which takes less time than the task itself:
- What's the model? Right-sized for the task's actual difficulty, not left on a high-cost default.
- What's already in context? A quick
/contextcheck catches an oversizedCLAUDE.mdor unused MCP servers before they tax every message in the task. - Does this need a plan first? For anything nontrivial, plan mode before implementation.
- What's verbose and delegable? Test runs, log processing, and documentation fetches that can go to a subagent instead of bloating the main session.
- Is there a hard cap?
--max-budget-usdfor an individual run, or an org-level spend limit for anything happening at team scale.
When after-the-fact checking still matters
None of this means /usage and /context mid-task are useless, they're just a different tool for a different job. Checking mid-task is for noticing drift: a session running longer than expected, a task that's ballooned past its original scope, a cache-miss spike after coming back from a break. That's worth catching early, and course-correcting immediately (pressing Escape to stop a session heading the wrong direction, or using /rewind to restore to an earlier checkpoint) is far cheaper than letting it run to completion and reviewing the bill afterward.
The distinction that matters: pre-task budgeting shapes how many tokens the task is going to need. Mid-task checking can only catch that it's using more than expected, useful, but reactive by nature. The highest-leverage moment to control cost on a big task is always before the first message, not the fifth check-in.
The underlying principle
Token cost on a big task is mostly determined by decisions made in the first few minutes, model choice, context hygiene, whether there's a plan before implementation, what gets delegated versus kept in the main conversation. Checking usage after the fact tells you whether those decisions were right. It doesn't make them for you. Do the budgeting up front, and the mid-task checks become a confirmation instead of a rescue.