Claude Code headless mode in CI without runaway costs
Claude Code headless mode in CI without runaway costs
A human running Claude Code interactively is a natural cost control: they notice when it goes off track and press Escape. In CI, nobody is watching. A prompt that loops, a trigger that fires on every push, or a job that hangs for an hour all bill the same as productive work.
This guide shows how to run claude -p (headless, or print, mode) in a pipeline with layered limits, so the worst case is a number you chose in advance. The layers are per-run caps, permission settings, job-level timeouts and concurrency, and a monthly spend limit on the API key's workspace. Flags and behaviors come from Anthropic's Claude Code documentation as retrieved on 2026-09-28. Claude Code changes quickly, so check your installed version against the linked pages.
Key Takeaways
--max-turnshas no default limit, and--max-budget-usdonly exists in print mode. Set both explicitly on every CI call.- The dollar cap is computed by Claude Code from token counts, and the docs call cost figures client-side estimates that can differ from your bill. Treat it as a per-run guard, not a billing guarantee.
- Use
--bareand a narrow permission setup so CI runs are reproducible and don't hang on prompts that nobody can answer.- Put the CI key in its own workspace with a monthly spend limit. That is the only cap enforced on Anthropic's side rather than by the client.
Why can headless runs get expensive?
Claude Code charges by API token consumption, and several properties of unattended use multiply it (Manage costs effectively, retrieved 2026-09-28):
- Every request carries the full conversation. Each time Claude uses tools, Claude Code sends another request with the accumulated context. Prompt caching re-reads that history at a discounted rate, but the volume still grows with every turn.
- Thinking counts as output. Extended thinking is on by default, and thinking tokens are billed as output tokens. Per the docs, you can't turn it off on Opus 5.5, Sonnet 5.5, or the Fable models, so lower effort is the main lever there.
- Turns are unbounded by default. The CLI reference says
--max-turnshas no limit by default. - Triggers fan out. One PR with ten pushes can start ten runs.
- Subagents add their own requests. Each one sends requests on top of the main conversation's.
None of these is a bug. They are simply unpriced until you decide what a run is allowed to cost.
What layers of protection should you use?
No single control covers every failure, so stack them. Each layer catches something the others miss.
| Layer | Control | What it bounds | Failure it catches |
|---|---|---|---|
| Run | --max-turns |
Agentic iterations | A loop that never converges |
| Run | --max-budget-usd |
Estimated dollars per run | Expensive turns, large context, subagents |
| Job | timeout-minutes (or timeout) |
Wall-clock time | Hangs, stuck background work |
| Fleet | concurrency, path filters |
Number of runs | Push storms, doc-only changes |
| Account | Workspace spend limit | Monthly dollars | Everything above failing at once |
The first two layers are Claude Code flags, the next two are your CI system, and the last is enforced by the API.
How do you set per-run limits?
Here is a locked-down call suited to any CI system:
timeout 600 claude --bare -p "$PROMPT" \
--model sonnet \
--effort medium \
--max-turns 8 \
--max-budget-usd 1.00 \
--permission-mode dontAsk \
--allowedTools "Read,Bash(git diff *),Bash(npm test)" \
--output-format json \
--no-session-persistence \
> result.json
status=$?
What each flag does, per the CLI reference:
--max-turns 8limits agentic turns and exits with an error when the limit is reached. Failing loudly is the point: a job that should take a handful of turns but loops should stop, not keep spending.--max-budget-usd 1.00stops the run before it spends more than that on API calls. Spend from subagents counts toward the cap. On v2.1.217 or later, once the cap is hit, spawning another subagent fails and background subagents are stopped.--model sonnetpicks a cheaper model explicitly. The costs docs say Sonnet handles most coding tasks well and costs less than Opus, so reserve Opus for jobs that need it.--effort mediumlowers reasoning effort, which reduces thinking tokens. Available levels depend on the model.--no-session-persistenceskips saving the session to disk. That is tidier on ephemeral runners, and it also means the run can't be resumed.timeout 600is an outer wall-clock bound from the shell. If you stopclaude -pwith SIGTERM, it exits with code 143, so treat 143 (or 124 from GNUtimeout) as a timeout rather than a crash.
Why use --bare in CI?
Without it, claude -p loads the same context an interactive session would: hooks, plugins, MCP servers, auto memory, and CLAUDE.md from the working directory and ~/.claude. The headless docs recommend --bare for scripted calls and say it will become the default for -p in a future release. Passing it explicitly gives you two benefits:
- Reproducibility and speed. Every runner behaves the same, and startup is faster.
- Less risk from the repo itself. Without
--bare, a-psession runs the hooks in a project's.claude/settings.jsonand connects the servers in its.mcp.json, even in a folder you've never trusted, with no trust dialog. That matters for pull requests from forks.
Bare mode doesn't read OAuth credentials or the keychain, so set ANTHROPIC_API_KEY (or an apiKeyHelper). Pass any context you do want, such as --append-system-prompt-file or --mcp-config, explicitly.
How trustworthy is the dollar cap?
Claude Code computes cost locally from token counts at list price, unless an admin has configured a modelPricing table. The docs describe these figures as estimates, and the JSON output's total_cost_usd is a client-side estimate that "can differ from your actual bill." Two more details affect how you use it:
- Resuming with
--continueor--resumereports the conversation's whole total, but earlier runs' totals don't count toward--max-budget-usd. Don't build a loop of resumed runs and assume the cap spans them. - For responses billed at the 1.1x data-residency rate, Claude Code applies the multiplier to its figure, on versions that include that fix.
So use the flag to stop an individual run early, and use the workspace limit below when you need a number you can put in a budget.
How do you avoid hangs and over-broad permissions?
In -p mode nobody can answer a permission prompt. The starting mode is Manual, so calls that would prompt are denied. That is safe, but it's also why a first unattended run often fails. You have four options:
--allowedToolspre-approves specific tools, using prefix patterns likeBash(git diff *). The space before the*matters, sinceBash(git diff*)would also matchgit diff-index.--permission-mode dontAskdenies anything that would otherwise prompt. It is the locked-down CI choice, and actions covered by your allow rules still run.--permission-mode acceptEditsauto-approves file edits and common filesystem commands, while other shell commands still need an allow entry.--permission-prompts none(v2.1.259 or later) makes Claude Code deny anything that would prompt and tells Claude not to retry it, which avoids wasted turns on a request that can never succeed.
Avoid --dangerously-skip-permissions unless the runner is genuinely isolated. A broad grant also widens what a runaway loop can do, in cost and in side effects.
How do you set up GitHub Actions?
Anthropic's GitHub Actions guidance names two separate meters: GitHub-hosted runner minutes and API tokens (GitHub Actions docs, retrieved 2026-09-28). Its cost tips are to use specific @claude commands rather than open-ended ones, configure --max-turns in claude_args, set workflow-level timeouts, and use GitHub's concurrency controls.
A review workflow that applies all of them:
name: Claude PR review
on:
pull_request:
types: [opened, synchronize]
paths-ignore: ["docs/**", "**/*.md"] # skip doc-only changes
concurrency:
group: claude-review-${{ github.event.pull_request.number }}
cancel-in-progress: true # a new push cancels the stale run
jobs:
review:
runs-on: ubuntu-latest
timeout-minutes: 10 # hard wall-clock backstop
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY_CI }}
prompt: "Review this PR for bugs and security issues. Be concise."
claude_args: "--model sonnet --max-turns 6 --max-budget-usd 1.00"
Three notes on this file:
- The
claude_argsstring forwards CLI flags to Claude Code. Check that your pinned version of the action passes--max-budget-usdthrough, and confirm what it does with--bareand permissions, since the action manages some tool access itself. cancel-in-progressis the cheapest saving in the file. If someone pushes twice in a minute, the first review is obsolete, and cancelling it stops the spend at once.- Use a dedicated secret for the CI key so it can live in its own workspace, as described below.
What is the worst-case bill?
A per-run cap turns an open-ended risk into arithmetic: worst-case monthly spend is roughly the number of runs times the cap. Suppose a team merges about 40 PRs a week with three pushes each. That is around 500 triggered runs a month.
To pick the cap, run a week with logging on, then set it at two to three times your median run. That leaves room for hard cases and still bounds the tail.
How do you add a hard spending backstop?
The flags above run on the client. For a limit that Anthropic enforces, create a dedicated workspace for CI in the Claude Console, issue the CI API key there, and set a monthly spend limit on that workspace. Workspace limits can be set at or below the organization's, and API keys are tied to the workspace where they were created (Workspaces, retrieved 2026-09-28).
Choose the workspace cap first, based on what you're willing to lose in a bad month, then derive the per-run cap: monthly cap divided by expected runs. With a $300 workspace limit and 500 runs, a $0.50 run cap keeps even the worst case inside the budget.
When the workspace limit is reached, the API returns HTTP 400 with an invalid_request_error whose message begins "You have reached your specified workspace API usage limits" (Rate limits). Your CI job will fail with that message, which is what you want. Make sure the failure is visible in the PR or a channel so someone raises the limit or investigates. A per-workspace rate limit also bounds how many concurrent runs can hit the API at once.
How do you measure what each run costs?
With --output-format json, the response includes total_cost_usd and a per-model breakdown, so scripts can track spend without opening the dashboard. Write it to the job summary and keep a running record:
cost=$(jq -r '.total_cost_usd' result.json)
echo "Claude run cost (estimate): \$${cost}" >> "$GITHUB_STEP_SUMMARY"
Then reconcile against real billing. The Console Usage page, and the Cost API grouped by workspace_id, show what Anthropic charged. A CI workspace with its own key makes that reconciliation a single filter.
What gotchas should you plan for?
- Background work keeps the process open. If Claude starts a background subagent or workflow,
claude -pstays open until it finishes. By default the wait ends after 10 minutes of continuous idle waiting. Your job timeout should sit above that or the two will race. - Cache lifetime is short on API keys. The prompt cache lasts five minutes by default on an API key, and caches are isolated per workspace, so don't budget for cache hits between runs that are minutes apart. Adding
--exclude-dynamic-system-prompt-sectionsmoves per-machine details out of the system prompt, which the docs say improves cache reuse across machines running the same task. - Flags depend on the version.
--permission-promptsneeds v2.1.259 or later, and the subagent cap behavior needs v2.1.217 or later. Pin the CLI version in CI (claude install <version>accepts a version number) and upgrade on purpose. - A cap hit is a failed run. Decide up front whether that should block a merge. For advisory reviews, many teams mark the step non-blocking but still surface the failure.
Frequently Asked Questions
Does --max-turns fail the job when it's reached?
Yes. The CLI reference says Claude Code exits with an error when the limit is reached, so your pipeline sees a non-zero status. That's useful for spotting loops, but plan how the workflow should treat it.
Should I use --dangerously-skip-permissions in CI?
Prefer --permission-mode dontAsk with a narrow --allowedTools list. The bypass flag removes prompts entirely, which also removes a limit on what a looping run can do. Reserve it for isolated, disposable environments.
Is --bare required?
No, but the docs recommend it for scripted calls and plan to make it the default for -p. Setting it now keeps runs consistent and avoids surprises when the default changes.
Conclusion
Runaway CI costs come from missing bounds, not from Claude Code itself. Bound each layer:
- Cap turns and estimated dollars on every
claude -pcall. - Use
--bare, a narrow permission setup, and a pinned CLI version. - Add job timeouts, concurrency cancellation, and path filters in CI.
- Put the CI key in a dedicated workspace with a monthly spend limit.
Start with the last item, since it takes minutes and caps your worst month immediately. Then add the per-run flags, log total_cost_usd for a week, and tighten the caps using real numbers.
Sources: Run Claude Code programmatically, CLI reference, Manage costs effectively, and GitHub Actions, all Claude Code Docs; Workspaces and Rate limits, Claude Platform Docs. All retrieved 2026-09-28. Flags, versions, and defaults change; verify against current documentation before relying on them.