Claude Code headless mode in CI without runaway costs

•By Blacdisk Team

Claude Code headless mode in CI without runaway costs

A human running Claude Code interactively is a natural cost control: they notice when it goes off track and press Escape. In CI, nobody is watching. A prompt that loops, a trigger that fires on every push, or a job that hangs for an hour all bill the same as productive work.

This guide shows how to run claude -p (headless, or print, mode) in a pipeline with layered limits, so the worst case is a number you chose in advance. The layers are per-run caps, permission settings, job-level timeouts and concurrency, and a monthly spend limit on the API key's workspace. Flags and behaviors come from Anthropic's Claude Code documentation as retrieved on 2026-09-28. Claude Code changes quickly, so check your installed version against the linked pages.

Key Takeaways

  • --max-turns has no default limit, and --max-budget-usd only exists in print mode. Set both explicitly on every CI call.
  • The dollar cap is computed by Claude Code from token counts, and the docs call cost figures client-side estimates that can differ from your bill. Treat it as a per-run guard, not a billing guarantee.
  • Use --bare and a narrow permission setup so CI runs are reproducible and don't hang on prompts that nobody can answer.
  • Put the CI key in its own workspace with a monthly spend limit. That is the only cap enforced on Anthropic's side rather than by the client.

Why can headless runs get expensive?

Claude Code charges by API token consumption, and several properties of unattended use multiply it (Manage costs effectively, retrieved 2026-09-28):

None of these is a bug. They are simply unpriced until you decide what a run is allowed to cost.

What layers of protection should you use?

No single control covers every failure, so stack them. Each layer catches something the others miss.

Layer Control What it bounds Failure it catches
Run --max-turns Agentic iterations A loop that never converges
Run --max-budget-usd Estimated dollars per run Expensive turns, large context, subagents
Job timeout-minutes (or timeout) Wall-clock time Hangs, stuck background work
Fleet concurrency, path filters Number of runs Push storms, doc-only changes
Account Workspace spend limit Monthly dollars Everything above failing at once

The first two layers are Claude Code flags, the next two are your CI system, and the last is enforced by the API.

How do you set per-run limits?

Here is a locked-down call suited to any CI system:

timeout 600 claude --bare -p "$PROMPT" \
  --model sonnet \
  --effort medium \
  --max-turns 8 \
  --max-budget-usd 1.00 \
  --permission-mode dontAsk \
  --allowedTools "Read,Bash(git diff *),Bash(npm test)" \
  --output-format json \
  --no-session-persistence \
  > result.json
status=$?

What each flag does, per the CLI reference:

Why use --bare in CI?

Without it, claude -p loads the same context an interactive session would: hooks, plugins, MCP servers, auto memory, and CLAUDE.md from the working directory and ~/.claude. The headless docs recommend --bare for scripted calls and say it will become the default for -p in a future release. Passing it explicitly gives you two benefits:

Bare mode doesn't read OAuth credentials or the keychain, so set ANTHROPIC_API_KEY (or an apiKeyHelper). Pass any context you do want, such as --append-system-prompt-file or --mcp-config, explicitly.

How trustworthy is the dollar cap?

Claude Code computes cost locally from token counts at list price, unless an admin has configured a modelPricing table. The docs describe these figures as estimates, and the JSON output's total_cost_usd is a client-side estimate that "can differ from your actual bill." Two more details affect how you use it:

So use the flag to stop an individual run early, and use the workspace limit below when you need a number you can put in a budget.

How do you avoid hangs and over-broad permissions?

In -p mode nobody can answer a permission prompt. The starting mode is Manual, so calls that would prompt are denied. That is safe, but it's also why a first unattended run often fails. You have four options:

Avoid --dangerously-skip-permissions unless the runner is genuinely isolated. A broad grant also widens what a runaway loop can do, in cost and in side effects.

How do you set up GitHub Actions?

Anthropic's GitHub Actions guidance names two separate meters: GitHub-hosted runner minutes and API tokens (GitHub Actions docs, retrieved 2026-09-28). Its cost tips are to use specific @claude commands rather than open-ended ones, configure --max-turns in claude_args, set workflow-level timeouts, and use GitHub's concurrency controls.

A review workflow that applies all of them:

name: Claude PR review
on:
  pull_request:
    types: [opened, synchronize]
    paths-ignore: ["docs/**", "**/*.md"]   # skip doc-only changes

concurrency:
  group: claude-review-${{ github.event.pull_request.number }}
  cancel-in-progress: true                 # a new push cancels the stale run

jobs:
  review:
    runs-on: ubuntu-latest
    timeout-minutes: 10                    # hard wall-clock backstop
    permissions:
      contents: read
      pull-requests: write
    steps:
      - uses: actions/checkout@v4
      - uses: anthropics/claude-code-action@v1
        with:
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY_CI }}
          prompt: "Review this PR for bugs and security issues. Be concise."
          claude_args: "--model sonnet --max-turns 6 --max-budget-usd 1.00"

Three notes on this file:

What is the worst-case bill?

A per-run cap turns an open-ended risk into arithmetic: worst-case monthly spend is roughly the number of runs times the cap. Suppose a team merges about 40 PRs a week with three pushes each. That is around 500 triggered runs a month.

Worst-case monthly cost, 500 runs, by per-run cap $250 $500 $1,000 $2,500 $0.50 cap $1 cap $2 cap $5 cap
Illustrative arithmetic by the author: runs per month multiplied by the per-run cap. Actual spend is usually far lower because most runs stop well short of the cap.

To pick the cap, run a week with logging on, then set it at two to three times your median run. That leaves room for hard cases and still bounds the tail.

How do you add a hard spending backstop?

The flags above run on the client. For a limit that Anthropic enforces, create a dedicated workspace for CI in the Claude Console, issue the CI API key there, and set a monthly spend limit on that workspace. Workspace limits can be set at or below the organization's, and API keys are tied to the workspace where they were created (Workspaces, retrieved 2026-09-28).

Choose the workspace cap first, based on what you're willing to lose in a bad month, then derive the per-run cap: monthly cap divided by expected runs. With a $300 workspace limit and 500 runs, a $0.50 run cap keeps even the worst case inside the budget.

When the workspace limit is reached, the API returns HTTP 400 with an invalid_request_error whose message begins "You have reached your specified workspace API usage limits" (Rate limits). Your CI job will fail with that message, which is what you want. Make sure the failure is visible in the PR or a channel so someone raises the limit or investigates. A per-workspace rate limit also bounds how many concurrent runs can hit the API at once.

How do you measure what each run costs?

With --output-format json, the response includes total_cost_usd and a per-model breakdown, so scripts can track spend without opening the dashboard. Write it to the job summary and keep a running record:

cost=$(jq -r '.total_cost_usd' result.json)
echo "Claude run cost (estimate): \$${cost}" >> "$GITHUB_STEP_SUMMARY"

Then reconcile against real billing. The Console Usage page, and the Cost API grouped by workspace_id, show what Anthropic charged. A CI workspace with its own key makes that reconciliation a single filter.

What gotchas should you plan for?

Frequently Asked Questions

Does --max-turns fail the job when it's reached?

Yes. The CLI reference says Claude Code exits with an error when the limit is reached, so your pipeline sees a non-zero status. That's useful for spotting loops, but plan how the workflow should treat it.

Should I use --dangerously-skip-permissions in CI?

Prefer --permission-mode dontAsk with a narrow --allowedTools list. The bypass flag removes prompts entirely, which also removes a limit on what a looping run can do. Reserve it for isolated, disposable environments.

Is --bare required?

No, but the docs recommend it for scripted calls and plan to make it the default for -p. Setting it now keeps runs consistent and avoids surprises when the default changes.

Conclusion

Runaway CI costs come from missing bounds, not from Claude Code itself. Bound each layer:

Start with the last item, since it takes minutes and caps your worst month immediately. Then add the per-run flags, log total_cost_usd for a week, and tighten the caps using real numbers.


Sources: Run Claude Code programmatically, CLI reference, Manage costs effectively, and GitHub Actions, all Claude Code Docs; Workspaces and Rate limits, Claude Platform Docs. All retrieved 2026-09-28. Flags, versions, and defaults change; verify against current documentation before relying on them.