7 Claude Code Plugins to stop Burning Through Tokens
7 Claude Code Plugins Worth Trying If You're Burning Through Tokens
If you use Claude Code regularly, you've probably watched a session eat through context faster than expected, Here are 7 plugins to help with that.
A quick note before diving in: these are all third-party, community-built tools, not official Anthropic products.
1. Superpowers
A skills/workflow plugin rather than a pure compression tool. An independent benchmark ran twelve otherwise-identical Claude Code sessions, six with Superpowers installed and six without Superpowers, and found runs with it were about 9% cheaper and used 14% fewer tokens, with better output quality.
🔗 Referenced in: MindStudio's benchmark writeup
2. token-saver
Takes a different philosophy than most tools in this space: instead of compressing tool output after the fact, it focuses on keeping large content out of the main context window in the first place. The project cites an analysis of 3.77 billion tokens in a single day showing that 95.7% of spend was reused input from prior turns, tool schemas, and files re-sent on every request, not large one-off output. It ships a skill plus a set of cheap-model subagents (Haiku for scouting and running commands, Sonnet for reading) that return conclusions instead of raw dumps, and a token-audit tool that measures real billing from Claude Code's own usage logs rather than counterfactual claims.
🔗 github.com/bryanvine/token-saver
3. token-squeeze
Indexes your codebase's symbols. Functions, classes, methods, types, using tree-sitter, then exposes an MCP server so Claude can pull just the source of the specific symbol it needs instead of reading whole files into context. Supports Python, JavaScript, TypeScript, C#, C, and C++.
🔗 github.com/mpool/token-squeeze
4. claude-code-token-saver
A grab-bag of small, targeted fixes rather than one big mechanism. Notably, it detects when Claude Code's one-hour prompt cache has expired (step away too long and your next message re-sends the entire context at full price) and blocks the expensive re-send calls. It can also restore previous session context by reading transcripts directly with no LLM calls involved, and it replaces Claude Code's built-in git instructions with a much smaller injected version, which the project estimates saves roughly 2,200 tokens per session plus around 1,700 tokens per call Claude makes.
🔗 claudepluginhub.com/plugins/ww-w-ai-cc-token-saver
5. Auto model-switching plugins
Several projects in this category do essentially the same thing: automatically route a prompt to Haiku, Sonnet, or Opus (or a custom tier you define) based on task complexity, so you're not burning an expensive model on mechanical or trivial work. This pairs naturally with tools like token-saver above, which uses a similar routing idea internally.
🔗 Search "auto model-switching Claude Code" on GitHub Topics: token-saving for current options, several exist and churn frequently, so it's worth comparing a couple before picking one.
6. Firecrawl
Not Claude-Code-specific, but very commonly paired with it for any workflow that involves scraping or reading web pages. Rather than dumping raw HTML into context, it returns cleaned, structured content, cutting token size significantly.
🔗 Referenced alongside Superpowers and others in MindStudio's Claude Code skills benchmark
7. Live token counters and cost dashboards
A handful of tools, bundled in some of the token-saver repos above, and standalone, add a live token counter to your CLI status bar or an interactive HTML dashboard breaking down cost, cache usage, and spend across sessions. These don't reduce token usage directly, but visibility is often what actually prompts a change in habits, which tends to save more than any single compression trick.
🔗 See the dashboard tools listed alongside the cache-expiry feature in claude-code-token-saver on ClaudePluginHub
The free alternative
Before installing anything, it's worth remembering that a lot of this can be done with zero plugins:
- Keep a lean
CLAUDE.mdat your project root instead of repeating context every session. - Don't let sessions sit idle for over an hour, Claude Code's prompt cache expires and your next message re-sends everything at full expensive price.
- Scope requests to specific files or directories in large repos instead of letting Claude explore everything.
- Use
/compactor start a fresh session once context balloons, rather than letting it grow indefinitely.