Reading Files Smartly: Why Letting Claude Grep First Beats Letting It Open Everything

•By Blacdisk Team

Reading Files Smartly: Why Letting Claude Grep First Beats Letting It Open Everything

Key Takeaways

  • Context is the real bottleneck for coding agents — not model size, not the size of the codebase, but how much of it gets dumped into the conversation.
  • Opening every file that might be relevant floods the context window with noise the model then has to reason around.
  • Grep-first search lets the agent form a hypothesis, test it cheaply, and only pay the token cost of a full file read when it actually needs one.
  • The right mental model isn't "grep vs. read" — it's grep, LSP, and subagents as a search stack, each used where it's cheapest.

If you've watched a coding agent work through an unfamiliar codebase, you've probably seen the pattern: it doesn't open every file. It greps for a symbol, skims the hits, opens two or three files, greps again. It looks almost hesitant compared to just reading the whole directory upfront. That hesitation is the point — and it's the difference between an agent that stays useful for an hour and one that runs out of room in twenty minutes.

The problem isn't the codebase size — it's what enters the context window

A large context window feels like it should solve this. Just read everything, right? But agents skilled at navigating large codebases using grep still become blind to any file they haven't fully tokenized, because grep only finds exact text matches and can't discover code that's conceptually related but named differently. Reading more doesn't fix that gap — it just adds volume without adding the right kind of understanding.

And volume has a real cost. Model performance doesn't stay flat as context fills up — it degrades meaningfully as context size increases, and benchmarks consistently improve when the model has less irrelevant material to sort through. This is the part that's easy to underestimate: it's not that a full context window makes the model slower or more expensive (though it does that too). It's that a cluttered context makes the model worse — more likely to miss the file that actually mattered because it's buried under twelve others that didn't.

One developer working on a subagent visibility issue measured this directly. Each subagent operation that surfaced its full internals — every grep call, every file read — added roughly 13,770 tokens to the conversation, compared to about 920 tokens for a clean summary. That's not a rounding error. That's the difference between a session that lasts all afternoon and one that hits its limit before lunch.

Grep-first search is a hypothesis-testing loop, not a shortcut

The instinct to call grep-then-read a "cheap trick" undersells what's actually happening. Coding agents spend more than 60% of their time searching for context, and the quality of that search — not the model's size, not the size of the context window — determines whether the agent succeeds or fails. Search is the bottleneck, so how the agent searches matters more than almost anything else it does.

Done well, it looks like this: the agent plans and executes multi-step searches with reasoning between each step, forming a hypothesis about where the relevant code might live, testing that hypothesis with a tool call, and following the resulting chains across files rather than retrieving everything in one pass. Grep is the cheapest way to test a hypothesis. A full file read is the expensive way to confirm one. Using grep first means the agent only pays the read cost when it already has good reason to believe the file matters.

This mirrors how a person actually works through unfamiliar code. Nobody opens every file in a repo before starting a task — you search for the function name, skim the handful of hits, and open the two files that look right. Claude Code was built to navigate a codebase the way a software engineer would: traversing the file system, reading files, using grep to find exactly what's needed, and following references across the codebase — without requiring an index to be built or maintained ahead of time.

Grep isn't free of trade-offs

It's worth being honest about where a pure text-search approach breaks down. Grep finds strings, not intent — a query like "how does the billing system handle failed payments" has no single string to search for, since the logic might be spread across several files connected by imports and function calls rather than shared keywords. And even when a grep query is well-chosen, it tends to be noisy: searching for a common function name returns the definition, every call site, test mocks, documentation, and changelog entries all at once, leaving the agent to reason about which results actually matter.

This is where pairing grep with other tools pays off rather than trying to make grep do everything. One practitioner's take, after testing both approaches side by side: run a language server for your main stack so the agent gets symbol-accurate definitions and references and catches broken call sites right after an edit, but keep grep for everything that isn't code — logs, configs, comments, string literals, feature flags — and for whole-tree sweeps where plain text search wins. Neither tool replaces the other; they cover different kinds of questions.

[Chart: Tokens added to context per operation — "verbose internals" (~13,770 tokens) vs. "clean summary" (~920 tokens), based on the subagent visibility measurement above]

Subagents push the savings further

Grep-first search controls what enters context within a single line of investigation. Subagents control it across investigations. The pattern is straightforward: without subagents, the main agent handles everything in one context window, and every grep, find, and read call stays there — after thirty minutes of exploration that can add up to eighty thousand tokens of accumulated noise. A subagent dispatched to explore a question works in its own window and returns only the result to the main conversation, so the searching happens off to the side and only the conclusion comes back.

[INTERNAL-LINK: how subagents keep long coding sessions from running out of context → deeper guide on delegation and context isolation in agentic workflows]

Anthropic's own tooling choices reflect this same philosophy at the model-provider level. Claude Code uses no embeddings at all — Anthropic chose grep over vector search for its own agent, betting that live, exact, zero-setup search beats a pre-built index that has to be maintained as the codebase changes. Agentic search avoids the failure modes of embedding pipelines entirely, since there's no centralized index to keep in sync as engineers commit new code — each developer's instance simply works from the live codebase.

What this means for how you prompt an agent

If you're directing a coding agent — through a system prompt, a project's conventions file, or just how you phrase a request — the practical takeaways are simple:

[INTERNAL-LINK: setting up a language server alongside Claude Code → walkthrough on configuring LSP for symbol-aware navigation]

None of this is about making the agent slower or more cautious for its own sake. It's the opposite: an agent that searches before it reads stays coherent for longer, burns fewer tokens per useful action, and is less likely to lose the file that mattered under a pile of ones that didn't.

Frequently Asked Questions

Does a bigger context window make grep-first search unnecessary?

Not really. A bigger window raises the ceiling, but it doesn't fix the underlying issue: model performance degrades as context size grows, and results consistently improve when the model has less irrelevant material to work through. Even with room to spare, dumping unnecessary files into context still costs accuracy, not just tokens.

Is grep search less accurate than reading everything?

It can miss conceptually related code that doesn't share matching keywords, which is a real limitation. Grep can only find exact text matches, so code that's related in purpose but named differently stays invisible to it — which is exactly why pairing it with a language server, or occasionally reading a file in full once grep has narrowed things down, closes the gap.

Does Claude Code still use ripgrep under the hood?

Not necessarily anymore. Around version 2.1.116/117, the built-in search reportedly swapped from ripgrep to ugrep — a reminder that the specific engine matters less than the search-first strategy sitting on top of it.