One Task per Session: The Case Against the Marathon Conversation
One Task per Session: The Case Against the Marathon Conversation
Key Takeaways
- A session that "still has room" isn't the same as a session that's still helping — accuracy degrades well before the context window fills up.
- Compaction doesn't restore a session to its earlier quality. It's a lossy summary, and specifics you typed an hour ago don't reliably survive it.
- Background tasks, scoped rules, and injected context can silently fail to carry across a compaction boundary, breaking coordination the model doesn't know it lost.
- The fix isn't a bigger context window. It's ending sessions on purpose, at task boundaries, instead of by accident when the window runs out.
There's a specific kind of session that feels productive in the moment and expensive in hindsight: the one that starts with a bug fix, drifts into a refactor, picks up a documentation update along the way, and is still open eight hours later when you finally notice the agent has started confusing details from the first task with the third. Nothing crashed. No error appeared. It just got worse, gradually, in a way that's easy to miss because the session never actually stopped working — it just started working less well.
The window filling up isn't the problem — what happens before it fills is
The instinct is to treat context capacity as the limiting factor: as long as there's room left in the window, keep going. But the quality of an agent's output starts declining well before that room runs out. One practical write-up on the topic puts a number on where this typically starts: (cite index="27-1">for a 1M-token context model, some level of context rot happens around 300-400k tokens, though it's highly dependent on the task rather than a fixed rule. That's a large window — and the degradation still shows up at a fraction of its capacity, not at the edge.
The mechanism behind that degradation is about attention, not storage. (cite index="27-1">Performance degrades as context grows because attention gets spread across more tokens, and older, irrelevant content starts to distract from the current task. A session carrying three unrelated tasks doesn't just contain more tokens — it contains more tokens the model has to actively work around to find the ones that matter right now. Every additional unrelated exchange makes that filtering job harder, even if none of it causes an outright error.
One engineer's summary of the practical cost is blunt: (cite index="26-1">mixed context degrades quality on every sub-task and inflates per-turn cost on every prompt afterward, because you're paying to re-process irrelevant noise. That's the marathon session in one sentence — you're not just risking a worse answer, you're paying full price for the model to wade through material that has nothing to do with what you just asked.
Compaction is a patch, not a reset
The obvious counter is: doesn't compaction handle this? Claude Code will summarize a long session and continue in a fresh window. It does — but it's worth being precise about what that actually preserves. (cite index="28-1">Compaction is not truncation — Claude Code doesn't just drop old messages, it writes a summary of the session covering decisions made, files touched, and current task state, then rebuilds the context from that summary plus a small set of always-loaded material. The key word is summary. A summary is a lossy compression of everything that came before it, and lossy compression means some things don't make it through.
What tends to get lost is exactly the kind of detail that matters mid-task: (cite index="28-1">mid-session instructions typed a couple of hours earlier exist only as whatever the summary kept of them, and injected context from earlier hook runs or tool output has its raw text discarded, leaving only the summary's paraphrase. If you gave the agent a specific constraint early in a long session, compaction doesn't guarantee that constraint survives intact — it survives as someone else's (the model's own) best paraphrase of it.
It gets more concrete than lost nuance, too. Compaction can break things that were actively running. (cite index="24-1">When compaction fires and creates a new session, background tasks started before it become permanently unreachable to the agent — the underlying processes keep running and stay visible in the UI, but the agent's own tools can no longer poll for their status, check results, or cancel them. From the model's perspective, work it dispatched simply vanishes. It can't tell you it finished, because it no longer knows it exists.
Users are already building around this, which is itself the signal
You can tell how real a problem is by how much manual infrastructure people build to route around it. Several open feature requests describe exactly that kind of workaround. One team runs a custom monitor because there's no built-in warning: (cite index="22-1">they built a context-monitor script that runs on every tool call, counting calls as a rough proxy for context usage, and prints a warning at an estimated 80% threshold that the agent is instructed to act on — while noting the mapping is a fragile heuristic and the agent still can't trigger compaction itself.
Another team maintains a hand-written session file specifically because compaction can't be trusted to preserve state on its own: (cite index="23-1">the only current mitigation is manually maintaining a session state file, which requires the model to proactively write it before compaction happens — something that's unreliable, since Claude often doesn't write that file until after compaction has already occurred, by which point critical context like the current task and pending work is already lost.
That's two separate teams building custom tooling to manage a single session across a long stretch of unrelated or extended work. The tooling isn't wrong to build — but its existence is itself the argument for the simpler alternative: don't let one session carry that much weight in the first place.
What ending a session on purpose actually looks like
The alternative isn't "never let a session get long." Plenty of long sessions are genuinely one continuous task and should stay open exactly as long as that task needs. The distinction is between a session that's long because the work is long, and one that's long because several unrelated pieces of work happened to share a window.
One practitioner's framing draws that line clearly: (cite index="26-1">work in focused chunks — spec it, plan it, execute it, verify it, then either keep going because the context is still directly useful, or leave a short handoff and start clean; don't keep a near-1M-token session alive through meetings, lunch, and unrelated work just because the window technically has room. The test isn't capacity. It's relevance — is everything currently in context still pulling weight for the thing you're doing right now?
The same source lays out three distinct tools for three distinct situations, which is worth keeping straight rather than reaching for whichever one is closest to hand: (cite index="26-1">/compact is a lossy LLM summary, right when you're still on the same task and the conversation is getting long; /clear is a hand-written fresh start that forces you to re-load only what matters, right for hard task switches; and a handoff document is /clear's premeditated cousin, right for pauses longer than the cache will survive. Using /compact at a task switch drags the old task's noise into the new one. Using /clear mid-task throws away state you still needed. Matching the tool to the actual situation — continuing versus switching — is most of what this comes down to.
The personal rule one engineer settled on captures the whole argument in a sentence: (cite index="26-1">"one coherent workstream per conversation, keep going until the context stops helping" — and the tell that it's stopped helping is almost always the moment you catch yourself typing "okay, switching topics" as a prompt prefix. That's not a note to the agent. It's a note to yourself that the session is over.
What this means in practice
- Treat a topic switch as a hard boundary, not a soft one. If the next thing you're about to ask has nothing to do with the last hour of conversation, that's the moment for
/clear, not another message. - Don't rely on compaction to preserve specifics. If a constraint, decision, or in-flight background task matters, write it down somewhere durable before the session compacts — don't assume the summary will carry it forward intact.
- Match the tool to what actually changed. Still on the same task but the conversation's gotten long →
/compact. Genuinely different task →/clear. Stepping away for a while → a short handoff doc. - Notice the drift, not just the limit. A session doesn't have to hit its ceiling to be past its useful point. If it's carrying three unrelated threads, it's already paying the cost — regardless of how much room is technically left.
None of this is about being precious with sessions. It's the opposite of precious, actually — a session is cheap to start and cheap to end. What's expensive is pretending one long conversation is still doing the job of several short, focused ones just because it never technically ran out of room.
Frequently Asked Questions
How do I know when a session has crossed from "still useful" to "context rot territory"?
There's no universal number — (cite index="27-1">degradation is highly dependent on the task rather than a fixed threshold — but a practical signal is the one already mentioned: if you're about to prefix your next message with something like "switching topics" or "unrelated question," that's usually the point where the session has stopped being one coherent workstream.
Does /compact fix context rot?
Partially, and imperfectly. It reduces token count, but (cite index="28-1">it works by writing a lossy summary of decisions, files touched, and task state, so specifics from earlier in the session aren't guaranteed to survive intact. It's a reasonable tool for staying on the same task through a long conversation — it's not a substitute for starting fresh when the task itself has changed.
Is this less of an issue now that context windows are much larger?
Not as much as it seems. A larger window raises the point where hard limits bite, but (cite index="27-1">some degree of context rot shows up well before the window's stated capacity, and a bigger window doesn't fix the more basic problem — that unrelated material in context still dilutes the model's attention and still costs tokens to reprocess on every subsequent turn, no matter how much spare capacity is left.