AI Coding & Agents

[AI Context #4] Why CLAUDE.md Should Be Short

In Part 3, we saw how to clear conversation history with /clear. But checking the remaining context immediately afterward shows it is not at 100%. Run /context in Claude Code and the reason becomes clear: the system prompt, tool definitions, CLAUDE.md, and tools registered by MCP servers already consume tens of thousands of tokens…

5 min read
Cover image for [AI Context #4] Why CLAUDE.md Should Be Short

In Part 3, we saw how to clear conversation history with /clear. But checking the remaining context immediately afterward shows it is not at 100%. Run /context in Claude Code and the reason becomes clear. The system prompt, tool definitions, CLAUDE.md, and tools registered by MCP servers already consume tens of thousands of tokens. Before the conversation even begins.

This part covers this “always-loaded context.” If conversation history is a variable cost, this is a fixed cost. It is carried on every turn without exception, survives /clear, and occupies part of the model’s attention throughout the session. Poor fixed-cost design means every session starts at a disadvantage, making it something to fix before conversation management.

What is already loaded before the session starts

An agent’s context window does not start as a blank page. Several layers are already in place before the first user input arrives.

At the bottom is the system prompt. It contains the baseline instructions installed by the agent harness—an execution environment such as Claude Code—including rules for tool use and response formats, totaling thousands of tokens. Users cannot modify this area.

Tool definitions are layered on top. To use a tool, the model needs its name, description, and parameter schema, all of which enter the context as text. The overhead is modest with only the built-in tools, but the situation changes once you connect MCP servers. A single server commonly registers dozens of tools, and just a few servers can consume tens of thousands of tokens through tool definitions alone. Tools you never use make the round trip on every turn.

The final layer is project instruction files such as CLAUDE.md. Global, project, and subdirectory settings are loaded automatically, and this is the only layer users can design directly.

The principle of CLAUDE.md: keep only what is always true

There is one criterion for deciding what belongs in CLAUDE.md: “Is this always true for every task in this project?” Because the file is carried on every turn and for every task, content that applies to some tasks but is irrelevant to others is not earning its space.

Build and test commands, the codebase’s broad structure, and a small number of rules that must not be violated pass the bar. By contrast, detailed specifications for a particular feature, full library usage instructions, and records of past work are needed only for specific tasks, so they do not.

There is a paradox about length. More instructions may seem likely to improve compliance, but in practice the opposite is closer to the truth. As we saw in Part 2, a longer context dilutes attention across individual items, and in a hundreds-of-lines instruction file, the rules that matter most get buried in the lost in the middle. With 50 rules, all 50 are followed vaguely; with 10, those 10 are followed clearly. If an agent keeps violating CLAUDE.md rules, shortening the file before adding more rules may be the right order of operations.

Of the three always-loaded layers, CLAUDE.md is the only one you can design directly
Of the three always-loaded layers, CLAUDE.md is the only one you can design directly

Do not load everything; load only pointers

So where should the detailed documents that did not pass the bar go—the ones needed only occasionally? The answer is: “Keep them outside the context and provide only their location.”

Instead of putting the full database migration procedure in CLAUDE.md, write one line: “See docs/migration.md for the migration procedure.” The agent reads that file only when performing a migration. Detailed content enters the context only for sessions that need it; unrelated sessions pay only the cost of a one-line pointer.

Claude Code’s skill is a systematic application of the same principle. A skill is a procedure for a specific task; normally, only its name and one-line description are present in the context, and the body loads when that task begins. If you turn the “deployment procedure” into a skill instead of keeping it resident in CLAUDE.md, you save those tokens on the 99% of turns that do not involve deployment.

Tools also need pruning. If an MCP server is connected but unused, disabling it removes tens of thousands of tokens in fixed costs. More harnesses now support lazy loading, keeping only tool names normally and loading full schemas when needed, but the direction is the same: load what is needed when it is needed instead of loading everything all the time.

Put only a few always-true rules on the wall; keep detailed manuals on a shelf and take them down when needed
Put only a few always-true rules on the wall; keep detailed manuals on a shelf and take them down when needed

A fixed-cost review routine

Always-loaded context is difficult to notice once it expands. The cost is the same every session, so there is nothing to compare it against. That is why it is worth checking deliberately from time to time.

In Claude Code, /context breaks down where and how much of the current context is being used. It displays tokens for the system prompt, tools, MCP, and memory files, so if tool definitions appear anomalously larger than the conversation, start by pruning MCP servers. Open CLAUDE.md roughly once a quarter and delete lines based on whether “this line actually helped last month.” Instruction files only grow when neglected, so pruning must become a routine to keep them short.

Summary

  • The system prompt, tool definitions, and CLAUDE.md are fixed costs carried on every turn. They do not disappear with /clear, so they must be designed before managing the conversation.
  • The criterion for CLAUDE.md is “Is this always true for every task?” Since length blurs individual rules, shortening the file should come first when rules are not being followed.
  • Use pointers instead of full detailed documents, split recurring procedures into skills, and disable unused MCP servers. The principle is to load what is needed only when it is needed.

At this point, one session’s context is fairly clean. But no matter how carefully you conserve it, there comes a moment when a large task will not fit in one session. The next part covers the structured solution for that moment: subagents. We will see how an isolation pattern—having another context explore and returning only the conclusion—protects context.