Skip to Content
ConceptsContext Engineering

Context Engineering

The thesis: bulk loading fails. When you load every project doc upfront — “just in case” — the context window fills with noise. The signal you actually need gets buried, the model starts forgetting code you just wrote, and the session crashes around exchange 5–7.

Navigator does the opposite. It treats your context window as a curated workspace and loads strategically: what you need, when you need it. This is the discipline Anthropic calls context engineering — the multi-turn cousin of prompt engineering. Prompt engineering optimizes a single query; context engineering curates what stays in the window across an entire session.

The Problem

The default AI workflow looks reasonable and quietly fails:

  1. Load all project docs at session start (“better safe than sorry”).
  2. Context fills with everything.
  3. The model tries to weigh all of it at once.
  4. Recent changes get lost — signal drowns in noise.
  5. Session dies. Start over. Repeat.

The failure isn’t the tool or the model. It’s the workflow. A window that’s 70%+ full before you write a line of code has no room left for the work itself, and quality degrades as the conversation grows.

The Core Principle

Strategic curation beats bulk loading.

Not “load everything just in case.” Not “better safe than sorry.” Load the index, load the task at hand, and let everything else arrive on demand. A typical session starts with the navigator (~2k) plus the current task doc (~3k) — and leaves the rest of the window free for actual reasoning and code.

Anti-Patterns That Waste Context

  • Upfront loading — pulling all .agent/ docs at session start. 70–90% of it is irrelevant to the task in front of you, and the relevant parts get diluted. This is the root anti-pattern; most others are variations of it.
  • “Better safe than sorry” — the mindset that justifies upfront loading. Availability feels free; it isn’t. Every loaded token is a token the model has to weigh against everything else.
  • Manual multi-file Reads — reading 15–20 files by hand to answer “where does X happen?”. Each file lands in full, most are irrelevant, and nothing is summarized. An exploration agent does the search and returns only the curated answer.
  • Loading full docs for a summary’s worth of need — pulling a 15k file to read one overview section. Fetch the overview first; drill down only if the task demands it.

Patterns That Win

  • Lazy loading — the foundation. Start minimal (navigator + task), then load system docs, SOPs, or integration notes only when a step actually requires them. Progressive, strategic, curated.
  • Navigator-first — always open with the index (DEVELOPMENT-README.md). It’s the map: what documentation exists and when to load it. Without it you navigate blind and fall back to manual searching.
  • Agents for exploration — delegate multi-file searches and “how does this work?” questions to a Task agent. It reads in its own context and returns a summary, so exploration never bloats your main window. Reserve manual Reads for known files you can name.
  • Progressive refinement — fetch metadata before details. Understand the structure, decide what’s relevant, then load only that.

A Token-Budget Illustration

The numbers below are an illustrative sample from one of Navigator’s own sessions — measured with OpenTelemetry, not estimated from file sizes. Your savings depend on your project’s docs and how you work; run /nav:stats to see your own.

200k token context window Bulk loading: ┌─────────────────────────────────────────┐ │ System prompt & tools: 50k (25%) │ fixed overhead │ All docs loaded: 100k (50%) │ ← mostly unused │ Available for work: 50k (25%) │ ← cramped └─────────────────────────────────────────┘ Result: 75% full before you start. Sessions die in 5–7 exchanges. Strategic curation (Navigator): ┌─────────────────────────────────────────┐ │ System prompt & tools: 50k (25%) │ fixed overhead │ Navigator + current task: 12k (6%) │ ← curated │ Available for work: 138k (69%) │ ← spacious └─────────────────────────────────────────┘ Result: ~31% used, room to spare. Sessions last 20+ exchanges.

In this sample, ~150k tokens of available documentation collapsed to ~12k actually loaded — roughly a 92% reduction, with the freed space going straight into longer, more capable sessions. Treat the figure as illustrative of the approach, not a guarantee; the principle holds regardless of the exact number.

  • nav-start — how a context-efficient session begins
  • Performance — metrics, scoring, and /nav:stats
  • Task Mode — phased orchestration for substantial work
  • Knowledge Graph — query everything without loading everything