Skip to Content
ReferencePerformance & Benchmarks

Performance & Benchmarks

Navigator’s value comes from loading only the documentation a task needs, instead of front-loading everything. This page explains the token-budget model, how Navigator measures real usage, and how to benchmark your own sessions.

The token-budget model

Navigator divides the context window into a few predictable buckets:

BucketTypical sizeNotes
System + tools~50kFixed by Claude Code; not controllable
CLAUDE.md~7kProject instructions, loaded once
Message historymanagedTrimmed via /nav:compact
Documentationon-demandThe bucket Navigator optimizes

The first three buckets are largely fixed. Navigator’s strategy targets the documentation bucket: instead of loading all of .agent/ upfront, it loads the navigator index first (~2k), then the current task doc (~3k), then system docs or SOPs only when a task requires them. See Context Engineering for the underlying principle.

Measured, not estimated

Navigator does not claim a fixed savings number. Plugin-level figures depend on your project’s doc footprint, the task, and how disciplined the lazy-loading is.

Instead, Navigator is instrumented with OpenTelemetry, so it reports cache efficiency and context usage from real session telemetry rather than file-size guesses. Token baselines come from actually measuring .agent/ doc sizes against what was loaded in the session.

To measure your own session, run:

/nav:stats

or say “show my stats”. The skill verifies Navigator is initialized, computes baseline vs. loaded tokens, and prints a report.

Example output (illustrative)

The panel below is illustrative formatting only — your numbers will differ:

╔══════════════════════════════════════════════════════╗ ║ NAVIGATOR EFFICIENCY REPORT ║ ╚══════════════════════════════════════════════════════╝ 📊 TOKEN USAGE Documentation loaded: 12,000 tokens Baseline (all docs): 150,000 tokens (example) Tokens saved: 138,000 tokens (example) 💾 CACHE PERFORMANCE Cache efficiency: (from OpenTelemetry) 📈 SESSION METRICS Context usage: (measured this session) Efficiency score: 0-100 (this session)

The efficiency score (0-100) weights token savings (40), cache efficiency (30), and context usage (30). It describes this session, not a guaranteed product benchmark.

Validation methodology

  • OTel-backed: cache and context metrics are read from Claude Code’s OpenTelemetry stream, not inferred.
  • Per-session: every figure is computed for the session you ran it in.
  • Reproducible: run /nav:stats again after a task to see how your own loading discipline moved the numbers.

If your score is low, the report suggests fixes (load fewer docs upfront, compact after sub-tasks, check caching).