Performance & Benchmarks
Navigator’s value comes from loading only the documentation a task needs, instead of front-loading everything. This page explains the token-budget model, how Navigator measures real usage, and how to benchmark your own sessions.
The token-budget model
Navigator divides the context window into a few predictable buckets:
| Bucket | Typical size | Notes |
|---|---|---|
| System + tools | ~50k | Fixed by Claude Code; not controllable |
CLAUDE.md | ~7k | Project instructions, loaded once |
| Message history | managed | Trimmed via /nav:compact |
| Documentation | on-demand | The bucket Navigator optimizes |
The first three buckets are largely fixed. Navigator’s strategy targets the documentation bucket: instead of loading all of .agent/ upfront, it loads the navigator index first (~2k), then the current task doc (~3k), then system docs or SOPs only when a task requires them. See Context Engineering for the underlying principle.
Measured, not estimated
Navigator does not claim a fixed savings number. Plugin-level figures depend on your project’s doc footprint, the task, and how disciplined the lazy-loading is.
Instead, Navigator is instrumented with OpenTelemetry, so it reports cache efficiency and context usage from real session telemetry rather than file-size guesses. Token baselines come from actually measuring .agent/ doc sizes against what was loaded in the session.
To measure your own session, run:
/nav:statsor say “show my stats”. The skill verifies Navigator is initialized, computes baseline vs. loaded tokens, and prints a report.
Example output (illustrative)
The panel below is illustrative formatting only — your numbers will differ:
╔══════════════════════════════════════════════════════╗
║ NAVIGATOR EFFICIENCY REPORT ║
╚══════════════════════════════════════════════════════╝
📊 TOKEN USAGE
Documentation loaded: 12,000 tokens
Baseline (all docs): 150,000 tokens (example)
Tokens saved: 138,000 tokens (example)
💾 CACHE PERFORMANCE
Cache efficiency: (from OpenTelemetry)
📈 SESSION METRICS
Context usage: (measured this session)
Efficiency score: 0-100 (this session)The efficiency score (0-100) weights token savings (40), cache efficiency (30), and context usage (30). It describes this session, not a guaranteed product benchmark.
Validation methodology
- OTel-backed: cache and context metrics are read from Claude Code’s OpenTelemetry stream, not inferred.
- Per-session: every figure is computed for the session you ran it in.
- Reproducible: run
/nav:statsagain after a task to see how your own loading discipline moved the numbers.
If your score is low, the report suggests fixes (load fewer docs upfront, compact after sub-tasks, check caching).