nav-deep-research
Web deep research: one question in, one cited report.md out, conclusions in the knowledge graph. Sources are fetched into fenced notes, a critic attacks the draft, a patcher applies hunks, and a deterministic ship gate decides whether the report may ship. New in v7.3.0; answer-first report layout since v7.4.0; source lens and the corroboration check since v7.6.0. Ships off.
For questions about your own codebase, use the navigator-research agent instead. This skill is for questions about the outside world: literature, vendors, standards, the state of a technology.
Enable
{ "deep_research": { "enabled": true } }Or say “enable deep_research”. It ships off because a run fetches dozens of third-party pages and spends several opus subagent calls.
Triggers
- “deep research on X”
- “research the web for X”
- “write a research report on X”
- “what does the literature say about X”
- “resume research
<slug>”
How a run works
The skill is a thin router. Each step’s procedure lives in its own file and is loaded at the moment the step starts, so a long run never depends on a procedure that context compaction has already evicted. Every heavy step runs in a subagent; your main session sees digests, findings, and the final report.
| Step | Who | Artifact |
|---|---|---|
| 1 Decompose | main session | atomic items + a search plan from three lenses: breadth, canonical/primary, adversarial |
| the lens travels with each URL through the whole run | ||
| 2 Sweep | WebSearch + parallel deep-research-fetcher agents | sources/NNN.md |
| 3 Draft | one deep-research-writer | report.md, written once |
| 4 Critique | one deep-research-critic | findings/critic.json |
| 5 Patch | one deep-research-patcher | patch-log.json |
| 6 Ship | main session | ship.json, graph memories, an index line in your docs navigator |
Everything lives under .agent/research/<slug>/. The verbatim query is persisted once in query.md and pasted, unparaphrased, into every subagent prompt. Fetched bodies under sources/ are gitignored; source_store.py refetch rebuilds them from the recorded URLs and reports checksum mismatches.
Three rules the pipeline enforces
Patch, never regenerate. After the writer produces the report, the only changes are surgical edits. The critic has no Edit tool; the patcher has Read and Edit only. A finding that needs a rewrite is forwarded as structural, not applied.
Fetched text is data. Every fetched body is wrapped in a <nav-untrusted-source url="…"> fence with a treat-as-data preamble. Forged fence tags inside a page are renamed rather than removed, and the URL attribute is escaped. A page that says “ignore previous instructions” is stored as content.
The gate is a script, not a judgment. ship_gate.py runs thirteen checks and exits non-zero on any failure: report exists, required sections, no citation ranges, every [n] in the body has a Sources row and vice versa, every row maps to an ok source note, no fence text leaked into the report, no unresolved critical finding, the Sources table never shrank after patching, minimum source count, no finding resting on a single breadth-lens source, typed Key findings, a scannable Summary, no wall-of-text paragraphs. Failures are fixed in the report, never by changing the gate.
Report layout (v7.4.0)
The report is written to be read from the top down: the answer first, then the evidence, then what is still open. The writer reads reference/REPORT-FORMAT.md before drafting, and two gate checks make the shape mechanical rather than a matter of the writer remembering it.
# <Title: the conclusion in at most 12 words>
> **Query:** <verbatim>
> **Date:** 2026-09-12 · **Sources:** 22 · **Register:** survey
## Summary
**Answer:** <one or two sentences> [n]
- **<Lead phrase>.** <one sentence> [n] (three to six bullets)
**Counter-position:** <one or two sentences> [n]
**Still open:** <one sentence>
## At a glance (comparisons only)
| entity | field | field |
## <one section per atomic item>
<lead of one to three sentences> [n]
**What the sources show**
- <one fact per bullet> [n]
**Caveats**
- <disagreement or gap> [n]
## Key findings · ## Open questions · ## Sources| Gate check | Rule |
|---|---|
summary-scannable | ## Summary opens with a **Answer:** line and has at least two bullets |
no-wall-of-text | no paragraph, blockquote, or single bullet before ## Sources exceeds 700 characters; tables, headings, and fenced code are exempt |
findings-corroborated | no Key finding cites a single breadth-lens source; see Source lens |
At ship time the session prints a four-line status block and then the Summary verbatim, so the first screen you see is the answer, not a paragraph.
Source lens (v7.6.0)
A quality score would be a model opinion dressed as a metric, and the gate only keeps checks a script can settle. So the pipeline records provenance instead: which of the search lenses found each URL. That is a fact about the search, not a judgment about the source, and it travels from the queue into the note, the writer’s digest, and the Sources table.
| Lens | Meaning |
|---|---|
canonical | a primary source: spec, vendor doc, paper, issue tracker |
breadth | a general search result |
adversarial | found by a search for limitations, failures, counter-evidence |
gap | added by the single gap wave when an atomic item was thin |
One gate check uses it. findings-corroborated fails when a Key finding’s only citation is a
breadth-lens source, naming the finding. A lone canonical or adversarial source passes:
a spec or a bug report can stand alone. The critic runs a matching pass first, so most cases
are fixed before the gate sees them — by citing a corroborating source already on disk, or by
moving the claim to Open questions. Never by relabelling a lens. Reports written before v7.6.0
have no lens column and are exempt.
Citation contract
Cite with [n] in first-use order from 1. Several sources: [3][5], never [3-5]. The last section is ## Sources:
| n | id | title | url | lens |
|---|---|---|---|---|
| 1 | 007 | Python support for free threading | https://docs.python.org/3/howto/free-threading-python.html | canonical |## Key findings bullets are typed so they can become knowledge-graph memories:
- (pitfall) Importing a C extension not marked free-threading-safe re-enables the GIL for the whole process [5][4]After the gate passes, each typed bullet becomes a memory with the cited URLs as evidence, at the same 0.7 confidence the codebase research agent uses. Since v7.5.1 each URL is tagged with the fetch date and a content-hash prefix — https://… (fetched 2026-09-10, sha256 aab8bb007ea7) — so a reader months later can tell which version of the page supported the claim, and source_store.py refetch compares against the same hash.
Configuration
| Key | Default | Meaning |
|---|---|---|
enabled | false | Master switch |
max_sources | 30 | Target upper bound for the sweep |
min_sources | 8 | Ship gate minimum |
fetchers | 4 | Parallel fetcher agents per wave |
max_full_reads | 10 | Sources the writer may read in full; the rest are used from digests |
critic_enabled | true | Run steps 4 and 5 |
models | fetcher sonnet, others opus | Per-role model |
Resume
python3 "$NDR/research_run.py" resume --run <slug>The manifest records every step. When it is stale (compaction ate the step calls), the artifact scan takes over: the highest-numbered artifact on disk decides the next step.
Limits
- WebFetch is only a fallback for pages the raw fetch cannot read. It returns model-summarized text, so such sources are tagged and may be paraphrased, never quoted. Since v7.8.0 the fallback note replaces the blocked stub under the same id, so a 403 followed by a successful WebFetch yields exactly one usable source.
- PDFs are recorded as skipped.
- No vault across runs, no source quality scoring, no academic APIs, no browser lane. If the project has the
hyperresearchharness installed (.hyperresearch/exists), the skill defers to it and only ingests its final report.