Typed Prompt Judge (judge)
Configures the typed judge that sits behind Navigator’s prompt-time classifiers. The loop-trigger, complexity and ambiguity detectors are keyword matchers: fourteen trigger phrases, weighted complexity keywords, an ambiguity base minus credits. They are cheap and deterministic, and they misfire on text that merely contains a trigger phrase — a pasted report header reading “Loop mode” is enough to put a session into Loop Mode.
With judge.enabled, one request per prompt goes to TypeSafe’s Jev , a model that answers typed questions with calibrated probabilities instead of text. Eight questions travel in that one request: is this a task, does it ask for autonomous iteration, complexity on a four-level rubric, ambiguity on three, and whether scope, limits, approach and verification are stated. Each axis overrides its keyword heuristic only when the answer is decisive. An undecided answer, a timeout, a missing key or any error leaves the heuristic in charge for that axis, byte for byte.
New in v7.7.0. Ships off: the prompt text is sent to the API when the judge is on.
Measured
60 labeled prompts, including the observed false-fires (the fixture and recorded responses ship in the plugin under hooks/nav_hook_lib/fixtures/):
| Axis | Keyword heuristic | With the judge |
|---|---|---|
| Tier (DIRECT / TASK / LOOP) | 38/60 | 53/60 |
| Task-shaped (brief gate) | 43/60 | 54/60 |
| Ambiguous (brief needed) | 47/60 | 50/60 |
Median latency 662 ms, about 700 input tokens per prompt, roughly $0.00003 each at list price. The “Loop mode” header scores 0.14 on the loop question and is silenced. Replay the sweep with python3 scripts/judge_eval.py --replay hooks/nav_hook_lib/fixtures/judge_eval_recorded.json from the plugin root.
Setup
- Key — create one at console.typesafe.ai/keys . Provide it either as
TYPESAFE_API_KEYin the environment Claude Code runs in, or in the file~/.config/typesafe/api_key(chmod 600). The environment variable wins when both exist. Never put the key in.agent/.nav-config.json; that file is committed. - Enable — say “enable judge” (
nav-features), or set the block below. The toggle prints the key instructions. - Verify — from the plugin root:
Prints the enable flag, the key source (never the key) and a live round trip with model, latency and tokens. Exit 0 = ready, 1 = no key, 2 = round trip failed.
python3 hooks/nav_hook_lib/judge.py --check
Session start then shows Typed judge: on (jev-latest, key from env:TYPESAFE_API_KEY), or a warning with the same instructions when no key resolves. The WORKFLOW CHECK block gains one line, Judged by jev-1.13.0 (662 ms), whenever a judgment was used.
Config block
{
"judge": {
"enabled": false,
"provider": "typesafe",
"endpoint": "https://api.typesafe.ai/v1/systemone",
"model": "jev-latest",
"timeout_ms": 1500,
"min_confidence": 0.4,
"noul_low": 0.4,
"noul_high": 0.6,
"api_key_env": "TYPESAFE_API_KEY",
"api_key_file": "~/.config/typesafe/api_key",
"max_state_chars": 4000
}
}Keys
enabled(defaultfalse) — Master switch.model(defaultjev-latest) —jev-latestmoves with TypeSafe releases, which can change answers under you; pinjev-1.13.0for stable behavior.timeout_ms(default1500) — Fuse for the request.UserPromptSubmithas a 5-second manifest budget andprompt_gateis deadline-exempt, so this fuse is the only protection.min_confidence(default0.4) — Score axes (complexity, ambiguity) count at or above this confidence.noul_low/noul_high(default0.4/0.6) — Yes/no axes count only outside this band; inside it the heuristic answers.api_key_env,api_key_file— Where the key is read from, in that order.max_state_chars(default4000) — Head cap on the prompt before it is sent.
Thresholds are the winners of a nine-cell sweep over the recorded responses, not cookbook defaults.
Telemetry (v7.7.1)
Runtime state keeps a judge section (30-day TTL) with calls, failures, last and max latency, and per axis — loop, complexity, task-shaped, ambiguity — whether the judge overrode the keyword decision, agreed with it, or stayed undecided inside the band. Counters only; no prompt text. nav stats shows them as two lines:
judge: 86 calls / 1 failed · 499 ms last, 1271 ms max
judge axes: 4 overridden / 240 agreed / 15 undecidedMany undecided means the band is too wide for your prompts; overrides on an axis that keeps being wrong means tighten it or turn it off. First live read across two projects, four days, 92 prompts: 4 overrides, all on the task axis, none on loop or complexity — the keywords are right on ordinary prompts and the judge’s value is in the tails.
A labeled set from your own prompts can be built with scripts/judge_label.py (extract, one-key label, sheet/import). Its output is gitignored: it holds real prompts and must stay on your machine. scripts/judge_eval.py --fixture <file> --live scores it, skipping unlabeled rows.
Privacy
Only the user prompt leaves the machine: the same strip_all()’d text the ops score, with key-shaped tokens redacted (apikey_…, sk-…, ghp_…, AWS, Slack, JWT, long hex, key= / token= pairs) and head-capped. Tool output, files and the transcript never do. Under PILOT_EXECUTOR the judge never runs.
What it does not touch
Tier-1 exact match (fuzzy matching was rejected on purpose), the deep-research ship gate (documented as “not an LLM judge”), the read guard and the Stop gates (hard blockers inside a 5-second budget).
Related
.nav-config.jsonSchema- Hooks as Runtime
- Task Mode — the complexity threshold the judge feeds
- nav-brief — the ambiguity threshold the judge feeds