Skip to Content
ConfigurationTyped Prompt Judge

Typed Prompt Judge (judge)

Configures the typed judge that sits behind Navigator’s prompt-time classifiers. The loop-trigger, complexity and ambiguity detectors are keyword matchers: fourteen trigger phrases, weighted complexity keywords, an ambiguity base minus credits. They are cheap and deterministic, and they misfire on text that merely contains a trigger phrase — a pasted report header reading “Loop mode” is enough to put a session into Loop Mode.

With judge.enabled, one request per prompt goes to TypeSafe’s Jev , a model that answers typed questions with calibrated probabilities instead of text. Eight questions travel in that one request: is this a task, does it ask for autonomous iteration, complexity on a four-level rubric, ambiguity on three, and whether scope, limits, approach and verification are stated. Each axis overrides its keyword heuristic only when the answer is decisive. An undecided answer, a timeout, a missing key or any error leaves the heuristic in charge for that axis, byte for byte.

New in v7.7.0. Ships off: the prompt text is sent to the API when the judge is on.

Measured

60 labeled prompts, including the observed false-fires (the fixture and recorded responses ship in the plugin under hooks/nav_hook_lib/fixtures/):

AxisKeyword heuristicWith the judge
Tier (DIRECT / TASK / LOOP)38/6053/60
Task-shaped (brief gate)43/6054/60
Ambiguous (brief needed)47/6050/60

Median latency 662 ms, about 700 input tokens per prompt, roughly $0.00003 each at list price. The “Loop mode” header scores 0.14 on the loop question and is silenced. Replay the sweep with python3 scripts/judge_eval.py --replay hooks/nav_hook_lib/fixtures/judge_eval_recorded.json from the plugin root.

Setup

  1. Key — create one at console.typesafe.ai/keys . Provide it either as TYPESAFE_API_KEY in the environment Claude Code runs in, or in the file ~/.config/typesafe/api_key (chmod 600). The environment variable wins when both exist. Never put the key in .agent/.nav-config.json; that file is committed.
  2. Enable — say “enable judge” (nav-features), or set the block below. The toggle prints the key instructions.
  3. Verify — from the plugin root:
    python3 hooks/nav_hook_lib/judge.py --check
    Prints the enable flag, the key source (never the key) and a live round trip with model, latency and tokens. Exit 0 = ready, 1 = no key, 2 = round trip failed.

Session start then shows Typed judge: on (jev-latest, key from env:TYPESAFE_API_KEY), or a warning with the same instructions when no key resolves. The WORKFLOW CHECK block gains one line, Judged by jev-1.13.0 (662 ms), whenever a judgment was used.

Config block

{ "judge": { "enabled": false, "provider": "typesafe", "endpoint": "https://api.typesafe.ai/v1/systemone", "model": "jev-latest", "timeout_ms": 1500, "min_confidence": 0.4, "noul_low": 0.4, "noul_high": 0.6, "api_key_env": "TYPESAFE_API_KEY", "api_key_file": "~/.config/typesafe/api_key", "max_state_chars": 4000 } }

Keys

  • enabled (default false) — Master switch.
  • model (default jev-latest) — jev-latest moves with TypeSafe releases, which can change answers under you; pin jev-1.13.0 for stable behavior.
  • timeout_ms (default 1500) — Fuse for the request. UserPromptSubmit has a 5-second manifest budget and prompt_gate is deadline-exempt, so this fuse is the only protection.
  • min_confidence (default 0.4) — Score axes (complexity, ambiguity) count at or above this confidence.
  • noul_low / noul_high (default 0.4 / 0.6) — Yes/no axes count only outside this band; inside it the heuristic answers.
  • api_key_env, api_key_file — Where the key is read from, in that order.
  • max_state_chars (default 4000) — Head cap on the prompt before it is sent.

Thresholds are the winners of a nine-cell sweep over the recorded responses, not cookbook defaults.

Telemetry (v7.7.1)

Runtime state keeps a judge section (30-day TTL) with calls, failures, last and max latency, and per axis — loop, complexity, task-shaped, ambiguity — whether the judge overrode the keyword decision, agreed with it, or stayed undecided inside the band. Counters only; no prompt text. nav stats shows them as two lines:

judge: 86 calls / 1 failed · 499 ms last, 1271 ms max judge axes: 4 overridden / 240 agreed / 15 undecided

Many undecided means the band is too wide for your prompts; overrides on an axis that keeps being wrong means tighten it or turn it off. First live read across two projects, four days, 92 prompts: 4 overrides, all on the task axis, none on loop or complexity — the keywords are right on ordinary prompts and the judge’s value is in the tails.

A labeled set from your own prompts can be built with scripts/judge_label.py (extract, one-key label, sheet/import). Its output is gitignored: it holds real prompts and must stay on your machine. scripts/judge_eval.py --fixture <file> --live scores it, skipping unlabeled rows.

Privacy

Only the user prompt leaves the machine: the same strip_all()’d text the ops score, with key-shaped tokens redacted (apikey_…, sk-…, ghp_…, AWS, Slack, JWT, long hex, key= / token= pairs) and head-capped. Tool output, files and the transcript never do. Under PILOT_EXECUTOR the judge never runs.

What it does not touch

Tier-1 exact match (fuzzy matching was rejected on purpose), the deep-research ship gate (documented as “not an LLM judge”), the read guard and the Stop gates (hard blockers inside a 5-second budget).