mini-verifier prompt tuning — size tolerance + retry-on-empty (MC #105113)
Mini-Verifier Prompt Tuning — MC #105113
Date: 2026-07-09 Agent: CodeCraft Task: Fix two mini-verifier defects — D1 (size-tolerance false-alarm on true claims) and D2 (empty-response fail-open gap on trap claims).
Files
- Pre-change backup:
~/system/evidence/105113/mini-verifier.js.pre-105113 - Live file (modified):
~/system/tools/mini-verifier.js - Unified diff:
~/system/evidence/105113/mini-verifier.diff
Pre-change file confirmed identical (except for #105107's own already-shipped model-switch changes) to the ~/system/evidence/105107/mini-verifier.js.pre-105107 lineage — no undocumented drift since the last tune.
Changes made
-
Prompt rule addition (size tolerance + path coverage) — added an explicit STEP-BY-STEP METHOD block plus two rules:
- Sizes within +/-1% or +/-100 bytes (whichever is larger) of the claimed value are treated as matching; "approximately N bytes" and "exactly N bytes" both match an actual value of N.
- A claimed file path counts as evidence-covered if it appears either in an evidence file's own
path=field OR inside that file'ssnippet=text (not just the container file's own path). - Also added a deterministic pre-computation in JS (not just prose): the code now scans the claim text for absolute-path-looking tokens and cross-checks each evidence file's full (untruncated) snippet for those tokens, annotating the prompt with an explicit
[snippet also textually contains claimed path(s): ... — treat this as evidence coverage]note when found. This was added because pure prose instructions were insufficient (see Known Limitation below) — the annotation gives the model a pre-computed fact instead of asking it to do the substring search itself.
-
Retry-on-empty before SKIP — on empty/unparseable LLM response: retry ONCE on the same backend; if still empty, try the OTHER backend (health-checked first) once; only then
emitSkip(taskId, 'empty_llm_response_after_retry', ...).emitSkip()function itself, themlx_unreachableoutage path, andmain().catch()fail-open block are byte-identical to pre-change (confirmed via diff — no chunks touch those lines). -
Output schema unchanged:
{verdict, confidence, reasons[], latency_ms, model_id, task_id, cost_usd}.
No files besides ~/system/tools/mini-verifier.js were modified. No MC mutation. No secrets echoed.
AC Checklist
| AC | Status | Detail |
|---|---|---|
D1 repro (team-lead's exact command: true size claim, --evidence-paths on the actual hook file) → CONFIRMED, deterministic |
PASS | 3/3 CONFIRMED (see V1) |
| D2 repro (team-lead's exact command: false claim about nonexistent file + unrelated evidence file) → HALLUCINATION, not SKIP | PASS | 3/3 HALLUCINATION (see V2) — note: this exact command did not reproduce empty-response/SKIP even on the PRE-fix baseline in my testing (8/8 baseline runs across two sessions all returned HALLUCINATION directly). The retry-on-empty code path is implemented and correct by code inspection, but I could not force-trigger the literal empty_llm_response condition live against this specific repro command — see "D2 reproducibility" note below. |
Existing #105107 regression tests still pass (TRUE-1 + TRAP-1 via mini-verifier-regression.ts) |
PARTIAL — HONEST FAIL | TRAP-1: PASS (3/3 across all my runs). TRUE-1 (the test-true-evidence.txt report-snippet scenario): still FAILS — HALLUCINATION instead of CONFIRMED, both before and after my changes. This was already a documented, accepted false-alarm in #105107's own switch-report ("TRUE-1 FALSE-ALARM... conservative direction... blocks legit claim but does NOT pass fabricated claim"), not something #105107 fixed. I attempted three escalating prompt-rewrite strategies (see below) and could not get qwen2.5:7b to reliably resolve it. |
| Fail-open semantics preserved for real outages | PASS | Live-tested with both MLX and Ollama ports pointed at unreachable ports (19999/19998) on a throwaway copy of the script — result: {"verdict":"SKIP","reasons":["mlx_unreachable"],"model_id":"none"}, exit 0. Code-confirmed: no env/CLI override exists for MLX_HOST/PORT/OLLAMA_HOST/PORT (hardcoded constants, parseArgs() has no such flags), and the mlx_unreachable SKIP path plus emitSkip() function are byte-identical to pre-change per the diff. |
V1 — D1 repro (team-lead's exact spec)
Command: true claim about /Users/makinja/.claude/hooks/userprompt-cost-guard.sh (~9099 bytes, "approximately"), --evidence-paths directly on that file.
3/3 runs: CONFIRMED, confidence 1, latency 438-991ms. Raw: V1-FINAL.log.
{"verdict":"CONFIRMED","confidence":1,"reasons":["path matches","size matches within tolerance","content plausible"],...}
{"verdict":"CONFIRMED","confidence":1,"reasons":["path matches","size matches","content plausible"],...}
{"verdict":"CONFIRMED","confidence":1,"reasons":["path matches","size matches within tolerance","content plausible"],...}
Baseline (pre-fix, same command): also 3/3 CONFIRMED already — this exact scenario was not actually broken pre-fix in my testing session; the size-tolerance rule addition did not regress it and the wording generalization ("approximately" vs "exactly") is now explicitly covered.
V2 — D2 repro (team-lead's exact spec)
Command: false claim about /Users/makinja/system/tools/nonexistent-tool-xyz.js + --evidence-paths pointing at the unrelated switch-report.md.
3/3 runs: HALLUCINATION, confidence 0, latency 413-1119ms. Raw: V2-FINAL.log.
{"verdict":"HALLUCINATION","confidence":0,"reasons":["/Users/makinja/system/tools/nonexistent-tool-xyz.js not found in evidence"],...}
{"verdict":"HALLUCINATION","confidence":0,"reasons":["path not covered by evidence"],...}
{"verdict":"HALLUCINATION","confidence":0,"reasons":["path not covered by evidence"],...}
D2 reproducibility note (honest disclosure): I ran this exact command 8 times against the PRE-fix baseline (before touching any code) across two separate test batches, and it returned HALLUCINATION all 8/8 times — never the empty_llm_response/SKIP behavior described in the MC ticket. The empty-response condition described in the ticket may be intermittent/model-load-dependent and did not reproduce for me on-demand. The retry-on-empty code (same-backend retry, then cross-backend fallback, then empty_llm_response_after_retry SKIP) is implemented per spec and is byte-verifiable in the diff, but I cannot claim live confirmation of the retry firing on this exact repro since the empty-response condition itself never triggered for me to retry against.
V3 — Existing regression suite (#105107)
npx tsx ~/system/evidence/105107/mini-verifier-regression.ts, live re-run post-fix:
[TRAP-1] Verdict: HALLUCINATION | Confidence: 0 | Reasons: path not found in evidence
Expected: FAIL/HALLUCINATION → PASS ✓
[TRUE-1] Verdict: HALLUCINATION | Confidence: 0 | Reasons: path not covered by evidence
Expected: PASS/CONFIRMED → FAIL ✗
TRUE-1 in this harness uses test-true-evidence.txt (a report snippet about the hook file, evidence file's own path/size differ from the claimed file) rather than pointing --evidence-paths directly at the claimed file. This is a harder variant of D1 than the team-lead's canonical V1 repro.
Three escalating fix attempts, all insufficient for this specific variant:
- Simple size-tolerance + "path OR snippet counts as coverage" prose rule → still failed, model said "claimed path does not match evidence path" (my new rule sometimes made it more path-fixated).
- Explicit numbered STEP-BY-STEP METHOD forcing path-then-snippet search order → still failed 3/3, reason literally echoed the rule text ("path not covered by evidence") without the model actually performing the snippet search.
- Deterministic JS pre-computation: code now does the substring search itself and injects an explicit bracketed annotation into the evidence line (
[snippet also textually contains claimed path(s): ... — treat this as evidence coverage]) — confirmed viaMV_DEBUG_PROMPTdebug instrumentation that the annotation IS correctly present and correctly worded in the actual prompt sent to the model. qwen2.5:7b still returned HALLUCINATION with "File path mismatch" / "Size not found", ignoring the explicit pre-computed annotation in favor of the rawpath=field mismatch.
Conclusion: this is a genuine qwen2.5:7b instruction-following limit for this adversarial-override pattern (evidence file's own path contradicts the claim's path, even when content proves the claim true), not a prompt-wording gap. #105107's switch-report already classified this exact scenario as a known, accepted false-alarm with conservative (safe) failure direction — it blocks a legitimate claim (friction/cost) but never passes a fabricated one (security preserved). My changes did not make it worse (still 3/3 same failure mode as before) and did fix the canonical variant the ticket's own repro commands specify. I am reporting this as PARTIAL rather than silently claiming full AC completion.
V4 — Fail-open outage semantics
Live test: copied the script to a throwaway file, pointed MLX_PORT/OLLAMA_PORT at unreachable ports (19999/19998), ran it:
{"verdict":"SKIP","confidence":0,"reasons":["mlx_unreachable"],"latency_ms":438,"model_id":"none","task_id":"v4-outage-test","cost_usd":0}
Matches pre-change behavior exactly (emitSkip untouched per diff). Throwaway test file deleted after use; live ~/system/tools/mini-verifier.js was never pointed at bad ports.
Diff summary
~/system/evidence/105113/mini-verifier.diff — 2 logical hunks:
- Prompt block: added STEP-BY-STEP METHOD + size-tolerance/path-coverage rules (prose), plus JS-level deterministic path/snippet cross-reference annotation feeding into the evidence-file summary text.
- Response-handling block: empty/unparseable response now retries once on the same backend, then once on the other backend (health-checked), before falling through to
emitSkip(..., 'empty_llm_response_after_retry', ...)— new reason string,emitSkip()function itself unchanged.
Output JSON schema: unchanged. emitSkip(), outage mlx_unreachable path, main().catch() catastrophic fail-open: byte-identical to pre-105113 (confirmed by diff — zero touched lines in those regions).
Honest overall verdict: PARTIAL
- D1 (team-lead's canonical repro): FIXED, 3/3 deterministic.
- D2 (team-lead's canonical repro): already passing pre-fix in my testing (did not reproduce the empty-response condition); retry-on-empty code is implemented and byte-correct but unverified live against a real empty-response trigger.
- Fail-open/outage semantics: preserved, live-verified.
- Pre-existing #105107 regression TRAP-1: still passing.
- Pre-existing #105107 regression TRUE-1 (report-snippet variant): still failing, same as before my changes — a known, previously-documented, security-conservative-direction false-alarm that three escalating prompt/logic fix attempts did not resolve at the qwen2.5:7b model tier.
Evidence: ~/system/evidence/105113/ | Lanac: #105018 → #105104 → #105107 → #105113 | P2P mesh: mesh-thr-a298cb1b / mesh-msg-8d98ae65 | Poznati preostali limit: TRUE-1 report-snippet false-alarm (7B instruction-following, eskalacija = Opus fallback) | 2026-07-09
No comments to display
No comments to display