llm-judge: record every judge stage, and report jobs that never reach the person - #818
Conversation
9473d50 to
ed08f60
Compare
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 2ca0013. Configure here.
|
Mergify repair stopped: GitHub reports merge conflict. The retry cap was reached for current head 2ca0013. |
… the person The runner's per-run table counted wrong-check-reflect as "silent 2,582, spoke 0" on Claude. That number cannot tell a mute hook from a working one: the Stop hook never prints, it queues a background judge job, and the verdict arrives on a later PostToolUse. The live cache showed the real losses were downstream: 39 hit verdicts for this hook sat undelivered, 25 of them on jobs queued with an empty transcript path, which no drain can ever find. Each stage now writes one catstack.hook_event.v1 row with a reason: - judge_skipped (wrong-check-reflect): stop_hook_active, gate_off, empty_reply, already_prompted, user_asked_reflect, judge_child, bad_payload - judge_queued (llm-judge, every hook): transcript or no_transcript - judge_finished (llm-judge): hit, clean or unchecked Delivery was already recorded by drain. Rows join on the job id. A job queued with no transcript also prints catstack-hook-error, so the runner classifies the run as caught_error and hook-health surfaces it on the next prompt instead of the verdict vanishing. report.py --judge prints the per-hook funnel plus leak counts (no_transcript, stuck, undelivered, undelivered_hits) older than --grace (default 1h); --check exits 1 on any leak. Stage rows carry no rule_id, so the existing rule table ignores them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Change-Id: Ic6eeb8c30dabc90c4d11270e98ef4d1a2eecddfd
2ca0013 to
96ff887
Compare
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_c3fec6fe-2fef-43c9-a261-148c741264e9) |
|
Queued — the merge queue status continues in this comment ↓. |
|
@Mergifyio queue |
Merge Queue Status
This pull request spent 27 minutes 37 seconds in the queue, including 26 minutes 46 seconds running CI. Required conditions to merge
|

Summary
Background review checks now record every step: skipped (and why), queued, judged, and delivered.
The old per-run table only saw the Stop hook, which never prints. So "silent" could not tell a broken check from a working one.
A new per-hook summary lines those steps up and flags jobs that were lost on the way. A job queued without a transcript now warns the person right away.
Review Claim
Every background review job leaves one row per step, joined by job id, and the per-hook summary command exits 1 when a job is stuck or never reaches the person.
Review Lane
behavior
Review Unit
engine-runtime
Safety Invariant
A job that cannot be delivered is reported, never silent: a missing transcript prints
catstack-hook-error, and a finished verdict with no delivery row counts as a leak. Recording failures print on stderr and do not change what the hook decides.Slice Rationale
The writer, the three places that call it, and the report that reads the rows are one feature. Any one alone records rows nobody reads, or reads rows nobody writes.
Non-goals
Does not fix the two leaks the report shows (jobs with no transcript, and verdicts from subagent transcripts that are never drained). Does not change any verdict or message.
Test Plan
Test Plan
The 7 new stage and report tests fail against the parent commit's code.
Real path, scratch metrics dir, real runner, real Stop and PostToolUse scripts, real transcript:
Revert Plan
Revert Plan
Revert this commit. Stage rows already written have no rule_id, so the rule table keeps ignoring them.
🤖 Generated with Claude Code
Note
Medium Risk
Adds metrics and reporting only; hook verdict behavior is unchanged, but new stage writes and
--checkexit codes could affect monitoring or CI that runs the report CLI.Overview
Background LLM judge work is now observable end-to-end: skipped, queued, finished, and delivered verdicts are written as metrics events (linked by job id), instead of disappearing into silent Stop-hook paths.
A new
write_stage_eventhelper records lifecycle rows withmode_source: stage. wrong-check-reflect logs every skip with a reason (gate off, already prompted, etc.) and passes harness (claude/codex/cursor) into queued jobs. llm-judge emitsjudge_queued/judge_finishedon enqueue and completion, and stderr-warns when a job is queued without a transcript path.report.py --judgeaggregates per-hook counts and detects leaks (no transcript, stuck past grace, finished but never delivered).--checkexits 1 when leaks exist. Stage rows stay out of the existing rule/event table (emptyrule_id).Reviewed by Cursor Bugbot for commit 96ff887. Bugbot is set up for automated code reviews on this repo. Configure here.