Skip to content

llm-judge: record every judge stage, and report jobs that never reach the person - #818

Merged
mergify[bot] merged 1 commit into
mainfrom
wrong-check-reflect/judge-stage-analytics
Sep 24, 2026
Merged

mergify[bot] merged 1 commit into
mainfrom
wrong-check-reflect/judge-stage-analytics

Conversation

@EdbertChan

@EdbertChan EdbertChan commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

Background review checks now record every step: skipped (and why), queued, judged, and delivered.

The old per-run table only saw the Stop hook, which never prints. So "silent" could not tell a broken check from a working one.

A new per-hook summary lines those steps up and flags jobs that were lost on the way. A job queued without a transcript now warns the person right away.

Review Claim

Every background review job leaves one row per step, joined by job id, and the per-hook summary command exits 1 when a job is stuck or never reaches the person.

Review Lane

behavior

Review Unit

engine-runtime

Safety Invariant

A job that cannot be delivered is reported, never silent: a missing transcript prints catstack-hook-error, and a finished verdict with no delivery row counts as a leak. Recording failures print on stderr and do not change what the hook decides.

Slice Rationale

The writer, the three places that call it, and the report that reads the rows are one feature. Any one alone records rows nobody reads, or reads rows nobody writes.

Non-goals

Does not fix the two leaks the report shows (jobs with no transcript, and verdicts from subagent transcripts that are never drained). Does not change any verdict or message.

Test Plan

Test Plan
python3 engine/hooks/_sdk/tests/test_events.py            Ran 8 tests   OK
python3 engine/hooks/llm-judge/tests/test_judge.py        Ran 36 tests  OK
python3 engine/hooks/wrong-check-reflect/tests/test_hooks.py  Ran 29 tests  OK
python3 engine/hooks/_runner/tests/test_report.py         Ran 9 tests   OK

The 7 new stage and report tests fail against the parent commit's code.

Real path, scratch metrics dir, real runner, real Stop and PostToolUse scripts, real transcript:

hook skipped queued finished delivered no_transcript stuck undelivered undelivered_hits
wrong-check-reflect already_prompted=1 2 clean=1,hit=0,unchecked=1 clean=1,hit=0,unchecked=0 1 0 1 0
LEAK: 2 judge job(s) queued with no transcript, stuck, or never delivered after 0s
check exit=1

Revert Plan

Revert Plan

Revert this commit. Stage rows already written have no rule_id, so the rule table keeps ignoring them.

🤖 Generated with Claude Code


Note

Medium Risk
Adds metrics and reporting only; hook verdict behavior is unchanged, but new stage writes and --check exit codes could affect monitoring or CI that runs the report CLI.

Overview
Background LLM judge work is now observable end-to-end: skipped, queued, finished, and delivered verdicts are written as metrics events (linked by job id), instead of disappearing into silent Stop-hook paths.

A new write_stage_event helper records lifecycle rows with mode_source: stage. wrong-check-reflect logs every skip with a reason (gate off, already prompted, etc.) and passes harness (claude / codex / cursor) into queued jobs. llm-judge emits judge_queued / judge_finished on enqueue and completion, and stderr-warns when a job is queued without a transcript path.

report.py --judge aggregates per-hook counts and detects leaks (no transcript, stuck past grace, finished but never delivered). --check exits 1 when leaks exist. Stage rows stay out of the existing rule/event table (empty rule_id).

Reviewed by Cursor Bugbot for commit 96ff887. Bugbot is set up for automated code reviews on this repo. Configure here.

@EdbertChan
EdbertChan force-pushed the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-one-shot-per-reply-scope--60c5103d branch from 9473d50 to ed08f60 Compare September 23, 2026 14:06

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2ca0013. Configure here.

Comment thread engine/hooks/wrong-check-reflect/detect.py Outdated
@EdbertChan

Copy link
Copy Markdown
Owner Author

Mergify repair stopped: GitHub reports merge conflict. The retry cap was reached for current head 2ca0013.

… the person

The runner's per-run table counted wrong-check-reflect as "silent 2,582,
spoke 0" on Claude. That number cannot tell a mute hook from a working one:
the Stop hook never prints, it queues a background judge job, and the verdict
arrives on a later PostToolUse. The live cache showed the real losses were
downstream: 39 hit verdicts for this hook sat undelivered, 25 of them on jobs
queued with an empty transcript path, which no drain can ever find.

Each stage now writes one catstack.hook_event.v1 row with a reason:

- judge_skipped (wrong-check-reflect): stop_hook_active, gate_off,
  empty_reply, already_prompted, user_asked_reflect, judge_child, bad_payload
- judge_queued (llm-judge, every hook): transcript or no_transcript
- judge_finished (llm-judge): hit, clean or unchecked

Delivery was already recorded by drain. Rows join on the job id.

A job queued with no transcript also prints catstack-hook-error, so the
runner classifies the run as caught_error and hook-health surfaces it on the
next prompt instead of the verdict vanishing.

report.py --judge prints the per-hook funnel plus leak counts (no_transcript,
stuck, undelivered, undelivered_hits) older than --grace (default 1h);
--check exits 1 on any leak. Stage rows carry no rule_id, so the existing
rule table ignores them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Change-Id: Ic6eeb8c30dabc90c4d11270e98ef4d1a2eecddfd
@EdbertChan
EdbertChan changed the base branch from stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-one-shot-per-reply-scope--60c5103d to main September 24, 2026 03:04
@EdbertChan
EdbertChan force-pushed the wrong-check-reflect/judge-stage-analytics branch from 2ca0013 to 96ff887 Compare September 24, 2026 03:06
@cursor

cursor Bot commented Sep 24, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_c3fec6fe-2fef-43c9-a261-148c741264e9)

@mergify

mergify Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Queued — the merge queue status continues in this comment ↓.

@EdbertChan

Copy link
Copy Markdown
Owner Author

@Mergifyio queue

@mergify

mergify Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merge Queue Status

This pull request spent 27 minutes 37 seconds in the queue, including 26 minutes 46 seconds running CI.

Required conditions to merge
  • check-success = lint
  • check-success = test
  • check-success = validate

@mergify mergify Bot added the queued label Sep 24, 2026
@mergify
mergify Bot merged commit ab34e20 into main Sep 24, 2026
6 checks passed
@mergify mergify Bot removed the queued label Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant