wrong-check-reflect: one shot per reply, and scope the reflect scan to the turn - #781
Conversation
…o the turn Two independent lockouts kept this hook from speaking on Claude Code. One: `already_prompted` was keyed on the transcript path, so the Stop of the reply BEFORE a correction spent the session's single shot. The correction itself then hit an already-prompted key. The key is now the transcript path plus a hash of the reply text, so each distinct reply gets its own chance and the same reply is still judged only once. Two: `user_already_asked_reflect` scanned the whole transcript. One `/reflect` typed at the start of a session switched the detector off for every later reply. The scan is now scoped to the user messages of the turn that produced the reply being judged, which is the case the skip was written for. An unreadable transcript now says so on stderr instead of silently reading as "the user did not ask". The harness-parity test asserts the same reply is judged under both `Stop` and `stop`. It passes before this change too: this hook compares no event name anywhere, so event-name case cannot be what split its hits by harness. The test pins that, rather than leaving the parity unstated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Change-Id: Idb90d55fa90496fcab44e1f0255b303b5cf0b514
|
This pull request is part of a Mergify stack:
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit dad472d. Configure here.
| if not text or text.lstrip().startswith(META_USER_PREFIXES): | ||
| continue | ||
| if ALREADY_REFLECT_RE.search(text): | ||
| return True |
There was a problem hiding this comment.
Reflect scan splits mid-turn replies
Medium Severity
user_already_asked_reflect opens the turn after the previous non-empty assistant line. One Claude turn often writes earlier assistant text (a preamble before tools) before the reply being judged, so a /reflect in that turn's user prompt sits outside the window and the skip does not run.
Reviewed by Cursor Bugbot for commit dad472d. Configure here.


Two independent lockouts kept this hook from speaking on Claude Code.
One:
already_promptedwas keyed on the transcript path, so the Stop of thereply BEFORE a correction spent the session's single shot. The correction
itself then hit an already-prompted key. The key is now the transcript path
plus a hash of the reply text, so each distinct reply gets its own chance and
the same reply is still judged only once.
Two:
user_already_asked_reflectscanned the whole transcript. One/reflecttyped at the start of a session switched the detector off for every later
reply. The scan is now scoped to the user messages of the turn that produced
the reply being judged, which is the case the skip was written for.
An unreadable transcript now says so on stderr instead of silently reading as
"the user did not ask".
The harness-parity test asserts the same reply is judged under both
Stopandstop. It passes before this change too: this hook compares no event nameanywhere, so event-name case cannot be what split its hits by harness. The
test pins that, rather than leaving the parity unstated.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Depends-On: #780
Note
Medium Risk
Changes when reflect enforcement fires across multi-turn sessions; behavior is hook-local with expanded tests but affects safety/reflection workflows when
CATSTACK_REFLECT_ENFORCEMENTis on.Overview
Fixes two wrong-check-reflect lockouts that stopped the hook from enqueueing LLM judge jobs when it should.
Once per reply: Prompted state is keyed by transcript path plus a hash of the assistant reply (
reply_key), not the transcript alone. A Stop on an earlier reply no longer consumes the session’s only shot, so a later correction reply can still be judged; identical reply text is still deduped./reflectskip scoped to the current turn:user_already_asked_reflectonly scans user messages in the turn that produced the reply being judged, so an early/reflectno longer disables the hook for the rest of the session. Same-turn/reflectstill suppresses double follow-up.Unreadable transcripts emit a stderr hook error (“unchecked”) via
_transcript_rolesinstead of failing open as “user did not ask.” README and tests cover correction-after-prior-reply, session-long/reflect, dedup, harnessStop/stopparity, and read failures.Reviewed by Cursor Bugbot for commit dad472d. Bugbot is set up for automated code reviews on this repo. Configure here.