wrong-check-reflect: only the person's own reflect request suppresses the hook - #787
Conversation
…o the turn Two independent lockouts kept this hook from speaking on Claude Code. One: `already_prompted` was keyed on the transcript path, so the Stop of the reply BEFORE a correction spent the session's single shot. The correction itself then hit an already-prompted key. The key is now the transcript path plus a hash of the reply text, so each distinct reply gets its own chance and the same reply is still judged only once. Two: `user_already_asked_reflect` scanned the whole transcript. One `/reflect` typed at the start of a session switched the detector off for every later reply. The scan is now scoped to the user messages of the turn that produced the reply being judged, which is the case the skip was written for. An unreadable transcript now says so on stderr instead of silently reading as "the user did not ask". The harness-parity test asserts the same reply is judged under both `Stop` and `stop`. It passes before this change too: this hook compares no event name anywhere, so event-name case cannot be what split its hits by harness. The test pins that, rather than leaving the parity unstated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Change-Id: Idb90d55fa90496fcab44e1f0255b303b5cf0b514
…n, not from wording Both retraction detectors need a wrongness word. Two real corrections carry none, because the claim was true and only its order was wrong: "Correcting one claim and arming the check I implied:" "I was right - but I said it a turn before I checked it" Primary fix, no regex and no model: a ledger row going outstanding -> discharged IS the event "a claim was asserted before its check ran, and the check has now run". The Stop that discharges a row now reports it as a reflect trigger. It reads state, so the reply's wording is not consulted at all. Secondary fix: the judge dictionary's meaning now covers "right, but said before it was checked", and both real strings are eval cases. Run against the real judge, all six cases land where the dictionary says they should. The regex scan keeps its hole on purpose, pinned by three tests so the next author widens it knowingly rather than by accident. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Change-Id: I99190104b68d88f78aa9c8a706c9a6d23055aecc
…e only The reminder was re-read to the user on every prompt for every outstanding claim, including claims whose proof had already been pasted. There was no way to turn it down: CATSTACK_HOOK_MODE_UNVERIFIED_TAG_LEDGER=off still printed the full list at exit 0, because this hook is not on run_hook and nothing reads the registry mode for it. CATSTACK_UNVERIFIED_TAG_REMINDER is read in-hook through _flags/flags.py, the one pattern proven to work here (wrong-check-reflect/detect.py:12-15, 191): stale default, and what unset means: only claims already 3+ turns old all the old behaviour off no injection Unset resolves to stale on purpose. A flag whose unset value is the old behaviour changes nothing for the person who asked for less. reminder() at :138 already computed that exact subset and threw it away at :145. The gate is on the injection only. Emission and recording are untouched on every setting, so `off` buys silence without emptying the ledger it is mined from. A value nobody understands, and an .env candidate that cannot be read, are both named on stderr instead of passing as "not set". hooks.toml gains an enabled_by line naming the flag. Nothing reads that field -- registry.py:15 is its only mention in the repo -- so it is documentation beside the gate, not the gate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Change-Id: I4f465761ee7f84489616ba1e83c54eb98ceefff0
FILE_LINE_RE matched a path and a line number and the paragraph loop skipped
the paragraph. Nothing checked that the file existed, that anyone had read it,
or at what ref. Pasting a made-up path beside the same claim turned the gate
off just as well:
$ printf '%s' '{"last_assistant_message":"The problem was the old value at
totally-made-up-file-that-does-not-exist.ts:99999.", ...}' \
| python3 engine/hooks/diu-stop/claude_stop_check.py
exit=0
A citation now counts when it names the ref it was read at
(`path:line @ origin/main`), or when the session transcript shows a tool call
that named that path. That is the rule corpus/CLAUDE.learned.md already ships
in prose and nothing enforced.
Third outcome, pinned: when the transcript cannot be read, the citation is
unchecked rather than clear. It does not buy silence, the reason goes to
stderr, and the block names the path it could not check.
tests/test_hooks.py:412 asserted the old behaviour. Its own PR, #479, listed
the file:line exemption under Non-goals -- it carried the exemption over so
fence normalization would not regress it, and did not decide that an unread
path should silence anything. The test keeps that intent and now backs its
citation with a read.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I9ece412c527ac3272bf86d31d8089db65fd685f5
find_marker_problems ran markers.TAG_RE and the legacy-marker check over the
raw message. find_unverified_claims, ten lines below it, already stripped
fenced blocks and inline code and never shared that with the marker check.
The result: explaining the tag tripped the gate, and so did quoting
cat-mode's own rule, and so did relaying this gate's refusal word for word --
which cat-mode/SKILL.md:55 asks for.
$ printf '%s' '{"last_assistant_message":"The escape hatch is
`{{CAT-UNVERIFIED}}` and it has to name a blocker."}' \
| python3 engine/hooks/diu-stop/claude_stop_check.py
exit=2
A `{{CAT-UNVERIFIED}}` tag here names no blocker. ...
After: exit=0, no output. A tag in running prose with no blocker named still
exits 2.
An unterminated fence drops everything after the marker it left behind,
rather than letting the unclosed block read as prose.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I8d19b3831299b0b18d5b31288017b1b0fff9cd21
The sentence that tells the reader to use the tag showed it as a bare token naming no blocker. That is the shape diu-stop rejects, so quoting cat-mode's own rule tripped the gate enforcing it. It now shows the same template engine/CLAUDE.core.md already uses. tests/test_cat_mode.py gains the structural catch: no tag anywhere in the skill may name no blocker, and the skill must still show the tag at all, so the fix cannot be "delete the example". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Change-Id: If47021e83bdf3d8771ccbdb879001f6effc65a3c
… the hook
user_already_asked_reflect matched any user-shaped line mentioning reflect.
The harness files its own injections as `type: "user"` rows, so two of them
matched:
- `diu-stop`'s Stop-hook feedback, whose text tells the agent to reflect
- the reflect skill's own injected body, so running /reflect disarmed the
hook that asks for it
Read-confirmed on a real transcript: both carry `isMeta: true` on a
`type: "user"` row, which is the field
engine/skills/reflect/scripts/token_audit.py:313 already keys off. This reads
the record rather than a prose prefix, which could only ever catch wordings
somebody had already seen. `agentId` and `isSidechain` are excluded the same
way.
Chesterton's fence checked. `git log -S "user_already_asked_reflect"
origin/main` returns three commits: 5b3b2fb, 30f62e5 "Add
wrong-check-reflect hook for first-person check admissions (#57)", and
e9a9750 (#462). 30f62e5 is the one that introduced the function, and its
message asks for the follow-up "once per transcript" -- do not nag twice.
Matching any mention anywhere in the file is far wider than that. This
narrows it back to the original intent; the turn scoping and the
once-per-reply key are the commit before this one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I31224f6a0f4626d5b7ca6590776a30f314ac50d8
|
This pull request is part of a Mergify stack:
|
Open question for the reviewer: one existing scenario now fails, and it contradicts itself
The failing case is "situation": "Pins the documented skip: ALREADY_REFLECT_RE suppresses the follow-up when the user's own message already asked for /reflect ...",
"user": "\"Claim I made was wrong\" is a trigger for /reflect",
"expect_no_enqueue": ["wrong-check-reflect"]The This is the same mention-vs-use defect that PR #785 fixes in the marker gate, showing up in a fixture instead of in code. Not resolved here, deliberately — a scenario whose label and data contradict each other is a judgment call for the owner, not something to silently re-point. Recommendation:
Separately: 🤖 Generated with Claude Code |
0201bdc to
fab6704
Compare
user_already_asked_reflect matched any user-shaped line mentioning reflect.
The harness files its own injections as
type: "user"rows, so two of themmatched:
diu-stop's Stop-hook feedback, whose text tells the agent to reflecthook that asks for it
Read-confirmed on a real transcript: both carry
isMeta: trueon atype: "user"row, which is the fieldengine/skills/reflect/scripts/token_audit.py:313 already keys off. This reads
the record rather than a prose prefix, which could only ever catch wordings
somebody had already seen.
agentIdandisSidechainare excluded the sameway.
Chesterton's fence checked.
git log -S "user_already_asked_reflect" origin/mainreturns three commits: 5b3b2fb, 30f62e5 "Addwrong-check-reflect hook for first-person check admissions (#57)", and
e9a9750 (#462). 30f62e5 is the one that introduced the function, and its
message asks for the follow-up "once per transcript" -- do not nag twice.
Matching any mention anywhere in the file is far wider than that. This
narrows it back to the original intent; the turn scoping and the
once-per-reply key are the commit before this one.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Depends-On: #786
Note
Medium Risk
Changes when the reflect-enforcement hook is suppressed on Stop; misclassification could either nag after the user already asked for reflect or skip needed follow-ups, though behavior is covered by new tests.
Overview
Fixes wrong-check-reflect so harness-injected transcript rows no longer count as the user asking for
/reflect.user_already_asked_reflectused to treat every user-shaped line in the current turn as the person speaking, so diu-stop Stop-hook feedback and the reflect skill’s injected body (both filed astype: "user"withisMeta) could mention “reflect” and silence the hook—including the case where running/reflectdisarmed the detector that should follow up.The change adds
_is_meta_line(aligned withtoken_auditviaisMeta, plusagentId/isSidechainand a"Stop hook feedback"prefix fallback), tags those rows asmetawhen parsing the transcript, and only scans real user lines for the reflect regex. README documents the behavior; tests cover hook feedback, skill body, and genuine user/reflectin the same turn.Reviewed by Cursor Bugbot for commit bd6a6fd. Bugbot is set up for automated code reviews on this repo. Configure here.