Skip to content

wrong-check-reflect: only the person's own reflect request suppresses the hook - #787

Closed
EdbertChan wants to merge 7 commits into
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/cat-mode-show-escape-hatch-tag-form-gate-accepts--f47021e8from
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-only-person-s-own-reflect-req--31224f6a
Closed

EdbertChan wants to merge 7 commits into
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/cat-mode-show-escape-hatch-tag-form-gate-accepts--f47021e8from
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-only-person-s-own-reflect-req--31224f6a

Conversation

@EdbertChan

@EdbertChan EdbertChan commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

user_already_asked_reflect matched any user-shaped line mentioning reflect.
The harness files its own injections as type: "user" rows, so two of them
matched:

  • diu-stop's Stop-hook feedback, whose text tells the agent to reflect
  • the reflect skill's own injected body, so running /reflect disarmed the
    hook that asks for it

Read-confirmed on a real transcript: both carry isMeta: true on a
type: "user" row, which is the field
engine/skills/reflect/scripts/token_audit.py:313 already keys off. This reads
the record rather than a prose prefix, which could only ever catch wordings
somebody had already seen. agentId and isSidechain are excluded the same
way.

Chesterton's fence checked. git log -S "user_already_asked_reflect" origin/main returns three commits: 5b3b2fb, 30f62e5 "Add
wrong-check-reflect hook for first-person check admissions (#57)", and
e9a9750 (#462). 30f62e5 is the one that introduced the function, and its
message asks for the follow-up "once per transcript" -- do not nag twice.
Matching any mention anywhere in the file is far wider than that. This
narrows it back to the original intent; the turn scoping and the
once-per-reply key are the commit before this one.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com

Depends-On: #786


Note

Medium Risk
Changes when the reflect-enforcement hook is suppressed on Stop; misclassification could either nag after the user already asked for reflect or skip needed follow-ups, though behavior is covered by new tests.

Overview
Fixes wrong-check-reflect so harness-injected transcript rows no longer count as the user asking for /reflect.

user_already_asked_reflect used to treat every user-shaped line in the current turn as the person speaking, so diu-stop Stop-hook feedback and the reflect skill’s injected body (both filed as type: "user" with isMeta) could mention “reflect” and silence the hook—including the case where running /reflect disarmed the detector that should follow up.

The change adds _is_meta_line (aligned with token_audit via isMeta, plus agentId / isSidechain and a "Stop hook feedback" prefix fallback), tags those rows as meta when parsing the transcript, and only scans real user lines for the reflect regex. README documents the behavior; tests cover hook feedback, skill body, and genuine user /reflect in the same turn.

Reviewed by Cursor Bugbot for commit bd6a6fd. Bugbot is set up for automated code reviews on this repo. Configure here.

EdbertChan and others added 7 commits September 22, 2026 00:51
…o the turn

Two independent lockouts kept this hook from speaking on Claude Code.

One: `already_prompted` was keyed on the transcript path, so the Stop of the
reply BEFORE a correction spent the session's single shot. The correction
itself then hit an already-prompted key. The key is now the transcript path
plus a hash of the reply text, so each distinct reply gets its own chance and
the same reply is still judged only once.

Two: `user_already_asked_reflect` scanned the whole transcript. One `/reflect`
typed at the start of a session switched the detector off for every later
reply. The scan is now scoped to the user messages of the turn that produced
the reply being judged, which is the case the skip was written for.

An unreadable transcript now says so on stderr instead of silently reading as
"the user did not ask".

The harness-parity test asserts the same reply is judged under both `Stop` and
`stop`. It passes before this change too: this hook compares no event name
anywhere, so event-name case cannot be what split its hits by harness. The
test pins that, rather than leaving the parity unstated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: Idb90d55fa90496fcab44e1f0255b303b5cf0b514
…n, not from wording

Both retraction detectors need a wrongness word. Two real corrections carry
none, because the claim was true and only its order was wrong:

  "Correcting one claim and arming the check I implied:"
  "I was right - but I said it a turn before I checked it"

Primary fix, no regex and no model: a ledger row going outstanding ->
discharged IS the event "a claim was asserted before its check ran, and the
check has now run". The Stop that discharges a row now reports it as a reflect
trigger. It reads state, so the reply's wording is not consulted at all.

Secondary fix: the judge dictionary's meaning now covers "right, but said
before it was checked", and both real strings are eval cases. Run against the
real judge, all six cases land where the dictionary says they should.

The regex scan keeps its hole on purpose, pinned by three tests so the next
author widens it knowingly rather than by accident.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I99190104b68d88f78aa9c8a706c9a6d23055aecc
…e only

The reminder was re-read to the user on every prompt for every outstanding
claim, including claims whose proof had already been pasted. There was no way
to turn it down: CATSTACK_HOOK_MODE_UNVERIFIED_TAG_LEDGER=off still printed
the full list at exit 0, because this hook is not on run_hook and nothing
reads the registry mode for it.

CATSTACK_UNVERIFIED_TAG_REMINDER is read in-hook through _flags/flags.py, the
one pattern proven to work here (wrong-check-reflect/detect.py:12-15, 191):

  stale  default, and what unset means: only claims already 3+ turns old
  all    the old behaviour
  off    no injection

Unset resolves to stale on purpose. A flag whose unset value is the old
behaviour changes nothing for the person who asked for less. reminder() at
:138 already computed that exact subset and threw it away at :145.

The gate is on the injection only. Emission and recording are untouched on
every setting, so `off` buys silence without emptying the ledger it is mined
from. A value nobody understands, and an .env candidate that cannot be read,
are both named on stderr instead of passing as "not set".

hooks.toml gains an enabled_by line naming the flag. Nothing reads that field
-- registry.py:15 is its only mention in the repo -- so it is documentation
beside the gate, not the gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I4f465761ee7f84489616ba1e83c54eb98ceefff0
FILE_LINE_RE matched a path and a line number and the paragraph loop skipped
the paragraph. Nothing checked that the file existed, that anyone had read it,
or at what ref. Pasting a made-up path beside the same claim turned the gate
off just as well:

  $ printf '%s' '{"last_assistant_message":"The problem was the old value at
    totally-made-up-file-that-does-not-exist.ts:99999.", ...}' \
      | python3 engine/hooks/diu-stop/claude_stop_check.py
    exit=0

A citation now counts when it names the ref it was read at
(`path:line @ origin/main`), or when the session transcript shows a tool call
that named that path. That is the rule corpus/CLAUDE.learned.md already ships
in prose and nothing enforced.

Third outcome, pinned: when the transcript cannot be read, the citation is
unchecked rather than clear. It does not buy silence, the reason goes to
stderr, and the block names the path it could not check.

tests/test_hooks.py:412 asserted the old behaviour. Its own PR, #479, listed
the file:line exemption under Non-goals -- it carried the exemption over so
fence normalization would not regress it, and did not decide that an unread
path should silence anything. The test keeps that intent and now backs its
citation with a read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I9ece412c527ac3272bf86d31d8089db65fd685f5
find_marker_problems ran markers.TAG_RE and the legacy-marker check over the
raw message. find_unverified_claims, ten lines below it, already stripped
fenced blocks and inline code and never shared that with the marker check.

The result: explaining the tag tripped the gate, and so did quoting
cat-mode's own rule, and so did relaying this gate's refusal word for word --
which cat-mode/SKILL.md:55 asks for.

  $ printf '%s' '{"last_assistant_message":"The escape hatch is
    `{{CAT-UNVERIFIED}}` and it has to name a blocker."}' \
      | python3 engine/hooks/diu-stop/claude_stop_check.py
  exit=2
  A `{{CAT-UNVERIFIED}}` tag here names no blocker. ...

After: exit=0, no output. A tag in running prose with no blocker named still
exits 2.

An unterminated fence drops everything after the marker it left behind,
rather than letting the unclosed block read as prose.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I8d19b3831299b0b18d5b31288017b1b0fff9cd21
The sentence that tells the reader to use the tag showed it as a bare token
naming no blocker. That is the shape diu-stop rejects, so quoting cat-mode's
own rule tripped the gate enforcing it. It now shows the same template
engine/CLAUDE.core.md already uses.

tests/test_cat_mode.py gains the structural catch: no tag anywhere in the
skill may name no blocker, and the skill must still show the tag at all, so
the fix cannot be "delete the example".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: If47021e83bdf3d8771ccbdb879001f6effc65a3c
… the hook

user_already_asked_reflect matched any user-shaped line mentioning reflect.
The harness files its own injections as `type: "user"` rows, so two of them
matched:

  - `diu-stop`'s Stop-hook feedback, whose text tells the agent to reflect
  - the reflect skill's own injected body, so running /reflect disarmed the
    hook that asks for it

Read-confirmed on a real transcript: both carry `isMeta: true` on a
`type: "user"` row, which is the field
engine/skills/reflect/scripts/token_audit.py:313 already keys off. This reads
the record rather than a prose prefix, which could only ever catch wordings
somebody had already seen. `agentId` and `isSidechain` are excluded the same
way.

Chesterton's fence checked. `git log -S "user_already_asked_reflect"
origin/main` returns three commits: 5b3b2fb, 30f62e5 "Add
wrong-check-reflect hook for first-person check admissions (#57)", and
e9a9750 (#462). 30f62e5 is the one that introduced the function, and its
message asks for the follow-up "once per transcript" -- do not nag twice.
Matching any mention anywhere in the file is far wider than that. This
narrows it back to the original intent; the turn scoping and the
once-per-reply key are the commit before this one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Change-Id: I31224f6a0f4626d5b7ca6590776a30f314ac50d8
@EdbertChan

EdbertChan commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

This pull request is part of a Mergify stack:

# Pull Request Link
1 unverified-tag-ledger: read the turn's tools from the transcript, not a key Claude Code never sends #780
2 wrong-check-reflect: one shot per reply, and scope the reflect scan to the turn #781
3 reflect: catch an evidence-order correction from the ledger transition, not from wording #782
4 unverified-tag-ledger: gate the next-prompt reminder, default to stale only #783
5 diu-stop: a file citation buys silence only when it is backed #784
6 diu-stop: the marker gate reads prose, so a mention is not a use #785
7 cat-mode: show the escape-hatch tag in the form the gate accepts #786
8 wrong-check-reflect: only the person's own reflect request suppresses the hook #787 👈
9 principle-subagent-inherits-scope: a brief carries decisions, not just facts and a question #788
10 wrong-check-reflect: only a real /reflect invocation counts as already asked #789
11 llm-judge: claude is the only default runner #793

@EdbertChan

Copy link
Copy Markdown
Owner Author

Open question for the reviewer: one existing scenario now fails, and it contradicts itself

bash scripts/test/run_all_tests.sh on this stack:

Ran 644 tests in 1472.932s
FAILED (failures=1)
  'wrong-check-reflect: expected NO judge job, queued one'

The failing case is tests/scenarios/self-correction.json, scenario admission-skipped-when-user-already-said-reflect:

"situation": "Pins the documented skip: ALREADY_REFLECT_RE suppresses the follow-up when the user's own message already asked for /reflect ...",
"user": "\"Claim I made was wrong\" is a trigger for /reflect",
"expect_no_enqueue": ["wrong-check-reflect"]

The situation and the user field disagree. The situation says the user asked for a reflect. The user text does not ask for anything — it quotes the phrase that triggers a reflect while describing the rule. That is a mention, not a request. This PR makes the hook fire on it, which is why the scenario goes red.

This is the same mention-vs-use defect that PR #785 fixes in the marker gate, showing up in a fixture instead of in code.

Not resolved here, deliberately — a scenario whose label and data contradict each other is a judgment call for the owner, not something to silently re-point. Recommendation:

  1. Change this scenario's user to an actual request (e.g. please /reflect on this) so it keeps pinning the legitimate skip: do not nag when a reflect is genuinely in flight.
  2. Add a second scenario using the current text verbatim, pinning that a mention does not suppress.

Separately: scripts/test/run_all_tests.sh printed FAILED (failures=1) and still exited 0. A check that failed reported success. Worth its own fix; not in this stack.

🤖 Generated with Claude Code

@EdbertChan
EdbertChan force-pushed the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/cat-mode-show-escape-hatch-tag-form-gate-accepts--f47021e8 branch from 0201bdc to fab6704 Compare September 22, 2026 18:38
@EdbertChan EdbertChan closed this Sep 22, 2026
@EdbertChan
EdbertChan deleted the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-only-person-s-own-reflect-req--31224f6a branch September 22, 2026 18:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant