Skip to content

wrong-check-reflect: only a real /reflect invocation counts as already asked - #789

Closed
EdbertChan wants to merge 9 commits into
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/principle-subagent-inherits-scope-brief-carries--37038435from
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-only-real-reflect-invocation--14642777
Closed

EdbertChan wants to merge 9 commits into
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/principle-subagent-inherits-scope-brief-carries--37038435from
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-only-real-reflect-invocation--14642777

Conversation

@EdbertChan

@EdbertChan EdbertChan commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

The suppression matched any prose containing the word, so a sentence about
the hook switched the hook off. "Claim I made was wrong" is a trigger for /reflect describes when the detector fires; it asks for nothing, and
nothing is in flight to avoid duplicating.

Match the harness's own record of an invocation -- the
<command-name>/reflect</command-name> envelope -- instead. That envelope
starts with <command-, which the meta filter drops, so the one
unambiguous signal was the one being discarded while vague prose was kept.
The window now includes meta rows for this check.

Also stop returning False when the transcript holds no assistant row. A
Stop payload can carry the reply before its row lands, and a session's
first turn has none at all; in both cases every row belongs to this turn,
and bailing out meant a real request went unhonoured.

Erring toward asking is the safe direction here: this detector recorded
1,682 Claude Code invocations and spoke 0 times.

tests/scenarios/self-correction.json: the scenario that pinned this said
the user "already asked for /reflect" while its user field only mentioned
the word. Its intent and its data disagreed. The request case now uses a
real invocation; the mention case is its own scenario and must still fire.

Ran 27 tests in 16.245s OK (wrong-check-reflect)
Ran 6 tests in 0.045s OK (scenario suite)
check_hook_test_coverage: OK (1 hook(s) checked)

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com

Depends-On: #788


Note

Low Risk
Narrower suppression logic in a stop-hook detector with expanded tests; may enqueue slightly more often but avoids missing real reflect requests and false silencing.

Overview
Fixes false suppression when users only mention /reflect in prose (e.g. describing hook triggers). The hook now treats “already asked for reflect” only when the transcript records a real command envelope (<command-name>/reflect</command-name> or <command-message>reflect</command-message>), not bare words like please /reflect in text.

Turn-window logic no longer bails out when there is no assistant row yet (first turn or Stop before the reply is written). The scan runs through the end of the file and includes meta rows so a genuine invocation is not missed while prose about the hook no longer disarms the detector.

Tests and self-correction.json scenarios are split: real invocation still suppresses enqueue; mention-only text must still enqueue wrong-check-reflect.

Reviewed by Cursor Bugbot for commit 721426e. Bugbot is set up for automated code reviews on this repo. Configure here.

EdbertChan and others added 9 commits September 22, 2026 00:51
…o the turn

Two independent lockouts kept this hook from speaking on Claude Code.

One: `already_prompted` was keyed on the transcript path, so the Stop of the
reply BEFORE a correction spent the session's single shot. The correction
itself then hit an already-prompted key. The key is now the transcript path
plus a hash of the reply text, so each distinct reply gets its own chance and
the same reply is still judged only once.

Two: `user_already_asked_reflect` scanned the whole transcript. One `/reflect`
typed at the start of a session switched the detector off for every later
reply. The scan is now scoped to the user messages of the turn that produced
the reply being judged, which is the case the skip was written for.

An unreadable transcript now says so on stderr instead of silently reading as
"the user did not ask".

The harness-parity test asserts the same reply is judged under both `Stop` and
`stop`. It passes before this change too: this hook compares no event name
anywhere, so event-name case cannot be what split its hits by harness. The
test pins that, rather than leaving the parity unstated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: Idb90d55fa90496fcab44e1f0255b303b5cf0b514
…n, not from wording

Both retraction detectors need a wrongness word. Two real corrections carry
none, because the claim was true and only its order was wrong:

  "Correcting one claim and arming the check I implied:"
  "I was right - but I said it a turn before I checked it"

Primary fix, no regex and no model: a ledger row going outstanding ->
discharged IS the event "a claim was asserted before its check ran, and the
check has now run". The Stop that discharges a row now reports it as a reflect
trigger. It reads state, so the reply's wording is not consulted at all.

Secondary fix: the judge dictionary's meaning now covers "right, but said
before it was checked", and both real strings are eval cases. Run against the
real judge, all six cases land where the dictionary says they should.

The regex scan keeps its hole on purpose, pinned by three tests so the next
author widens it knowingly rather than by accident.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I99190104b68d88f78aa9c8a706c9a6d23055aecc
…e only

The reminder was re-read to the user on every prompt for every outstanding
claim, including claims whose proof had already been pasted. There was no way
to turn it down: CATSTACK_HOOK_MODE_UNVERIFIED_TAG_LEDGER=off still printed
the full list at exit 0, because this hook is not on run_hook and nothing
reads the registry mode for it.

CATSTACK_UNVERIFIED_TAG_REMINDER is read in-hook through _flags/flags.py, the
one pattern proven to work here (wrong-check-reflect/detect.py:12-15, 191):

  stale  default, and what unset means: only claims already 3+ turns old
  all    the old behaviour
  off    no injection

Unset resolves to stale on purpose. A flag whose unset value is the old
behaviour changes nothing for the person who asked for less. reminder() at
:138 already computed that exact subset and threw it away at :145.

The gate is on the injection only. Emission and recording are untouched on
every setting, so `off` buys silence without emptying the ledger it is mined
from. A value nobody understands, and an .env candidate that cannot be read,
are both named on stderr instead of passing as "not set".

hooks.toml gains an enabled_by line naming the flag. Nothing reads that field
-- registry.py:15 is its only mention in the repo -- so it is documentation
beside the gate, not the gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I4f465761ee7f84489616ba1e83c54eb98ceefff0
FILE_LINE_RE matched a path and a line number and the paragraph loop skipped
the paragraph. Nothing checked that the file existed, that anyone had read it,
or at what ref. Pasting a made-up path beside the same claim turned the gate
off just as well:

  $ printf '%s' '{"last_assistant_message":"The problem was the old value at
    totally-made-up-file-that-does-not-exist.ts:99999.", ...}' \
      | python3 engine/hooks/diu-stop/claude_stop_check.py
    exit=0

A citation now counts when it names the ref it was read at
(`path:line @ origin/main`), or when the session transcript shows a tool call
that named that path. That is the rule corpus/CLAUDE.learned.md already ships
in prose and nothing enforced.

Third outcome, pinned: when the transcript cannot be read, the citation is
unchecked rather than clear. It does not buy silence, the reason goes to
stderr, and the block names the path it could not check.

tests/test_hooks.py:412 asserted the old behaviour. Its own PR, #479, listed
the file:line exemption under Non-goals -- it carried the exemption over so
fence normalization would not regress it, and did not decide that an unread
path should silence anything. The test keeps that intent and now backs its
citation with a read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I9ece412c527ac3272bf86d31d8089db65fd685f5
find_marker_problems ran markers.TAG_RE and the legacy-marker check over the
raw message. find_unverified_claims, ten lines below it, already stripped
fenced blocks and inline code and never shared that with the marker check.

The result: explaining the tag tripped the gate, and so did quoting
cat-mode's own rule, and so did relaying this gate's refusal word for word --
which cat-mode/SKILL.md:55 asks for.

  $ printf '%s' '{"last_assistant_message":"The escape hatch is
    `{{CAT-UNVERIFIED}}` and it has to name a blocker."}' \
      | python3 engine/hooks/diu-stop/claude_stop_check.py
  exit=2
  A `{{CAT-UNVERIFIED}}` tag here names no blocker. ...

After: exit=0, no output. A tag in running prose with no blocker named still
exits 2.

An unterminated fence drops everything after the marker it left behind,
rather than letting the unclosed block read as prose.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I8d19b3831299b0b18d5b31288017b1b0fff9cd21
The sentence that tells the reader to use the tag showed it as a bare token
naming no blocker. That is the shape diu-stop rejects, so quoting cat-mode's
own rule tripped the gate enforcing it. It now shows the same template
engine/CLAUDE.core.md already uses.

tests/test_cat_mode.py gains the structural catch: no tag anywhere in the
skill may name no blocker, and the skill must still show the tag at all, so
the fix cannot be "delete the example".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: If47021e83bdf3d8771ccbdb879001f6effc65a3c
… the hook

user_already_asked_reflect matched any user-shaped line mentioning reflect.
The harness files its own injections as `type: "user"` rows, so two of them
matched:

  - `diu-stop`'s Stop-hook feedback, whose text tells the agent to reflect
  - the reflect skill's own injected body, so running /reflect disarmed the
    hook that asks for it

Read-confirmed on a real transcript: both carry `isMeta: true` on a
`type: "user"` row, which is the field
engine/skills/reflect/scripts/token_audit.py:313 already keys off. This reads
the record rather than a prose prefix, which could only ever catch wordings
somebody had already seen. `agentId` and `isSidechain` are excluded the same
way.

Chesterton's fence checked. `git log -S "user_already_asked_reflect"
origin/main` returns three commits: 5b3b2fb, 30f62e5 "Add
wrong-check-reflect hook for first-person check admissions (#57)", and
e9a9750 (#462). 30f62e5 is the one that introduced the function, and its
message asks for the follow-up "once per transcript" -- do not nag twice.
Matching any mention anywhere in the file is far wider than that. This
narrows it back to the original intent; the turn scoping and the
once-per-reply key are the commit before this one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Change-Id: I31224f6a0f4626d5b7ca6590776a30f314ac50d8
…t facts and a question

The parent's side of the skill offered two slots: facts and the question. A
decision the user already made is neither, so it gets filed under the
question -- and the "don't prime the answer" rule then rewards re-opening it,
because that reads as rigour.

Third slot added: decisions travel as constraint sentences marked settled.
"Don't prime" is scoped to findings, never decisions; leaving a decision out
is dropping a constraint, not staying neutral.

The skill says plainly that repeating the user's words is not enough: a
delegation can contain the requirement verbatim, pass a containment check,
and still lose it by appending one open question beside it. No mechanical
catch is claimed, and the fires fixture records why -- the containment check
returns PASS on the exact call that drifted.

Prior art and its caveats are in the skill: Gotel & Finkelstein on
pre-requirements-specification traceability, marked single-source because
Crossref's issued field is null and its author rendering differs, with no
publisher page read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I370384357ee26e01c4b4a2d8ad1d65808c496ab3
…y asked

The suppression matched any prose containing the word, so a sentence about
the hook switched the hook off. `"Claim I made was wrong" is a trigger for
/reflect` describes when the detector fires; it asks for nothing, and
nothing is in flight to avoid duplicating.

Match the harness's own record of an invocation -- the
`<command-name>/reflect</command-name>` envelope -- instead. That envelope
starts with `<command-`, which the meta filter drops, so the one
unambiguous signal was the one being discarded while vague prose was kept.
The window now includes meta rows for this check.

Also stop returning False when the transcript holds no assistant row. A
Stop payload can carry the reply before its row lands, and a session's
first turn has none at all; in both cases every row belongs to this turn,
and bailing out meant a real request went unhonoured.

Erring toward asking is the safe direction here: this detector recorded
1,682 Claude Code invocations and spoke 0 times.

tests/scenarios/self-correction.json: the scenario that pinned this said
the user "already asked for /reflect" while its user field only mentioned
the word. Its intent and its data disagreed. The request case now uses a
real invocation; the mention case is its own scenario and must still fire.

Ran 27 tests in 16.245s OK (wrong-check-reflect)
Ran 6 tests in 0.045s OK (scenario suite)
check_hook_test_coverage: OK (1 hook(s) checked)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I1464277797323c004f488a6de5676d7430c10d21
@EdbertChan

EdbertChan commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

This pull request is part of a Mergify stack:

# Pull Request Link
1 unverified-tag-ledger: read the turn's tools from the transcript, not a key Claude Code never sends #780
2 wrong-check-reflect: one shot per reply, and scope the reflect scan to the turn #781
3 reflect: catch an evidence-order correction from the ledger transition, not from wording #782
4 unverified-tag-ledger: gate the next-prompt reminder, default to stale only #783
5 diu-stop: a file citation buys silence only when it is backed #784
6 diu-stop: the marker gate reads prose, so a mention is not a use #785
7 cat-mode: show the escape-hatch tag in the form the gate accepts #786
8 wrong-check-reflect: only the person's own reflect request suppresses the hook #787
9 principle-subagent-inherits-scope: a brief carries decisions, not just facts and a question #788
10 wrong-check-reflect: only a real /reflect invocation counts as already asked #789 👈
11 llm-judge: claude is the only default runner #793

@EdbertChan
EdbertChan force-pushed the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/principle-subagent-inherits-scope-brief-carries--37038435 branch from 83d83f9 to 843cbad Compare September 22, 2026 18:38
@EdbertChan EdbertChan closed this Sep 22, 2026
@EdbertChan
EdbertChan deleted the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/wrong-check-reflect-only-real-reflect-invocation--14642777 branch September 22, 2026 18:38
mergify Bot pushed a commit that referenced this pull request Sep 23, 2026
* pr-schema-gate: find the validator in catstack, report UNCHECKED when a repo has none

Scope now comes from the .git boundary, not scripts/create-pr.mjs. The
validator is scripts/validate-pr-body.mjs, then
engine/skills/draft-pr/scripts/validate-pr-body.mjs. A PR-publishing
command in a git repo with neither reports UNCHECKED instead of staying
silent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790182316514-5/implement-hook-finds-catstack-checker — Review claim: pr-schema-gate checks PR descriptions in a repo whose validator lives at engine/skills/draft-pr/scripts/validate-pr-body.mjs, and reports UNCHECKED instead of staying silent when a PR-publishing command runs in a repo where no validator can be found.
Review lane: behavior
Safety invariant: The change can only add a failure or an UNCHECKED notice; a PR description that passes engine/skills/draft-pr/scripts/validate-pr-body.mjs today is never newly blocked.
Effectiveness measurement: Hook tests replay the three commands that stayed silent in the incident: `mergify stack push` in a catstack-shaped repo (fires or UNCHECKED, never silent), `gh pr create --body-file <body failing the validator>` in a catstack-shaped repo (fires with the validator's errors), and a passing body in the same repo (clean). The existing Invoker-shaped cases keep their current results.
Slice rationale: One hook, one claim: where the hook looks for the validator and what it says when it finds none.
Architectural effect: pr-schema-gate stops depending on scripts/create-pr.mjs to decide whether a repo is in scope.
Goal: Make the PR-description guard fire in catstack.
Motivation: A reflect pass found PR descriptions on catstack PRs #780-#789, #793 and #795 failing the required PR Body check after publication, and the user had to ask for a manual fix of every PR. In catstack the guard stayed silent: fed `mergify stack push` and `gh pr create --body-file <failing body>` it exited 0 with no output, while the same stack push in the Invoker checkout fired. Read at origin/main: engine/hooks/pr-schema-gate/detect.py scopes itself by walking up for scripts/create-pr.mjs (repo_root_with_create_pr_tool) and expects the validator at scripts/validate-pr-body.mjs (VALIDATOR_RELATIVE_PATH); catstack has neither.
Alternative considerations: Adding a scripts/create-pr.mjs shim to catstack was set aside because it makes scope depend on an unrelated file again. Copying the validator to scripts/ was set aside because there are already too many copies.
Implementation details: In engine/hooks/pr-schema-gate/detect.py, find the repo root from the .git boundary, then look for the validator in a short ordered list: scripts/validate-pr-body.mjs, then engine/skills/draft-pr/scripts/validate-pr-body.mjs. A repo with either is in scope. When a PR-publishing command (gh pr create/edit with a body, gh api PATCH on pulls with a body, mergify stack push) runs in a git repo where no validator is found, return the existing unchecked outcome with a one-line reason instead of None. Keep the existing hit / clean / unchecked outcomes and messages for Invoker-shaped repos unchanged.
Non-goals: No change to the validator, preflight, install.sh, or any other hook. Do not edit engine/skills/draft-pr/scripts/validate-pr-body.mjs, scripts/pr/validate-pr-body-local.mjs, or any other file open PR #742 changes; call the validator as it is. If a change would overlap an open PR, stop and report instead.
Layer: domain
Feature state: active
Files:
- engine/hooks/pr-schema-gate/detect.py
- engine/hooks/pr-schema-gate/tests/test_hooks.py
- engine/hooks/pr-schema-gate/README.md
Change types:
- engine/hooks/pr-schema-gate/detect.py: modify
- engine/hooks/pr-schema-gate/tests/test_hooks.py: modify
- engine/hooks/pr-schema-gate/README.md: modify
Acceptance criteria:
- `python3 -m unittest discover -s engine/hooks/pr-schema-gate/tests -v` exits 0.
- `python3 scripts/check_hook_test_coverage.py` exits 0.
- `python3 scripts/check_no_silent_hook_except.py` exits 0.
- `python3 scripts/check_no_new_comments.py` exits 0.

Exit code: 0

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-3 — Review claim: `python3 scripts/check_no_silent_hook_except.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_silent_hook_except.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-1 — Review claim: `python3 -m unittest discover -s engine/hooks/pr-schema-gate/tests -v` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 -m unittest discover -s engine/hooks/pr-schema-gate/tests -v`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-2 — Review claim: `python3 scripts/check_hook_test_coverage.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_hook_test_coverage.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-2 — Review claim: `python3 scripts/check_hook_test_coverage.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_hook_test_coverage.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-2 — Review claim: `python3 scripts/check_hook_test_coverage.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_hook_test_coverage.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Solution:
  Review claim: `python3 scripts/check_hook_test_coverage.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_hook_test_coverage.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Solution:
  Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-2 — Review claim: `python3 scripts/check_hook_test_coverage.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_hook_test_coverage.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182316514-5/verify-hook-finds-catstack-checker-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182316514-5/scrub-handoff-artifacts — Review claim: No ephemeral inter-task handoff files remain in the worktree before the merge gate.
Review lane: cleanup
Safety invariant: The scrub script only checks for known handoff artifact names and never touches source, tests, or other repository files.
Effectiveness measurement: The script exits non-zero if any handoff artifact remains.
Slice rationale: Required terminal scrub for every implementation workflow.
Architectural effect: None; hygiene only.
Goal: Leave the branch free of handoff artifacts.
Motivation: Handoff files must not reach the PR.
Alternative considerations: Manual cleanup was set aside as non-deterministic.
Implementation details: Run scripts/scrub-handoff-artifacts.sh.
Non-goals: No product edits.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* wrong-check-reflect: one shot per reply, and scope the reflect scan to the turn (#796)

* wrong-check-reflect: one shot per reply, and scope the reflect scan to the turn

Two independent lockouts kept this hook from speaking on Claude Code.

One: `already_prompted` was keyed on the transcript path, so the Stop of the
reply BEFORE a correction spent the session's single shot. The correction
itself then hit an already-prompted key. The key is now the transcript path
plus a hash of the reply text, so each distinct reply gets its own chance and
the same reply is still judged only once.

Two: `user_already_asked_reflect` scanned the whole transcript. One `/reflect`
typed at the start of a session switched the detector off for every later
reply. The scan is now scoped to the user messages of the turn that produced
the reply being judged, which is the case the skip was written for.

An unreadable transcript now says so on stderr instead of silently reading as
"the user did not ask".

The harness-parity test asserts the same reply is judged under both `Stop` and
`stop`. It passes before this change too: this hook compares no event name
anywhere, so event-name case cannot be what split its hits by harness. The
test pins that, rather than leaving the parity unstated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: Idb90d55fa90496fcab44e1f0255b303b5cf0b514

Three: the suppression matched any prose containing the word. A sentence
ABOUT the hook switched the hook off. `"Claim I made was wrong" is a trigger
for /reflect` describes when the detector fires; it asks for nothing, and
nothing is in flight to avoid duplicating. It now matches the harness's own
record of an invocation, the `<command-name>/reflect</command-name>` envelope.
That envelope starts with `<command-`, which the meta filter drops, so the one
unambiguous signal was being discarded while vague prose was kept; the window
now includes meta rows for this check.

Four: a missing assistant row no longer reads as "no request found". A Stop
payload can carry the reply before its row lands, and a session's first turn
has none at all; in both cases every row belongs to this turn, and bailing out
meant a real request went unhonoured.

tests/scenarios/self-correction.json: the scenario pinning this said the user
"already asked for /reflect" while its user field only mentioned the word --
its stated intent and its data disagreed. The request case now uses a real
invocation; the mention case is its own scenario and must still fire. Both
were folded into this commit rather than a later one so no commit in the stack
leaves the scenario suite red.

Erring toward asking is the safe direction: this detector recorded 1,682
Claude Code invocations and spoke 0 times.

  Ran 27 tests in 18.811s  OK   (wrong-check-reflect)
  Ran  6 tests in  0.098s  OK   (scenario suite)

Change-Id: I60c5103d0541efd34f49e2401d50f4c5e312ddf0

* wrong-check-reflect: anchor the reflect scan on the person's message, not the last assistant row

The same-turn /reflect skip read the wrong turn. It took the last
text-bearing assistant row as the reply under judgment and scanned only
back to the assistant row before it.

A Stop payload carries the reply before its transcript row is written, so
that "last assistant row" was usually the previous turn's reply. The
window then sat one turn behind: a /reflect from the turn before
suppressed this reply (the session lockout, one turn wide), and the
/reflect the person had just typed sat past the window and was ignored, so
the hook nagged for a reflect already in flight. A turn also writes more
than one assistant row -- mid-turn narration, a subagent's sidechain rows
-- and each one pushed the window's start past the message that opened
the turn.

The window now runs from the person's own last message to the end of the
file, so it does not depend on assistant rows at all. A typed slash
command counts as that message, so <command-name> leaves the harness
prefix list; <local-command stdout, which is harness echo, joins it.

Four tests, each failing before this change: reflect this turn before the
reply row lands, last turn's reflect while the reply is still being
written, mid-turn narration, and a subagent's rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790174651006-232/repair — Resolve bot review thread on PR #796

Exit code: 0

* wrong-check-reflect: drop the banned code comment, keep the why in the README

scripts/ci/check_no_new_comments.py fails the test job on the three comment
lines this branch added above META_USER_PREFIXES. The point they made -- a
typed slash command is the person, not harness text -- now sits in the hook's
README next to the harness-rows paragraph it belongs with.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* draft-pr: pin that a changed file's stem is a code name even when it reads like English

The PR body validator failed this branch on "self-correction" in the Summary,
because tests/scenarios/self-correction.json is a changed file. The suite
covered changed folder names and plain hyphenated English separately, but not
the overlap, which is what a person hits: an everyday phrase that happens to
be a file's name. Both directions are now tested -- it fails, and the
reworded Summary passes with the same changed files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790175683747-249/repair — Repair PR #796 (failed check validate)

Exit code: 0

* wrong-check-reflect: a tool result is not the person opening a turn

Claude Code files every tool result as a `type: "user"` row, and unlike
its other injections it marks that row with neither `isMeta` nor
`isSidechain`. The turn window anchors on the newest real user row, so
the first tool call of a turn became the anchor: a `/reflect` typed at
the top of the turn fell outside the window and the hook nagged for a
reflect already asked for. A tool result quoting the word the other way
round could also suppress the nudge.

`_is_tool_result_line` reads the record the harness does write -- a
`tool_result` content block, or the `toolUseResult` field beside the
message -- the same way `engine/hooks/agent-relay-attribution/detect.py`
does, and `_is_meta_line` now files those rows under the harness rather
than the person.

Two tests: a tool result mid-turn no longer hides the turn's own
`/reflect`, and `/reflect` printed inside a grep result does not count
as a request. Both fail without the detect.py change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790177409884-260/repair — Resolve bot review thread on PR #796

Exit code: 0

---------

Co-authored-by: CI Bot <ci@invoker.dev>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* prove-it-ship-gate: the user's machine is a live surface, a PR link is not proof (#779)

* Make the done-gate the real path, not the layers under it

cat-mode's two e2e bullets were deleted in #68 as a duplicate of the
global Named constraints text. That text is narrower: it fires on
"UI/layout work" or on a test the user asked for. Work that was neither
-- an agent running a CI Playwright shard on the user's own Mac and
opening Electron windows on their desktop -- had no trigger left, and a
"Shipped" claim went out with no end-to-end run behind it.

Restore the gate in a form that names the surface rather than the kind
of work: any surface the repo's fixtures cannot stand in for, including
the user's own machine, session, or screen. Add the clause that failed
hardest in the incident -- "I chose not to run it" is not a blocker --
and make "admit what was not exercised" an enumeration against the
done-gate instead of a recollection.

The full text and the end-to-end argument it rests on (Saltzer, Reed &
Clark, ACM TOCS 2(4) 1984) live in the reference; SKILL.md keeps every
trigger condition, because a pointer narrower than the text it replaces
is exactly what failed here. Four tests pin the surfaces, the
not-a-blocker clause, the enumeration, and the citation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: Ie690ba84bc28698b6e528c9598479e1d541633ec

* Count the user's own machine as a live surface, and a PR link as no proof

Replayed against this gate unchanged, both of the incident's own ship
messages return SILENT. Two defects stack, and either one alone keeps
them silent.

The live-noun list held only external services, so "the popups are
stopped on your machine" matched nothing and the gate returned before it
ever looked at evidence. Add the surfaces an agent can disturb without
leaving the desk: the user's machine, Mac, laptop, desktop, screen or
session, plus end-to-end, e2e, Playwright, Electron and popup.

The evidence scan then accepted any URL, so the two links to this
change's own pull requests discharged the gate -- while the gate's own
block text already said the PR number of this change does not prove the
live path ran. Blank a pull-request link's span before the scan, so the
two agree. Only that span: a sha, an exit code, or a fenced block beside
the link still counts, and an Actions-run URL is still a receipt.

Both incident messages are fixtures, verbatim. The negatives keep the
near neighbours silent: a mention with no claim, a link beside a real
sha, an Actions-run URL, and the follow-up message that pasted the
suite's own output.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Change-Id: I3a816e04b2241f38ba7069f61176807bbf50b1fc

* Stop the end-to-end idiom from counting as a live surface

The gate needs two parts within 240 characters: a ship claim and a live
noun. Listing bare `end-to-end` and `e2e` as live nouns broke that, because
`working end-to-end` and `confirmed end-to-end` are already claim phrases --
so the claim always sat on top of a live noun and the two parts collapsed
into one. "Done. The parser now works end-to-end" was blocked for showing no
live evidence it never needed.

Drop the bare idiom. It says how a check ran, not where. The surface in the
incident was the desktop the windows opened on, and `your machine`, `your
Mac`, playwright, electron and popup already name it -- every incident
fixture still fires on one of those, and `end to end on your laptop` still
fires on the laptop.

Two tests pin it: the idiom-only messages stay silent, and no claim phrase
may consist entirely of live nouns.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790174635085-231/repair — Resolve bot review thread on PR #779

Exit code: 0

* Keep the end-to-end rationale out of code comments

The no-comments CI twin fails any diff that adds `#` comment lines to a
code file. The two blocks explaining why bare "end-to-end" and "e2e" are
not live nouns went in as plain comments, so the test job's last step
failed with "10 new comment line(s)".

Move the detect.py rationale into the module docstring, which the
detector skips by design (engine/hooks/no-comments/detect.py), and drop
the test-file comment: test_silent_when_only_the_idiom_names_the_surface
already carries the same reasoning in its own docstring. No behaviour
change -- LIVE_NOUN_RE and every fixture are untouched.

The gate is its own repro:
  python3 scripts/ci/check_no_new_comments.py --base origin/reflect/done-gate-real-path
before: fail  10 new comment line(s) ... (exit 1)
after:  ok      no new comments        (exit 0)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790177084312-254/repair — Repair PR #779 (failed check test)

Exit code: 0

* invoker: wf-1790182337423-299/repair — Repair PR #779 (failed check validate)

Exit code: 0

* prove-it-ship-gate: let a qualifier sit between the possessive and the surface noun

The local-surface pattern required the noun to sit immediately after
`your`, `their`, or `the user's`, so `your own machine` and `the
user's own screen` never matched. That is the exact wording the block
message, the README, and the skill all use, which meant the gate stayed
silent on a done-claim phrased the way the gate itself recommends: same
sentence, one extra word, opposite verdict.

The noun scan now accepts one optional qualifier (own, real, actual,
personal, local) after the possessive. Tests cover each qualifier and
assert the block message's own phrasing trips the scan, so the gate can
never again describe a surface it cannot detect.

* invoker: wf-1790184149094-309/repair — Resolve bot review thread on PR #779

Exit code: 0

---------

Co-authored-by: Edbert Chan <chanedbert@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Bot <ci@invoker.dev>

* llm-judge: leave out a runner that cannot answer (#799)

* llm-judge: leave a runner that cannot answer out of the table for 6 hours

A runner that is not installed or exits non-zero is skipped by later
asks until the window ends. Timeouts and non-JSON replies do not count.
If every runner is skipped the whole table is tried, and an answer
clears the marker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Change-Id: I9db64c4f2d23e801a3f0b7ca729f39fca7e6adaf

* draft-pr: pin the Review Claim half of the code-name check

CODE_NAME_SECTIONS covers Summary and Review Claim, but every code-name
test edited the Summary, so nothing exercised Review Claim. PR #799's
validate job failed there: "A judge runner that cannot answer ..." names
judge.py, a changed file, even though it reads as plain English.

Add both directions on a Review Claim -- the stem fails and names itself,
and the reworded claim passes -- so a later change cannot quietly exempt
file stems that happen to be English words.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790179585073-282/repair — Repair PR #799 (failed check validate)

Exit code: 0

* draft-pr tests: name the fixture instead of commenting it

The no-comments CI gate rejected the `# A changed-file list whose stem
("judge") is also an ordinary English word.` line added above JUDGE_FILES.
Comments are banned in code in this repo, so the note moves into the
constant's own name: FILES_WHOSE_STEM_IS_AN_ORDINARY_ENGLISH_WORD. The
two tests that use it read the same, and the test that needed the note
already states the reasoning in its docstring, which the gate allows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790182343006-301/repair — Repair PR #799 (failed check test)

Exit code: 0

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Bot <ci@example.com>
Co-authored-by: CI Bot <ci@invoker.dev>

* pr-schema-gate: a skipped sub-check is not a vacuous pass

A PR body the repo's validator accepted was reported to the agent as
"could not check", and an owed stack follow-up stayed armed.

The hook treats exit 0 alongside an UNCHECKED/SKIPPED/not-installed line
as a vacuous pass, so the validator gets no credit for a run that never
judged the body. The catstack validator prints exactly such a line for a
sub-check it skipped -- "Summary reading grade unchecked: Summary has N
words; under 30 the score is too noisy to trust" -- while still accepting
the body and exiting 0. Every accepted body with a short Summary was
therefore reported unchecked.

check_body_file now lifts the vacuous reading when the run states its own
verdict on the body (VALIDATOR_PASS_VERDICT_RE, "PR body validation
passed"). Only a pass verdict counts, so a "failed" banner beside exit 0
stays unchecked. The scan also reads the whole output rather than the
first VALIDATOR_OUTPUT_MAX_LINES: truncation shortens what the agent is
shown, never what is judged.

Two tests in tests/test_advisory.py drive a stub validator with the
catstack shape (skip note on stderr, verdict on stdout, exit 0): one
asserts the write is silent, one asserts it clears the pending follow-up.
Both fail before this change. test_validator_exit_zero_with_an_unchecked_
line_is_not_clean still passes, so the guard against a validator that
really did not run is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790193889047-345/repair — Resolve bot review thread on PR #840

Exit code: 0

* draft-pr tests: pin the changed-folder-name half of the code-name check

PR #840's own body gate failed on `Review Claim: "draft-pr" (changed
folder name)`, and no test covered that kind. Every existing code-name
test uses a changed *file* stem, so deleting the folder loop in
changedFileNames kept the whole suite green.

Adds the real failing Review Claim as a repro plus its reworded,
passing twin. Removing `for (const folder of parts) add(folder,
'changed folder name')` now fails the new test with `0 != 1`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* invoker: wf-1790200429814-416/repair — Repair PR #840 (failed check validate)

Exit code: 0

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Bot <ci@invoker.dev>
Co-authored-by: Edbert Chan <chanedbert@gmail.com>
Co-authored-by: CI Bot <ci@example.com>
mergify Bot pushed a commit that referenced this pull request Sep 24, 2026
…841)

* make-pr: preflight runs the PR-body validator on --body-file

description_check only read the description's prose for claims about the
repo's past, so preflight printed "ok preflight passed" for a body the
required PR Body check then rejected after publication (#780-#789, #793,
#795). describe() now also runs
engine/skills/draft-pr/scripts/validate-pr-body.mjs on the same file and
prints its errors under the gate line.

Exit 1 from the validator fails preflight. No node, a missing validator,
a timeout, or any other exit code prints "description unchecked: ..." and
fails too -- a check that could not run is not a pass.

The #795 description lands as a fixture: it exits 1 with Test Plan and
Revert Plan outside <details> and code names in the Summary, and the new
tests show preflight failing on it and still passing a valid body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* invoker: wf-1790182452439-6/implement-preflight-runs-validator — Review claim: make-pr preflight runs the PR-description validator on --body-file and fails when it fails, so it no longer prints 'preflight passed' for a description the required PR Body check rejects.
Review lane: policy
Safety invariant: The change can only add a failure or an UNCHECKED notice; a PR description that passes engine/skills/draft-pr/scripts/validate-pr-body.mjs today is never newly blocked.
Effectiveness measurement: A preflight test feeds the original #795 description (Test Plan and Revert Plan not inside <details>, code names in Summary), which the validator rejects with exit 1, and asserts preflight exits non-zero and prints the validator's errors; a passing description still gives 'ok preflight passed'; a validator that cannot run gives an unchecked failure, not a pass.
Slice rationale: One script, one claim: preflight's description step also runs the schema validator.
Architectural effect: preflight's describe() step covers the same rules as the required PR Body check.
Goal: Stop preflight from approving a PR description the required check will reject.
Motivation: A reflect pass found PR descriptions on catstack PRs #780-#789, #793 and #795 failing the required PR Body check after publication, and the user had to ask for a manual fix of every PR. In catstack the guard stayed silent: fed `mergify stack push` and `gh pr create --body-file <failing body>` it exited 0 with no output, while the same stack push in the Invoker checkout fired. Read at origin/main: engine/skills/make-pr/scripts/preflight.py describe() only runs description_check, then main() prints 'ok preflight passed'; the session that opened #795 got that line and then ran gh pr create.
Alternative considerations: Leaving the validator to CI was set aside because the failure then shows up only after publication. Reimplementing the rules in Python was set aside because it would drift from the validator.
Implementation details: In engine/skills/make-pr/scripts/preflight.py describe(), after description_check, run `node engine/skills/draft-pr/scripts/validate-pr-body.mjs --body-file <file>` from the repo root with a timeout. Print its output lines indented like the other gates. Exit 1 from the validator fails preflight; a missing node binary, a timeout, or any other exit code prints 'description unchecked: <reason>' and fails preflight. Add the tests to engine/skills/make-pr/tests/test_preflight.py with the #795 body as a fixture file under engine/skills/make-pr/tests/fixtures/.
Non-goals: No change to the validator, the hook, or the make-pr SKILL.md rules. Do not edit engine/skills/draft-pr/scripts/validate-pr-body.mjs, scripts/pr/validate-pr-body-local.mjs, or any other file open PR #742 changes; call the validator as it is. If a change would overlap an open PR, stop and report instead.
Layer: domain
Feature state: active
Files:
- engine/skills/make-pr/scripts/preflight.py
- engine/skills/make-pr/tests/test_preflight.py
- engine/skills/make-pr/tests/fixtures/pr795-failing-body.md
Change types:
- engine/skills/make-pr/scripts/preflight.py: modify
- engine/skills/make-pr/tests/test_preflight.py: modify
- engine/skills/make-pr/tests/fixtures/pr795-failing-body.md: create
Acceptance criteria:
- `python3 -m unittest discover -s engine/skills/make-pr/tests -v` exits 0.
- `python3 scripts/check_skill_test_coverage.py --base origin/main --head HEAD` exits 0.
- `python3 scripts/check_skill_file_refs.py` exits 0.
- `python3 scripts/check_no_new_comments.py` exits 0.

Exit code: 0

* invoker: wf-1790182452439-6/verify-preflight-runs-validator-1 — Review claim: `python3 -m unittest discover -s engine/skills/make-pr/tests -v` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 -m unittest discover -s engine/skills/make-pr/tests -v`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1

* invoker: wf-1790182452439-6/verify-preflight-runs-validator-2 — Review claim: `python3 scripts/check_skill_test_coverage.py --base origin/main --head HEAD` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_skill_test_coverage.py --base origin/main --head HEAD`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182452439-6/verify-preflight-runs-validator-3 — Review claim: `python3 scripts/check_skill_file_refs.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_skill_file_refs.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1

* invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Solution:
  Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

* invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* invoker: wf-1790182452439-6/scrub-handoff-artifacts — Review claim: No ephemeral inter-task handoff files remain in the worktree before the merge gate.
Review lane: cleanup
Safety invariant: The scrub script only checks for known handoff artifact names and never touches source, tests, or other repository files.
Effectiveness measurement: The script exits non-zero if any handoff artifact remains.
Slice rationale: Required terminal scrub for every implementation workflow.
Architectural effect: None; hygiene only.
Goal: Leave the branch free of handoff artifacts.
Motivation: Handoff files must not reach the PR.
Alternative considerations: Manual cleanup was set aside as non-deterministic.
Implementation details: Run scripts/scrub-handoff-artifacts.sh.
Non-goals: No product edits.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0

* make-pr: preflight hands the validator the changed files

The required PR Body check passes --changed-files-file, which is how the
validator rejects changed file and folder names in the Summary and a
Review Unit that does not match the diff. Preflight passed only
--body-file, so those descriptions passed locally and failed after
publish. describe() now takes the changed paths preflight already
computed and writes them to a temp file for the validator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Change-Id: Id36a64869b7cfb2b76b912fd3cf81a88a39dd344

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Edbert Chan <edbertchantech@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant