[Hook architecture](50) Install every hook from the registry through the metrics runner and report installed hooks the registry does not name - #1204
Merged
Merged
Conversation
…ut the unverified-tag-ledger hook onto the shared hook code. Review claim: This hook reports findings to the shared hook code, which applies its registry mode and writes event rows. It keeps mode stop. Review lane: behavior Safety invariant: The hook gives the same stop, warn, or silent result on every case in its current test folder, except the mode change named in this claim, and its test folder keeps exiting 0. Effectiveness measurement: `python3 -m unittest discover -s engine/hooks/unverified-tag-ledger/tests` exits 0, and the new mode-override case fails before this change. Slice rationale: One hook per workflow, as the user asked, so each migration is reviewed on its own. Architectural effect: The unverified-tag-ledger entry scripts become thin calls into the shared runtime; its detection returns findings. Goal: Stops when unchecked claims pile up. Keep that behavior while its mode moves into the registry. Motivation: Mode and output shape live inside each hook today; the shared code makes a mode change a one-line registry edit. Alternative considerations: Migrating several hooks per workflow was set aside because the user asked for one hook per workflow. Implementation details: Turn this hook's detection into detect(event) returning Finding objects with stable rule ids, and make each harness entry script call run_hook from engine/hooks/_sdk/runtime.py. It keeps mode stop. Non-goals: No change to what the hook detects. No other hook changes. Layer: domain Feature state: active Files: engine/hooks/unverified-tag-ledger/claude_prompt_reminder.py, engine/hooks/unverified-tag-ledger/claude_stop_check.py, engine/hooks/unverified-tag-ledger/detect.py, engine/hooks/unverified-tag-ledger/install_claude_hook.py, engine/hooks/unverified-tag-ledger/tests/test_hooks_sdk_mode.py Change types: - engine/hooks/unverified-tag-ledger/claude_prompt_reminder.py: modify - engine/hooks/unverified-tag-ledger/claude_stop_check.py: modify - engine/hooks/unverified-tag-ledger/detect.py: modify - engine/hooks/unverified-tag-ledger/install_claude_hook.py: modify - engine/hooks/unverified-tag-ledger/tests/test_hooks_sdk_mode.py: create Acceptance criteria: - `python3 -m unittest discover -s engine/hooks/unverified-tag-ledger/tests` exits 0. - With CATSTACK_HOOK_MODE_UNVERIFIED_TAG_LEDGER set to warn, a case that stops today produces a warning instead, proving the registry mode drives the response. - Each finding writes one event row with the hook's rule_id. Exit code: 0 Invoker-Finalize-Id: 48e433fc-e1d3-4b76-94e5-ffd3eb3fb67c
…the deterministic proof for put the unverified-tag-ledger hook onto the shared hook code. Review claim: The proof exits 0 only when put the unverified-tag-ledger hook onto the shared hook code holds. Review lane: proof Safety invariant: Proof only; it changes no product behavior. Effectiveness measurement: The exit status of `python3 -m unittest discover -s engine/hooks/unverified-tag-ledger/tests` is the signal for this slice. Slice rationale: One proof unit for this workflow. Architectural effect: None; verification only. Goal: Prove put the unverified-tag-ledger hook onto the shared hook code with one deterministic run. Motivation: Each workflow carries its own proof so a reviewer can trust the slice alone. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Execute the proof as a terminal gate. Non-goals: No product edits here. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: d33259c1-d2cd-4a03-9752-f1898bf787ec
…only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: 88a9a276-8784-4563-bf65-7aeaa4174240
…-ad68f99c4-2ac9fea1 — Terminal read-only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active
…the verdict-flip-watch hook onto the shared hook code. Review claim: This hook reports findings to the shared hook code, which applies its registry mode and writes event rows. It keeps mode warn. Review lane: behavior Safety invariant: The hook gives the same stop, warn, or silent result on every case in its current test folder, except the mode change named in this claim, and its test folder keeps exiting 0. Effectiveness measurement: `python3 -m unittest discover -s engine/hooks/verdict-flip-watch/tests` exits 0, and the new mode-override case fails before this change. Slice rationale: One hook per workflow, as the user asked, so each migration is reviewed on its own. Architectural effect: The verdict-flip-watch entry scripts become thin calls into the shared runtime; its detection returns findings. Goal: Notes a check that passed earlier and failed later. Keep that behavior while its mode moves into the registry. Motivation: Mode and output shape live inside each hook today; the shared code makes a mode change a one-line registry edit. Alternative considerations: Migrating several hooks per workflow was set aside because the user asked for one hook per workflow. Implementation details: Turn this hook's detection into detect(event) returning Finding objects with stable rule ids, and make each harness entry script call run_hook from engine/hooks/_sdk/runtime.py. It keeps mode warn. Non-goals: No change to what the hook detects. No other hook changes. Layer: domain Feature state: active Files: engine/hooks/verdict-flip-watch/claude_stop_check.py, engine/hooks/verdict-flip-watch/detect.py, engine/hooks/verdict-flip-watch/install_claude_hook.py, engine/hooks/verdict-flip-watch/tests/test_hooks_sdk_mode.py Change types: - engine/hooks/verdict-flip-watch/claude_stop_check.py: modify - engine/hooks/verdict-flip-watch/detect.py: modify - engine/hooks/verdict-flip-watch/install_claude_hook.py: modify - engine/hooks/verdict-flip-watch/tests/test_hooks_sdk_mode.py: create Acceptance criteria: - `python3 -m unittest discover -s engine/hooks/verdict-flip-watch/tests` exits 0. - With CATSTACK_HOOK_MODE_VERDICT_FLIP_WATCH set to warn, a case that stops today produces a warning instead, proving the registry mode drives the response. - Each finding writes one event row with the hook's rule_id. Exit code: 0 Invoker-Finalize-Id: 5c05c0bc-f01b-490e-a0ef-a9903bd7be6f
… deterministic proof for put the verdict-flip-watch hook onto the shared hook code. Review claim: The proof exits 0 only when put the verdict-flip-watch hook onto the shared hook code holds. Review lane: proof Safety invariant: Proof only; it changes no product behavior. Effectiveness measurement: The exit status of `python3 -m unittest discover -s engine/hooks/verdict-flip-watch/tests` is the signal for this slice. Slice rationale: One proof unit for this workflow. Architectural effect: None; verification only. Goal: Prove put the verdict-flip-watch hook onto the shared hook code with one deterministic run. Motivation: Each workflow carries its own proof so a reviewer can trust the slice alone. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Execute the proof as a terminal gate. Non-goals: No product edits here. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: 426060cc-ee4e-4853-a0ed-01c4ee43290c
…only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: 0015d6b6-90c4-4314-8ea7-c5eee86a6bdf
…-ae4f65fb4-1d718c84 — Terminal read-only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active
…he wait-needs-wakeup hook onto the shared hook code. Review claim: This hook reports findings to the shared hook code, which applies its registry mode and writes event rows. It keeps mode stop. Review lane: behavior Safety invariant: The hook gives the same stop, warn, or silent result on every case in its current test folder, except the mode change named in this claim, and its test folder keeps exiting 0. Effectiveness measurement: `python3 -m unittest discover -s engine/hooks/wait-needs-wakeup/tests` exits 0, and the new mode-override case fails before this change. Slice rationale: One hook per workflow, as the user asked, so each migration is reviewed on its own. Architectural effect: The wait-needs-wakeup entry scripts become thin calls into the shared runtime; its detection returns findings. Goal: Stops waiting language with no time or wake-up set. Keep that behavior while its mode moves into the registry. Motivation: Mode and output shape live inside each hook today; the shared code makes a mode change a one-line registry edit. Alternative considerations: Migrating several hooks per workflow was set aside because the user asked for one hook per workflow. Implementation details: Turn this hook's detection into detect(event) returning Finding objects with stable rule ids, and make each harness entry script call run_hook from engine/hooks/_sdk/runtime.py. It keeps mode stop. Non-goals: No change to what the hook detects. No other hook changes. Layer: domain Feature state: active Files: engine/hooks/wait-needs-wakeup/backtest.py, engine/hooks/wait-needs-wakeup/claude_pretooluse.py, engine/hooks/wait-needs-wakeup/claude_stop_check.py, engine/hooks/wait-needs-wakeup/detect.py, engine/hooks/wait-needs-wakeup/install_claude_hook.py, engine/hooks/wait-needs-wakeup/tests/test_hooks_sdk_mode.py Change types: - engine/hooks/wait-needs-wakeup/backtest.py: modify - engine/hooks/wait-needs-wakeup/claude_pretooluse.py: modify - engine/hooks/wait-needs-wakeup/claude_stop_check.py: modify - engine/hooks/wait-needs-wakeup/detect.py: modify - engine/hooks/wait-needs-wakeup/install_claude_hook.py: modify - engine/hooks/wait-needs-wakeup/tests/test_hooks_sdk_mode.py: create Acceptance criteria: - `python3 -m unittest discover -s engine/hooks/wait-needs-wakeup/tests` exits 0. - With CATSTACK_HOOK_MODE_WAIT_NEEDS_WAKEUP set to warn, a case that stops today produces a warning instead, proving the registry mode drives the response. - Each finding writes one event row with the hook's rule_id. Solution: Put the wait-needs-wakeup hook onto the shared hook code. Review claim: This hook reports findings to the shared hook code, which applies its registry mode and writes event rows. It keeps mode stop. Review lane: behavior Safety invariant: The hook gives the same stop, warn, or silent result on every case in its current test folder, except the mode change named in this claim, and its test folder keeps exiting 0. Effectiveness measurement: `python3 -m unittest discover -s engine/hooks/wait-needs-wakeup/tests` exits 0, and the new mode-override case fails before this change. Slice rationale: One hook per workflow, as the user asked, so each migration is reviewed on its own. Architectural effect: The wait-needs-wakeup entry scripts become thin calls into the shared runtime; its detection returns findings. Goal: Stops waiting language with no time or wake-up set. Keep that behavior while its mode moves into the registry. Motivation: Mode and output shape live inside each hook today; the shared code makes a mode change a one-line registry edit. Alternative considerations: Migrating several hooks per workflow was set aside because the user asked for one hook per workflow. Implementation details: Turn this hook's detection into detect(event) returning Finding objects with stable rule ids, and make each harness entry script call run_hook from engine/hooks/_sdk/runtime.py. It keeps mode stop. Non-goals: No change to what the hook detects. No other hook changes. Layer: domain Feature state: active Files: engine/hooks/wait-needs-wakeup/backtest.py, engine/hooks/wait-needs-wakeup/claude_pretooluse.py, engine/hooks/wait-needs-wakeup/claude_stop_check.py, engine/hooks/wait-needs-wakeup/detect.py, engine/hooks/wait-needs-wakeup/install_claude_hook.py, engine/hooks/wait-needs-wakeup/tests/test_hooks_sdk_mode.py Change types: - engine/hooks/wait-needs-wakeup/backtest.py: modify - engine/hooks/wait-needs-wakeup/claude_pretooluse.py: modify - engine/hooks/wait-needs-wakeup/claude_stop_check.py: modify - engine/hooks/wait-needs-wakeup/detect.py: modify - engine/hooks/wait-needs-wakeup/install_claude_hook.py: modify - engine/hooks/wait-needs-wakeup/tests/test_hooks_sdk_mode.py: create Acceptance criteria: - `python3 -m unittest discover -s engine/hooks/wait-needs-wakeup/tests` exits 0. - With CATSTACK_HOOK_MODE_WAIT_NEEDS_WAKEUP set to warn, a case that stops today produces a warning instead, proving the registry mode drives the response. - Each finding writes one event row with the hook's rule_id. Invoker-Finalize-Id: 5e2fd291-ec78-4aac-80cd-49d3d3bd5097
…deterministic proof for put the wait-needs-wakeup hook onto the shared hook code. Review claim: The proof exits 0 only when put the wait-needs-wakeup hook onto the shared hook code holds. Review lane: proof Safety invariant: Proof only; it changes no product behavior. Effectiveness measurement: The exit status of `python3 -m unittest discover -s engine/hooks/wait-needs-wakeup/tests` is the signal for this slice. Slice rationale: One proof unit for this workflow. Architectural effect: None; verification only. Goal: Prove put the wait-needs-wakeup hook onto the shared hook code with one deterministic run. Motivation: Each workflow carries its own proof so a reviewer can trust the slice alone. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Execute the proof as a terminal gate. Non-goals: No product edits here. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: c17b931c-37ec-4fd8-bc25-fdd6dcdbc2d2
…only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: b4107492-8edb-4b2e-a3d0-4fb2b17e0a87
…-ae10bb318-7be25f7d — Terminal read-only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active
… the wrong-check-reflect hook onto the shared hook code. Review claim: This hook reports findings to the shared hook code, which applies its registry mode and writes event rows. It keeps mode warn. Review lane: behavior Safety invariant: The hook gives the same stop, warn, or silent result on every case in its current test folder, except the mode change named in this claim, and its test folder keeps exiting 0. Effectiveness measurement: `python3 -m unittest discover -s engine/hooks/wrong-check-reflect/tests` exits 0, and the new mode-override case fails before this change. Slice rationale: One hook per workflow, as the user asked, so each migration is reviewed on its own. Architectural effect: The wrong-check-reflect entry scripts become thin calls into the shared runtime; its detection returns findings. Goal: Suggests reflect after a claim is taken back. Keep that behavior while its mode moves into the registry. Motivation: Mode and output shape live inside each hook today; the shared code makes a mode change a one-line registry edit. Alternative considerations: Migrating several hooks per workflow was set aside because the user asked for one hook per workflow. Implementation details: Turn this hook's detection into detect(event) returning Finding objects with stable rule ids, and make each harness entry script call run_hook from engine/hooks/_sdk/runtime.py. It keeps mode warn. Non-goals: No change to what the hook detects. No other hook changes. Layer: domain Feature state: active Files: engine/hooks/wrong-check-reflect/claude_stop_check.py, engine/hooks/wrong-check-reflect/codex_notify.py, engine/hooks/wrong-check-reflect/cursor_session.py, engine/hooks/wrong-check-reflect/detect.py, engine/hooks/wrong-check-reflect/eval_dictionary.py, engine/hooks/wrong-check-reflect/install_claude_hook.py, engine/hooks/wrong-check-reflect/install_codex_notify.py, engine/hooks/wrong-check-reflect/install_cursor_hook.py, engine/hooks/wrong-check-reflect/tests/test_hooks_sdk_mode.py Change types: - engine/hooks/wrong-check-reflect/claude_stop_check.py: modify - engine/hooks/wrong-check-reflect/codex_notify.py: modify - engine/hooks/wrong-check-reflect/cursor_session.py: modify - engine/hooks/wrong-check-reflect/detect.py: modify - engine/hooks/wrong-check-reflect/eval_dictionary.py: modify - engine/hooks/wrong-check-reflect/install_claude_hook.py: modify - engine/hooks/wrong-check-reflect/install_codex_notify.py: modify - engine/hooks/wrong-check-reflect/install_cursor_hook.py: modify - engine/hooks/wrong-check-reflect/tests/test_hooks_sdk_mode.py: create Acceptance criteria: - `python3 -m unittest discover -s engine/hooks/wrong-check-reflect/tests` exits 0. - With CATSTACK_HOOK_MODE_WRONG_CHECK_REFLECT set to warn, a case that stops today produces a warning instead, proving the registry mode drives the response. - Each finding writes one event row with the hook's rule_id. Exit code: 0 Invoker-Finalize-Id: 236d39ea-e79a-49eb-ad7c-d4c42784c1e1
…e deterministic proof for put the wrong-check-reflect hook onto the shared hook code. Review claim: The proof exits 0 only when put the wrong-check-reflect hook onto the shared hook code holds. Review lane: proof Safety invariant: Proof only; it changes no product behavior. Effectiveness measurement: The exit status of `python3 -m unittest discover -s engine/hooks/wrong-check-reflect/tests` is the signal for this slice. Slice rationale: One proof unit for this workflow. Architectural effect: None; verification only. Goal: Prove put the wrong-check-reflect hook onto the shared hook code with one deterministic run. Motivation: Each workflow carries its own proof so a reviewer can trust the slice alone. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Execute the proof as a terminal gate. Non-goals: No product edits here. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: 72a70a0a-cd89-466e-8737-3f2002218b29
…only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: 4d7d5267-ffe5-4df7-b21c-46d480f766db
…-ac3aef88e-b1f9d695 — Terminal read-only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active
…l every hook from the registry through the metrics runner and report installed hooks the registry does not name. Review claim: Install writes Claude, Cursor, and Codex hook entries from the registry through the metrics runner, and the install check reports installed hooks the registry does not name. Review lane: policy Safety invariant: The installed hook set equals the registry's hooks whose mode is not off, each through the metrics runner. Effectiveness measurement: `python3 -m unittest tests.test_install` exits 0, with a case where an installed hook absent from the registry is reported. Slice rationale: One unit: install driven by the registry, after every hook has migrated. Architectural effect: One installer reads the registry and writes all three harness configs; the per-hook installer scripts stop being called. Goal: Stop installs drifting from the hook set in the repo. Motivation: This Mac ran two hooks that were deleted upstream, and hook registration is spread across dozens of installer scripts. Alternative considerations: Keeping per-hook installers and only adding the drift check was set aside because the registration would still be spread out. Implementation details: Create engine/hooks/_sdk/install_from_registry.py, call it from install.sh in place of the per-hook installer calls, and extend scripts/check_install_effective.py to report installed catstack hooks the registry does not name. Non-goals: No hook behavior changes. Layer: domain Feature state: active Files: engine/hooks/_sdk/install_from_registry.py, install.sh, scripts/check_install_effective.py, tests/test_install.py Change types: - engine/hooks/_sdk/install_from_registry.py: create - install.sh: modify - scripts/check_install_effective.py: modify - tests/test_install.py: modify Acceptance criteria: - `python3 -m unittest tests.test_install` exits 0. - Installing into a temporary home writes entries only for registry hooks whose mode is not off, each through the metrics runner. - An installed catstack hook the registry does not name is reported by the install check. Solution: Install every hook from the registry through the metrics runner and report installed hooks the registry does not name. Review claim: Install writes Claude, Cursor, and Codex hook entries from the registry through the metrics runner, and the install check reports installed hooks the registry does not name. Review lane: policy Safety invariant: The installed hook set equals the registry's hooks whose mode is not off, each through the metrics runner. Effectiveness measurement: `python3 -m unittest tests.test_install` exits 0, with a case where an installed hook absent from the registry is reported. Slice rationale: One unit: install driven by the registry, after every hook has migrated. Architectural effect: One installer reads the registry and writes all three harness configs; the per-hook installer scripts stop being called. Goal: Stop installs drifting from the hook set in the repo. Motivation: This Mac ran two hooks that were deleted upstream, and hook registration is spread across dozens of installer scripts. Alternative considerations: Keeping per-hook installers and only adding the drift check was set aside because the registration would still be spread out. Implementation details: Create engine/hooks/_sdk/install_from_registry.py, call it from install.sh in place of the per-hook installer calls, and extend scripts/check_install_effective.py to report installed catstack hooks the registry does not name. Non-goals: No hook behavior changes. Layer: domain Feature state: active Files: engine/hooks/_sdk/install_from_registry.py, install.sh, scripts/check_install_effective.py, tests/test_install.py Change types: - engine/hooks/_sdk/install_from_registry.py: create - install.sh: modify - scripts/check_install_effective.py: modify - tests/test_install.py: modify Acceptance criteria: - `python3 -m unittest tests.test_install` exits 0. - Installing into a temporary home writes entries only for registry hooks whose mode is not off, each through the metrics runner. - An installed catstack hook the registry does not name is reported by the install check. Invoker-Finalize-Id: 82f89819-8a38-4511-ac5d-c33273d64cf5
…eterministic proof for install every hook from the registry through the metrics runner and report installed hooks the registry does not name. Review claim: The proof exits 0 only when install every hook from the registry through the metrics runner and report installed hooks the registry does not name holds. Review lane: proof Safety invariant: Proof only; it changes no product behavior. Effectiveness measurement: The exit status of `python3 -m unittest tests.test_install` is the signal for this slice. Slice rationale: One proof unit for this workflow. Architectural effect: None; verification only. Goal: Prove install every hook from the registry through the metrics runner and report installed hooks the registry does not name with one deterministic run. Motivation: Each workflow carries its own proof so a reviewer can trust the slice alone. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Execute the proof as a terminal gate. Non-goals: No product edits here. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: 46360da2-1609-4d19-80bb-69fb3a91c870
…only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active Exit code: 0 Invoker-Finalize-Id: 13f07187-d6a5-46ed-ba78-caf31725c2f8
…-a5b497e7b-b8c4a67f — Terminal read-only gate confirming no ephemeral handoff files were left behind. Review claim: The workflow leaves no ephemeral handoff files in the tree. Review lane: proof Safety invariant: Read-only; it never deletes files, alters the index, or commits caller work. Effectiveness measurement: A non-zero exit when ephemeral handoff files remain is the signal. Slice rationale: One unit: the hygiene gate. Architectural effect: None. Goal: Confirm no ephemeral handoff files remain after every other task finishes. Motivation: Ephemeral inter-task files leak into the diff and read as part of the change. Alternative considerations: Manual inspection was set aside as non-deterministic. Implementation details: Run scripts/scrub-handoff-artifacts.sh in check mode. Non-goals: No deletion, no index changes, no commits. Layer: e2e_regression Feature state: active
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_95c9f795-f010-41f2-944d-ab2014bbb08d) |
Contributor
|
Queued — the merge queue status continues in this comment ↓. |
added 2 commits
September 27, 2026 06:48
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_5256a7cc-2c3e-4979-806f-f70720f9a3df) |
Owner
Author
|
@Mergifyio queue |
Contributor
Merge Queue Status
This pull request spent 31 minutes 22 seconds in the queue, including 30 minutes 58 seconds running CI. Required conditions to merge
|
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Setup now follows the shared registry for Claude, Cursor, and Codex.
Separate lists could leave deleted entries active or omit new ones.
The setup path now selects enabled entries from the registry and routes them through the shared runner. Its drift check also reports active entries the registry does not name.
Review Claim
Approve registry-driven setup that keeps each enabled entry on the shared metrics path and reports active entries missing from the registry.
Review Lane
behavior
Review Unit
engine-runtime
Safety Invariant
The installed hook set equals the registry’s hooks whose mode is not off, and every installed hook continues through the shared metrics runner.
Slice Rationale
This slice keeps the setup path, drift check, and regression coverage together so they verify one source of truth end to end.
Non-goals
Architecture
Before
Separate per-harness hook lists drove installation, while the registry independently informed drift checks.
After
The hook registry selects enabled entries, the shared metrics runner drives Claude, Cursor, and Codex setup, and the drift check compares active entries against that registry.
Test Plan
Test Plan
python3 -m unittest tests.test_installbash scripts/scrub-handoff-artifacts.shRevert Plan
Revert Plan
git revert <sha>python3 -m unittest tests.test_install.Note
Medium Risk
Install behavior now depends on registry listing and mode filtering; a bug in
harness_hook_namesor the new loops could omit symlinks or leave stale links until prune runs.Overview
Replaces hand-maintained per-harness
link_itemlists ininstall.shwith loops that symlink hook directories frominstall_from_registry.py --list-harness-hooksfor Claude, Cursor, and Codex. Infrastructure dirs (_flags,_sdk,_markers,_runner) stay explicit; everything else follows the registry.Adds
harness_hook_names()and a CLI flag so tests and the shell installer share one definition of which hooks belong on each harness (e.g.cat-mode-defaultonly on Claude). Layout tests no longer regex-matchinstall.sh.Aligns “active hook” semantics with registry
mode: config merge, SubagentStop mirroring, CI drift checks, and install tests now treat only hooks withmode != "off"as wired—whilemode = "off"hooks can still be symlinked so scripts exist on disk but are not registered in settings/hooks.json (covered forskill-usage-logandunverified-tag-check).Reviewed by Cursor Bugbot for commit dd755a3. Bugbot is set up for automated code reviews on this repo. Configure here.