Skip to content

merge queue: checking #841 on main (90dffd5), stacked on #890 and #782 - #905

Closed
mergify[bot] wants to merge 21 commits into
mainfrom
mergify/merge-queue/1012dbad59
Closed

mergify[bot] wants to merge 21 commits into
mainfrom
mergify/merge-queue/1012dbad59

Conversation

@mergify

@mergify mergify Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

🎉 This pull request has been checked successfully and will be merged soon. 🎉

#841 is queued for merge on branch main (90dffd5).

Stacked behind 2 pull requests queued ahead of this batch, not part of it. These checks run on a tip that also carries their commits, so a failure here can come from them as much as from #841.

Queued ahead of this batch:

This pull request has been created by Mergify to speculatively check the mergeability of #841.
You don't need to do anything. Mergify will close this pull request automatically when it is complete.

Required conditions of queue rule admin-bypass for merge:

  • check-success = lint
  • check-success = test
  • check-success = validate

Required conditions to stay in the queue:

---
checking_base_sha: 5ec556a829cac997f501a4b46d2d6a0968125d92
previous_check_retries: []
previous_failed_batches: []
pull_requests:
  - number: 841
    scopes: []
scopes: []
...

EdbertChan and others added 21 commits September 24, 2026 00:58
description_check only read the description's prose for claims about the
repo's past, so preflight printed "ok preflight passed" for a body the
required PR Body check then rejected after publication (#780-#789, #793,
#795). describe() now also runs
engine/skills/draft-pr/scripts/validate-pr-body.mjs on the same file and
prints its errors under the gate line.

Exit 1 from the validator fails preflight. No node, a missing validator,
a timeout, or any other exit code prints "description unchecked: ..." and
fails too -- a check that could not run is not a pass.

The #795 description lands as a fixture: it exits 1 with Test Plan and
Revert Plan outside <details> and code names in the Summary, and the new
tests show preflight failing on it and still passing a valid body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ew claim: make-pr preflight runs the PR-description validator on --body-file and fails when it fails, so it no longer prints 'preflight passed' for a description the required PR Body check rejects.

Review lane: policy
Safety invariant: The change can only add a failure or an UNCHECKED notice; a PR description that passes engine/skills/draft-pr/scripts/validate-pr-body.mjs today is never newly blocked.
Effectiveness measurement: A preflight test feeds the original #795 description (Test Plan and Revert Plan not inside <details>, code names in Summary), which the validator rejects with exit 1, and asserts preflight exits non-zero and prints the validator's errors; a passing description still gives 'ok preflight passed'; a validator that cannot run gives an unchecked failure, not a pass.
Slice rationale: One script, one claim: preflight's description step also runs the schema validator.
Architectural effect: preflight's describe() step covers the same rules as the required PR Body check.
Goal: Stop preflight from approving a PR description the required check will reject.
Motivation: A reflect pass found PR descriptions on catstack PRs #780-#789, #793 and #795 failing the required PR Body check after publication, and the user had to ask for a manual fix of every PR. In catstack the guard stayed silent: fed `mergify stack push` and `gh pr create --body-file <failing body>` it exited 0 with no output, while the same stack push in the Invoker checkout fired. Read at origin/main: engine/skills/make-pr/scripts/preflight.py describe() only runs description_check, then main() prints 'ok preflight passed'; the session that opened #795 got that line and then ran gh pr create.
Alternative considerations: Leaving the validator to CI was set aside because the failure then shows up only after publication. Reimplementing the rules in Python was set aside because it would drift from the validator.
Implementation details: In engine/skills/make-pr/scripts/preflight.py describe(), after description_check, run `node engine/skills/draft-pr/scripts/validate-pr-body.mjs --body-file <file>` from the repo root with a timeout. Print its output lines indented like the other gates. Exit 1 from the validator fails preflight; a missing node binary, a timeout, or any other exit code prints 'description unchecked: <reason>' and fails preflight. Add the tests to engine/skills/make-pr/tests/test_preflight.py with the #795 body as a fixture file under engine/skills/make-pr/tests/fixtures/.
Non-goals: No change to the validator, the hook, or the make-pr SKILL.md rules. Do not edit engine/skills/draft-pr/scripts/validate-pr-body.mjs, scripts/pr/validate-pr-body-local.mjs, or any other file open PR #742 changes; call the validator as it is. If a change would overlap an open PR, stop and report instead.
Layer: domain
Feature state: active
Files:
- engine/skills/make-pr/scripts/preflight.py
- engine/skills/make-pr/tests/test_preflight.py
- engine/skills/make-pr/tests/fixtures/pr795-failing-body.md
Change types:
- engine/skills/make-pr/scripts/preflight.py: modify
- engine/skills/make-pr/tests/test_preflight.py: modify
- engine/skills/make-pr/tests/fixtures/pr795-failing-body.md: create
Acceptance criteria:
- `python3 -m unittest discover -s engine/skills/make-pr/tests -v` exits 0.
- `python3 scripts/check_skill_test_coverage.py --base origin/main --head HEAD` exits 0.
- `python3 scripts/check_skill_file_refs.py` exits 0.
- `python3 scripts/check_no_new_comments.py` exits 0.

Exit code: 0
…w claim: `python3 -m unittest discover -s engine/skills/make-pr/tests -v` passes on the finished branch.

Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 -m unittest discover -s engine/skills/make-pr/tests -v`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0
…w claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.

Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1
…w claim: `python3 scripts/check_skill_test_coverage.py --base origin/main --head HEAD` passes on the finished branch.

Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_skill_test_coverage.py --base origin/main --head HEAD`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0
…w claim: `python3 scripts/check_skill_file_refs.py` passes on the finished branch.

Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_skill_file_refs.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0
…w claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.

Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 1
…w claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.

Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Solution:
  Review claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.
Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.
…w claim: `python3 scripts/check_no_new_comments.py` passes on the finished branch.

Review lane: proof
Safety invariant: Verification is read-only and alters no repository file.
Effectiveness measurement: The command's exit code is the direct measurement.
Slice rationale: One check per proof task.
Architectural effect: None; verification only.
Goal: Prove the slice.
Motivation: Running the check is the proof.
Alternative considerations: The full suite was not required because the slice touches one component with its own tests.
Implementation details: Run `python3 scripts/check_no_new_comments.py`.
Non-goals: No mutations.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0
…o ephemeral inter-task handoff files remain in the worktree before the merge gate.

Review lane: cleanup
Safety invariant: The scrub script only checks for known handoff artifact names and never touches source, tests, or other repository files.
Effectiveness measurement: The script exits non-zero if any handoff artifact remains.
Slice rationale: Required terminal scrub for every implementation workflow.
Architectural effect: None; hygiene only.
Goal: Leave the branch free of handoff artifacts.
Motivation: Handoff files must not reach the PR.
Alternative considerations: Manual cleanup was set aside as non-deterministic.
Implementation details: Run scripts/scrub-handoff-artifacts.sh.
Non-goals: No product edits.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.

Exit code: 0
…c4e653a5-a4933bc3 — Review claim: No ephemeral inter-task handoff files remain in the worktree before the merge gate.

Review lane: cleanup
Safety invariant: The scrub script only checks for known handoff artifact names and never touches source, tests, or other repository files.
Effectiveness measurement: The script exits non-zero if any handoff artifact remains.
Slice rationale: Required terminal scrub for every implementation workflow.
Architectural effect: None; hygiene only.
Goal: Leave the branch free of handoff artifacts.
Motivation: Handoff files must not reach the PR.
Alternative considerations: Manual cleanup was set aside as non-deterministic.
Implementation details: Run scripts/scrub-handoff-artifacts.sh.
Non-goals: No product edits.
Layer: app_regression
Feature state: active
Layer exception: allowed. Verification and the terminal scrub run after the docs commit so they check the final branch the PR will carry.
Acceptance criteria:
- The command exits 0.
…rially

route_execution(units=N) never returns local for N>1 publishing units:
Invoker first, else one worktree subagent per unit (also when the user
directs subagents).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Change-Id: I9d5797fbed454079561fd2f4b36379ab7b7df02d
…n, not from wording

Both retraction detectors need a wrongness word. Two real corrections carry
none, because the claim was true and only its order was wrong:

  "Correcting one claim and arming the check I implied:"
  "I was right - but I said it a turn before I checked it"

Primary fix, no regex and no model: a ledger row going outstanding ->
discharged IS the event "a claim was asserted before its check ran, and the
check has now run". The Stop that discharges a row now reports it as a reflect
trigger. It reads state, so the reply's wording is not consulted at all.

Secondary fix: the judge dictionary's meaning now covers "right, but said
before it was checked", and both real strings are eval cases. Run against the
real judge, all six cases land where the dictionary says they should.

The regex scan keeps its hole on purpose, pinned by three tests so the next
author widens it knowingly rather than by accident.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I99190104b68d88f78aa9c8a706c9a6d23055aecc
The required PR Body check passes --changed-files-file, which is how the
validator rejects changed file and folder names in the Summary and a
Review Unit that does not match the diff. Preflight passed only
--body-file, so those descriptions passed locally and failed after
publish. describe() now takes the changed paths preflight already
computed and writes them to a temp file for the validator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Change-Id: Id36a64869b7cfb2b76b912fd3cf81a88a39dd344
@mergify mergify Bot closed this Sep 24, 2026
@mergify
mergify Bot deleted the mergify/merge-queue/1012dbad59 branch September 24, 2026 04:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants