Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
25622aa
make-pr: preflight runs the PR-body validator on --body-file
EdbertChan Sep 23, 2026
aeb2d1c
invoker: wf-1790182452439-6/implement-preflight-runs-validator — Revi…
EdbertChan Sep 23, 2026
e896182
invoker: wf-1790182452439-6/verify-preflight-runs-validator-1 — Revie…
EdbertChan Sep 23, 2026
fe1887b
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
394d321
invoker: wf-1790182452439-6/verify-preflight-runs-validator-2 — Revie…
EdbertChan Sep 23, 2026
5950111
invoker: wf-1790182452439-6/verify-preflight-runs-validator-3 — Revie…
EdbertChan Sep 23, 2026
e59f975
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
0c52061
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
54e76b3
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
19ead97
Invoker: merge experiment/wf-1790182452439-6/verify-preflight-runs-va…
EdbertChan Sep 23, 2026
cb1089e
Invoker: merge experiment/wf-1790182452439-6/verify-preflight-runs-va…
EdbertChan Sep 23, 2026
ed02a2a
Invoker: merge experiment/wf-1790182452439-6/verify-preflight-runs-va…
EdbertChan Sep 23, 2026
cc4dd39
invoker: wf-1790182452439-6/scrub-handoff-artifacts — Review claim: N…
EdbertChan Sep 23, 2026
9a9c913
Merge experiment/wf-1790182452439-6/scrub-handoff-artifacts/g0.t1.a-a…
EdbertChan Sep 23, 2026
9bb7258
cat-mode: many PR stacks run as one parallel unit per stack, never se…
EdbertChan Sep 24, 2026
13ab621
reflect: catch an evidence-order correction from the ledger transitio…
EdbertChan Sep 22, 2026
8ae48c3
invoker: wf-1790219068137-115/repair — Repair PR #890: failed_checks:…
Sep 24, 2026
1c598ed
make-pr: preflight hands the validator the changed files
edbertchantech-ai Sep 24, 2026
4688b8d
Merge of #890
mergify[bot] Sep 24, 2026
5ec556a
Merge of #782
mergify[bot] Sep 24, 2026
2124b89
Merge of #841
mergify[bot] Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions corpus/skills/cat-mode/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -138,6 +138,7 @@ isolated subagents and report back async rather than blocking on each one.
durable artifact.** Separable and parallel is not authorization to fan
out; a fan-out default cannot hand a subagent publishing authority the
routing table never granted. Route that work through Execution routing.
- **Many PR stacks: one parallel unit per stack, never serial** (Invoker, else a worktree subagent each).
- **A fork/subagent told to touch files must run in its own worktree, not
the live checkout** — even when told "read-only." Scope wording is not
filesystem isolation.
Expand Down
1 change: 1 addition & 0 deletions corpus/skills/cat-mode/references/execution-routing.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ Catstack owns judgment and local fallback. Invoker owns durable plan submission,

## Decision

0. **More than one independent publishing unit** (several PR stacks to land or repair, several workflows): never serial in the parent chat. Invoker first, one workflow per unit; `subagent_worktree_per_unit` when Invoker is unavailable or the user directs subagents. `route_execution(units=N)` returns it; see [subagents.md](subagents.md).
1. **Invoker unavailable** (no `invoker_prepare_plan_review` / `invoker_submit_plan` tools): stay local — subagents, `loop-generator`, `land-stack`, current chat execution.
2. **Small local work** (one-file fix, short edit, read-only question): stay local even if Invoker is installed. Post-land wait until `MERGED`, merge-queue babysit, and already-named execution Backlog are **not** this bucket — they are `durable_parallel`.
3. **Approved plan or durable/parallel work** and Invoker MCP is available: delegate. If Invoker is missing, use a separate git worktree + PR stack. Do not park that work in the parent chat.
Expand Down
23 changes: 23 additions & 0 deletions corpus/skills/cat-mode/references/subagents.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,29 @@ definition whatever `produces` claims; and an empty or unrecognized `produces`
raises rather than falling through to fan-out, because an output nobody
declared is unchecked, not clean.

## Many stacks: parallel per unit, never serial

Routing picks *who* runs publishing work; it never licenses running several
independent units one after another in the parent thread. When the work is N
independent PR stacks (landing, conflict repair, review fixes), each stack is
its own unit with its own worktree, and the units run in parallel: an Invoker
workflow per stack first, and one worktree-isolated subagent per stack as the
fallback when Invoker is unavailable or the user directs it
(`subagent_worktree_per_unit`). The per-unit subagent still inherits only the
scope the parent names, and its transcript is still grepped for writes.

The failure shape: asked to land dozens of admin-bypass PRs across two repos,
the parent recommended working the ~20 rebases and review fixes "one at a
time" and started serially in one worktree, until the user asked for a
worktree subagent per stack. The routing table allowed it: `route_execution`
returned `local` for publishing work without Invoker at any unit count.

Prior art: Amdahl's law — Gene M. Amdahl, "Validity of the single processor
approach to achieving large scale computing capabilities", AFIPS 1967,
https://doi.org/10.1145/1465482.1465560 — the serial fraction bounds the
whole job, so independent units forced through one thread set the finish
time.

## Defer to the harness's routing skill

The precedence above is catstack's fallback, not the owner. When a harness
Expand Down
55 changes: 46 additions & 9 deletions corpus/skills/cat-mode/scripts/route_execution.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,8 @@
from typing import Literal

WorkKind = Literal["readonly", "small_local", "approved_plan", "durable_parallel"]
Route = Literal["local", "delegate_invoker"]
Delegation = Literal["local", "delegate_invoker", "subagent_fanout"]
Route = Literal["local", "delegate_invoker", "subagent_worktree_per_unit"]
Delegation = Literal["local", "delegate_invoker", "subagent_fanout", "subagent_worktree_per_unit"]

DURABLE_ALIASES = frozenset({"post_land_babysit", "named_execution_backlog"})

Expand All @@ -36,6 +36,13 @@
"invoker_wait_for_workflow_or_status",
)

SUBAGENT_PER_UNIT_STEPS = (
"one_worktree_per_unit",
"spawn_one_subagent_per_unit_in_parallel",
"collect_reports_async",
"grep_transcripts_for_writes",
)

SUBAGENT_FANOUT_STEPS = (
"spawn_worktree_isolated_subagents",
"collect_reports_async",
Expand Down Expand Up @@ -72,14 +79,32 @@ def normalize_work_kind(work_kind: str) -> WorkKind:
raise ValueError(f"unknown work_kind: {work_kind!r}")


def route_execution(*, tools: set[str] | frozenset[str] | list[str], work_kind: str) -> Route:
def route_execution(
*,
tools: set[str] | frozenset[str] | list[str],
work_kind: str,
units: int = 1,
user_directed_subagents: bool = False,
) -> Route:
"""Return where execution should run for this request.

1. Invoker MCP missing → local
2. Small / read-only work → local even if Invoker exists
3. Approved plan or durable/parallel → delegate_invoker
`units` counts independent publishing units (PR stacks, workflows).
More than one never runs serially in the parent thread:
Invoker first, else one worktree-isolated subagent per unit.

1. units > 1 → delegate_invoker, or subagent_worktree_per_unit when
Invoker is missing or the user directed subagents
2. Invoker MCP missing → local
3. Small / read-only work → local even if Invoker exists
4. Approved plan or durable/parallel → delegate_invoker
"""
kind = normalize_work_kind(work_kind)
if units < 1:
raise ValueError(f"units must be >= 1, got {units!r}")
if units > 1 and kind != "readonly":
if invoker_mcp_available(tools) and not user_directed_subagents:
return "delegate_invoker"
return "subagent_worktree_per_unit"
if not invoker_mcp_available(tools):
return "local"
if kind in ("readonly", "small_local"):
Expand Down Expand Up @@ -112,6 +137,8 @@ def route_delegation(
tools: set[str] | frozenset[str] | list[str],
work_kind: str,
produces: set[str] | frozenset[str] | list[str] | tuple[str, ...],
units: int = 1,
user_directed_subagents: bool = False,
) -> Delegation:
"""Resolve the Subagents default against the execution-routing table.

Expand All @@ -122,7 +149,9 @@ def route_delegation(
"""
normalize_work_kind(work_kind)
if publishes(work_kind=work_kind, produces=produces):
return route_execution(tools=tools, work_kind=work_kind)
return route_execution(
tools=tools, work_kind=work_kind, units=units, user_directed_subagents=user_directed_subagents,
)
return "subagent_fanout"


Expand All @@ -131,6 +160,8 @@ def handoff_steps_for(route: Route | Delegation) -> tuple[str, ...]:
return ("stay_local",)
if route == "subagent_fanout":
return SUBAGENT_FANOUT_STEPS
if route == "subagent_worktree_per_unit":
return SUBAGENT_PER_UNIT_STEPS
return DELEGATE_HANDOFF_STEPS


Expand All @@ -142,9 +173,15 @@ def handoff_steps_for(route: Route | Delegation) -> tuple[str, ...]:
tools = payload.get("tools", [])
work_kind = payload.get("work_kind", "small_local")
produces = payload.get("produces")
units = int(payload.get("units", 1))
directed = bool(payload.get("user_directed_subagents", False))
if produces is None:
route: Route | Delegation = route_execution(tools=tools, work_kind=work_kind)
route: Route | Delegation = route_execution(
tools=tools, work_kind=work_kind, units=units, user_directed_subagents=directed,
)
else:
route = route_delegation(tools=tools, work_kind=work_kind, produces=produces)
route = route_delegation(
tools=tools, work_kind=work_kind, produces=produces, units=units, user_directed_subagents=directed,
)
defer_to = installed_harness_routing_skill(payload.get("home"))
print(json.dumps({"route": route, "steps": list(handoff_steps_for(route)), "defer_to": defer_to}))
6 changes: 4 additions & 2 deletions engine/hooks/llm-judge/phrases/wrong-check-reflect.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"checker": "wrong-check-reflect",
"reads": "reply",
"meaning": "The latest assistant reply admits that something it told the user earlier was wrong, misread, or answered the wrong question.",
"meaning": "The latest assistant reply admits that something it told the user earlier was wrong, misread, or answered the wrong question, OR that an earlier claim of its own went out before the check backing it had run. A reply saying the claim turned out correct but was stated before it was checked IS a match: the failure is the order, not the answer, so no word in the reply has to say anything was wrong.",
"match": [
"my earlier check was wrong",
"You're right, I misread that",
Expand All @@ -10,7 +10,9 @@
"I was wrong about the path",
"Correction: the file I pointed you to earlier is not the one in use; the real one is src/b.py.",
"Good catch. The earlier number was off; the real count is 12.",
"My earlier read of the config was wrong \u2014 the default is 4, not 8."
"My earlier read of the config was wrong \u2014 the default is 4, not 8.",
"Correcting one claim and arming the check I implied:",
"I was right - but I said it a turn before I checked it"
],
"not_match": [
"You're right. Let's go with option B.",
Expand Down
6 changes: 6 additions & 0 deletions engine/hooks/unverified-tag-ledger/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,12 @@ fixtures in `tests/test_hooks.py`.
reason is written to stderr.
- **Escalation** — a claim outstanding `ESCALATE_AFTER_TURNS` (3) turns or more
is reported as a reflect trigger rather than accumulating quietly.
- **Discharge is itself a reflect trigger** — a row going outstanding ->
discharged is the record of a claim that went out first and was checked
after. That is an evidence-order miss, and it carries no wrongness word, so
the phrase scanners (`engine/skills/reflect/scripts/self_retraction_scan.py`,
and the `wrong-check-reflect` dictionary) cannot see it from the text. This
hook sees it from state instead, and says so on the Stop that discharges.

Malformed tags are deliberately ignored here; `diu-stop` already rejects those.

Expand Down
35 changes: 26 additions & 9 deletions engine/hooks/unverified-tag-ledger/detect.py
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,15 @@
ESCALATE_AFTER_TURNS = 3
MAX_LISTED = 5

DISCHARGE_REFLECT = (
"unverified-tag-ledger: {count} claim(s) went from unverified to checked this turn: "
"{claims}. That transition is the whole event: the claim went out first and the check "
"ran after. No wording has to admit anything for this to be true, which is why the "
"phrase scanners miss it -- an evidence-order miss carries no wrongness word. "
"Treat it as a reflect trigger, not a milestone: run reflect on this transcript, or "
"say plainly why this one does not need it."
)

CLAIM_RE = re.compile(
r"\{\{\s*CAT-UNVERIFIED\s*:?\s*(?P<claim>.*?)(?:--|—)\s*cannot\s+verify\s*:\s*(?P<reason>[^}]*)\}\}",
re.IGNORECASE | re.DOTALL,
Expand Down Expand Up @@ -177,18 +186,26 @@ def evaluate(payload: dict) -> dict:
session_id = str(payload.get("session_id") or "")
message = _last_assistant_text(payload)
tools = tools_used_this_turn(payload)
was_open = {row["claim"] for row in outstanding(read_ledger(session_id))}
rows = record_turn(session_id, message, tools)
notes = []
discharged = sorted(
row["claim"] for row in rows if row.get("resolved") and row["claim"] in was_open)
if discharged:
notes.append(DISCHARGE_REFLECT.format(
count=len(discharged), claims="; ".join(discharged[:MAX_LISTED])))

new_claims = {tag["claim"] for tag in parse_tags(message)}
if not new_claims:
return {"note": "", "block": ""}
return {"note": "\n".join(notes), "block": ""}

if tools is None:
return {"note": (
notes.append(
f"unverified-tag-ledger: logged {len(new_claims)} CAT-UNVERIFIED claim(s), but this "
"turn's tool calls could not be read from transcript_path (see the line above), so "
"whether a check was attempted is UNCHECKED, not clean. Nothing was discharged and "
"the turn was not refused."), "block": ""}
"the turn was not refused.")
return {"note": "\n".join(notes), "block": ""}

if not tools & VERIFY_TOOLS and not payload.get("stop_hook_active"):
claims = "; ".join(sorted(new_claims)[:MAX_LISTED])
Expand All @@ -201,12 +218,12 @@ def evaluate(payload: dict) -> dict:

fresh = [row for row in rows
if row["claim"] in new_claims and not row.get("resolved") and row.get("turns", 0) == 0]
if not fresh:
return {"note": "", "block": ""}
return {"note": (
f"unverified-tag-ledger: logged {len(fresh)} CAT-UNVERIFIED claim(s) against this session. "
"They are deferred, not discharged, and will be raised again next turn "
"(cat-mode/SKILL.md:269)."), "block": ""}
if fresh:
notes.append(
f"unverified-tag-ledger: logged {len(fresh)} CAT-UNVERIFIED claim(s) against this "
"session. They are deferred, not discharged, and will be raised again next turn "
"(cat-mode/SKILL.md:269).")
return {"note": "\n".join(notes), "block": ""}


def decide_stop(payload: dict) -> str:
Expand Down
30 changes: 30 additions & 0 deletions engine/hooks/unverified-tag-ledger/tests/test_hooks.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,9 @@
"-- cannot verify: my own reasoning isn't observable by any command}}")
MALFORMED = "{{CAT-UNVERIFIED: something I did not check}}"

EVIDENCE_ORDER_1 = "Correcting one claim and arming the check I implied:"
EVIDENCE_ORDER_2 = "I was right - but I said it a turn before I checked it"


def _real_transcript_lines() -> list[str]:
with open(REAL_TRANSCRIPT, encoding="utf-8") as handle:
Expand Down Expand Up @@ -209,6 +212,33 @@ def test_verified_and_dropped_tag_is_discharged_through_the_real_payload(self) -
self.assertEqual(self.detect.outstanding(self.detect.read_ledger("s1")), [])
self.assertEqual(self.detect.reminder("s1"), "")

def test_a_discharged_claim_fires_the_reflect_trigger(self) -> None:
self.detect.evaluate(self.payload(REAL_TAG_1, tools=True))
verdict = self.detect.evaluate(
self.payload("Here is the pasted output proving it.", tools=True))
self.assertIn("reflect trigger", verdict["note"])
self.assertIn("widened scope", verdict["note"])

def test_an_evidence_order_correction_triggers_with_no_wrongness_word(self) -> None:
"""The transition fires; the reply's wording is not consulted at all."""
for index, reply in enumerate((EVIDENCE_ORDER_1, EVIDENCE_ORDER_2)):
session = f"evidence-order-{index}"
self.detect.evaluate(self.payload(REAL_TAG_1, tools=True, session_id=session))
verdict = self.detect.evaluate(
self.payload(reply, tools=True, session_id=session))
self.assertIn("reflect trigger", verdict["note"])
self.assertIn("carries no wrongness word", verdict["note"])

def test_a_turn_that_discharges_nothing_stays_silent_about_reflect(self) -> None:
verdict = self.detect.evaluate(
self.payload("Ran the tests, all green.", tools=True))
self.assertEqual(verdict["note"], "")

def test_a_reemitted_tag_is_not_reported_as_discharged(self) -> None:
self.detect.evaluate(self.payload(REAL_TAG_1, tools=True))
verdict = self.detect.evaluate(self.payload(REAL_TAG_1, tools=True))
self.assertNotIn("reflect trigger", verdict["note"])

def test_unchecked_turn_does_not_discharge_a_row(self) -> None:
self.detect.evaluate(self.payload(REAL_TAG_1, tools=True))
self.detect.evaluate({
Expand Down
3 changes: 3 additions & 0 deletions engine/hooks/wrong-check-reflect/eval_dictionary.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,9 @@
(HIT_TEXT, True),
("You're right. Let's go with option B.", False),
("I double-checked my earlier count and it holds; nothing in it was wrong.", False),
("Correcting one claim and arming the check I implied:", True),
("I was right - but I said it a turn before I checked it", True),
("I ran the check first and then said it, so the order was right.", False),
)


Expand Down
48 changes: 45 additions & 3 deletions engine/skills/make-pr/scripts/preflight.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@
import shutil
import subprocess
import sys
import tempfile

sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))

Expand All @@ -39,6 +40,9 @@

DEFAULT_CONFIG = os.path.join(REPO_ROOT, "drafter.config.json")
UNCHECKED_EXIT = 3
VALIDATOR = "engine/skills/draft-pr/scripts/validate-pr-body.mjs"
VALIDATOR_NAME = os.path.basename(VALIDATOR)
VALIDATOR_TIMEOUT = 120


class UnitRulesUnreadable(Exception):
Expand Down Expand Up @@ -222,7 +226,39 @@ def changed_paths(base: str, repo: str = REPO_ROOT) -> list[str]:
return sorted({p for p in (out + untracked).splitlines() if p.strip()})


def describe(body_file: str | None) -> int:
def validate_body(body_file: str, run=None, which=None, changed_paths: list[str] | None = None) -> tuple[int, list[str]]:
"""Run the PR-body schema validator that the required PR Body check runs.

description_check only reads the prose for history claims, so a body whose
Test Plan is not inside a <details> block passed preflight and then failed
the required check after publication (PRs #780-#789, #793, #795). Exit 1
from the validator is a real rejection; anything else -- no node, a
timeout, an unexpected exit code -- is unchecked, which also fails.
"""
run = run if run is not None else subprocess.run
which = which if which is not None else shutil.which
if not which("node"):
return 1, [f"description unchecked: node is not on PATH, so {VALIDATOR_NAME} did not run"]
if not os.path.isfile(os.path.join(REPO_ROOT, VALIDATOR)):
return 1, [f"description unchecked: no {VALIDATOR} under {REPO_ROOT}"]
with tempfile.TemporaryDirectory(prefix="preflight-") as tmp:
files_file = os.path.join(tmp, "changed-files.txt")
with open(files_file, "w", encoding="utf-8") as handle:
handle.writelines(p + "\n" for p in changed_paths or [])
cmd = ["node", VALIDATOR, "--body-file", os.path.abspath(body_file), "--changed-files-file", files_file]
try:
res = run(cmd, cwd=REPO_ROOT, capture_output=True, text=True, timeout=VALIDATOR_TIMEOUT)
except subprocess.TimeoutExpired:
return 1, [f"description unchecked: {VALIDATOR_NAME} did not finish in {VALIDATOR_TIMEOUT}s"]
except OSError as exc:
return 1, [f"description unchecked: cannot run {VALIDATOR_NAME}: {exc}"]
lines = (res.stdout + res.stderr).strip().splitlines()
if res.returncode in (0, 1):
return res.returncode, lines
return 1, lines + [f"description unchecked: {VALIDATOR_NAME} exited {res.returncode}"]


def describe(body_file: str | None, changed_paths: list[str]) -> int:
if not body_file:
print("fail description unchecked: pass --body-file with the PR description")
return 1
Expand All @@ -239,7 +275,13 @@ def describe(body_file: str | None) -> int:
for line in lines:
print(" " + line)
print(f" description {outcome}")
return 0 if outcome == "clean" else 1
status = 0 if outcome == "clean" else 1

print(f"gate node {VALIDATOR} --body-file {body_file} --changed-files-file <{len(changed_paths)} changed path(s)>")
schema_status, schema_lines = validate_body(body_file, changed_paths=changed_paths)
for line in schema_lines:
print(" " + line)
return status or schema_status


def main(argv: list[str] | None = None) -> int:
Expand Down Expand Up @@ -301,7 +343,7 @@ def main(argv: list[str] | None = None) -> int:
if res.returncode != 0:
status = 1
if args.paths is None and not args.dry_run:
if describe(args.body_file) != 0:
if describe(args.body_file, paths) != 0:
status = 1
print("ok preflight passed" if status == 0 else "fail preflight: fix the above before gh pr create")
return status
Expand Down
Loading
Loading