Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
25622aa
make-pr: preflight runs the PR-body validator on --body-file
EdbertChan Sep 23, 2026
aeb2d1c
invoker: wf-1790182452439-6/implement-preflight-runs-validator — Revi…
EdbertChan Sep 23, 2026
e896182
invoker: wf-1790182452439-6/verify-preflight-runs-validator-1 — Revie…
EdbertChan Sep 23, 2026
fe1887b
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
394d321
invoker: wf-1790182452439-6/verify-preflight-runs-validator-2 — Revie…
EdbertChan Sep 23, 2026
5950111
invoker: wf-1790182452439-6/verify-preflight-runs-validator-3 — Revie…
EdbertChan Sep 23, 2026
e59f975
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
0c52061
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
54e76b3
invoker: wf-1790182452439-6/verify-preflight-runs-validator-4 — Revie…
EdbertChan Sep 23, 2026
19ead97
Invoker: merge experiment/wf-1790182452439-6/verify-preflight-runs-va…
EdbertChan Sep 23, 2026
cb1089e
Invoker: merge experiment/wf-1790182452439-6/verify-preflight-runs-va…
EdbertChan Sep 23, 2026
ed02a2a
Invoker: merge experiment/wf-1790182452439-6/verify-preflight-runs-va…
EdbertChan Sep 23, 2026
cc4dd39
invoker: wf-1790182452439-6/scrub-handoff-artifacts — Review claim: N…
EdbertChan Sep 23, 2026
9a9c913
Merge experiment/wf-1790182452439-6/scrub-handoff-artifacts/g0.t1.a-a…
EdbertChan Sep 23, 2026
9bb7258
cat-mode: many PR stacks run as one parallel unit per stack, never se…
EdbertChan Sep 24, 2026
13ab621
reflect: catch an evidence-order correction from the ledger transitio…
EdbertChan Sep 22, 2026
8ae48c3
invoker: wf-1790219068137-115/repair — Repair PR #890: failed_checks:…
Sep 24, 2026
1c598ed
make-pr: preflight hands the validator the changed files
edbertchantech-ai Sep 24, 2026
ccddeb1
event-wait: read a final unterminated line, kill a stuck source, hono…
EdbertChan Sep 24, 2026
aa9412e
llm-judge: find Codex's transcript by thread-id, and deliver subagent…
EdbertChan Sep 23, 2026
9df0641
llm-judge: look for Codex rollouts under CODEX_HOME when it is set
Sep 23, 2026
619f890
invoker: wf-1790173197097-223/repair — Repair PR #819: failed_checks:…
Sep 23, 2026
3500da6
Merge of #890
mergify[bot] Sep 24, 2026
71f78a5
Merge of #782
mergify[bot] Sep 24, 2026
86149ac
Merge of #841
mergify[bot] Sep 24, 2026
021b3bd
Merge of #819
mergify[bot] Sep 24, 2026
5b5e6c2
Merge of #791
mergify[bot] Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions corpus/skills/cat-mode/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -138,6 +138,7 @@ isolated subagents and report back async rather than blocking on each one.
durable artifact.** Separable and parallel is not authorization to fan
out; a fan-out default cannot hand a subagent publishing authority the
routing table never granted. Route that work through Execution routing.
- **Many PR stacks: one parallel unit per stack, never serial** (Invoker, else a worktree subagent each).
- **A fork/subagent told to touch files must run in its own worktree, not
the live checkout** — even when told "read-only." Scope wording is not
filesystem isolation.
Expand Down
1 change: 1 addition & 0 deletions corpus/skills/cat-mode/references/execution-routing.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ Catstack owns judgment and local fallback. Invoker owns durable plan submission,

## Decision

0. **More than one independent publishing unit** (several PR stacks to land or repair, several workflows): never serial in the parent chat. Invoker first, one workflow per unit; `subagent_worktree_per_unit` when Invoker is unavailable or the user directs subagents. `route_execution(units=N)` returns it; see [subagents.md](subagents.md).
1. **Invoker unavailable** (no `invoker_prepare_plan_review` / `invoker_submit_plan` tools): stay local — subagents, `loop-generator`, `land-stack`, current chat execution.
2. **Small local work** (one-file fix, short edit, read-only question): stay local even if Invoker is installed. Post-land wait until `MERGED`, merge-queue babysit, and already-named execution Backlog are **not** this bucket — they are `durable_parallel`.
3. **Approved plan or durable/parallel work** and Invoker MCP is available: delegate. If Invoker is missing, use a separate git worktree + PR stack. Do not park that work in the parent chat.
Expand Down
23 changes: 23 additions & 0 deletions corpus/skills/cat-mode/references/subagents.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,29 @@ definition whatever `produces` claims; and an empty or unrecognized `produces`
raises rather than falling through to fan-out, because an output nobody
declared is unchecked, not clean.

## Many stacks: parallel per unit, never serial

Routing picks *who* runs publishing work; it never licenses running several
independent units one after another in the parent thread. When the work is N
independent PR stacks (landing, conflict repair, review fixes), each stack is
its own unit with its own worktree, and the units run in parallel: an Invoker
workflow per stack first, and one worktree-isolated subagent per stack as the
fallback when Invoker is unavailable or the user directs it
(`subagent_worktree_per_unit`). The per-unit subagent still inherits only the
scope the parent names, and its transcript is still grepped for writes.

The failure shape: asked to land dozens of admin-bypass PRs across two repos,
the parent recommended working the ~20 rebases and review fixes "one at a
time" and started serially in one worktree, until the user asked for a
worktree subagent per stack. The routing table allowed it: `route_execution`
returned `local` for publishing work without Invoker at any unit count.

Prior art: Amdahl's law — Gene M. Amdahl, "Validity of the single processor
approach to achieving large scale computing capabilities", AFIPS 1967,
https://doi.org/10.1145/1465482.1465560 — the serial fraction bounds the
whole job, so independent units forced through one thread set the finish
time.

## Defer to the harness's routing skill

The precedence above is catstack's fallback, not the owner. When a harness
Expand Down
55 changes: 46 additions & 9 deletions corpus/skills/cat-mode/scripts/route_execution.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,8 @@
from typing import Literal

WorkKind = Literal["readonly", "small_local", "approved_plan", "durable_parallel"]
Route = Literal["local", "delegate_invoker"]
Delegation = Literal["local", "delegate_invoker", "subagent_fanout"]
Route = Literal["local", "delegate_invoker", "subagent_worktree_per_unit"]
Delegation = Literal["local", "delegate_invoker", "subagent_fanout", "subagent_worktree_per_unit"]

DURABLE_ALIASES = frozenset({"post_land_babysit", "named_execution_backlog"})

Expand All @@ -36,6 +36,13 @@
"invoker_wait_for_workflow_or_status",
)

SUBAGENT_PER_UNIT_STEPS = (
"one_worktree_per_unit",
"spawn_one_subagent_per_unit_in_parallel",
"collect_reports_async",
"grep_transcripts_for_writes",
)

SUBAGENT_FANOUT_STEPS = (
"spawn_worktree_isolated_subagents",
"collect_reports_async",
Expand Down Expand Up @@ -72,14 +79,32 @@ def normalize_work_kind(work_kind: str) -> WorkKind:
raise ValueError(f"unknown work_kind: {work_kind!r}")


def route_execution(*, tools: set[str] | frozenset[str] | list[str], work_kind: str) -> Route:
def route_execution(
*,
tools: set[str] | frozenset[str] | list[str],
work_kind: str,
units: int = 1,
user_directed_subagents: bool = False,
) -> Route:
"""Return where execution should run for this request.

1. Invoker MCP missing → local
2. Small / read-only work → local even if Invoker exists
3. Approved plan or durable/parallel → delegate_invoker
`units` counts independent publishing units (PR stacks, workflows).
More than one never runs serially in the parent thread:
Invoker first, else one worktree-isolated subagent per unit.

1. units > 1 → delegate_invoker, or subagent_worktree_per_unit when
Invoker is missing or the user directed subagents
2. Invoker MCP missing → local
3. Small / read-only work → local even if Invoker exists
4. Approved plan or durable/parallel → delegate_invoker
"""
kind = normalize_work_kind(work_kind)
if units < 1:
raise ValueError(f"units must be >= 1, got {units!r}")
if units > 1 and kind != "readonly":
if invoker_mcp_available(tools) and not user_directed_subagents:
return "delegate_invoker"
return "subagent_worktree_per_unit"
if not invoker_mcp_available(tools):
return "local"
if kind in ("readonly", "small_local"):
Expand Down Expand Up @@ -112,6 +137,8 @@ def route_delegation(
tools: set[str] | frozenset[str] | list[str],
work_kind: str,
produces: set[str] | frozenset[str] | list[str] | tuple[str, ...],
units: int = 1,
user_directed_subagents: bool = False,
) -> Delegation:
"""Resolve the Subagents default against the execution-routing table.

Expand All @@ -122,7 +149,9 @@ def route_delegation(
"""
normalize_work_kind(work_kind)
if publishes(work_kind=work_kind, produces=produces):
return route_execution(tools=tools, work_kind=work_kind)
return route_execution(
tools=tools, work_kind=work_kind, units=units, user_directed_subagents=user_directed_subagents,
)
return "subagent_fanout"


Expand All @@ -131,6 +160,8 @@ def handoff_steps_for(route: Route | Delegation) -> tuple[str, ...]:
return ("stay_local",)
if route == "subagent_fanout":
return SUBAGENT_FANOUT_STEPS
if route == "subagent_worktree_per_unit":
return SUBAGENT_PER_UNIT_STEPS
return DELEGATE_HANDOFF_STEPS


Expand All @@ -142,9 +173,15 @@ def handoff_steps_for(route: Route | Delegation) -> tuple[str, ...]:
tools = payload.get("tools", [])
work_kind = payload.get("work_kind", "small_local")
produces = payload.get("produces")
units = int(payload.get("units", 1))
directed = bool(payload.get("user_directed_subagents", False))
if produces is None:
route: Route | Delegation = route_execution(tools=tools, work_kind=work_kind)
route: Route | Delegation = route_execution(
tools=tools, work_kind=work_kind, units=units, user_directed_subagents=directed,
)
else:
route = route_delegation(tools=tools, work_kind=work_kind, produces=produces)
route = route_delegation(
tools=tools, work_kind=work_kind, produces=produces, units=units, user_directed_subagents=directed,
)
defer_to = installed_harness_routing_skill(payload.get("home"))
print(json.dumps({"route": route, "steps": list(handoff_steps_for(route)), "defer_to": defer_to}))
33 changes: 33 additions & 0 deletions engine/hooks/_sdk/transcripts.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
from __future__ import annotations

import glob
import os
import re

CODEX_SESSIONS_ENV = "CATSTACK_CODEX_SESSIONS_DIR"
CODEX_HOME_ENV = "CODEX_HOME"
THREAD_ID = re.compile(r"^[0-9A-Za-z-]{8,64}$")


def codex_sessions_root() -> str:
override = os.environ.get(CODEX_SESSIONS_ENV)
if override:
return override
home = os.environ.get(CODEX_HOME_ENV) or os.path.join(os.path.expanduser("~"), ".codex")
return os.path.join(home, "sessions")


def codex_rollout(payload: dict) -> str:
thread = payload.get("thread-id") or payload.get("thread_id")
if not isinstance(thread, str) or not THREAD_ID.match(thread):
return ""
pattern = os.path.join(codex_sessions_root(), "*", "*", "*", f"rollout-*-{thread}.jsonl")
matches = sorted(glob.glob(pattern))
return matches[-1] if matches else ""


def subagent_transcripts(transcript: str) -> list[str]:
if not transcript.endswith(".jsonl"):
return []
folder = os.path.join(transcript[: -len(".jsonl")], "subagents")
return sorted(glob.glob(os.path.join(folder, "*.jsonl")))
8 changes: 6 additions & 2 deletions engine/hooks/llm-judge/inbox.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
import os

import judge
from transcripts import codex_rollout, subagent_transcripts

NO_TRANSCRIPT = "llm-judge: {harness} payload has no transcript path, so finished verdicts were not checked"
REPORT_LIMIT = 600
Expand All @@ -31,7 +32,7 @@ def resolve_transcript(payload: dict) -> str:
candidate = os.path.join(root, project, "agent-transcripts", conv, f"{conv}.jsonl")
if os.path.isfile(candidate):
return candidate
return ""
return codex_rollout(payload)


def _turn_text(data: dict) -> str:
Expand Down Expand Up @@ -167,7 +168,10 @@ def unchecked_message(item: dict) -> str:

def messages(transcript: str) -> list[str]:
out = []
for item in judge.drain(transcript):
drained = []
for path in [transcript, *subagent_transcripts(transcript)]:
drained.extend(judge.drain(path))
for item in drained:
outcome = item.get("outcome")
if outcome == "clean":
continue
Expand Down
6 changes: 4 additions & 2 deletions engine/hooks/llm-judge/phrases/wrong-check-reflect.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"checker": "wrong-check-reflect",
"reads": "reply",
"meaning": "The latest assistant reply admits that something it told the user earlier was wrong, misread, or answered the wrong question.",
"meaning": "The latest assistant reply admits that something it told the user earlier was wrong, misread, or answered the wrong question, OR that an earlier claim of its own went out before the check backing it had run. A reply saying the claim turned out correct but was stated before it was checked IS a match: the failure is the order, not the answer, so no word in the reply has to say anything was wrong.",
"match": [
"my earlier check was wrong",
"You're right, I misread that",
Expand All @@ -10,7 +10,9 @@
"I was wrong about the path",
"Correction: the file I pointed you to earlier is not the one in use; the real one is src/b.py.",
"Good catch. The earlier number was off; the real count is 12.",
"My earlier read of the config was wrong \u2014 the default is 4, not 8."
"My earlier read of the config was wrong \u2014 the default is 4, not 8.",
"Correcting one claim and arming the check I implied:",
"I was right - but I said it a turn before I checked it"
],
"not_match": [
"You're right. Let's go with option B.",
Expand Down
73 changes: 73 additions & 0 deletions engine/hooks/llm-judge/tests/test_inbox.py
Original file line number Diff line number Diff line change
Expand Up @@ -355,5 +355,78 @@ def test_no_transcript_says_unchecked_on_stderr(self):
self.assertIn("no transcript path", self.run_codex([json.dumps({"type": "agent-turn-complete", "thread-id": "t1"})]))



REAL_CODEX_NOTIFY_KEYS = ["client", "cwd", "input-messages", "last-assistant-message", "thread-id", "turn-id", "type"]
THREAD = "01a0ce9a-301b-7d42-bb5c-b9a9a4cfe8c6"


class TestTranscriptResolution(InboxTestCase):
def setUp(self):
super().setUp()
sessions = os.path.join(self.work.name, "codex-sessions")
day = os.path.join(sessions, "2026", "09", "23")
os.makedirs(day)
self.rollout = os.path.join(day, f"rollout-2026-09-23T22-10-06-{THREAD}.jsonl")
with open(self.rollout, "w", encoding="utf-8") as handle:
handle.write("{}\n")
self.sessions_env = patch.dict(os.environ, {"CATSTACK_CODEX_SESSIONS_DIR": sessions})
self.sessions_env.start()

def tearDown(self):
self.sessions_env.stop()
super().tearDown()

def real_codex_payload(self):
payload = {
"type": "agent-turn-complete",
"thread-id": THREAD,
"turn-id": "01a0ce9a-30dc-73f1-bfe6-ccce644446e1",
"cwd": self.work.name,
"client": "codex_exec",
"input-messages": ["Reply with the single word ok."],
"last-assistant-message": "ok",
}
self.assertEqual(sorted(payload), REAL_CODEX_NOTIFY_KEYS)
return payload

def test_real_codex_notify_payload_resolves_to_its_rollout_file(self):
self.assertEqual(inbox.resolve_transcript(self.real_codex_payload()), self.rollout)

def test_hit_for_a_codex_rollout_reaches_the_next_codex_notify(self):
self.transcript = self.rollout
self.seed(ANSWERS_TRUE)
self.assertEqual(self.run_codex([json.dumps(self.real_codex_payload())]), ON_HIT + "\n")

def test_a_codex_home_override_is_where_the_rollout_is_looked_up(self):
home = os.path.join(self.work.name, "codex-home")
day = os.path.join(home, "sessions", "2026", "09", "23")
os.makedirs(day)
rollout = os.path.join(day, f"rollout-2026-09-23T22-40-00-{THREAD}.jsonl")
with open(rollout, "w", encoding="utf-8") as handle:
handle.write("{}\n")
self.sessions_env.stop()
try:
with patch.dict(os.environ, {"CODEX_HOME": home}, clear=False):
os.environ.pop("CATSTACK_CODEX_SESSIONS_DIR", None)
self.assertEqual(inbox.resolve_transcript(self.real_codex_payload()), rollout)
finally:
self.sessions_env.start()

def test_unknown_or_unsafe_thread_id_resolves_to_nothing(self):
for thread in ("0000000-0000-not-there", "../../etc", "*"):
with self.subTest(thread=thread):
self.assertEqual(inbox.resolve_transcript({"thread-id": thread}), "")

def test_a_subagent_verdict_is_delivered_to_the_parent_session(self):
subagent = os.path.join(self.transcript[: -len(".jsonl")], "subagents", "agent-a1.jsonl")
os.makedirs(os.path.dirname(subagent))
with open(subagent, "w", encoding="utf-8") as handle:
handle.write("{}\n")
parent = self.transcript
self.transcript = subagent
self.seed(ANSWERS_TRUE)
self.assertEqual(inbox.messages(parent), [ON_HIT])
self.assertEqual(inbox.messages(parent), [])

if __name__ == "__main__":
unittest.main()
6 changes: 6 additions & 0 deletions engine/hooks/unverified-tag-ledger/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,12 @@ fixtures in `tests/test_hooks.py`.
reason is written to stderr.
- **Escalation** — a claim outstanding `ESCALATE_AFTER_TURNS` (3) turns or more
is reported as a reflect trigger rather than accumulating quietly.
- **Discharge is itself a reflect trigger** — a row going outstanding ->
discharged is the record of a claim that went out first and was checked
after. That is an evidence-order miss, and it carries no wrongness word, so
the phrase scanners (`engine/skills/reflect/scripts/self_retraction_scan.py`,
and the `wrong-check-reflect` dictionary) cannot see it from the text. This
hook sees it from state instead, and says so on the Stop that discharges.

Malformed tags are deliberately ignored here; `diu-stop` already rejects those.

Expand Down
35 changes: 26 additions & 9 deletions engine/hooks/unverified-tag-ledger/detect.py
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,15 @@
ESCALATE_AFTER_TURNS = 3
MAX_LISTED = 5

DISCHARGE_REFLECT = (
"unverified-tag-ledger: {count} claim(s) went from unverified to checked this turn: "
"{claims}. That transition is the whole event: the claim went out first and the check "
"ran after. No wording has to admit anything for this to be true, which is why the "
"phrase scanners miss it -- an evidence-order miss carries no wrongness word. "
"Treat it as a reflect trigger, not a milestone: run reflect on this transcript, or "
"say plainly why this one does not need it."
)

CLAIM_RE = re.compile(
r"\{\{\s*CAT-UNVERIFIED\s*:?\s*(?P<claim>.*?)(?:--|—)\s*cannot\s+verify\s*:\s*(?P<reason>[^}]*)\}\}",
re.IGNORECASE | re.DOTALL,
Expand Down Expand Up @@ -177,18 +186,26 @@ def evaluate(payload: dict) -> dict:
session_id = str(payload.get("session_id") or "")
message = _last_assistant_text(payload)
tools = tools_used_this_turn(payload)
was_open = {row["claim"] for row in outstanding(read_ledger(session_id))}
rows = record_turn(session_id, message, tools)
notes = []
discharged = sorted(
row["claim"] for row in rows if row.get("resolved") and row["claim"] in was_open)
if discharged:
notes.append(DISCHARGE_REFLECT.format(
count=len(discharged), claims="; ".join(discharged[:MAX_LISTED])))

new_claims = {tag["claim"] for tag in parse_tags(message)}
if not new_claims:
return {"note": "", "block": ""}
return {"note": "\n".join(notes), "block": ""}

if tools is None:
return {"note": (
notes.append(
f"unverified-tag-ledger: logged {len(new_claims)} CAT-UNVERIFIED claim(s), but this "
"turn's tool calls could not be read from transcript_path (see the line above), so "
"whether a check was attempted is UNCHECKED, not clean. Nothing was discharged and "
"the turn was not refused."), "block": ""}
"the turn was not refused.")
return {"note": "\n".join(notes), "block": ""}

if not tools & VERIFY_TOOLS and not payload.get("stop_hook_active"):
claims = "; ".join(sorted(new_claims)[:MAX_LISTED])
Expand All @@ -201,12 +218,12 @@ def evaluate(payload: dict) -> dict:

fresh = [row for row in rows
if row["claim"] in new_claims and not row.get("resolved") and row.get("turns", 0) == 0]
if not fresh:
return {"note": "", "block": ""}
return {"note": (
f"unverified-tag-ledger: logged {len(fresh)} CAT-UNVERIFIED claim(s) against this session. "
"They are deferred, not discharged, and will be raised again next turn "
"(cat-mode/SKILL.md:269)."), "block": ""}
if fresh:
notes.append(
f"unverified-tag-ledger: logged {len(fresh)} CAT-UNVERIFIED claim(s) against this "
"session. They are deferred, not discharged, and will be raised again next turn "
"(cat-mode/SKILL.md:269).")
return {"note": "\n".join(notes), "block": ""}


def decide_stop(payload: dict) -> str:
Expand Down
Loading
Loading