Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 26 additions & 1 deletion corpus/skills/principle-subagent-inherits-scope/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,8 +62,33 @@ the work in a single round trip.
until verified — `engine/hooks/agent-relay-attribution` flags the shape.
- **State the scope in the prompt, not in your head.** An unstated boundary
is not inherited. Name the files, the write authority, and the question.
- **A brief carries three things, not two: facts, the question, and
decisions.** A decision the user already made is neither a fact to weigh
nor a question to answer. Relay it as a constraint sentence that says it is
settled — "the toggle gates the re-injection, not the emission; that is
decided, do not re-open it." Filed under the question, it reads as
something to work out, and a subagent that re-opens it looks rigorous while
discarding the only part the user owned outright. Repeating the user's
words is not enough on its own: a brief can carry the requirement verbatim
and still lose it by appending one open question beside it.
- **Don't prime the answer.** Hand over the facts and the question. A
subagent told what you expect finds roughly that.
subagent told what you expect finds roughly that. This governs findings,
never decisions. Leaving out a decision the user already made is not
neutrality — it is dropping a constraint, and the anti-priming rule then
rewards re-opening it.

Prior art for the third slot: Orlena Gotel and Anthony Finkelstein, "An
analysis of the requirements traceability problem", Proc. IEEE International
Conference on Requirements Engineering, https://doi.org/10.1109/ICRE.1994.292398
— pre-requirements-specification traceability exists so a requirement keeps
its link to the stakeholder who set it; without that link it gets
renegotiated by people who do not own it. Read this citation as
single-source: Crossref's `issued` field for the record is null, so the 1994
date is inferred from the DOI string and the conference rather than confirmed
by metadata, and Crossref renders the second author as "C.W. Finkelstein"
while the paper is normally cited as Anthony Finkelstein. IEEE Xplore
returned an empty body and ACM DL returned 403, so no publisher page was
read.

## Related

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -23,3 +23,18 @@ and conventions are not concurrency control, so a read-only brief is not
filesystem isolation. A subagent told to write files runs in its own worktree
even when the rest of the prompt defaults to read-only, and the parent that
omitted one has not stated the boundary at all.

A third shape, the one the decisions slot exists for. The user has already
decided which stage a new toggle gates. The parent relays that requirement
word for word and then appends "work out what a toggle would actually gate."
A mechanical containment check on the delegation passes — the requirement is
verbatim-contained and every content word is present — and the subagent still
ranks the user's own requirement fourth of six and argues its premise away.
The skill fires here because the decision was relayed as part of the
question instead of as a settled constraint, and because the anti-priming
rule reads re-opening it as rigour.

No mechanical catch is claimed for that third shape, and the obvious one is
known not to work: the containment check returns PASS on the exact
delegation that drifted, because the words were all there. The gate is the
parent's wording, so this stays an `unchecked` case pinned by prose.
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,11 @@ whose authority could exceed the parent's. The principle has no target: its
four limits all describe what a delegate may do, and the contradiction
contract describes how a delegate reports back. A single agent editing one
file on its own behalf is the case this principle is not about.

A second silent shape, against the decisions slot. A parent hands an explorer
a read-only brief that names the files, the question, and one settled
decision relayed as a constraint the explorer may not re-open. The explorer
reads those files, answers the question, and reports one contradiction with a
`file:line` and the ref it was read at, leaving the decision alone. Nothing
here is a widening: the boundary was stated, the constraint travelled as a
constraint, and the contradiction went back up rather than being acted on.
19 changes: 15 additions & 4 deletions engine/hooks/llm-judge/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -180,14 +180,25 @@ one into a line of text:
- **hit**: the job's `on_hit` text, word for word, followed by a space and the
answer's `report` string when the answer has a non-blank one, clipped to 600
characters.
- **unchecked**: `llm-judge: <hook> could not judge the last reply: ` then
`<runner>: <reason>` for each try, joined by `; `. If there were no tries
(the judge broke, or the verdict file was unreadable), the verdict's own
`reason` is used instead.
- **unchecked**: `llm-judge UNCHECKED: <hook> could not judge the last
reply, so that reply is unchecked, not clean.` It then says the check fails
open, tells the agent to tell the user the check did not run, and ends with
`Tried: ` and `<runner>: <reason>` for each try, joined by `; `. If there
were no tries (the judge broke, or the verdict file was unreadable), the
verdict's own `reason` is used instead.
- **clean**: nothing.

Each verdict is delivered once. Draining deletes it.

An unchecked verdict fails open: the judge runs after the reply is sent, so it
never holds that reply up, and a judge that cannot run cannot hold it up
either. What it must not do is read as clean. So on Claude, when any drained
verdict is unchecked, the hook also sets `systemMessage` to
`llm-judge: <n> check(s) did not run and failed open, so the replies they cover
are unchecked, not clean: <hooks>`. That line is shown to the user directly, so
it reaches them even if the agent drops the context line. The verdict event is
also written with action `unchecked`.

One small script per harness calls it:

| Harness | Script | Event | How the text reaches the agent |
Expand Down
11 changes: 8 additions & 3 deletions engine/hooks/llm-judge/claude_post_tool_use.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,12 +20,17 @@ def main() -> None:
print(inbox.NO_TRANSCRIPT.format(harness=HARNESS), file=sys.stderr)
return
try:
found = inbox.messages(transcript)
found, unchecked = inbox.report(transcript)
except Exception as exc:
print(f"llm-judge: {HARNESS} could not drain verdicts for {transcript}: {type(exc).__name__}: {exc}", file=sys.stderr)
return
if found:
print(json.dumps({"hookSpecificOutput": {"hookEventName": "PostToolUse", "additionalContext": "\n\n".join(found)}}))
if not found:
return
output = {"hookSpecificOutput": {"hookEventName": "PostToolUse", "additionalContext": "\n\n".join(found)}}
notice = inbox.user_notice(unchecked)
if notice:
output["systemMessage"] = notice
print(json.dumps(output))


if __name__ == "__main__":
Expand Down
11 changes: 8 additions & 3 deletions engine/hooks/llm-judge/claude_prompt_submit.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,12 +18,17 @@ def main() -> None:
print(inbox.NO_TRANSCRIPT.format(harness="Claude UserPromptSubmit"), file=sys.stderr)
return
try:
found = inbox.messages(transcript)
found, unchecked = inbox.report(transcript)
except Exception as exc:
print(f"llm-judge: could not drain verdicts for {transcript}: {type(exc).__name__}: {exc}", file=sys.stderr)
return
if found:
print(json.dumps({"hookSpecificOutput": {"hookEventName": "UserPromptSubmit", "additionalContext": "\n\n".join(found)}}))
if not found:
return
output = {"hookSpecificOutput": {"hookEventName": "UserPromptSubmit", "additionalContext": "\n\n".join(found)}}
notice = inbox.user_notice(unchecked)
if notice:
output["systemMessage"] = notice
print(json.dumps(output))


if __name__ == "__main__":
Expand Down
35 changes: 29 additions & 6 deletions engine/hooks/llm-judge/inbox.py
Original file line number Diff line number Diff line change
Expand Up @@ -159,15 +159,33 @@ def enqueue_judge(job_fields: dict, payload: dict) -> str | None:
return judge.enqueue(job)


UNCHECKED = (
"llm-judge UNCHECKED: {hook} could not judge the last reply, so that reply is unchecked, "
"not clean. This check fails open: the reply was already sent and was not held. "
"Tell the user this check did not run before relying on that reply. Tried: {tried}"
)
USER_NOTICE = (
"llm-judge: {count} check(s) did not run and failed open, so the replies they cover "
"are unchecked, not clean: {hooks}"
)


def unchecked_message(item: dict) -> str:
hook = item.get("hook") or "unknown hook"
attempts = item.get("attempts") or []
tried = "; ".join(f"{a.get('runner')}: {a.get('reason')}" for a in attempts if isinstance(a, dict))
return f"llm-judge: {hook} could not judge the last reply: {tried or item.get('reason') or 'no reason recorded'}"
return UNCHECKED.format(hook=hook, tried=tried or item.get("reason") or "no reason recorded")


def messages(transcript: str) -> list[str]:
def user_notice(unchecked_hooks: list[str]) -> str | None:
if not unchecked_hooks:
return None
return USER_NOTICE.format(count=len(unchecked_hooks), hooks=", ".join(sorted(set(unchecked_hooks))))


def report(transcript: str) -> tuple[list[str], list[str]]:
out = []
unchecked = []
drained = []
for path in [transcript, *subagent_transcripts(transcript)]:
drained.extend(judge.drain(path))
Expand All @@ -181,10 +199,15 @@ def messages(transcript: str) -> list[str]:
text = f"llm-judge: {item.get('hook') or 'unknown hook'} flagged the last reply: {item.get('reason')}"
answer = item.get("answer")
if isinstance(answer, dict):
report = answer.get("report")
if isinstance(report, str) and report.strip():
text = f"{text} {report.strip()[:REPORT_LIMIT]}"
detail = answer.get("report")
if isinstance(detail, str) and detail.strip():
text = f"{text} {detail.strip()[:REPORT_LIMIT]}"
out.append(text)
continue
out.append(unchecked_message(item))
return out
unchecked.append(str(item.get("hook") or "unknown hook"))
return out, unchecked


def messages(transcript: str) -> list[str]:
return report(transcript)[0]
50 changes: 43 additions & 7 deletions engine/hooks/llm-judge/tests/test_inbox.py
Original file line number Diff line number Diff line change
Expand Up @@ -236,9 +236,30 @@ def test_clean_with_report_yields_nothing(self):

def test_unchecked_yields_one_reason_per_runner(self):
self.assertEqual(self.seed(MISSING, CRASHES)["outcome"], "unchecked")
found = inbox.messages(self.transcript)
self.assertEqual(len(found), 1)
self.assertTrue(found[0].endswith("Tried: ghost: not installed; crashes: exit 3: model quota exhausted"), found)

def test_unchecked_says_unchecked_not_clean_and_names_the_fail_direction(self):
self.seed(MISSING, CRASHES)
found = inbox.messages(self.transcript)
self.assertTrue(found[0].startswith("llm-judge UNCHECKED: demo-hook could not judge the last reply"), found)
self.assertIn("unchecked, not clean", found[0])
self.assertIn("fails open", found[0])
self.assertIn("Tell the user this check did not run", found[0])

def test_report_lists_unchecked_hooks_apart_from_hits(self):
self.seed(ANSWERS_TRUE, job_id="a")
self.seed(MISSING, job_id="b")
found, unchecked = inbox.report(self.transcript)
self.assertEqual(len(found), 2)
self.assertEqual(unchecked, ["demo-hook"])

def test_user_notice_is_none_when_nothing_was_unchecked(self):
self.assertIsNone(inbox.user_notice([]))
self.assertEqual(
inbox.messages(self.transcript),
["llm-judge: demo-hook could not judge the last reply: ghost: not installed; crashes: exit 3: model quota exhausted"],
inbox.user_notice(["b-hook", "a-hook", "b-hook"]),
"llm-judge: 3 check(s) did not run and failed open, so the replies they cover are unchecked, not clean: a-hook, b-hook",
)

def test_clean_yields_nothing(self):
Expand All @@ -256,7 +277,8 @@ def test_unreadable_verdict_file_is_unchecked_not_clean(self):
handle.write("{not json")
found = inbox.messages(self.transcript)
self.assertEqual(len(found), 1)
self.assertTrue(found[0].startswith("llm-judge: unknown hook could not judge the last reply: unreadable verdict file"), found)
self.assertTrue(found[0].startswith("llm-judge UNCHECKED: unknown hook could not judge the last reply"), found)
self.assertIn("Tried: unreadable verdict file", found[0])

def test_verdicts_for_another_transcript_are_not_delivered(self):
self.seed(ANSWERS_TRUE)
Expand All @@ -270,12 +292,26 @@ def test_hit_is_delivered_as_additional_context_once(self):
self.seed(MISSING, job_id="b")
out, err = self.run_claude(self.claude_payload())
self.assertEqual(err, "")
self.assertEqual(json.loads(out), {"hookSpecificOutput": {
"hookEventName": "UserPromptSubmit",
"additionalContext": ON_HIT + "\n\nllm-judge: demo-hook could not judge the last reply: ghost: not installed",
}})
data = json.loads(out)
self.assertEqual(data["hookSpecificOutput"]["hookEventName"], "UserPromptSubmit")
context = data["hookSpecificOutput"]["additionalContext"]
self.assertTrue(context.startswith(ON_HIT + "\n\nllm-judge UNCHECKED: demo-hook"), context)
self.assertTrue(context.endswith("Tried: ghost: not installed"), context)
self.assertEqual(self.run_claude(self.claude_payload()), ("", ""))

def test_unchecked_is_also_shown_to_the_user_as_a_system_message(self):
self.seed(MISSING)
out, err = self.run_claude(self.claude_payload())
self.assertEqual(err, "")
data = json.loads(out)
self.assertIn("llm-judge UNCHECKED: demo-hook", data["hookSpecificOutput"]["additionalContext"])
self.assertEqual(data["systemMessage"], inbox.user_notice(["demo-hook"]))

def test_hit_alone_adds_no_system_message(self):
self.seed(ANSWERS_TRUE)
data = json.loads(self.run_claude(self.claude_payload())[0])
self.assertNotIn("systemMessage", data)

def test_clean_prints_nothing(self):
self.seed(ANSWERS_FALSE)
self.assertEqual(self.run_claude(self.claude_payload()), ("", ""))
Expand Down
12 changes: 12 additions & 0 deletions engine/hooks/llm-judge/tests/test_post_tool_use.py
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,18 @@ def test_malformed_stdin_exits_zero_with_a_stderr_line(self):
self.assert_bad_json(self.script, self.harness)


def test_unchecked_verdict_adds_a_system_message_for_the_user(self):
judge.write_json_atomic(
os.path.join(judge.verdict_dir(self.transcript), "job-u.json"),
{"id": "job-u", "hook": "demo-hook", "outcome": "unchecked", "reason": "no runner answered", "attempts": []},
)
result = self.run_script(self.script, self.payload())
self.assert_success(result)
data = json.loads(result.stdout)
self.assertIn("llm-judge UNCHECKED: demo-hook", data["hookSpecificOutput"]["additionalContext"])
self.assertIn("failed open", data["systemMessage"])


class TestCodexPostToolUse(PostToolUseTestCase):
script = "codex_post_tool_use.py"
harness = "Codex PostToolUse"
Expand Down
Loading