Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion corpus/skills/cat-mode/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -265,7 +265,7 @@ agent switch, or resubmit is a fix, and none comes before the repro.

**A factual or technical claim gets a real repro script, not a history search.** Judging an old comment or a "probably confabulated" suspicion needs an actual attempt under the claimed conditions, not a `git log` sweep. No citation means "never verified," not "false."

**Unhedged root-cause or fix claims about live system behavior need instrument-level proof in the same message, or a `{{CAT-UNVERIFIED}}` tag naming the blocker.** The gate is the claim type, not a hedge word. Invoking `/prove-it` once does not arm it for later claims. Any hedge auto-runs prove-it in the same turn — a hedge is a trigger to verify, never a place to stop.
**Unhedged root-cause or fix claims about live system behavior need instrument-level proof in the same message, or a `{{CAT-UNVERIFIED: <claim> -- cannot verify: <reason>}}` tag naming the blocker.** The gate is the claim type, not a hedge word. Invoking `/prove-it` once does not arm it for later claims. Any hedge auto-runs prove-it in the same turn — a hedge is a trigger to verify, never a place to stop.

Outputs carry failures explicitly (a status column, an error row), never
dropped — [[principle-explicit-errors]].
Expand Down
27 changes: 26 additions & 1 deletion corpus/skills/principle-subagent-inherits-scope/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,8 +61,33 @@ the work in a single round trip.
until verified — `engine/hooks/agent-relay-attribution` flags the shape.
- **State the scope in the prompt, not in your head.** An unstated boundary
is not inherited. Name the files, the write authority, and the question.
- **A brief carries three things, not two: facts, the question, and
decisions.** A decision the user already made is neither a fact to weigh
nor a question to answer. Relay it as a constraint sentence that says it is
settled — "the toggle gates the re-injection, not the emission; that is
decided, do not re-open it." Filed under the question, it reads as
something to work out, and a subagent that re-opens it looks rigorous while
discarding the only part the user owned outright. Repeating the user's
words is not enough on its own: a brief can carry the requirement verbatim
and still lose it by appending one open question beside it.
- **Don't prime the answer.** Hand over the facts and the question. A
subagent told what you expect finds roughly that.
subagent told what you expect finds roughly that. This governs findings,
never decisions. Leaving out a decision the user already made is not
neutrality — it is dropping a constraint, and the anti-priming rule then
rewards re-opening it.

Prior art for the third slot: Orlena Gotel and Anthony Finkelstein, "An
analysis of the requirements traceability problem", Proc. IEEE International
Conference on Requirements Engineering, https://doi.org/10.1109/ICRE.1994.292398
— pre-requirements-specification traceability exists so a requirement keeps
its link to the stakeholder who set it; without that link it gets
renegotiated by people who do not own it. Read this citation as
single-source: Crossref's `issued` field for the record is null, so the 1994
date is inferred from the DOI string and the conference rather than confirmed
by metadata, and Crossref renders the second author as "C.W. Finkelstein"
while the paper is normally cited as Anthony Finkelstein. IEEE Xplore
returned an empty body and ACM DL returned 403, so no publisher page was
read.

## Related

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -20,3 +20,18 @@ grounds that it was told not to. The link to
and conventions are not concurrency control, so a read-only brief is not
filesystem isolation. A subagent that may write gets its own worktree, and the
parent that omitted one has not stated the boundary at all.

A third shape, the one the decisions slot exists for. The user has already
decided which stage a new toggle gates. The parent relays that requirement
word for word and then appends "work out what a toggle would actually gate."
A mechanical containment check on the delegation passes — the requirement is
verbatim-contained and every content word is present — and the subagent still
ranks the user's own requirement fourth of six and argues its premise away.
The skill fires here because the decision was relayed as part of the
question instead of as a settled constraint, and because the anti-priming
rule reads re-opening it as rigour.

No mechanical catch is claimed for that third shape, and the obvious one is
known not to work: the containment check returns PASS on the exact
delegation that drifted, because the words were all there. The gate is the
parent's wording, so this stays an `unchecked` case pinned by prose.
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,11 @@ whose authority could exceed the parent's. The principle has no target: its
four limits all describe what a delegate may do, and the contradiction
contract describes how a delegate reports back. A single agent editing one
file on its own behalf is the case this principle is not about.

A second silent shape, against the decisions slot. A parent hands an explorer
a read-only brief that names the files, the question, and one settled
decision relayed as a constraint the explorer may not re-open. The explorer
reads those files, answers the question, and reports one contradiction with a
`file:line` and the ref it was read at, leaving the decision alone. Nothing
here is a widening: the boundary was stated, the constraint travelled as a
constraint, and the contradiction went back up rather than being acted on.
20 changes: 20 additions & 0 deletions engine/hooks/diu-stop/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,26 @@ last appeared before however many compactions have happened since.

The claim check reads only the main agent's turn-final message: of 337 unproven claims found in stored transcripts, 196 were mid-turn or subagent text it never saw. See [`COVERAGE.md`](COVERAGE.md) before reading its silence as clearance.

## What buys a paragraph its silence

A fenced block of output, inline code that looks like output, a well-formed
`{{CAT-UNVERIFIED: ... -- cannot verify: ...}}` tag, or a file citation --
and a citation has to be backed. `path:line` on its own used to silence a
paragraph with no check that the file existed, that anyone read it, or at
what ref; a made-up path silenced the gate exactly as well as a real one. A
citation now counts when it names the ref it was read at (`path:line @
origin/main`, the form `corpus/CLAUDE.learned.md` already asks for in prose),
or when the session's transcript shows a tool call that named that path.

Three outcomes, not two. When the transcript cannot be read, whether the path
was read is *unchecked*: the citation does not buy silence, and the block says
which path it could not check and that the ref would settle it.

The marker check reads prose only. A marker inside a fence or a pair of
backticks is being shown, not used, so explaining the tag, quoting the rule
that defines it, or relaying this gate's own refusal word for word all stay
silent. A marker in running prose is a use and still counts.

Not one file per harness, because there is no single "stop" mechanism
shared by every harness -- each one has a genuinely different amount of
power at that point:
Expand Down
141 changes: 129 additions & 12 deletions engine/hooks/diu-stop/claude_stop_check.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@
Every block names every flagged sentence, so one rewrite that fixes them
all gets through.
"""
import json
import os
import re
import sys
Expand Down Expand Up @@ -90,6 +91,9 @@
FENCED_BODY_RE = re.compile(r"```[^\n]*\n(.*?)```", re.DOTALL)
INLINE_CODE_RE = re.compile(r"`([^`\n]+)`")
FILE_LINE_RE = re.compile(r"(?<![\w/.-])(?:[\w.-]+/)*[\w-]+\.[A-Za-z]\w*(?::\d+\b|#L\d+\b)")
CITATION_REF_RE = re.compile(
r"(?<![\w/.-])(?:[\w.-]+/)*[\w-]+\.[A-Za-z]\w*(?::\d+\b|#L\d+\b)\s*@\s*\S+")
LINE_SUFFIX_RE = re.compile(r":\d+\b|#L\d+\b")
OUTPUT_SHAPE_RE = re.compile(
r"^(?:\$ |> |\+\+\+ |--- |@@ |diff --git|commit [0-9a-f]{7,}|[0-9a-f]{7,10} )"
r"|Traceback|^\s*at [\w.$<>]+ \(.*:\d+:\d+\)"
Expand Down Expand Up @@ -144,15 +148,37 @@ def _opening_word(message):
return match.group(0).lower() if match else ""


def prose_only(message):
"""`message` with fenced blocks and inline code removed.

A marker inside a fence or a pair of backticks is being shown, not used:
explaining the tag, quoting the rule that defines it, or pasting a gate's
own message back to the user all put the token on screen without claiming
anything. `find_unverified_claims` has stripped both for a while; the
marker check read the raw message, so the gate fired on the sentence that
taught the reader how not to trip it.

An unterminated fence leaves a `\u0060\u0060\u0060` behind after the
substitution. Everything from that marker on is inside a code block that
never closed, so it is dropped too.
"""
prose = INLINE_CODE_RE.sub("", FENCED_BODY_RE.sub("", message or ""))
if FENCE_MARKER in prose:
prose = prose[:prose.rindex(FENCE_MARKER)]
return prose


def find_marker_problems(message):
"""Return the marker complaints this message earns, in report order.

A tag that names no blocker, and the retired bare `UNVERIFIED:`, each
draw their own message. Both can be present at once."""
draw their own message. Both can be present at once. Only prose counts --
see `prose_only`."""
prose = prose_only(message)
problems = []
if markers.malformed_tags(message):
if markers.malformed_tags(prose):
problems.append(markers.MALFORMED_TAG_MESSAGE)
if markers.has_legacy_marker(message):
if markers.has_legacy_marker(prose):
problems.append(markers.LEGACY_MARKER_MESSAGE)
return problems

Expand Down Expand Up @@ -189,7 +215,76 @@ def _paragraph_claim(para):
return None


def find_unverified_claims(message):
def cited_paths(text):
"""The path part of every file:line in `text`, longest first."""
seen = []
for match in FILE_LINE_RE.finditer(text):
path = LINE_SUFFIX_RE.split(match.group(0))[0]
if path and path not in seen:
seen.append(path)
return sorted(seen, key=len, reverse=True)


def read_evidence(event):
"""Everything this session handed a tool, as one string, or None.

None means the check could not run -- no transcript to read, or the file
would not open. That is a third outcome, not a clean one: a citation whose
read cannot be checked does not buy silence, and the finding says why.
"""
path = event.get("transcript_path") if isinstance(event, dict) else None
if not isinstance(path, str) or not path:
return None
parts = []
try:
with open(path, encoding="utf-8") as handle:
for line in handle:
line = line.strip()
if not line:
continue
try:
entry = json.loads(line)
except json.JSONDecodeError:
continue
if not isinstance(entry, dict):
continue
inner = entry.get("message")
content = inner.get("content") if isinstance(inner, dict) else entry.get("content")
if not isinstance(content, list):
continue
for block in content:
if isinstance(block, dict) and block.get("type") == "tool_use":
parts.append(json.dumps(block.get("input"), default=str))
except (OSError, UnicodeError) as exc:
print(
f"catstack-hook-error diu-stop: cannot read {path}, so which files "
f"were read this session is unchecked: {exc}",
file=sys.stderr,
)
return None
return "\n".join(parts)


def citation_earns_silence(para, read_blob):
"""(exempt, unchecked) for the file:line citations in one paragraph.

A bare `file.ts:99` used to silence a paragraph on its own, with no check
that the file exists, that anyone read it, or at what ref. A fabricated
path silenced the gate exactly as well as a real one. It now has to carry
the ref it was read at (`path:line @ origin/main`), which is what
corpus/CLAUDE.learned.md already asks for in prose, or the session has to
show a tool call that named that path.
"""
if not FILE_LINE_RE.search(para):
return False, False
if CITATION_REF_RE.search(para):
return True, False
if read_blob is None:
return False, True
return any(path in read_blob for path in cited_paths(para)), False


def find_unverified_claims(message, read_blob=None):
"""Return one (trigger phrase, sentence) pair for every paragraph that
makes an unverified-shaped claim with no evidence marker in that same
paragraph, in message order.
Expand All @@ -209,7 +304,8 @@ def find_unverified_claims(message):
continue
if markers.excuses_paragraph(para):
continue
if FILE_LINE_RE.search(para):
exempt, _unchecked = citation_earns_silence(para, read_blob)
if exempt:
continue
inline = INLINE_CODE_RE.findall(para)
if inline and (fenced_output or any(OUTPUT_SHAPE_RE.search(code) for code in inline)):
Expand All @@ -221,13 +317,26 @@ def find_unverified_claims(message):
return claims


def find_unverified_claim(message):
def find_unverified_claim(message, read_blob=None):
"""Return the first offending phrase find_unverified_claims reports, or
None."""
claims = find_unverified_claims(message)
claims = find_unverified_claims(message, read_blob)
return claims[0][0] if claims else None


def unchecked_citations(message, read_blob):
"""Paths cited in a flagged paragraph whose read could not be checked."""
if read_blob is not None:
return []
found = []
for para in re.split(r"\n\s*\n", message):
para = FENCED_BODY_RE.sub("", para)
_exempt, unchecked = citation_earns_silence(para, read_blob)
if unchecked:
found.extend(path for path in cited_paths(para) if path not in found)
return found


def detect(event):
if event.get("agent_id"):
return []
Expand All @@ -239,7 +348,8 @@ def detect(event):

word_count = counted_words(message)
over_limit = word_count > WORD_LIMIT and not retry
claims = find_unverified_claims(message)
read_blob = read_evidence(event)
claims = find_unverified_claims(message, read_blob)
marker_problems = find_marker_problems(message)

findings = []
Expand All @@ -260,11 +370,18 @@ def detect(event):
for number, (phrase, sentence) in enumerate(claims, 1):
lines.append(f"{number}. \"{sentence}\" (trigger: \"{' '.join(phrase.split())}\")")
lines.append(
"A backticked name or command alone is not output. Per "
"skills/prove-it/SKILL.md: for each one, either paste the output "
"of what was actually run/checked in its paragraph, or -- only if "
"the check cannot run -- tag the claim there and say why."
"A backticked name or command alone is not output, and neither is a "
"bare file:line. Per skills/prove-it/SKILL.md: for each one, either "
"paste the output of what was actually run/checked in its paragraph, "
"cite it as `path:line @ <ref>`, or -- only if the check cannot run "
"-- tag the claim there and say why."
)
unchecked = unchecked_citations(message, read_blob)
if unchecked:
lines.append(
"This turn's transcript could not be read, so whether "
f"{', '.join(unchecked)} was read this session is UNCHECKED, not "
"clear. Add the ref it was read at to the citation.")
claim_message = "\n".join(lines)
findings.append(Finding(
rule_id=RULE_UNVERIFIED_CLAIM,
Expand Down
Loading
Loading