Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions corpus/skills/cat-mode/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,13 @@ CLAUDE.md's evidence rules already apply here. Also, don't declare something fix

**UI testing must not disrupt the user's own session.** Prove a UI or surface change somewhere disposable — a test channel or workspace, a throwaway profile, a second display, a VM, a headless run. Driving the user's real keyboard, mouse, or screen is a last resort needing an explicit hands-off window first: state the acceptance test in one line, get the yes, `touch /tmp/.ui-input-window`, and remove it when the window closes; a PreToolUse hook (`engine/hooks/ui-input-guard/`) blocks synthetic input and screen recording while no window is open, the screen is locked, or the user is still typing. Stop at the first sign the session is theirs again (idle time drops, the frontmost app changes, the screen locks), and leave no residue: undo stray messages, pins, or reactions, or say what was left behind.

**Visual Proof authenticity** (extends [[visual-proof]] / [[principle-prove-it]]):

- **The Visual Proof surface must match the Review Claim surface.** A claim about one product surface needs pixels from that surface (e.g. a Slack-thread claim → Slack-thread pixels). A different product's screen, a provider login page, or an adjacent flow is not that proof.
- **Declare Expected surface and Expected predicates before capture.** Write what must be visible and what must not appear; only then capture. `Manually inspected:` checks claim↔pixels against that Expected list by reading the image — a marker-only line is not a check.
- **Never submit synthesized UI as Visual Proof** unless the user asked for a mockup: generated text slides, HTML mock surfaces, reconstructed controls, or redrawn UI do not count.
- **When a Review Claim covers multiple major behavioral cases, Visual Proof is not done until each major case has its own UI proof media — or an explicit waiver naming the skipped case.** OR claims need one capture per disjunct; one case's pixels do not prove another. Declare Expected cases and Expected predicates before capture.

**A factual or technical claim gets a real repro script, not a history search.** Judging an old comment or a "probably confabulated" suspicion needs an actual attempt under the claimed conditions, not a `git log` sweep. No citation means "never verified," not "false."

**Unhedged root-cause or fix claims about live system behavior need instrument-level proof in the same message, or a `{{CAT-UNVERIFIED: <claim> -- cannot verify: <reason>}}` tag naming the blocker.** The gate is the claim type, not a hedge word. Invoking `/prove-it` once does not arm it for later claims. Any hedge auto-runs prove-it in the same turn — a hedge is a trigger to verify, never a place to stop.
Expand Down
41 changes: 41 additions & 0 deletions corpus/skills/cat-mode/references/verify.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,3 +89,44 @@ hang. Invoking `/prove-it` once does not arm it for later claims — each new
causal claim needs its own same-message evidence. Any hedge — "I think,"
"probably," a retired bare `UNVERIFIED:` — auto-runs prove-it in the same turn; a hedge is a
trigger to verify, never a place to stop.

## Visual Proof authenticity

Personal standing rules on top of [[visual-proof]] and [[principle-prove-it]].
These extend those skills; they harden authenticity when a near-neighbor
capture would otherwise stand in for the claimed surface.

- **The Visual Proof surface must match the Review Claim surface.** The claim
names which product surface must be proved. Capture that surface's pixels —
not a different product, not an upstream provider login, not an adjacent
step that "looks related." A Slack-thread claim needs Slack-thread pixels; a Claude or OpenAI login page is not Slack UI proof.
- **Declare Expected surface and Expected predicates before capture.** Before
any screenshot or frame grab, write (1) the Expected surface and (2) the
Expected predicates: what must be visible, and what must not appear. Capture
only after that list exists. Then `Manually inspected:` walks claim↔pixels
against that list by reading the image (or extracted frames). A bare
`Manually inspected:` marker with no predicate check is not a check.
- **Never submit synthesized UI as Visual Proof** unless the user explicitly
asked for a mockup. Generated text slides (e.g. ffmpeg lavfi/drawtext), HTML mock surfaces,
reconstructed controls, and redrawn UI prove only that the generator ran —
not that the claimed surface showed the claimed state.

## Visual Proof case coverage

Personal standing rules on top of [[visual-proof]] and [[principle-prove-it]].
When a Review Claim or feature names multiple major behavioral cases —
disjuncts joined by OR — Visual Proof is incomplete until every named case
has its own UI proof media, or an explicit waiver that names the skipped
case.

- **One capture covers one case.** Pixels that prove one behavioral case
(e.g. a usage-limit Slack thread) do not prove a different case the claim
also covers (e.g. authentication-needed). Treat each major case as its
own done-gate.
- **Declare Expected cases and Expected predicates before capture.** List
every major case the claim covers, and for each case what must be visible
and what must not appear. Capture only after that list exists. Then
inspect claim↔pixels per case against that list.
- **An incomplete set is not done.** Shipping with proof for a subset of
the claim's major cases, without a waiver naming each missing case, is
an unfinished Visual Proof — not a partial success.
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
`disable-model-invocation: true` means the model never reads this
skill's `description:` to decide whether to apply it. cat-mode applies
only when the `CATSTACK_CAT_MODE_DEFAULT=on` hook fires, or on an
explicit `/cat-mode` invocation. Here the hook is on.

The Review Claim covers two major behavioral cases joined by OR:
usage-limit and authentication-needed. The agent never writes Expected
cases or Expected predicates. It captures only a usage-limit Slack
screenshot, uploads that as Visual Proof, and treats the claim as
proved.

This rule fires. One case's pixels do not prove another; Visual Proof
is not done until each major case has its own UI proof media or an
explicit waiver naming the skipped case. The correct next step is
declare Expected cases + predicates for both disjuncts, capture (or
waive) each, then inspect claim↔pixels per case.
16 changes: 16 additions & 0 deletions corpus/skills/cat-mode/tests/fires_visual_proof_wrong_surface.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
`disable-model-invocation: true` means the model never reads this
skill's `description:` to decide whether to apply it. cat-mode applies
only when the `CATSTACK_CAT_MODE_DEFAULT=on` hook fires, or on an
explicit `/cat-mode` invocation. Here the hook is on.

The Review Claim is about a Slack-thread UI change. The agent never
writes Expected surface or Expected predicates. It captures a Claude
(or OpenAI) provider login page, uploads that image as Visual Proof,
and writes a marker-only `Manually inspected:` line with no claim↔pixels
check against any Expected list.

This rule fires. The Visual Proof surface does not match the Review
Claim surface, Expected predicates were never declared before capture,
and a marker-only inspection line is not a check. The correct next
step is declare Expected surface + predicates for the Slack thread,
capture that thread's pixels, then inspect against the Expected list.
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
`disable-model-invocation: true` means the model never reads this
skill's `description:` to decide whether to apply it. cat-mode applies
only when the `CATSTACK_CAT_MODE_DEFAULT=on` hook fires, or on an
explicit `/cat-mode` invocation. Here the hook is on.

The Review Claim covers two major behavioral cases joined by OR:
usage-limit and authentication-needed. Before capture the agent writes
Expected cases (both disjuncts) and Expected predicates for each. It
then captures UI proof media for each case from the live surface — or
records an explicit waiver naming any skipped case — and inspects
claim↔pixels per case against that Expected list.

This skill's Visual Proof case-coverage rules stay silent: every major
case is covered or waived by name, Expected was declared before
capture, and inspection checked the Expected predicates per case.
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
`disable-model-invocation: true` means the model never reads this
skill's `description:` to decide whether to apply it. cat-mode applies
only when the `CATSTACK_CAT_MODE_DEFAULT=on` hook fires, or on an
explicit `/cat-mode` invocation. Here the hook is on.

The Review Claim is about a Slack-thread UI change. Before capture the
agent writes Expected surface (the Slack thread) and Expected
predicates (what must be visible / must not appear). It then captures
that Slack thread's pixels from the live surface — not a mock HTML UI,
not a generated text slide — reads the image, and writes
`Manually inspected:` that walks claim↔pixels against the Expected
list.

This skill's Visual Proof authenticity rules stay silent: surface
matches claim, Expected was declared before capture, the artifact is
real pixels from that surface, and inspection checked the Expected
predicates.
39 changes: 39 additions & 0 deletions tests/test_cat_mode.py
Original file line number Diff line number Diff line change
Expand Up @@ -129,6 +129,45 @@ def test_requires_cleanup_of_what_the_run_left_behind(self):
self.assertTrue("residue" in text or "undo stray" in text, "cleanup rule missing")



class TestVisualProofAuthenticity(unittest.TestCase):
"""Standing authenticity defaults for Visual Proof under Verify.

Agents must match proof surface to claim, declare Expected predicates
before capture, reject synthesized UI as proof, and treat a marker-only
Manually inspected line as unchecked.
"""

VERIFY_REF = os.path.join(
REPO_ROOT, "corpus", "skills", "cat-mode", "references", "verify.md"
)

def test_skill_names_surface_match_and_expected_before_capture(self):
text = read_skill_text()
self.assertIn("Visual Proof authenticity", text)
self.assertIn("Visual Proof surface must match the Review Claim surface", text)
self.assertIn("Declare Expected surface and Expected predicates before capture", text)
self.assertIn("marker-only line is not a check", text)
self.assertIn("Never submit synthesized UI as Visual Proof", text)
self.assertIn("multiple major behavioral cases", text)
self.assertIn("each major case has its own UI proof media", text)

def test_verify_reference_carries_full_predicates(self):
with open(self.VERIFY_REF, encoding="utf-8") as handle:
text = handle.read()
self.assertIn("## Visual Proof authenticity", text)
self.assertIn("Expected surface", text)
self.assertIn("## Visual Proof case coverage", text)
self.assertIn("One capture covers one case", text)
self.assertIn("Expected predicates", text)
self.assertIn("Manually inspected:", text)
self.assertIn("ffmpeg lavfi/drawtext", text)
self.assertIn("HTML", text)
self.assertIn("mock surfaces", text)
self.assertIn("unless the user explicitly", text)
self.assertIn("asked for a mockup", text)


class TestCatModeReferences(unittest.TestCase):
def test_every_referenced_skill_still_exists(self):
text = read_skill_text()
Expand Down
Loading