Skip to content

cat-mode: Visual Proof authenticity defaults (Expected before capture) - #1207

Merged
mergify[bot] merged 2 commits into
mainfrom
stack/EdbertChan/automate-me/visual-proof-expected-20260928/cat-mode-require-expected-predicates-visual-proof--b9421a35
Sep 29, 2026
Merged

mergify[bot] merged 2 commits into
mainfrom
stack/EdbertChan/automate-me/visual-proof-expected-20260928/cat-mode-require-expected-predicates-visual-proof--b9421a35

Conversation

@EdbertChan

@EdbertChan EdbertChan commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Summary

Agents prove UI claims with screenshots. Reviewers trust those shots only when they show the same surface the claim names.

The failure mode is a near-neighbor capture: another product's login page, a generated text slide, or a mock screen treated as real UI.

This change adds standing authenticity rules: match proof to claim, write expected predicates before capture, reject synthesized UI, and require real inspection.

Review Claim

Visual proof for a claimed product surface must come from that surface's real pixels, after expected predicates are written, and a marker-only inspection line does not count as a check.

Review Lane

behavior

Review Unit

corpus-lesson

Safety Invariant

This slice only adds standing verification rules and colocated trigger scenarios to a personal skill. It does not change runtime hooks, install wiring, or product capture tooling. Agents that already follow visual-proof correctly keep the same path; the new text forbids substituting a different surface or a synthetic image as proof.

Slice Rationale

One review unit under the personal skill. The authenticity defaults belong with the user's verify posture, not a rewrite of the shared visual-proof skill.

Non-goals

  • Do not change the shared visual-proof or prove-it skills beyond cross-links.
  • Do not add hooks or PR-body validators in this slice.
  • Do not implement Slack product UI or copy-button behavior.
  • Do not merge this PR.

Assumptions

  • Safety Invariant above is proposed under mandated non-interactive automate-me from reflect; treat as unconfirmed until a human edits or approves.
  • Lane is behavior (product-unit corpus-lesson) rather than policy/docs.

Test Plan

Test Plan
  • python3 scripts/ci/check_no_dated_provenance.py --base origin/main
  • python3 -m unittest tests.test_cat_mode.TestVisualProofAuthenticity -v
  • python3 scripts/ci/check_skill_test_coverage.py --base origin/main --head HEAD
  • python3 scripts/ci/check_skills_three_harnesses.py
  • python3 scripts/ci/check_ecosystem_boundaries.py
  • python3 engine/skills/make-pr/scripts/preflight.py --base origin/main --body-file /tmp/catstack-automate-me-visual-proof-pr-description.md → ok preflight passed

Revert Plan

Revert Plan
  • Safe to revert? Yes
  • Revert command: git revert <sha>
  • Post-revert steps: None (re-run install if a machine had already pulled the skill)
  • Data migration? No

Note

Low Risk
Corpus and unittest changes only; no hooks, install wiring, or capture tooling—behavior guidance for agents already following visual-proof correctly.

Overview
Adds standing Visual Proof authenticity rules to personal cat-mode (Verify section in SKILL.md plus expanded references/verify.md), extending [[visual-proof]] without changing shared skills or hooks.

Agents must match capture surface to the Review Claim, declare Expected surface/predicates (and per-case lists for OR claims) before capture, treat marker-only Manually inspected: as not a check, reject synthesized UI (slides, HTML mocks, etc.) unless the user asked for a mockup, and cover each major behavioral case with its own proof or a named waiver.

Four colocated skill test scenarios document fire vs. stay-silent cases; TestVisualProofAuthenticity in tests/test_cat_mode.py locks key phrases in the skill and reference.

Reviewed by Cursor Bugbot for commit 86a58ad. Bugbot is set up for automated code reviews on this repo. Configure here.

Codify authenticity defaults so agents match proof surface to claim,
declare Expected before capture, and reject synthesized UI as proof.
Pin the standing wording with structural tests in test_cat_mode.

Co-authored-by: Cursor <cursoragent@cursor.com>
Change-Id: Ib9421a359f38c44ab8f1f342f83dd783711b8dfc
@EdbertChan EdbertChan changed the title cat-mode: require Expected predicates before Visual Proof capture cat-mode: Visual Proof authenticity defaults (Expected before capture) Sep 27, 2026
OR Review Claims need one capture or named waiver per major behavioral
case; one case pixels do not prove another. Pin with structural tests
and fires/stays-silent fixtures.

Co-authored-by: Cursor <cursoragent@cursor.com>
Change-Id: If0239eceaf455a3bc23cc07bc8b5bfc21c3d7862
@EdbertChan
EdbertChan marked this pull request as ready for review September 29, 2026 00:37
@cursor

cursor Bot commented Sep 29, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_0b70c110-1980-4e27-ae45-e93a0a2b2b93)

@mergify

mergify Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Queued — the merge queue status continues in this comment ↓.

@EdbertChan

Copy link
Copy Markdown
Owner Author

@Mergifyio queue

@mergify

mergify Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Merge Queue Status

This pull request spent 45 minutes 6 seconds in the queue, including 44 minutes 39 seconds running CI.

Required conditions to merge
  • check-success = lint
  • check-success = test
  • check-success = validate

@mergify mergify Bot added the queued label Sep 29, 2026
@mergify mergify Bot mentioned this pull request Sep 29, 2026
6 tasks done
@mergify
mergify Bot merged commit 3902aa5 into main Sep 29, 2026
7 checks passed
@mergify mergify Bot removed the queued label Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant