cat-mode: Visual Proof authenticity defaults (Expected before capture) - #1207
Merged
Conversation
Codify authenticity defaults so agents match proof surface to claim, declare Expected before capture, and reject synthesized UI as proof. Pin the standing wording with structural tests in test_cat_mode. Co-authored-by: Cursor <cursoragent@cursor.com> Change-Id: Ib9421a359f38c44ab8f1f342f83dd783711b8dfc
OR Review Claims need one capture or named waiver per major behavioral case; one case pixels do not prove another. Pin with structural tests and fires/stays-silent fixtures. Co-authored-by: Cursor <cursoragent@cursor.com> Change-Id: If0239eceaf455a3bc23cc07bc8b5bfc21c3d7862
EdbertChan
marked this pull request as ready for review
September 29, 2026 00:37
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_0b70c110-1980-4e27-ae45-e93a0a2b2b93) |
Contributor
|
Queued — the merge queue status continues in this comment ↓. |
Owner
Author
|
@Mergifyio queue |
Contributor
Merge Queue Status
This pull request spent 45 minutes 6 seconds in the queue, including 44 minutes 39 seconds running CI. Required conditions to merge
|
6 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Agents prove UI claims with screenshots. Reviewers trust those shots only when they show the same surface the claim names.
The failure mode is a near-neighbor capture: another product's login page, a generated text slide, or a mock screen treated as real UI.
This change adds standing authenticity rules: match proof to claim, write expected predicates before capture, reject synthesized UI, and require real inspection.
Review Claim
Visual proof for a claimed product surface must come from that surface's real pixels, after expected predicates are written, and a marker-only inspection line does not count as a check.
Review Lane
behavior
Review Unit
corpus-lesson
Safety Invariant
This slice only adds standing verification rules and colocated trigger scenarios to a personal skill. It does not change runtime hooks, install wiring, or product capture tooling. Agents that already follow visual-proof correctly keep the same path; the new text forbids substituting a different surface or a synthetic image as proof.
Slice Rationale
One review unit under the personal skill. The authenticity defaults belong with the user's verify posture, not a rewrite of the shared visual-proof skill.
Non-goals
Assumptions
Test Plan
Test Plan
python3 scripts/ci/check_no_dated_provenance.py --base origin/mainpython3 -m unittest tests.test_cat_mode.TestVisualProofAuthenticity -vpython3 scripts/ci/check_skill_test_coverage.py --base origin/main --head HEADpython3 scripts/ci/check_skills_three_harnesses.pypython3 scripts/ci/check_ecosystem_boundaries.pypython3 engine/skills/make-pr/scripts/preflight.py --base origin/main --body-file /tmp/catstack-automate-me-visual-proof-pr-description.md→ok preflight passedRevert Plan
Revert Plan
git revert <sha>Note
Low Risk
Corpus and unittest changes only; no hooks, install wiring, or capture tooling—behavior guidance for agents already following visual-proof correctly.
Overview
Adds standing Visual Proof authenticity rules to personal cat-mode (Verify section in
SKILL.mdplus expandedreferences/verify.md), extending[[visual-proof]]without changing shared skills or hooks.Agents must match capture surface to the Review Claim, declare Expected surface/predicates (and per-case lists for OR claims) before capture, treat marker-only
Manually inspected:as not a check, reject synthesized UI (slides, HTML mocks, etc.) unless the user asked for a mockup, and cover each major behavioral case with its own proof or a named waiver.Four colocated skill test scenarios document fire vs. stay-silent cases;
TestVisualProofAuthenticityintests/test_cat_mode.pylocks key phrases in the skill and reference.Reviewed by Cursor Bugbot for commit 86a58ad. Bugbot is set up for automated code reviews on this repo. Configure here.