Skip to content

unverified-tag-ledger: gate the next-prompt reminder, default to stale only - #783

Merged
mergify[bot] merged 1 commit into
mainfrom
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/unverified-tag-ledger-gate-next-prompt-reminder--4f465761
Sep 24, 2026
Merged

mergify[bot] merged 1 commit into
mainfrom
stack/EdbertChan/reflect/subagent-decisions-slot-20260922/unverified-tag-ledger-gate-next-prompt-reminder--4f465761

Conversation

@EdbertChan

@EdbertChan EdbertChan commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

The reminder about open claims repeats every single turn, forever.

In one session it appeared eight times for one claim, and seven of those came after the proof had already been shown.

The code already works out which claims are old enough to matter, then ignores that list and prints them all.

This adds a setting with three values: off, stale only, or all. Stale only is the default.

Review Claim

The reminder repeats only for claims past the escalation age, and can be switched off entirely.

Review Lane

behavior

Review Unit

engine-runtime

Safety Invariant

Every value still records claims, so the ledger stays complete and minable. The setting gates display only. An unreadable setting file reports itself rather than silently choosing a value.

Slice Rationale

Changes one function's output and the switch that gates it. Kept apart from the change that lets a claim close, because either alone is useful and each can be reverted without the other.

Non-goals

Does not stop a claim being written into a reply, which would hide unchecked claims instead of reducing repeats. Does not change the Stop-time block.

Test Plan

Test Plan

Run at this commit:

python3 -m unittest discover -s engine/hooks/unverified-tag-ledger/tests
Ran 45 tests in 1.299s
OK

Cases: off prints nothing; unset shows a fresh claim nothing and an aged claim once; all keeps today's behaviour.

Revert Plan

Revert Plan

Revert this commit. The reminder returns to printing every open claim every turn. No stored data changes, so nothing needs migrating.


Note

Medium Risk
Touches hook install/wrapping, background judge delivery, and PR preflight gates used before publish; most changes are guarded by tests but mis-routing or a broken notify wrap could affect Codex sessions fleet-wide.

Overview
This PR is much wider than the title suggests: it tightens catstack hooks, routing, PR gates, and reflect/automation guidance in one pass.

Unverified-tag ledger: CATSTACK_UNVERIFIED_TAG_REMINDER now gates injection only (off / stale default / all); ledger recording and Stop blocking stay on. Discharging a claim now emits a reflect trigger for evidence-order misses (claim before check), complementing wrong-check-reflect phrase updates for the same pattern.

Hooks & metrics: Codex notify chains get wrapped by the metrics runner (wrap_installed, check_install_effective, run.py --notify). LLM-judge work emits stage events (judge_skipped / queued / finished) with a new report.py --judge --check leak detector. Codex rollouts resolve from thread-id; inbox reads last turns from the end of huge transcripts; subagent verdicts drain into parent sessions.

Execution routing: route_execution(units=N) routes multiple independent PR stacks to Invoker in parallel or subagent_worktree_per_unit, never serial parent-thread publishing.

PR / reflect / cost tooling: make-pr preflight runs the same Node PR-body schema validator as CI (fixes bodies that passed history-only checks). New fleet session scanners (scan_session_tokens.py, context_growth.py), token_audit marks Bash-only file access as unchecked for read/edit detectors, and reflect adds a User-did-it lens wired to automate-me. Smaller fixes: update_fleet.sh app-replace recovery, ci_logs note dedup, cat-mode ops rules (superseded-branch check, Bash cd anchoring in CLAUDE.learned.md).

Reviewed by Cursor Bugbot for commit ba95eda. Bugbot is set up for automated code reviews on this repo. Configure here.

@EdbertChan

EdbertChan commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

This pull request is part of a Mergify stack:

# Pull Request Link
1 wrong-check-reflect: one shot per reply, and scope the reflect scan to the turn #796
2 reflect: catch an evidence-order correction from the ledger transition, not from wording #782
3 unverified-tag-ledger: gate the next-prompt reminder, default to stale only #783 👈
4 diu-stop: a file citation buys silence only when it is backed #784
5 diu-stop: the marker gate reads prose, so a mention is not a use #785
6 cat-mode: show the escape-hatch tag in the form the gate accepts #786
7 principle-subagent-inherits-scope: a brief carries decisions, not just facts and a question #788
8 llm-judge: claude is the only default runner #793

@EdbertChan

EdbertChan commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Revision history

# Type Changes Reason Date
1 initial ee35240 2026-09-22 18:38 UTC
2 rebase ee35240 → 4d9f2f3 (rebase only) 2026-09-22 18:38 UTC
3 rebase 4d9f2f3 → 1d1cad5 (rebase only) 2026-09-23 14:06 UTC

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4d9f2f3. Configure here.

Comment thread engine/hooks/unverified-tag-ledger/detect.py
@EdbertChan
EdbertChan force-pushed the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/catch-evidence-order-correction-ledger-transition--99190104 branch from 65902c9 to 21fb367 Compare September 23, 2026 14:06
@EdbertChan
EdbertChan force-pushed the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/unverified-tag-ledger-gate-next-prompt-reminder--4f465761 branch from 4d9f2f3 to 1d1cad5 Compare September 23, 2026 14:06
@EdbertChan
EdbertChan force-pushed the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/catch-evidence-order-correction-ledger-transition--99190104 branch from 21fb367 to 13ab621 Compare September 24, 2026 03:33
@EdbertChan

Copy link
Copy Markdown
Owner Author

Mergify repair stopped: GitHub reports merge conflict. The retry cap was reached for current head 1d1cad5.

…e only

The reminder was re-read to the user on every prompt for every outstanding
claim, including claims whose proof had already been pasted. There was no way
to turn it down: CATSTACK_HOOK_MODE_UNVERIFIED_TAG_LEDGER=off still printed
the full list at exit 0, because this hook is not on run_hook and nothing
reads the registry mode for it.

CATSTACK_UNVERIFIED_TAG_REMINDER is read in-hook through _flags/flags.py, the
one pattern proven to work here (wrong-check-reflect/detect.py:12-15, 191):

  stale  default, and what unset means: only claims already 3+ turns old
  all    the old behaviour
  off    no injection

Unset resolves to stale on purpose. A flag whose unset value is the old
behaviour changes nothing for the person who asked for less. reminder() at
:138 already computed that exact subset and threw it away at :145.

The gate is on the injection only. Emission and recording are untouched on
every setting, so `off` buys silence without emptying the ledger it is mined
from. A value nobody understands, and an .env candidate that cannot be read,
are both named on stderr instead of passing as "not set".

hooks.toml gains an enabled_by line naming the flag. Nothing reads that field
-- registry.py:15 is its only mention in the repo -- so it is documentation
beside the gate, not the gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Change-Id: I4f465761ee7f84489616ba1e83c54eb98ceefff0
@EdbertChan
EdbertChan force-pushed the stack/EdbertChan/reflect/subagent-decisions-slot-20260922/unverified-tag-ledger-gate-next-prompt-reminder--4f465761 branch from 1d1cad5 to ba95eda Compare September 24, 2026 08:02
@cursor

cursor Bot commented Sep 24, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_b530443d-9b4a-4228-b84f-abd818f77705)

@EdbertChan
EdbertChan changed the base branch from stack/EdbertChan/reflect/subagent-decisions-slot-20260922/catch-evidence-order-correction-ledger-transition--99190104 to main September 24, 2026 08:03
@EdbertChan

Copy link
Copy Markdown
Owner Author

Mergify repair stopped: missing required check lint

@EdbertChan

Copy link
Copy Markdown
Owner Author

@Mergifyio queue

@mergify

mergify Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merge Queue Status

This pull request spent 43 minutes 15 seconds in the queue, including 30 minutes 33 seconds running CI.

Required conditions to merge
  • check-success = lint
  • check-success = test
  • check-success = validate

@EdbertChan

Copy link
Copy Markdown
Owner Author

@Mergifyio queue

@mergify

mergify Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

queue

☑️ Command queue ignored because it is already running from a previous command.

@mergify
mergify Bot merged commit 78678c2 into main Sep 24, 2026
4 checks passed
@mergify mergify Bot removed the queued label Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant