unverified-tag-ledger: gate the next-prompt reminder, default to stale only - #783
Conversation
|
This pull request is part of a Mergify stack:
|
97efd7e to
65902c9
Compare
ee35240 to
4d9f2f3
Compare
Revision history
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 4d9f2f3. Configure here.
65902c9 to
21fb367
Compare
4d9f2f3 to
1d1cad5
Compare
21fb367 to
13ab621
Compare
|
Mergify repair stopped: GitHub reports merge conflict. The retry cap was reached for current head 1d1cad5. |
…e only The reminder was re-read to the user on every prompt for every outstanding claim, including claims whose proof had already been pasted. There was no way to turn it down: CATSTACK_HOOK_MODE_UNVERIFIED_TAG_LEDGER=off still printed the full list at exit 0, because this hook is not on run_hook and nothing reads the registry mode for it. CATSTACK_UNVERIFIED_TAG_REMINDER is read in-hook through _flags/flags.py, the one pattern proven to work here (wrong-check-reflect/detect.py:12-15, 191): stale default, and what unset means: only claims already 3+ turns old all the old behaviour off no injection Unset resolves to stale on purpose. A flag whose unset value is the old behaviour changes nothing for the person who asked for less. reminder() at :138 already computed that exact subset and threw it away at :145. The gate is on the injection only. Emission and recording are untouched on every setting, so `off` buys silence without emptying the ledger it is mined from. A value nobody understands, and an .env candidate that cannot be read, are both named on stderr instead of passing as "not set". hooks.toml gains an enabled_by line naming the flag. Nothing reads that field -- registry.py:15 is its only mention in the repo -- so it is documentation beside the gate, not the gate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Change-Id: I4f465761ee7f84489616ba1e83c54eb98ceefff0
1d1cad5 to
ba95eda
Compare
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_b530443d-9b4a-4228-b84f-abd818f77705) |
|
Mergify repair stopped: missing required check lint |
|
@Mergifyio queue |
Merge Queue Status
This pull request spent 43 minutes 15 seconds in the queue, including 30 minutes 33 seconds running CI. Required conditions to merge
|
|
@Mergifyio queue |
☑️ Command
|

Summary
The reminder about open claims repeats every single turn, forever.
In one session it appeared eight times for one claim, and seven of those came after the proof had already been shown.
The code already works out which claims are old enough to matter, then ignores that list and prints them all.
This adds a setting with three values: off, stale only, or all. Stale only is the default.
Review Claim
The reminder repeats only for claims past the escalation age, and can be switched off entirely.
Review Lane
behavior
Review Unit
engine-runtime
Safety Invariant
Every value still records claims, so the ledger stays complete and minable. The setting gates display only. An unreadable setting file reports itself rather than silently choosing a value.
Slice Rationale
Changes one function's output and the switch that gates it. Kept apart from the change that lets a claim close, because either alone is useful and each can be reverted without the other.
Non-goals
Does not stop a claim being written into a reply, which would hide unchecked claims instead of reducing repeats. Does not change the Stop-time block.
Test Plan
Test Plan
Run at this commit:
Cases: off prints nothing; unset shows a fresh claim nothing and an aged claim once; all keeps today's behaviour.
Revert Plan
Revert Plan
Revert this commit. The reminder returns to printing every open claim every turn. No stored data changes, so nothing needs migrating.
Note
Medium Risk
Touches hook install/wrapping, background judge delivery, and PR preflight gates used before publish; most changes are guarded by tests but mis-routing or a broken notify wrap could affect Codex sessions fleet-wide.
Overview
This PR is much wider than the title suggests: it tightens catstack hooks, routing, PR gates, and reflect/automation guidance in one pass.
Unverified-tag ledger:
CATSTACK_UNVERIFIED_TAG_REMINDERnow gates injection only (off/staledefault /all); ledger recording and Stop blocking stay on. Discharging a claim now emits a reflect trigger for evidence-order misses (claim before check), complementingwrong-check-reflectphrase updates for the same pattern.Hooks & metrics: Codex
notifychains get wrapped by the metrics runner (wrap_installed,check_install_effective,run.py --notify). LLM-judge work emits stage events (judge_skipped/queued/finished) with a newreport.py --judge --checkleak detector. Codex rollouts resolve fromthread-id; inbox reads last turns from the end of huge transcripts; subagent verdicts drain into parent sessions.Execution routing:
route_execution(units=N)routes multiple independent PR stacks to Invoker in parallel orsubagent_worktree_per_unit, never serial parent-thread publishing.PR / reflect / cost tooling:
make-prpreflight runs the same Node PR-body schema validator as CI (fixes bodies that passed history-only checks). New fleet session scanners (scan_session_tokens.py,context_growth.py),token_auditmarks Bash-only file access asuncheckedfor read/edit detectors, and reflect adds a User-did-it lens wired toautomate-me. Smaller fixes:update_fleet.shapp-replace recovery,ci_logsnote dedup, cat-mode ops rules (superseded-branch check, Bashcdanchoring inCLAUDE.learned.md).Reviewed by Cursor Bugbot for commit ba95eda. Bugbot is set up for automated code reviews on this repo. Configure here.