merge queue: checking #890 on main (5b4a082), stacked on #818, #822 and #875 - #903
Closed
mergify[bot] wants to merge 11 commits into
Closed
mergify[bot] wants to merge 11 commits into
mergify[bot] wants to merge 11 commits into
Conversation
model-tier-candidates, redundant-reads and no-verify-edit-streak all key on
structured tool names: LOOKUP_TOOLS = ("Read", "Grep", "Glob") and the
Edit/Write scan. Under a bash-first harness every file read and edit happens
inside a Bash command (cat, sed -n, sed -i), so all three counted zero and
reported "no" with a confident rationale ("0 redundant re-read(s) of an
identical file+offset/limit window").
That is absence of evidence reported as evidence of absence, and it is read
downstream as a clean bill of health. Found while auditing a 2.8M-token Invoker
repair session that ran 33 Bash calls and zero structured tool calls: every flag
in the report said "no", including on a turn pair where the agent re-ran a
command it had already truncated, and on its only file edit (a sed -i).
These three now report "unchecked" with count None when the session used Bash
and made no Read/Grep/Glob/Edit/Write call at all. "unchecked" is the state the
report already uses for brevity-hook-blocks, frustration-signals and
subagent-thrash, so no consumer shape changes.
The gate is deliberately the whole-session bash-only shape rather than
per-detector. A first attempt keyed each detector on its own tool
(edit_blind = no Edit/Write seen) and the negative fixture caught it calling a
Read-using session blind for edits -- an agent that had Edit available and did
not use it really did not edit, and that measurement is true.
audit_subagents sums flags["redundant-reads"]["count"], which can now be None,
so that aggregation coalesces to 0.
Fixture pair, per the repro gate:
- positive: a Bash-only session (two identical sed -n reads plus a sed -i)
asserts all three flags are unchecked with a None count
- negative: a session with two identical Read calls plus a Bash call asserts
all three still measure, and redundant-reads still counts 1
Backtested against the motivating transcript: all three flipped no -> unchecked.
Full suite: Ran 276 tests, OK.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… the person The runner's per-run table counted wrong-check-reflect as "silent 2,582, spoke 0" on Claude. That number cannot tell a mute hook from a working one: the Stop hook never prints, it queues a background judge job, and the verdict arrives on a later PostToolUse. The live cache showed the real losses were downstream: 39 hit verdicts for this hook sat undelivered, 25 of them on jobs queued with an empty transcript path, which no drain can ever find. Each stage now writes one catstack.hook_event.v1 row with a reason: - judge_skipped (wrong-check-reflect): stop_hook_active, gate_off, empty_reply, already_prompted, user_asked_reflect, judge_child, bad_payload - judge_queued (llm-judge, every hook): transcript or no_transcript - judge_finished (llm-judge): hit, clean or unchecked Delivery was already recorded by drain. Rows join on the job id. A job queued with no transcript also prints catstack-hook-error, so the runner classifies the run as caught_error and hook-health surfaces it on the next prompt instead of the verdict vanishing. report.py --judge prints the per-hook funnel plus leak counts (no_transcript, stuck, undelivered, undelivered_hits) older than --grace (default 1h); --check exits 1 on any leak. Stage rows carry no rule_id, so the existing rule table ignores them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Change-Id: Ic6eeb8c30dabc90c4d11270e98ef4d1a2eecddfd
…rially route_execution(units=N) never returns local for N>1 publishing units: Invoker first, else one worktree subagent per unit (also when the user directs subagents). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Change-Id: I9d5797fbed454079561fd2f4b36379ab7b7df02d
…uard CI's no-new-comments gate failed on three added `#` lines in engine/skills/reflect/scripts/token_audit.py. The variable names (bash_only, lookup_blind/read_blind/edit_blind) and BLIND_RATIONALE already say what the comment said, so the lines are removed rather than reworded. Repro: python3 scripts/ci/check_no_new_comments.py --base origin/main failed with 3 new comment line(s) before, exits 0 after. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…est) Exit code: 0
… validate; Exit code: 0
capture_github stored its notes in the manifest and then passed the same list to artifact_response, which appends the manifest's notes, so every note printed twice. It now passes no extra notes, like the cached and local paths. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Change-Id: Ib46d51d0ddda60167c6742233458c9e603adb973
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🎉 This pull request has been checked successfully and will be merged soon. 🎉
#890 is queued for merge on branch main (5b4a082).
Stacked behind 3 pull requests queued ahead of this batch, not part of it. These checks run on a tip that also carries their commits, so a failure here can come from them as much as from #890.
Queued ahead of this batch:
This pull request has been created by Mergify to speculatively check the mergeability of #890.
You don't need to do anything. Mergify will close this pull request automatically when it is complete.
Required conditions of queue rule
admin-bypassfor merge:check-success = lintcheck-success = testcheck-success = validateRequired conditions to stay in the queue:
-draftbase=mainlabel=admin-bypass