Repository navigation
ci(merge-queue): a shard that never got a runner is not a test failure #21933
Description
Activity
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsPath: fleet decision — the merge queue lands green PRs (#4859) | 缺项 | none
Triage: first grade —
tooling·priority:p2·domain:devx·area:devpath·pm:queue. A cancelled-with-no-runner shard is classified as infrastructure and re-queued once, and the attestation rule is unchangedTriage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-06T01:54Z. ⛔ Not a claim, ⛔ not a dispatch.Triage: lands in the merge-queue triage workflow (#4859) and its classifier ⇒
domain:devx, as filed; rationale: it is CI wiring and classification only, and nothing a published package ships.Unblocks: every PR the merge queue drops on a runner outage, which today needs a person to open the jobs before re-queuing. On 2026-10-05, PR #21872 was dropped twice this way. The same window put 7 jobs of the hourly
CIrun and 4 relaypr_createdispatches without a runner (#21909, objectui#11660).- Why p2: this is an operational cost measured on real PRs within one day, with a person in the loop each time. It is not p1, because nothing red is hidden and nothing is wrongly landed.
- Kept, and pinned:
- ⛔ A shard that never ran still does not attest as passing (CI 聚合门禁把合并队列重建的
aggregate result: abandoned判成红 —— 在队 PR 零测试失败被踢出(ci.yml 两处白名单缺abandoned) #6082). - ⛔ A re-queue happens at most once per head, and never on a real failure.
- A second no-runner red goes to a person.
- ⛔ A shard that never ran still does not attest as passing (CI 聚合门禁把合并队列重建的
- The classifier keys on the job record (
runner_nameempty with no steps run, or the "not acquired by Runner" annotation), not on log text.
Generated by Claude Code
- addedarea:devpathThe road — create, dev, verify, publish/install, connect an agent, iterateThe road — create, dev, verify, publish/install, connect an agent, iteratepriority:p2Medium: important, M3Medium: important, M3and removed
on Oct 6, 2026 objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsClaim: PM loop round 64
Session:session_01VDtqoecgES7ScQYGbFVDRv
Branch:claude/issue-21933-mq-no-runner-requeue(cut fromorigin/main9dce635337)
Worktree:objectstack-issue-21933
Domain:domain:devx
Seat:domain:devx#1
File surface:.github/workflows/merge-queue-triage.ymland the script it runs;scripts/check-merge-queue-triage-outcome.mjs;scripts/fixtures/merge-queue-triage/**;- the test or dry-run fixtures the card's Done-when names.
The card's three steps:
- Classify. A required job that ends
cancelledwith an emptyrunner_nameand no steps, or that carries the "not acquired by Runner" annotation, is classedinfra:no-runner. The classifier reads the job record, not log text. - Re-queue. A build whose only red is
infra:no-runneris re-queued at most once per head. A real failure is never re-queued, and a second no-runner red goes to a person. - The triage comment names the class.
- ⛔ The attestation rule is unchanged: a shard that never ran never attests as passing (CI 聚合门禁把合并队列重建的
aggregate result: abandoned判成红 —— 在队 PR 零测试失败被踢出(ci.yml 两处白名单缺abandoned) #6082). - ⛔ No change to whether a build is green.
- ⛔ No governed surface.
Stop on a breach and explain it in the report. That includes a re-queue mechanism that needs a credential or permission the workflow does not hold; report the exact blocker, not a partial PASS.
Container & model:M(one workflow, its script, fixtures and test),mode:subagent,model: opus
Clause-②: no
Thread-read: 6007722545
Serial constraints cleared: board read at 2026-10-06T02:53Z. No open PR touches.github/workflows/**,scripts/ci/**or the merge-queue triage files (each file list read byfilename). The surface is disjoint from round 63.Priority rule 3 reading: this is a p2, dispatched alongside the lane's P1 (round 63) on a disjoint surface. A workflow or script is not a review face. No package publishes, so
skip-changesetapplies.
Generated by Claude Code
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsos-dev-report
{ "issue": 21933, "status": "blocked", "branch": "claude/issue-21933-mq-no-runner-requeue", "pr": "https://github.com/objectstack-ai/objectstack/pull/21946", "head_sha": "499d0ae32e2259ee14eb10406e24c67013788548", "session": "session_01VDtqoecgES7ScQYGbFVDRv", "premise_still_valid": true, "summary": "Delivered limbs 1 and 3, plus the decision half of limb 2, as draft PR #21946 (Part of #21933, not Closes). The triage script classifies infra:no-runner from the job record: a cancelled job with an empty runner_name and zero steps, or the not-acquired-by-Runner check-run annotation for a cancelled job the record does not decide. A build is infra:no-runner only when every other red is a Verify-results aggregate explained by a no-runner member of its own shard family. The comment names the class and states a re-queue verdict of once per PR head (pulls.get head; durable per-head marker read back from the PR comments); the second red on the same head goes to a person, and an unreadable head or comment list gives unknown. BLOCKED: the re-queue ACT. The workflow holds no credential that can enqueue. The only measured path is the fleet App token with contents: write (fleet-write.yml says it has ONE consumer). Granting contents: write to GITHUB_TOKEN would widen permissions, and a merge group that GITHUB_TOKEN creates is not known to trigger CI. So the comment tells a person to re-queue once, and an enqueue call fails the harness (M28). The attestation rule and green semantics are unchanged.", "files_changed": [ ".github/workflows/merge-queue-triage.yml", "scripts/check-merge-queue-triage-outcome.mjs", "scripts/fixtures/merge-queue-triage/README.md", "scripts/fixtures/merge-queue-triage/run-37374282440.jobs.json", "scripts/fixtures/merge-queue-triage/run-37371558473.jobs.json", "scripts/fixtures/merge-queue-triage/check-run-111979038621.annotations.json" ], "done_when": { "synthetic no-runner cancellation classified and re-queued once": "Classified: proven by N1 (real records of run 37374282440), N2 (run 37371558473: 8 no-runner jobs and 2 aggregates) and N7 (real annotation of check run 111979038621; control without the annotation reads as failure). Re-queued once: only the VERDICT is proven (N1 grants once; the N3 pair has run 2 read run 1 own comment and sends the second red to a person; N4 shows the budget is per head). The ACT is blocked, see blockers.", "real shard failure not re-queued": "N5 (real failure), N6 (no-runner beside a real failing shard) and N9 (an aggregate from another family) all read as failure, with no verdict, no marker and no pulls.get.", "triage comment names the class": "N1-N10 assert the line naming the class: infra:no-runner or failure.", "test or dry-run fixture for both cases": "10 new scenarios (N1-N10) over 3 trimmed real fixtures, plus 9 new self-test mutations (M20-M28). Each goes red at its named scenarios, and its control scenario stays green." }, "tests": "At 499d0ae32e: node scripts/check-merge-queue-triage-outcome.mjs -> OK (133 assertions over 34 scenarios); the base had 90 over 24. --self-test -> 181 assertions, 32 mutations each driven to red, exit 0. dispatch-gates --commands derived 51 commands; 50 ran with exit 0 (including check:workflow-status-functions, check:nul-bytes, check:required-contexts, check:shard-attestation and check:self-test-wired). --ran reconciliation: 51 accounted, 1 NOT-MEASURED, 0 UNRUN. NOT MEASURED: pnpm check:pm-dispatch-gates. Its --self-test half alone exceeded the foreground cap (timeout 570s, exit 124; 1785 pass marks and 1 fail mark, on packages/qa/dogfood/test/per-file-cwd.setup.ts, outside the diff); the plain half also exceeded 500s. Declared to CI. Narrowed eslint: the changed .mjs file, --no-inline-config, json format: 1 file, 0 errors, 0 warnings; no parserOptions.project, so untouched files cannot move. No package touched, so no build, unit-test or typecheck leg applies.", "validations": "see tests", "deviations": [ "PR body says Part of #21933, not Closes #21933 as briefed: the re-queue act is not delivered, and merging must not close the card.", "Permission widened, and declared: job triage gains checks: read, used only to read the not-acquired-by-Runner annotation of a cancelled job that has a runner. No other grant changes; the harness now pins checks: read and no contents grant.", "The workflow header limb-2 boundary (DECIDES NOTHING, no auto-requeue) still holds. The new limb 3 decides a verdict and states it in the comment, but performs no act.", "Zero label writes: the brief says the PM adds skip-changeset. PR assignee set to the card assignee (baozhoutao) through label-write." ], "blockers": [ "Re-queue act: needs an enqueue credential the workflow does not hold. Exact facts: (1) the only measured enqueue path is enablePullRequestAutoMerge with the fleet App token minted with permission-contents: write (fleet-write.yml lines 115-120: measured, without contents write it answered Resource not accessible by integration; that grant has ONE consumer). (2) The workflow GITHUB_TOKEN holds actions read, pull-requests write, issues write and checks read, with no contents grant. (3) Even with contents: write added, an enqueue by GITHUB_TOKEN creates the merge group as github-actions; GitHub does not start workflow runs for GITHUB_TOKEN-caused events except workflow_dispatch and repository_dispatch, so whether CI would build that group is NOT MEASURED." ], "mcp_calls": "0", "api_writes": "3 - three POST /repos/objectstack-ai/objectstack/dispatches through the fleet-write relay: pr_create (PR #21946), PR assignee through label-write, and this os-dev-report comment. Also 2 git pushes (not REST).", "open_questions": [ { "question": "How should the one automatic re-queue be performed?", "options": [ "A: mint the fleet App token in merge-queue-triage.yml (vars.OS_FLEET_APP_ID + secrets.OS_FLEET_PRIVATE_KEY, permission-contents: write and permission-pull-requests: write, repositories: objectstack). Call enablePullRequestAutoMerge on a once verdict. This makes it a second consumer of the contents grant; a maintainer must approve that.", "B: grant GITHUB_TOKEN contents: write and call enqueuePullRequest. Cheaper, but first measure whether a GITHUB_TOKEN-created merge group triggers CI at all; if not, the entry sits until the queue timeout and is ejected again.", "C: keep the act manual (this PR as it stands). The comment already names the class and the verdict, so the person no longer has to open the jobs." ], "recommendation": "C now, A as a follow-up: A reuses the one measured path, but adds a consumer to a contents: write credential, which is the maintainer call the brief reserves." } ], "out_of_scope_findings": [ "carrier: none. Noted in PR Acceptance notes and not filed. merge-queue-triage listens to workflows: [CI] only, so a no-runner red in the Governed Surface Guard / Closing-Target Claim Guard / Spec Liveness Check merge_group runs (for example 37371558492, which ejected #21872) is never triaged. Dedupe words: merge-queue-triage workflows CI only, guard workflow no-runner, workflow_run workflows list." ] }
Generated by Claude Code
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actions决策:合并队列"没分到 runner"的自动重新排队,用哪种凭据执行
PM 工位
domain:devx#1·session_01VDtqoecgES7ScQYGbFVDRv· 2026-10-06T03:37Z已完成的部分(PR #21946,
Part of #21933,不关卡):- 按 job 记录(不读日志文本)把"没分到 runner 就被取消"分类为
infra:no-runner。判断条件是runner_name为空且一步没跑,或带 "not acquired by Runner" 注解。 - triage 评论写明这个类别,并给出"本 head 可重新排队一次 / 第二次交给人"的判定。
- 分片没跑不算通过的规则(CI 聚合门禁把合并队列重建的
aggregate result: abandoned判成红 —— 在队 PR 零测试失败被踢出(ci.yml 两处白名单缺abandoned) #6082)不变。
卡住的部分: 重新排队这个动作需要一个能把 PR 放进队列的凭据,
merge-queue-triage.yml现在没有。Governing text:
.github/workflows/fleet-write.yml(铸 App token 那一步)里写着:permission-contents: write这项授权 "ONE consumer",只给enablePullRequestAutoMerge/disablePullRequestAutoMerge用。选项 做法 代价 / 风险 A triage workflow 自己铸 fleet App token( contents: write+pull-requests: write),判定"可以一次"时调用enablePullRequestAutoMerge这条路实测可用,但 contents: write会多出第二个使用方,正好违反上面那句"只有一个使用方"的约定B 给 workflow 的 GITHUB_TOKEN加contents: write,再调用enqueuePullRequest扩大了 workflow 的写权限。另外,由 GITHUB_TOKEN建出来的 merge group 会不会触发 CI 没有实测过;如果不会,这个条目会在队列里等到超时再被踢出C 保持手动:评论已经写明类别和判定,由人点一次重新排队 不扩权限;但仍然每次要有人动手 席位意见: 先按 C 落地 #21946,它在 A、B、C 三种选择下都有用。重新排队的动作要等你选 A 或 B;在你决定之前不做。如果选 B,要先实测
GITHUB_TOKEN建的 merge group 会不会触发 CI。你要做的(一个动作): 回复 A、B 或 C。
Generated by Claude Code
- 按 job 记录(不读日志文本)把"没分到 runner 就被取消"分类为
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsRuled, recorded by the PM seat (
session_018zT8d8NpiQ1ExhuNd5TxY6) from the maintainer's answer in that Claude Code session, 2026-10-06: C, keep the re-queue manual (「C 保持手动」).No credential is added.
merge-queue-triage.ymlgets nocontents: write, and the fleet App token keeps its single consumer. PR #21946 already delivers the classification, the triage comment and the once-per-head re-queue verdict. A person, or a seat, acts on that verdict by re-queuing once. Nothing more is owed on this card.
Generated by Claude Code
objectstack-fleet commented
on Oct 7, 2026 ContributorAuthorMore actionsTriage (owned 24 h re-check): closed
completed. Ruling C leaves nothing owed, and PR #21946 delivered the restTriage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-07T06:09Z. ⛔ Not a claim. The assignee is untouched.- Idle for 24 h with an owner: the last write is the ruling record
6010502906at 2026-10-06T06:12Z. The card carries nopm:state label. - What landed: PR ci(merge-queue): name a no-runner red as infra:no-runner, with a once-per-head re-queue verdict #21946 merged as
f57627b585, an ancestor oforigin/main. It delivered limbs 1 and 3, plus the verdict half of limb 2: theinfra:no-runnerclass from the job record, the triage comment naming it, and the once-per-head re-queue verdict. It left the card open only for the automatic re-queue act. - What the ruling settled: C, 「C 保持手动」. No credential is added, the re-queue stays a person's act on the verdict, and the record says "Nothing more is owed on this card."
- ⇒ Nothing remains. If someone later wants the re-queue act automated, that is a new card with a new ruling, because it would add a credential consumer.
- Idle for 24 h with an owner: the last write is the ruling record
- added a commit that references this issue
on Oct 7, 2026
Requested by the maintainer in Claude Code session
session_018zT8d8NpiQ1ExhuNd5TxY6, 2026-10-06 (「两张卡都开」).What happened
On 2026-10-05, PR #21872 was removed from the merge queue twice with
CI_FAILURE, and nothing in it was failing:Test Core (5/6)was cancelled with no runner assigned (emptyrunner_name, zero steps). The aggregator then reported "1 of 6 declared shard(s) of test published no positive attestation".The same day, PR #21877's head saw 13 cancelled jobs before any test body ran.
Each time, a person had to open the jobs to tell a runner outage from a test failure before re-queuing. The merge-queue triage comment (merge-queue-triage workflow, #4859) could not tell them apart either ("日志不可读").
Proposal
cancelledwith no runner assigned (runner_nameempty, no steps run), or with the "not acquired by Runner" annotation, label the buildinfra:no-runnerinstead of a test failure. The triage comment then names it as such.infra:no-runner, re-queue the PR once. Never re-queue a second time on the same head, and never re-queue a real failure. A second no-runner failure goes to a person.aggregate result: abandoned判成红 —— 在队 PR 零测试失败被踢出(ci.yml 两处白名单缺abandoned) #6082). This only changes how the red is classified and handled. It does not change whether the build is green.Done when
Generated by Claude Code