Skip to content

ci(merge-queue): a shard that never got a runner is not a test failure #21933

Description

@objectstack-fleet

Requested by the maintainer in Claude Code session session_018zT8d8NpiQ1ExhuNd5TxY6, 2026-10-06 (「两张卡都开」).

What happened

On 2026-10-05, PR #21872 was removed from the merge queue twice with CI_FAILURE, and nothing in it was failing:

  • Run 37371558492 and its siblings: the guard jobs were cancelled with "The job was not acquired by Runner of type hosted even after multiple attempts".
  • Run 37374282440: Test Core (5/6) was cancelled with no runner assigned (empty runner_name, zero steps). The aggregator then reported "1 of 6 declared shard(s) of test published no positive attestation".

The same day, PR #21877's head saw 13 cancelled jobs before any test body ran.

Each time, a person had to open the jobs to tell a runner outage from a test failure before re-queuing. The merge-queue triage comment (merge-queue-triage workflow, #4859) could not tell them apart either ("日志不可读").

Proposal

  1. Classify. When a required job ends cancelled with no runner assigned (runner_name empty, no steps run), or with the "not acquired by Runner" annotation, label the build infra:no-runner instead of a test failure. The triage comment then names it as such.
  2. Re-queue once, automatically. For a build whose only red is infra:no-runner, re-queue the PR once. Never re-queue a second time on the same head, and never re-queue a real failure. A second no-runner failure goes to a person.
  3. Keep the attestation rule. A shard that never ran still does not count as passing (CI 聚合门禁把合并队列重建的 aggregate result: abandoned 判成红 —— 在队 PR 零测试失败被踢出(ci.yml 两处白名单缺 abandoned) #6082). This only changes how the red is classified and handled. It does not change whether the build is green.

Done when

  • A synthetic no-runner cancellation is classified and re-queued once. A real shard failure is not re-queued.
  • The triage comment names the class.
  • The workflow has a test or a dry-run fixture for both cases.

Generated by Claude Code

Activity

  1. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Path: fleet decision — the merge queue lands green PRs (#4859) | 缺项 | none

    Triage: first grade — tooling · priority:p2 · domain:devx · area:devpath · pm:queue. A cancelled-with-no-runner shard is classified as infrastructure and re-queued once, and the attestation rule is unchanged

    Triage seat (objectstack-wide, seat post #6015) · session_01AavokzJ5DndAwitDXvKy4U · 2026-10-06T01:54Z. ⛔ Not a claim, ⛔ not a dispatch.

    Triage: lands in the merge-queue triage workflow (#4859) and its classifier ⇒ domain:devx, as filed; rationale: it is CI wiring and classification only, and nothing a published package ships.

    Unblocks: every PR the merge queue drops on a runner outage, which today needs a person to open the jobs before re-queuing. On 2026-10-05, PR #21872 was dropped twice this way. The same window put 7 jobs of the hourly CI run and 4 relay pr_create dispatches without a runner (#21909, objectui#11660).


    Generated by Claude Code

  2. added
    area:devpathThe road — create, dev, verify, publish/install, connect an agent, iterate
    and removed on Oct 6, 2026
  3. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Claim: PM loop round 64
    Session: session_01VDtqoecgES7ScQYGbFVDRv
    Branch: claude/issue-21933-mq-no-runner-requeue (cut from origin/main 9dce635337)
    Worktree: objectstack-issue-21933
    Domain: domain:devx
    Seat: domain:devx#1
    File surface:

    • .github/workflows/merge-queue-triage.yml and the script it runs;
    • scripts/check-merge-queue-triage-outcome.mjs;
    • scripts/fixtures/merge-queue-triage/**;
    • the test or dry-run fixtures the card's Done-when names.

    The card's three steps:

    1. Classify. A required job that ends cancelled with an empty runner_name and no steps, or that carries the "not acquired by Runner" annotation, is classed infra:no-runner. The classifier reads the job record, not log text.
    2. Re-queue. A build whose only red is infra:no-runner is re-queued at most once per head. A real failure is never re-queued, and a second no-runner red goes to a person.
    3. The triage comment names the class.

    Stop on a breach and explain it in the report. That includes a re-queue mechanism that needs a credential or permission the workflow does not hold; report the exact blocker, not a partial PASS.
    Container & model: M (one workflow, its script, fixtures and test), mode:subagent, model: opus
    Clause-②: no
    Thread-read: 6007722545
    Serial constraints cleared: board read at 2026-10-06T02:53Z. No open PR touches .github/workflows/**, scripts/ci/** or the merge-queue triage files (each file list read by filename). The surface is disjoint from round 63.

    Priority rule 3 reading: this is a p2, dispatched alongside the lane's P1 (round 63) on a disjoint surface. A workflow or script is not a review face. No package publishes, so skip-changeset applies.


    Generated by Claude Code

  4. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    os-dev-report

    {
      "issue": 21933,
      "status": "blocked",
      "branch": "claude/issue-21933-mq-no-runner-requeue",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/21946",
      "head_sha": "499d0ae32e2259ee14eb10406e24c67013788548",
      "session": "session_01VDtqoecgES7ScQYGbFVDRv",
      "premise_still_valid": true,
      "summary": "Delivered limbs 1 and 3, plus the decision half of limb 2, as draft PR #21946 (Part of #21933, not Closes). The triage script classifies infra:no-runner from the job record: a cancelled job with an empty runner_name and zero steps, or the not-acquired-by-Runner check-run annotation for a cancelled job the record does not decide. A build is infra:no-runner only when every other red is a Verify-results aggregate explained by a no-runner member of its own shard family. The comment names the class and states a re-queue verdict of once per PR head (pulls.get head; durable per-head marker read back from the PR comments); the second red on the same head goes to a person, and an unreadable head or comment list gives unknown. BLOCKED: the re-queue ACT. The workflow holds no credential that can enqueue. The only measured path is the fleet App token with contents: write (fleet-write.yml says it has ONE consumer). Granting contents: write to GITHUB_TOKEN would widen permissions, and a merge group that GITHUB_TOKEN creates is not known to trigger CI. So the comment tells a person to re-queue once, and an enqueue call fails the harness (M28). The attestation rule and green semantics are unchanged.",
      "files_changed": [
        ".github/workflows/merge-queue-triage.yml",
        "scripts/check-merge-queue-triage-outcome.mjs",
        "scripts/fixtures/merge-queue-triage/README.md",
        "scripts/fixtures/merge-queue-triage/run-37374282440.jobs.json",
        "scripts/fixtures/merge-queue-triage/run-37371558473.jobs.json",
        "scripts/fixtures/merge-queue-triage/check-run-111979038621.annotations.json"
      ],
      "done_when": {
        "synthetic no-runner cancellation classified and re-queued once": "Classified: proven by N1 (real records of run 37374282440), N2 (run 37371558473: 8 no-runner jobs and 2 aggregates) and N7 (real annotation of check run 111979038621; control without the annotation reads as failure). Re-queued once: only the VERDICT is proven (N1 grants once; the N3 pair has run 2 read run 1 own comment and sends the second red to a person; N4 shows the budget is per head). The ACT is blocked, see blockers.",
        "real shard failure not re-queued": "N5 (real failure), N6 (no-runner beside a real failing shard) and N9 (an aggregate from another family) all read as failure, with no verdict, no marker and no pulls.get.",
        "triage comment names the class": "N1-N10 assert the line naming the class: infra:no-runner or failure.",
        "test or dry-run fixture for both cases": "10 new scenarios (N1-N10) over 3 trimmed real fixtures, plus 9 new self-test mutations (M20-M28). Each goes red at its named scenarios, and its control scenario stays green."
      },
      "tests": "At 499d0ae32e: node scripts/check-merge-queue-triage-outcome.mjs -> OK (133 assertions over 34 scenarios); the base had 90 over 24. --self-test -> 181 assertions, 32 mutations each driven to red, exit 0. dispatch-gates --commands derived 51 commands; 50 ran with exit 0 (including check:workflow-status-functions, check:nul-bytes, check:required-contexts, check:shard-attestation and check:self-test-wired). --ran reconciliation: 51 accounted, 1 NOT-MEASURED, 0 UNRUN. NOT MEASURED: pnpm check:pm-dispatch-gates. Its --self-test half alone exceeded the foreground cap (timeout 570s, exit 124; 1785 pass marks and 1 fail mark, on packages/qa/dogfood/test/per-file-cwd.setup.ts, outside the diff); the plain half also exceeded 500s. Declared to CI. Narrowed eslint: the changed .mjs file, --no-inline-config, json format: 1 file, 0 errors, 0 warnings; no parserOptions.project, so untouched files cannot move. No package touched, so no build, unit-test or typecheck leg applies.",
      "validations": "see tests",
      "deviations": [
        "PR body says Part of #21933, not Closes #21933 as briefed: the re-queue act is not delivered, and merging must not close the card.",
        "Permission widened, and declared: job triage gains checks: read, used only to read the not-acquired-by-Runner annotation of a cancelled job that has a runner. No other grant changes; the harness now pins checks: read and no contents grant.",
        "The workflow header limb-2 boundary (DECIDES NOTHING, no auto-requeue) still holds. The new limb 3 decides a verdict and states it in the comment, but performs no act.",
        "Zero label writes: the brief says the PM adds skip-changeset. PR assignee set to the card assignee (baozhoutao) through label-write."
      ],
      "blockers": [
        "Re-queue act: needs an enqueue credential the workflow does not hold. Exact facts: (1) the only measured enqueue path is enablePullRequestAutoMerge with the fleet App token minted with permission-contents: write (fleet-write.yml lines 115-120: measured, without contents write it answered Resource not accessible by integration; that grant has ONE consumer). (2) The workflow GITHUB_TOKEN holds actions read, pull-requests write, issues write and checks read, with no contents grant. (3) Even with contents: write added, an enqueue by GITHUB_TOKEN creates the merge group as github-actions; GitHub does not start workflow runs for GITHUB_TOKEN-caused events except workflow_dispatch and repository_dispatch, so whether CI would build that group is NOT MEASURED."
      ],
      "mcp_calls": "0",
      "api_writes": "3 - three POST /repos/objectstack-ai/objectstack/dispatches through the fleet-write relay: pr_create (PR #21946), PR assignee through label-write, and this os-dev-report comment. Also 2 git pushes (not REST).",
      "open_questions": [
        {
          "question": "How should the one automatic re-queue be performed?",
          "options": [
            "A: mint the fleet App token in merge-queue-triage.yml (vars.OS_FLEET_APP_ID + secrets.OS_FLEET_PRIVATE_KEY, permission-contents: write and permission-pull-requests: write, repositories: objectstack). Call enablePullRequestAutoMerge on a once verdict. This makes it a second consumer of the contents grant; a maintainer must approve that.",
            "B: grant GITHUB_TOKEN contents: write and call enqueuePullRequest. Cheaper, but first measure whether a GITHUB_TOKEN-created merge group triggers CI at all; if not, the entry sits until the queue timeout and is ejected again.",
            "C: keep the act manual (this PR as it stands). The comment already names the class and the verdict, so the person no longer has to open the jobs."
          ],
          "recommendation": "C now, A as a follow-up: A reuses the one measured path, but adds a consumer to a contents: write credential, which is the maintainer call the brief reserves."
        }
      ],
      "out_of_scope_findings": [
        "carrier: none. Noted in PR Acceptance notes and not filed. merge-queue-triage listens to workflows: [CI] only, so a no-runner red in the Governed Surface Guard / Closing-Target Claim Guard / Spec Liveness Check merge_group runs (for example 37371558492, which ejected #21872) is never triaged. Dedupe words: merge-queue-triage workflows CI only, guard workflow no-runner, workflow_run workflows list."
      ]
    }

    Generated by Claude Code

  5. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    决策:合并队列"没分到 runner"的自动重新排队,用哪种凭据执行

    PM 工位 domain:devx#1 · session_01VDtqoecgES7ScQYGbFVDRv · 2026-10-06T03:37Z

    已完成的部分(PR #21946,Part of #21933,不关卡):

    卡住的部分: 重新排队这个动作需要一个能把 PR 放进队列的凭据,merge-queue-triage.yml 现在没有。

    Governing text: .github/workflows/fleet-write.yml(铸 App token 那一步)里写着:permission-contents: write 这项授权 "ONE consumer",只给 enablePullRequestAutoMerge / disablePullRequestAutoMerge 用。

    选项 做法 代价 / 风险
    A triage workflow 自己铸 fleet App token(contents: write + pull-requests: write),判定"可以一次"时调用 enablePullRequestAutoMerge 这条路实测可用,但 contents: write 会多出第二个使用方,正好违反上面那句"只有一个使用方"的约定
    B 给 workflow 的 GITHUB_TOKEN 加 contents: write,再调用 enqueuePullRequest 扩大了 workflow 的写权限。另外,由 GITHUB_TOKEN 建出来的 merge group 会不会触发 CI 没有实测过;如果不会,这个条目会在队列里等到超时再被踢出
    C 保持手动:评论已经写明类别和判定,由人点一次重新排队 不扩权限;但仍然每次要有人动手

    席位意见: 先按 C 落地 #21946,它在 A、B、C 三种选择下都有用。重新排队的动作要等你选 A 或 B;在你决定之前不做。如果选 B,要先实测 GITHUB_TOKEN 建的 merge group 会不会触发 CI。

    你要做的(一个动作): 回复 A、B 或 C。


    Generated by Claude Code

  6. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Ruled, recorded by the PM seat (session_018zT8d8NpiQ1ExhuNd5TxY6) from the maintainer's answer in that Claude Code session, 2026-10-06: C, keep the re-queue manual (「C 保持手动」).

    No credential is added. merge-queue-triage.yml gets no contents: write, and the fleet App token keeps its single consumer. PR #21946 already delivers the classification, the triage comment and the once-per-head re-queue verdict. A person, or a seat, acts on that verdict by re-queuing once. Nothing more is owed on this card.


    Generated by Claude Code

  7. objectstack-fleet commented on Oct 7, 2026

    @objectstack-fleet
    ContributorAuthor

    Triage (owned 24 h re-check): closed completed. Ruling C leaves nothing owed, and PR #21946 delivered the rest

    Triage seat (objectstack-wide, seat post #6015) · session_01AavokzJ5DndAwitDXvKy4U · 2026-10-07T06:09Z. ⛔ Not a claim. The assignee is untouched.

    • Idle for 24 h with an owner: the last write is the ruling record 6010502906 at 2026-10-06T06:12Z. The card carries no pm: state label.
    • What landed: PR ci(merge-queue): name a no-runner red as infra:no-runner, with a once-per-head re-queue verdict #21946 merged as f57627b585, an ancestor of origin/main. It delivered limbs 1 and 3, plus the verdict half of limb 2: the infra:no-runner class from the job record, the triage comment naming it, and the once-per-head re-queue verdict. It left the card open only for the automatic re-queue act.
    • What the ruling settled: C, 「C 保持手动」. No credential is added, the re-queue stays a person's act on the verdict, and the record says "Nothing more is owed on this card."
    • ⇒ Nothing remains. If someone later wants the re-queue act automated, that is a new card with a new ruling, because it would add a credential consumer.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:devpathThe road — create, dev, verify, publish/install, connect an agent, iteratedomain:devxpriority:p2Medium: important, M3tooling

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions