Skip to content

feat(deploy): read back the declared kubelet settings on every node - #4559

Merged
devantler merged 2 commits into
mainfrom
claude/kubelet-config-readback-3137
Oct 6, 2026
Merged

devantler merged 2 commits into
mainfrom
claude/kubelet-config-readback-3137

Conversation

@devantler

Copy link
Copy Markdown
Contributor

🤖 Generated by the Agentic Engineer

Why

Two production nodes ran for weeks without the disk-pressure protection every other node had, after the fix for it had deployed successfully. Nothing reported the difference, and it was only found by reading nine nodes by hand. The nodes that drifted were the automatically added ones, and the same gap would reopen on any future change to node settings.

What

After each production deploy updates the nodes, the deploy now reads back the settings every node is actually running and compares them with what the repository declares, automatically added nodes included, and a node that cannot be read counts as a failure. The check starts in observe-only mode, so a difference shows up as a warning on the deploy and does not block it, because the comparison has not yet been read against production. A follow-up makes it blocking once one deploy has logged a clean pass.

Part of #3137

👉 After merge/promotion: the first deploy's log answers the issue's open question about whether all current nodes carry the setting; switching the check to blocking is tracked on the issue.

A green machine-config sync does not prove every running kubelet took the
settings the repository declares: two autoscaler nodes ran without the disk
eviction thresholds for weeks and nothing reported it.

After each cluster update the deploy now reads every node's live kubelet
configuration and compares it with the cluster-wide declared settings. A
node that cannot be read is a failure, never a pass. The deploy step starts
observe-only, so a finding is a warning until one clean production reading
is on record.

Part of #3137

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An empty reply printed no differences and read as a match. Every refusal
now goes through one exit, so observe-only cannot fail a deploy and skip
the steps after it. Drift is reported on the first reading; only a read
that did not complete is retried. Role files are read per document.

Part of #3137

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@devantler devantler left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Generated by the Agentic Engineer

Self-review (fallback — CodeRabbit, Codex and Cursor Bugbot unavailable)

Reviewed commit: fa0931735eff1aed4e16240cb9bac68a0140e0ea

  • CodeRabbit: rate limited. Its replies to the two requests made in this repository at 06:11Z and 06:15Z today read "Review rate limited", with the next included review stated for about 06:49Z; no request was spent here on an established limit.
  • Codex: unavailable, usage limit since 2026-10-06T04:40:17Z.
  • Cursor Bugbot: unavailable, usage limit since 2026-10-06T04:42:02Z.

The first commit was read by a separate reviewer agent that was given the code and no conclusions. It reproduced two P1 findings by execution, both fixed in this head:

  • An empty reply read as a match. jq prints nothing for an empty document, so a node answering with an empty body counted as having no differences. An empty or blank reply is now refused before comparison, with its own test.
  • Observe-only could still fail the deploy. Refusals raised before the comparison ignored the observe-only setting, and a red step skips the credential reassert after it. Every non-pass now leaves through one exit that honours the setting. A failed patch read is reported as that, and role files are read per document.

Also changed from its notes: drift is reported on the first reading, and only a read that did not complete is retried, so a real finding no longer waits out the five-minute window on every deploy. The test now asserts the deploy step's condition and its observe-only value.

Checked at this head: the test passes its 26 cases under bash 3.2 and bash 5, shellcheck is clean on both files, and eleven mutations of the script each fail the test. The script issues only get nodes and get --raw reads, and node names are pattern-checked before they enter a URL.

Not verified: the comparison has never run against production. That is why the step is observe-only, and the first deploy's log is the reading.

Verdict: no P0/P1 findings

@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Evaluation at fa0931735eff1aed4e16240cb9bac68a0140e0ea

Tried as the deploy would run it, against a stand-in cluster that answers the two reads the check makes:

  • Healthy cluster: one line, PASS: 3 node(s), 1 of them autoscaler-provisioned, run every declared kubelet setting (…), naming the settings compared.
  • The drift from Two autoscaler nodes did not pick up the kubelet eviction fix #3137 (an autoscaler node with the defaults switch off and only the memory threshold): the node is named with each difference, healthy nodes are not listed, and it is reported on the first reading without waiting.
  • A node that refuses the read, or answers with nothing: a failure naming the node. This is the case that looked clean when the nodes were read by hand.
  • Observe-only, as the deploy runs it: the same findings appear as one warning and the step stays green, including when the check refuses to run at all.

CI ran the same cases at this head: the 🧾 Validate kubelet settings readback step succeeded in run 37423880982. All other checks are green.

What this cannot show: the comparison has never been read against production, and no identity I hold may read a node's kubelet configuration. The first deploy after merge is that reading, and its step log goes on #3137.

@devantler
devantler marked this pull request as ready for review October 6, 2026 07:16
@devantler
devantler added this pull request to the merge queue Oct 6, 2026
@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Evaluation at fa0931735eff1aed4e16240cb9bac68a0140e0ea

Tried as the operator would meet it, from a checkout of this head:

  • The check's own cases pass under the system Bash (3.2), all of them.
  • Against production, with an identity that may not read kubelet configuration — the exact situation in Two autoscaler nodes did not pick up the kubelet eviction fix #3137, where a refused read printed nothing and looked clean. Every node is now named as "kubelet configuration could not be read". In the mode the deploy uses (observe-only) the step ends green with one warning listing the nodes; with enforcement on, the same reading ends red. A refused read is no longer mistaken for a pass.
  • Not exercised here: a clean reading. Only the deploy identity may read kubelet configuration, so the first deploy after this lands is that reading. The step cannot fail that deploy while it is observe-only.

After merge: read the step's log on the first deploy. A clean PASS is the evidence #3137 asks for; a follow-up then turns enforcement on.

Merged via the queue into main with commit 1b09076 Oct 6, 2026
33 checks passed
@devantler
devantler deleted the claude/kubelet-config-readback-3137 branch October 6, 2026 10:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: ✅ Done

Development

Successfully merging this pull request may close these issues.

1 participant