Repository navigation
feat(deploy): read back the declared kubelet settings on every node - #4559
Conversation
A green machine-config sync does not prove every running kubelet took the settings the repository declares: two autoscaler nodes ran without the disk eviction thresholds for weeks and nothing reported it. After each cluster update the deploy now reads every node's live kubelet configuration and compares it with the cluster-wide declared settings. A node that cannot be read is a failure, never a pass. The deploy step starts observe-only, so a finding is a warning until one clean production reading is on record. Part of #3137 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An empty reply printed no differences and read as a match. Every refusal now goes through one exit, so observe-only cannot fail a deploy and skip the steps after it. Drift is reported on the first reading; only a read that did not complete is retried. Role files are read per document. Part of #3137 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
devantler
left a comment
There was a problem hiding this comment.
🤖 Generated by the Agentic Engineer
Self-review (fallback — CodeRabbit, Codex and Cursor Bugbot unavailable)
Reviewed commit: fa0931735eff1aed4e16240cb9bac68a0140e0ea
- CodeRabbit: rate limited. Its replies to the two requests made in this repository at 06:11Z and 06:15Z today read "Review rate limited", with the next included review stated for about 06:49Z; no request was spent here on an established limit.
- Codex: unavailable, usage limit since 2026-10-06T04:40:17Z.
- Cursor Bugbot: unavailable, usage limit since 2026-10-06T04:42:02Z.
The first commit was read by a separate reviewer agent that was given the code and no conclusions. It reproduced two P1 findings by execution, both fixed in this head:
- An empty reply read as a match.
jqprints nothing for an empty document, so a node answering with an empty body counted as having no differences. An empty or blank reply is now refused before comparison, with its own test. - Observe-only could still fail the deploy. Refusals raised before the comparison ignored the observe-only setting, and a red step skips the credential reassert after it. Every non-pass now leaves through one exit that honours the setting. A failed patch read is reported as that, and role files are read per document.
Also changed from its notes: drift is reported on the first reading, and only a read that did not complete is retried, so a real finding no longer waits out the five-minute window on every deploy. The test now asserts the deploy step's condition and its observe-only value.
Checked at this head: the test passes its 26 cases under bash 3.2 and bash 5, shellcheck is clean on both files, and eleven mutations of the script each fail the test. The script issues only get nodes and get --raw reads, and node names are pattern-checked before they enter a URL.
Not verified: the comparison has never run against production. That is why the step is observe-only, and the first deploy's log is the reading.
Verdict: no P0/P1 findings
Evaluation at
|
Evaluation at
|
Why
Two production nodes ran for weeks without the disk-pressure protection every other node had, after the fix for it had deployed successfully. Nothing reported the difference, and it was only found by reading nine nodes by hand. The nodes that drifted were the automatically added ones, and the same gap would reopen on any future change to node settings.
What
After each production deploy updates the nodes, the deploy now reads back the settings every node is actually running and compares them with what the repository declares, automatically added nodes included, and a node that cannot be read counts as a failure. The check starts in observe-only mode, so a difference shows up as a warning on the deploy and does not block it, because the comparison has not yet been read against production. A follow-up makes it blocking once one deploy has logged a clean pass.
Part of #3137
👉 After merge/promotion: the first deploy's log answers the issue's open question about whether all current nodes carry the setting; switching the check to blocking is tracked on the issue.