Repository navigation
fix(health): move the probe's schedule off GitHub cron to a systemd timer - #138
Merged
Merged
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configuration
📒 Files selected for processing (8)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…imer GitHub ran the hourly schedule once in nine hours: public repositories are scheduled best-effort and dropped ticks are not delayed, so the detection delay the cron bought was unbounded - a quieter version of the six-day silence the probe was built to end. Measured on main, with cancellation and the 60-day public-repo deactivation excluded. The scheduled prober is now a systemd timer on the Pi (~/ops, every 10 minutes, pages on the second consecutive failure) that probes robmini over the LAN, so it also catches the host being off - which the self-hosted loopback job structurally cannot report, since the run queues instead of failing. It materializes this repo's probe and pager from origin/main at each tick, so the code path stays reviewed here. This workflow keeps workflow_dispatch and its pull-request tests and no longer schedules anything; both probe jobs now page drill markers only, and the production marker belongs to the timer. The pager takes PC_SOURCE_LINE and PC_REPROBE_HINT so a non-Actions caller can replace the two Actions-specific halves of the alert body - telling a Pi-paged agent to dispatch a workflow whose runner is on the dead host is advice it cannot follow - and PC_CALL_ID so a caller with no run id can bucket the recovery idempotency key instead of repeating one key forever. Docs: ADR-0022 records the mechanism choice and what it costs; the runbook drops "hourly" and "the same day" as guarantees, names the three vantage points and their markers, and documents the credential and the drill from the Pi. Verified: pager tests 57/0, probe tests 11/0. Co-Authored-By: Daedalus <daedalus@agents.flopbut.local> Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fl0p
force-pushed
the
flo-1001-pi-healthz-timer
branch
from
October 5, 2026 10:07
2e570b8 to
cce73b3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why
GitHub ran the hourly health-probe schedule once in nine hours.
cron: "17 * * * *"landed onmainat 00:44Z and by 09:46Z had produced one run, andgh run list --workflow=health-probe.yml --limit 200shows exactly oneevent=schedulerun in the workflow's whole history - not a cancellation (those stay listed ascancelled) and not the 60-day public-repo deactivation (the repo was active that day). Public repositories are scheduled best-effort and dropped ticks are not delayed, so the detection delay the cron bought was unbounded - a quieter version of the six-day silence this probe was built to end.The scheduled prober is now a systemd timer on the Pi (
~/ops/cotel-healthz.sh, every 10 minutes, pages on the second consecutive failure), probing robmini over the LAN. That also closes the second gap: a job on the deploy host cannot report its own host being off - the run queues, andtimeout-minutesdoes not bound queue time - while from the Pi an absent host is a refused connection.The timer materializes this repo's
scripts/probe-healthz.shandscripts/page-cotel-health.shfromorigin/mainat each tick, so the code path stays reviewed here and reaches the timer with no sync step.Changes
.github/workflows/health-probe.yml- theschedule:trigger is gone.workflow_dispatchand the pull-request/push tests stay; both probe jobs now pin the drill markers unconditionally, and the production marker[cotel-health-probe]belongs to the timer.scripts/page-cotel-health.sh-PC_SOURCE_LINEandPC_REPROBE_HINTlet a non-Actions caller replace the two Actions-specific halves of the alert body (telling a Pi-paged agent to dispatch a workflow whose runner sits on the possibly-dead host is advice it cannot follow);PC_CALL_IDlets a caller with no run id bucket the recovery idempotency key instead of repeating one key forever. Defaults unchanged, so the workflow behaves exactly as before. The GitHub run URL moved into the context block, which removes an empty paragraph for callers that have none.docs/decisions/0022-...- the mechanism choice, the rejected options (Paperclip routine: ~24 agent runs a day; leave as-is: unbounded delay) and what it costs.docs/operations/health-probe.md- "hourly" and "catching an outage the same day" are gone as guarantees. Three vantage points and their markers, the 10-minute cadence and 2-strike rule, how a tick's exit code splits "watcher broken" from "production red", the credentials table, and the drill run from the Pi.Verified
bash scripts/page-cotel-health_test.sh-> 57/0 (12 new assertions: the overrides land, the defaults survive,PC_CALL_IDprecedence)bash scripts/probe-healthz_test.sh-> 11/0Summary by CodeRabbit