Skip to content

fix(health): move the probe's schedule off GitHub cron to a systemd timer - #138

Merged
Fl0p merged 1 commit into
mainfrom
flo-1001-pi-healthz-timer
Oct 5, 2026
Merged

Fl0p merged 1 commit into
mainfrom
flo-1001-pi-healthz-timer

Conversation

@Fl0p

@Fl0p Fl0p commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What and why

GitHub ran the hourly health-probe schedule once in nine hours. cron: "17 * * * *" landed on main at 00:44Z and by 09:46Z had produced one run, and gh run list --workflow=health-probe.yml --limit 200 shows exactly one event=schedule run in the workflow's whole history - not a cancellation (those stay listed as cancelled) and not the 60-day public-repo deactivation (the repo was active that day). Public repositories are scheduled best-effort and dropped ticks are not delayed, so the detection delay the cron bought was unbounded - a quieter version of the six-day silence this probe was built to end.

The scheduled prober is now a systemd timer on the Pi (~/ops/cotel-healthz.sh, every 10 minutes, pages on the second consecutive failure), probing robmini over the LAN. That also closes the second gap: a job on the deploy host cannot report its own host being off - the run queues, and timeout-minutes does not bound queue time - while from the Pi an absent host is a refused connection.

The timer materializes this repo's scripts/probe-healthz.sh and scripts/page-cotel-health.sh from origin/main at each tick, so the code path stays reviewed here and reaches the timer with no sync step.

Changes

  • .github/workflows/health-probe.yml - the schedule: trigger is gone. workflow_dispatch and the pull-request/push tests stay; both probe jobs now pin the drill markers unconditionally, and the production marker [cotel-health-probe] belongs to the timer.
  • scripts/page-cotel-health.sh - PC_SOURCE_LINE and PC_REPROBE_HINT let a non-Actions caller replace the two Actions-specific halves of the alert body (telling a Pi-paged agent to dispatch a workflow whose runner sits on the possibly-dead host is advice it cannot follow); PC_CALL_ID lets a caller with no run id bucket the recovery idempotency key instead of repeating one key forever. Defaults unchanged, so the workflow behaves exactly as before. The GitHub run URL moved into the context block, which removes an empty paragraph for callers that have none.
  • docs/decisions/0022-... - the mechanism choice, the rejected options (Paperclip routine: ~24 agent runs a day; leave as-is: unbounded delay) and what it costs.
  • docs/operations/health-probe.md - "hourly" and "catching an outage the same day" are gone as guarantees. Three vantage points and their markers, the 10-minute cadence and 2-strike rule, how a tick's exit code splits "watcher broken" from "production red", the credentials table, and the drill run from the Pi.

Verified

  • bash scripts/page-cotel-health_test.sh -> 57/0 (12 new assertions: the overrides land, the defaults survive, PC_CALL_ID precedence)
  • bash scripts/probe-healthz_test.sh -> 11/0
  • Live drill on the Pi against the drill marker: two red ticks (first paged nobody, second opened FLO-1002 with the Pi-specific re-probe instruction), then a green tick that routed the recovery wake. Streak file back to 0.

Summary by CodeRabbit

  • New Features
    • Health monitoring now runs every 10 minutes, with alerts sent after two consecutive failures.
    • Alerts can provide caller-specific source and recheck instructions, and recovery notifications can use a caller-provided identifier.
  • Bug Fixes
    • GitHub-triggered health probes no longer send production alerts; manual runs use separate drill alerts. Edge probe exit code 4 is reported as a warning rather than a failed verdict.
  • Documentation
    • Updated health probe operations guidance and added a decision record describing the scheduler and alerting behavior.

@coderabbitai

coderabbitai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 5feda64e-0282-4cc4-8c77-97e7a2787a19
📥 Commits

Reviewing files that changed from the base of the PR and between 4703042 and cce73b3.

📒 Files selected for processing (8)
  • .github/workflows/health-probe.yml
  • docs/.vitepress/config.js
  • docs/decisions/0022-health-probe-scheduler-outside-github.md
  • docs/decisions/index.md
  • docs/index.md
  • docs/operations/health-probe.md
  • scripts/page-cotel-health.sh
  • scripts/page-cotel-health_test.sh
 _______________________________________________
< Now streaming live: defusing your code bombs. >
 -----------------------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…imer

GitHub ran the hourly schedule once in nine hours: public repositories are
scheduled best-effort and dropped ticks are not delayed, so the detection
delay the cron bought was unbounded - a quieter version of the six-day
silence the probe was built to end. Measured on main, with cancellation and
the 60-day public-repo deactivation excluded.

The scheduled prober is now a systemd timer on the Pi (~/ops, every 10
minutes, pages on the second consecutive failure) that probes robmini over
the LAN, so it also catches the host being off - which the self-hosted
loopback job structurally cannot report, since the run queues instead of
failing. It materializes this repo's probe and pager from origin/main at
each tick, so the code path stays reviewed here.

This workflow keeps workflow_dispatch and its pull-request tests and no
longer schedules anything; both probe jobs now page drill markers only, and
the production marker belongs to the timer. The pager takes PC_SOURCE_LINE
and PC_REPROBE_HINT so a non-Actions caller can replace the two
Actions-specific halves of the alert body - telling a Pi-paged agent to
dispatch a workflow whose runner is on the dead host is advice it cannot
follow - and PC_CALL_ID so a caller with no run id can bucket the recovery
idempotency key instead of repeating one key forever.

Docs: ADR-0022 records the mechanism choice and what it costs; the runbook
drops "hourly" and "the same day" as guarantees, names the three vantage
points and their markers, and documents the credential and the drill from
the Pi.

Verified: pager tests 57/0, probe tests 11/0.

Co-Authored-By: Daedalus <daedalus@agents.flopbut.local>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@Fl0p
Fl0p force-pushed the flo-1001-pi-healthz-timer branch from 2e570b8 to cce73b3 Compare October 5, 2026 10:07
@Fl0p
Fl0p merged commit ace2a50 into main Oct 5, 2026
7 of 8 checks passed
@Fl0p
Fl0p deleted the flo-1001-pi-healthz-timer branch October 5, 2026 10:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant