Skip to content

feat(node): add opt-in local health endpoint (V2-1380) - #244

Draft
JimCollinson wants to merge 3 commits into
mainfrom
jimcollinson/v2-1380-ant-node-health-endpoint
Draft

JimCollinson wants to merge 3 commits into
mainfrom
jimcollinson/v2-1380-ant-node-health-endpoint

Conversation

@JimCollinson

@JimCollinson JimCollinson commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

The node already knows about its connectivity, stored chunks and uptime, but local tools cannot ask for that information. Operators and applications instead see only whether the process is running, or must inspect logs. Logging is optional, and log messages are not a stable interface for other software.

This PR makes the existing metrics-port setting useful: /health gives applications a small, structured answer, and /metrics gives monitoring tools the same information. Both are accessible only on the same machine. A separate ant-client follow-up can use this to show richer node status in the CLI.

It reuses information the node already holds, adds no dependencies, and is off by default. Existing deployments stay unchanged unless an operator enables it; an unavailable port never prevents the node from running.

Linear issue

Closes V2-1380

Issue and specification

Risk tier

  • T0 — docs / tooling / CI / pure UX-output. Repo CI only.
  • T1 — client-only, no network-facing behavior change. CI + prod compat smoke.
  • T2 — node/client logic with behavioral surface, no protocol/format/economics change. Dev testnet + ADR.
  • T3 — protocol / storage format / payments / routing. T2 evidence + adversarial testing.

Activate the existing metrics-port setting for local health queries and Prometheus scraping, using metric names compatible with existing deployment tools where their meanings match. The endpoint is opt-in, read-only, and bound only to IPv4 loopback. No listener is opened by default; the existing deployment scripts already pass explicit metrics ports. Risk tier is proposed for maintainer confirmation.

Compatibility

  • Wire: none; no changes to node-to-node or client protocols.
  • Storage: no stored-data format changes. One new discovery file, {root_dir}/metrics.port, published after successful binding and cleaned up best-effort on shutdown, disabled startup, or bind failure. It is a discovery hint, not proof of health.
  • API: new local GET /health JSON and GET /metrics Prometheus text endpoints on the existing configured metrics port. Config and CLI defaults change from 9100 to 0 (off); enable explicitly by config, --metrics-port, or ANT_METRICS_PORT. Changes apply at restart. Explicit CLI/environment values override the metrics config setting; an omitted override preserves it.

Both forms expose existing in-memory state only: identity/version, uptime, initial bootstrap completion, connected peers, routing-table size, storage enabled, current chunk count, and process-lifetime write/read activity. chunks_current is occupancy; _total fields reset on restart. Storage-read totals are not counts of completed client downloads. Live bytes are deliberately deferred.

Requests are handled sequentially, with bounded headers, a whole-request timeout, and active-request shutdown cancellation. Bind or discovery-file failures do not fail node startup. Remote worker-IP scrapers cannot reach a loopback listener; dashboard panels for deferred values remain empty. No deployment, client, storage instrumentation, CI, or dependency changes.

Follow-up polish requires one local Host header (127.0.0.1 or localhost, optionally the bound port) before returning telemetry, adds Allow: GET to 405 replies, and tests actual environment-variable precedence in isolated subprocesses. Untrusted/missing/duplicate hosts return 403. Normal curl and Prometheus requests satisfy the header contract. Connection-accept retry behaviour is deliberately unchanged.

Semver impact

  • breaking
  • feature
  • fix

Test evidence

Local evidence for follow-up commit 3c67abe7bc77b08cfda35161ce13b3da81d6ba9e, using Rust 1.98.1 and unchanged build settings, independently repeated by a fresh clean-context reviewer:

Command Result
cargo fmt --all -- --check Passed
cargo clippy --all-targets --all-features -- -D warnings Passed
cargo test --lib --features test-utils 1,232 passed
cargo test --lib --no-default-features 1,187 passed
cargo test --lib --features test-utils health_ 4 passed, 1,228 filtered
cargo test --bin ant-node 8 passed
cargo test --bin ant-node metrics_port_ 3 passed, 5 filtered
GITHUB_BASE_REF=main python3 scripts/adr-governance.py Passed

The real-node library tests cover both representations and response headers, field types, disabled mode, stale discovery cleanup, two-node port collision without startup failure, fragmented requests, active-request cancellation, joined cleanup, and socket release. Golden strings verify JSON and Prometheus naming/types. CLI tests cover default-off behavior, preserving the configured port, and explicit enabling/disabling overrides.

Polish tests cover trusted/rejected/duplicate Host values on both endpoints, no telemetry in rejection responses, Allow: GET, and actual environment enabling/explicit-zero disabling/CLI precedence. The subprocess test runs locally; existing CI does not execute binary tests, so CI green alone would not establish cross-platform execution of that test.

Executable development-mode smoke: one node with storage enabled answered both endpoints on loopback; unknown GET returned 404, POST returned 405, and the advertised port matched. A second separately launched node without a metrics setting opened no TCP listener. Orderly shutdown removed discovery and released the ports. This used an intentionally unavailable local bootstrap, so it demonstrates endpoint/lifecycle behavior, not connected-network or paid-storage operation. Connected dev-testnet evidence remains pending for T2 review.

Independent Craft Review found no findings; adversarial review findings were recorded and dispositioned; the final Claude clean-context review returned Pass. An earlier cached-workspace test build failed during linking; the unchanged committed candidate subsequently passed the full suites in a fresh checkout and independent review. The exact cause of that earlier local failure is unproven; no test expectations, build flags, or environment workarounds were changed to obtain passing results.

Draft status: the updated base 41b9848 passed all 20 checks, including Windows/btrfs and Rust 1.91. That result does not cover this follow-up. Current CI has completed with failures after its stable compiler updated to Rust 1.99 between runs; the pinned Rust 1.91 check passed.

CI attribution checked against unchanged main: a clean checkout of GitHub's base d407edd421d1f30198f804bd089db8ebd3f47add, containing none of this PR's changes, reproduces the same test-build, documentation and Clippy failures using Rust/Clippy 1.99 and CI's warnings-denied settings. All three report fetch_update deprecations at src/web_rtc.rs:170/202; Clippy also reproduces the same existing empty-assert complaints (29 library-test errors in total). This establishes that these inspected failures also occur without our contribution. The comparison was run on macOS, not every CI platform. No compiler pin, warning suppression or unrelated source fix is included here; current CI remains failed and is not being waived. Not requesting merge until required evidence and maintainer review are complete.

New dependency

None.

ADR

ADR-0017: Local node health endpoint (ant-node). Status: Proposed; human acceptance remains required.

Mitigation / rollback

off by default; set metrics_port = 0 and restart; revert one module plus the default.

grumbach pushed a commit to grumbach/ant-node that referenced this pull request Oct 2, 2026
Rust 1.99 broke CI on `main`: `Clippy`, `Documentation`, all three `Build`/`Test` matrices, the
three storage-filesystem jobs and the WebRTC devnet job fail, because the workflow sets
`RUSTFLAGS: -D warnings` and the new toolchain reports two new lint families here. The same tree
passed every job on 1.98.1, and PR WithAutonomi#244 is blocked by it rather than by anything in that PR.

`Atomic::fetch_update` is deprecated, renamed `try_update` for consistency with the new infallible
`update`. It is the same method under a new name, but only stable from 1.95, so the rename takes
the MSRV with it.

`assert_is_empty` is a new pedantic lint, picked up through the blanket `pedantic = "warn"`. Its
point is that a bare `assert!` prints nothing useful on failure; clippy's suggestion is of the form
`assert_eq!(dirs, [] as [String; 0])`, so these use a message carrying the value instead, which
gives the same diagnostic and reads better in a test.

- rename the two `fetch_update` calls in `web_rtc`
- raise `rust-version` to 1.95 and the MSRV job to 1.95.0, and update the MSRV quoted in
  `README.md` and `docs/WEBRTC_DIRECT_TESTNET.md`
- give the 37 flagged assertions a failure message, binding a local first where the subject was a
  method call so the value can be printed

All assertion conditions are unchanged, so every test passes or fails exactly as before.

Verified on 1.99.0: `cargo fmt --all -- --check`, `cargo clippy --all-targets --all-features
--keep-going -- -D warnings` and `RUSTDOCFLAGS="-D warnings" cargo doc --all-features --no-deps`,
all clean. The sweep used `--keep-going` deliberately: clippy aborts at the first failing target,
so the lints hid behind one another across `src`, `tests/e2e` and the devnet tests.

Closes V2-1399

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant