Repository navigation
feat(node): add opt-in local health endpoint (V2-1380) - #244
Draft
JimCollinson wants to merge 3 commits into
Draft
JimCollinson wants to merge 3 commits into
JimCollinson wants to merge 3 commits into
Conversation
2 of 7 tasks
grumbach
pushed a commit
to grumbach/ant-node
that referenced
this pull request
Oct 2, 2026
Rust 1.99 broke CI on `main`: `Clippy`, `Documentation`, all three `Build`/`Test` matrices, the three storage-filesystem jobs and the WebRTC devnet job fail, because the workflow sets `RUSTFLAGS: -D warnings` and the new toolchain reports two new lint families here. The same tree passed every job on 1.98.1, and PR WithAutonomi#244 is blocked by it rather than by anything in that PR. `Atomic::fetch_update` is deprecated, renamed `try_update` for consistency with the new infallible `update`. It is the same method under a new name, but only stable from 1.95, so the rename takes the MSRV with it. `assert_is_empty` is a new pedantic lint, picked up through the blanket `pedantic = "warn"`. Its point is that a bare `assert!` prints nothing useful on failure; clippy's suggestion is of the form `assert_eq!(dirs, [] as [String; 0])`, so these use a message carrying the value instead, which gives the same diagnostic and reads better in a test. - rename the two `fetch_update` calls in `web_rtc` - raise `rust-version` to 1.95 and the MSRV job to 1.95.0, and update the MSRV quoted in `README.md` and `docs/WEBRTC_DIRECT_TESTNET.md` - give the 37 flagged assertions a failure message, binding a local first where the subject was a method call so the value can be printed All assertion conditions are unchanged, so every test passes or fails exactly as before. Verified on 1.99.0: `cargo fmt --all -- --check`, `cargo clippy --all-targets --all-features --keep-going -- -D warnings` and `RUSTDOCFLAGS="-D warnings" cargo doc --all-features --no-deps`, all clean. The sweep used `--keep-going` deliberately: clippy aborts at the first failing target, so the lints hid behind one another across `src`, `tests/e2e` and the devnet tests. Closes V2-1399 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2 of 7 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The node already knows about its connectivity, stored chunks and uptime, but local tools cannot ask for that information. Operators and applications instead see only whether the process is running, or must inspect logs. Logging is optional, and log messages are not a stable interface for other software.
This PR makes the existing metrics-port setting useful:
/healthgives applications a small, structured answer, and/metricsgives monitoring tools the same information. Both are accessible only on the same machine. A separate ant-client follow-up can use this to show richer node status in the CLI.It reuses information the node already holds, adds no dependencies, and is off by default. Existing deployments stay unchanged unless an operator enables it; an unavailable port never prevents the node from running.
Linear issue
Closes V2-1380
Issue and specification
Risk tier
Activate the existing metrics-port setting for local health queries and Prometheus scraping, using metric names compatible with existing deployment tools where their meanings match. The endpoint is opt-in, read-only, and bound only to IPv4 loopback. No listener is opened by default; the existing deployment scripts already pass explicit metrics ports. Risk tier is proposed for maintainer confirmation.
Compatibility
{root_dir}/metrics.port, published after successful binding and cleaned up best-effort on shutdown, disabled startup, or bind failure. It is a discovery hint, not proof of health.GET /healthJSON andGET /metricsPrometheus text endpoints on the existing configured metrics port. Config and CLI defaults change from9100to0(off); enable explicitly by config,--metrics-port, orANT_METRICS_PORT. Changes apply at restart. Explicit CLI/environment values override the metrics config setting; an omitted override preserves it.Both forms expose existing in-memory state only: identity/version, uptime, initial bootstrap completion, connected peers, routing-table size, storage enabled, current chunk count, and process-lifetime write/read activity.
chunks_currentis occupancy;_totalfields reset on restart. Storage-read totals are not counts of completed client downloads. Live bytes are deliberately deferred.Requests are handled sequentially, with bounded headers, a whole-request timeout, and active-request shutdown cancellation. Bind or discovery-file failures do not fail node startup. Remote worker-IP scrapers cannot reach a loopback listener; dashboard panels for deferred values remain empty. No deployment, client, storage instrumentation, CI, or dependency changes.
Follow-up polish requires one local
Hostheader (127.0.0.1orlocalhost, optionally the bound port) before returning telemetry, addsAllow: GETto 405 replies, and tests actual environment-variable precedence in isolated subprocesses. Untrusted/missing/duplicate hosts return 403. Normal curl and Prometheus requests satisfy the header contract. Connection-accept retry behaviour is deliberately unchanged.Semver impact
Test evidence
Local evidence for follow-up commit
3c67abe7bc77b08cfda35161ce13b3da81d6ba9e, using Rust 1.98.1 and unchanged build settings, independently repeated by a fresh clean-context reviewer:cargo fmt --all -- --checkcargo clippy --all-targets --all-features -- -D warningscargo test --lib --features test-utilscargo test --lib --no-default-featurescargo test --lib --features test-utils health_cargo test --bin ant-nodecargo test --bin ant-node metrics_port_GITHUB_BASE_REF=main python3 scripts/adr-governance.pyThe real-node library tests cover both representations and response headers, field types, disabled mode, stale discovery cleanup, two-node port collision without startup failure, fragmented requests, active-request cancellation, joined cleanup, and socket release. Golden strings verify JSON and Prometheus naming/types. CLI tests cover default-off behavior, preserving the configured port, and explicit enabling/disabling overrides.
Polish tests cover trusted/rejected/duplicate Host values on both endpoints, no telemetry in rejection responses,
Allow: GET, and actual environment enabling/explicit-zero disabling/CLI precedence. The subprocess test runs locally; existing CI does not execute binary tests, so CI green alone would not establish cross-platform execution of that test.Executable development-mode smoke: one node with storage enabled answered both endpoints on loopback; unknown GET returned 404, POST returned 405, and the advertised port matched. A second separately launched node without a metrics setting opened no TCP listener. Orderly shutdown removed discovery and released the ports. This used an intentionally unavailable local bootstrap, so it demonstrates endpoint/lifecycle behavior, not connected-network or paid-storage operation. Connected dev-testnet evidence remains pending for T2 review.
Independent Craft Review found no findings; adversarial review findings were recorded and dispositioned; the final Claude clean-context review returned Pass. An earlier cached-workspace test build failed during linking; the unchanged committed candidate subsequently passed the full suites in a fresh checkout and independent review. The exact cause of that earlier local failure is unproven; no test expectations, build flags, or environment workarounds were changed to obtain passing results.
Draft status: the updated base
41b9848passed all 20 checks, including Windows/btrfs and Rust 1.91. That result does not cover this follow-up. Current CI has completed with failures after its stable compiler updated to Rust 1.99 between runs; the pinned Rust 1.91 check passed.CI attribution checked against unchanged main: a clean checkout of GitHub's base
d407edd421d1f30198f804bd089db8ebd3f47add, containing none of this PR's changes, reproduces the same test-build, documentation and Clippy failures using Rust/Clippy 1.99 and CI's warnings-denied settings. All three reportfetch_updatedeprecations atsrc/web_rtc.rs:170/202; Clippy also reproduces the same existing empty-assert complaints (29 library-test errors in total). This establishes that these inspected failures also occur without our contribution. The comparison was run on macOS, not every CI platform. No compiler pin, warning suppression or unrelated source fix is included here; current CI remains failed and is not being waived. Not requesting merge until required evidence and maintainer review are complete.New dependency
None.
ADR
ADR-0017: Local node health endpoint (ant-node). Status: Proposed; human acceptance remains required.
Mitigation / rollback
off by default; set metrics_port = 0 and restart; revert one module plus the default.