Skip to content

feat(cluster): respin instances with stalled nodes - #3712

Merged
mkoura merged 1 commit into
masterfrom
stalled_node_respin
Sep 25, 2026
Merged

mkoura merged 1 commit into
masterfrom
stalled_node_respin

Conversation

@mkoura

@mkoura mkoura commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator

A node whose tip is older than the forecast horizon (3k/f slots) has no ledger view for the current slot and cannot forge anymore. When it happens to all nodes, the chain is halted for good; a single node in that state is usually stuck on a fork deeper than k. Tests on such instance either fail on the missing UTxOs, or hang waiting for new blocks.

Treat such instance as unhealthy, so it is respun once the running tests finish and no new tests start on it meanwhile.

The time of the last block is taken from the newest write to the node's volatile DB, which covers also the blocks downloaded during a sync. Only running nodes are checked, as tests stop nodes on purpose, and a node is measured from the start of its process at the earliest, so a node restarted after a long stop has time to catch up. A node whose volatile DB can't be found or read is not reported. The genesis parameters are cached, as the check runs under the global cluster lock.

@mkoura
mkoura requested a review from saratomaz as a code owner September 25, 2026 11:09
@mkoura
mkoura requested a lite review from Copilot and removed request for saratomaz September 25, 2026 11:11

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Genesis parsing errors remain unhandled, and two tests need stronger mtime isolation.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity · 2 Low severity

Open (3)
What changed in this PR

Adds stalled-node detection based on genesis parameters, process uptime, and volatile database activity, triggering respins for unhealthy instances.

Changes:

  • Detects nodes that stop advancing.
  • Integrates detection into cluster health and respin decisions.
  • Adds detection and health-check tests.
File Changes Review notes
framework_tests/​test_cluster_nodes.py Tests uptime and volatile DB activity. Test directory mtimes should be reset to isolate file-mtime behavior.
framework_tests/​test_cluster_getter.py Tests health-check integration. No findings.
cardano_node_tests/​utils/​cluster_nodes.py Implements stalled-node detection and state resolution. Handle TypeError and OverflowError when parsing malformed genesis data.
cardano_node_tests/​cluster_management/​cluster_getter.py Integrates stalled nodes into respin decisions. No findings.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread cardano_node_tests/utils/cluster_nodes.py
Comment thread framework_tests/test_cluster_nodes.py
Comment thread framework_tests/test_cluster_nodes.py
@mkoura
mkoura force-pushed the stalled_node_respin branch from 2297395 to bcb3468 Compare September 25, 2026 11:21
A node whose tip is older than the forecast horizon (3k/f slots) has
no ledger view for the current slot and cannot forge anymore. When it
happens to all nodes, the chain is halted for good; a single node in
that state is usually stuck on a fork deeper than k. Tests on such
instance either fail on the missing UTxOs, or hang waiting for new
blocks.

Treat such instance as unhealthy, so it is respun once the running
tests finish and no new tests start on it meanwhile.

The time of the last block is taken from the newest write to the
node's volatile DB, which covers also the blocks downloaded during a
sync. Only running nodes are checked, as tests stop nodes on purpose,
and a node is measured from the start of its process at the earliest,
so a node restarted after a long stop has time to catch up. A node
whose volatile DB can't be found or read is not reported. The genesis
parameters are cached, as the check runs under the global cluster
lock.
@mkoura
mkoura force-pushed the stalled_node_respin branch from bcb3468 to ac1ee64 Compare September 25, 2026 11:27
@mkoura
mkoura merged commit 45a2021 into master Sep 25, 2026
3 checks passed
@mkoura
mkoura deleted the stalled_node_respin branch September 25, 2026 11:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants