feat(cluster): respin instances with stalled nodes - #3712
Merged
Merged
Conversation
mkoura
requested
a lite review from Copilot
and removed request for
saratomaz
September 25, 2026 11:11
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Genesis parsing errors remain unhandled, and two tests need stronger mtime isolation.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (3)
What changed in this PR
Adds stalled-node detection based on genesis parameters, process uptime, and volatile database activity, triggering respins for unhealthy instances.
Changes:
- Detects nodes that stop advancing.
- Integrates detection into cluster health and respin decisions.
- Adds detection and health-check tests.
| File | Changes | Review notes |
|---|---|---|
framework_tests/test_cluster_nodes.py |
Tests uptime and volatile DB activity. | Test directory mtimes should be reset to isolate file-mtime behavior. |
framework_tests/test_cluster_getter.py |
Tests health-check integration. | No findings. |
cardano_node_tests/utils/cluster_nodes.py |
Implements stalled-node detection and state resolution. | Handle TypeError and OverflowError when parsing malformed genesis data. |
cardano_node_tests/cluster_management/cluster_getter.py |
Integrates stalled nodes into respin decisions. | No findings. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
mkoura
force-pushed
the
stalled_node_respin
branch
from
September 25, 2026 11:21
2297395 to
bcb3468
Compare
A node whose tip is older than the forecast horizon (3k/f slots) has no ledger view for the current slot and cannot forge anymore. When it happens to all nodes, the chain is halted for good; a single node in that state is usually stuck on a fork deeper than k. Tests on such instance either fail on the missing UTxOs, or hang waiting for new blocks. Treat such instance as unhealthy, so it is respun once the running tests finish and no new tests start on it meanwhile. The time of the last block is taken from the newest write to the node's volatile DB, which covers also the blocks downloaded during a sync. Only running nodes are checked, as tests stop nodes on purpose, and a node is measured from the start of its process at the earliest, so a node restarted after a long stop has time to catch up. A node whose volatile DB can't be found or read is not reported. The genesis parameters are cached, as the check runs under the global cluster lock.
mkoura
force-pushed
the
stalled_node_respin
branch
from
September 25, 2026 11:27
bcb3468 to
ac1ee64
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


A node whose tip is older than the forecast horizon (3k/f slots) has no ledger view for the current slot and cannot forge anymore. When it happens to all nodes, the chain is halted for good; a single node in that state is usually stuck on a fork deeper than k. Tests on such instance either fail on the missing UTxOs, or hang waiting for new blocks.
Treat such instance as unhealthy, so it is respun once the running tests finish and no new tests start on it meanwhile.
The time of the last block is taken from the newest write to the node's volatile DB, which covers also the blocks downloaded during a sync. Only running nodes are checked, as tests stop nodes on purpose, and a node is measured from the start of its process at the earliest, so a node restarted after a long stop has time to catch up. A node whose volatile DB can't be found or read is not reported. The genesis parameters are cached, as the check runs under the global cluster lock.