Repository navigation
[Bug][HStore] Server exits 1 on cold start when Stores miss the 38s partition-lookup retry ceiling #3203
Description
Activity
Reproduced outside Kubernetes, same outcome
I reproduced this on 3 VMs (1 PD + 3 Store + 1 Server, tarball built from master
60c8803, no Docker) by starting the Server right after PD and delaying the Store registration by a fixed number of seconds. The results are deterministic:Server config stores register (after PD) result usePD=true, first boot of the cluster1 store at +5 s, the other two at +71 s 10 attempts in 38 s (sleeps 1,1,1,2,3,4,5,6,7,8 s), upper limit : 10, exit 1usePD=true, first boot of the cluster1 store at +5 s, the other two at +26 s 9 attempts, the 9th succeeds, Server starts local conf/graphs, nousePDall stores at +20 / +45 / +60 s Server up after 8 s, no error 105 at all; the missing stores only show up as code 102 in the task-db-workerthreadWith zero registered stores PD answers 102 (There is no any online store), with one or two it answers 105; the client treats both the same way.
Script and logs:
cluster/repro_coldstart.sh,results/issue-3203/.What the partition lookup on the main thread actually is
The path that ends in
NodeTxSessionProxy.createTableis the creation of the system graphDEFAULT-~sys_graph:GraphManager.loadMetaFromPD → kvStoreInit → createSysGraphIfNeed. That method callscreateGraph(..., init = true)only when PD holds no system-graph config yet, andinitBackend()on hstore isHstoreStore.init → createTable, i.e. the first partition lookup. This explains both observations at once: it can only happen on the very first boot of a cluster, and never on pod replacement, because by then the config is in PD andinitis false. The trigger isusePD=trueinrest-server.properties, which the entrypoint sets fromHG_SERVER_USE_PD.Two small corrections to the "Related" section:
init-storeis not involved here, it skips hstore graphs sinceb9a3dd9; and on the PD side code 105 comes fromStoreNodeService.allocShardsand is thrown only while the shard group of a partition does not exist yet, so also only on a cold cluster.Stdout
Confirmed: without
STDOUT_MODE=true,hugegraph-server.shredirects the JVM's stdout and stderr tologs/hugegraph-server-stdout.log, and the image's entrypoint does not enable that mode. All that reachesdocker logsis the line fromstart-hugegraph.sh.Direction I would consider for a fix
- Since this is a one-off bootstrap path, a gate "wait until PD reports at least
minStoreCountactive stores" right beforecreateSysGraphIfNeed(or inHstoreSessionsImpl.open), with a configurable timeout and a log line every few seconds.PDClient.getActiveStores()andgetPDConfig()already expose what is needed. - I would not change
NODE_MAX_RETRYING_TIMESor the general retry loop for this: the same counter handles failures during normal operation, and fix(store-client): enhance thread interrupts in NodeTxExecutor #3204 touches that loop right now; two PRs changing its semantics at the same time would be hard to review. STDOUT_MODE=truein the image's entrypoint as a separate one-line change.
If you want, I can try to implement such a gate with a test and measure it with the same script; if you would rather do it yourself, I am happy to test your patch against the same reproduction.
- Since this is a one-off bootstrap path, a gate "wait until PD reports at least
Server-side fix in #3210: with
usePD=true,GraphManagerwaits forpd.initial-store-countactive stores in PD before it opens the first hstore graph, bounded by a new optionpd.stores_wait_timeout(default 300 s, 0 restores today's behaviour). Measured on the same reproduction as above (1 PD + 3 Store + 1 Server from a tarball, stores registering +5 s and +71 s after PD,usePD=true, first boot):server result master 1a15e76218 × error code = 105, backoff 1,1,1,2,3,4,5,6,7,8,upper limit : 10, exit 1 after 38 s1a15e762+ #3210PD needs 3 active store(s) … waiting up to 300s, 14 progress lines (0/3 → 1/3 → 3/3),3 active store(s) in PD after 70s, 0 × 105, exit 0 after 81 s, REST 200+ #3210, pd.stores_wait_timeout=20, stores at +120 s5 progress lines, then Timed out after 20s waiting for 3 active store(s) in PD (0 registered); start the stores first or raise pd.stores_wait_timeout, exit 1 after 31 sTwo things from the measurement. The wait has to sit before any graph is opened, not only before the system graph is created: with a local
conf/graphshstore graph, that local graph is the first to hit the 10-retry ceiling. AndgetPDConfig()does not sendmin_store_count(ConfigServicebuilds the response frompartition_countandshard_countonly), so the server takesshard_count; on a cluster wherepd.initial-store-countdiffers from the replica count PD should fill that field, a one-line follow-up on the PD side. Logs and script:results/issue-3203/fix/in https://github.com/SebastianGruza/hugegraph-validation. The[wait-storage]gate from #3132 stays the right thing for images; this change covers the tarball and any start without an orchestrator.Reacted by imbajin
Bug Type (问题类型)
server status (启动/运行异常)
Before submit
Environment (环境信息)
hugegraph/server:latestpublished image, built 2026-09-08,hg-store-client-1.7.0.jar. Also seen on images built from master60c8803dand from6ec19838.Expected & Actual behavior (期望与实际表现)
Expected. On a cold start of a distributed cluster, a Server that finds the Stores have
not yet registered with PD waits for them and then starts.
Actual. The Server asks PD for partition information, gets
error code = 105, The number of active stores is less then 3, retries 10 times over 38seconds, hits a hard retry ceiling, and the JVM exits 1. Under Kubernetes the container
is restarted and usually succeeds on the second attempt, so the cluster ends up healthy and
the only lasting trace is
RESTARTS: 1. Outside an orchestrator that restarts it, the Serverstays down.
This is a startup race, not a timeout. It is not related to the entrypoint startup budget
from #3186: the log contains no
The operation timed out(...)line, and the value was 450 shere.
Case B is the reproduction detailed below. Case A stands for the successful path, which is
what the other two Servers on that same install did; its T+64 s is illustrative, while the
fixed timings (T+42, T+80, T+85) are measured.
Timeline from one reproduction
error code = 105x10, backoff 1,1,1,2,3,4,5,6,7,8 s = 38 sNodeTxExecutor - the number of retries reached the upper limit : 10Log and stack trace
The restarted container hits the same error 9 more times and then succeeds, because by
then the Stores have registered.
Why it is intermittent
It is a race between how long the Stores take to register with PD and a fixed 38 second
budget that begins about 42 seconds after the Server container starts. Rates observed across
four campaigns on two different hosts:
So 3 of 12 fresh Server starts failed, and 3 of the 4 installs produced at least one
crashed Server. The per-install figure is the one an operator meets.
Earlier campaigns saw a similar symptom that is deliberately not counted above. On
2026-08-29, two of six Server starts exited 1 on install, but each ran about 2 min 45 s
before dying, which fits the 120 s entrypoint self-kill of #3186 rather than this 38 s retry
ceiling, and the logs were lost so neither can be attributed. #3187 has since fixed that
one. If those two turn out to be this bug as well the rate is higher, not lower.
(Corrected after first posting. The first two rows originally read 6/2 and 10/0. The 6/2
was carried over from #3186, which is the entrypoint's 120 s self-kill and a different
failure, and the 10 was a pod count rather than a count of Server starts. Every campaign
deployed exactly 3 Servers.)
It does not reproduce on pod replacement. 66 Server replacements against an already
running cluster produced zero failures, because a replacement Server asks PD and gets a
straight answer in about 16 seconds. Only a cold cluster has the window. Anyone trying to
reproduce this must recreate the whole cluster, not restart a Server.
Steps to reproduce
at least one Server with
RESTARTS: 1andlastState.terminated.exitCode: 1.A note on diagnosing this
The message an operator actually sees is only:
That file is inside the container and the container is gone, so
kubectl logs --previousreturns the line above and nothing else. The stack trace in this report was only obtainable
by mounting a volume at
/hugegraph-server/logsbefore the first boot, so the crashed run'slog survived the restart. Logging the fatal startup cause to stdout would make this class
of failure diagnosable without that trick.
Suggested direction
These are suggestions rather than a diagnosis of the right fix:
HgStoreClientConst.NODE_MAX_RETRYING_TIMES = 10(
hugegraph-store/hg-store-client/.../util/HgStoreClientConst.java:48), used only byNodeTxExecutorat lines 376 and 383. Nothing reads it from configuration or theenvironment, so 38 seconds is the whole budget an operator gets. A cold cluster bootstrap
can reasonably take longer than that. Either make the budget configurable, or use a longer
one on the bootstrap path where "stores not registered yet" is an expected transient
rather than a fault.
error code = 105during startup as "not ready yet" rather than as a retryablefailure with a short ceiling, since it is precisely the condition that resolves on its own.
Orchestration can also side-step this by not starting the Server until PD reports the
expected number of active stores. I will carry that as a gate in the Helm chart in #3131 /
#3132, but it is a workaround for the retry ceiling, not a fix for it.
Related
HgStoreNodePartitionerImpl,NodeTxExecutor) but adifferent trigger: stale DNS after Store pod replacement, not a cold-start race.
Vertex/Edge example (问题点 / 边数据举例)
Not applicable, the cluster is empty.
Schema [VertexLabel, EdgeLabel, IndexLabel] (元数据结构)
Not applicable, the failure happens before any schema exists.