feat(pointer): create, update, read and resolve pointers - #200
Conversation
bb5f76b to
10ae243
Compare
Pointers (ADR-0015 in ant-node) are mutable, owner-signed references at BLAKE3(domain || owner_key). Public-key addressed and self-verifying: the key is inside the record, so a client can derive its own pointer's address offline and check any answer on the spot. Three things are the client's job, and this is all of them. The counter. A create is counter 0 and every update is exactly one more, so one payment buys one increment. pointer_update reads the current record and steps it rather than guessing, because a guessed counter is precisely what the network refuses. Paying at the state. The quote names the record's state_id while the close group that issues it is the one around the address — pay_for_storage_split carries both, where a chunk passes one address twice. Paying at the address would buy every future update at once, which is the defect this design exists to fix. Following targets. A node stores 33 opaque bytes and never reads them, so chain resolution lives here: pointer_resolve walks to the chunk a chain ends at, refusing to loop and refusing to walk past a bounded depth. Both guards have to be here because the node cannot enforce either. Answers are verified before use: the signature must check out and the record must belong at the address that was asked for, so a storer cannot answer with somebody else's pointer.
…ranch as_bytes became to_bytes upstream — the record no longer caches its encoding — so the client re-encodes at the call sites that send it. The devnet and test dependency on ant-node moved from the published 0.19.0 to the pointer branch. The published node pins ant-protocol 2.4.0, which put two incompatible copies of the wire types in the graph: one for the client's runtime and one for the node the harness spawns. Both now resolve to the same protocol branch, so there is a single ant-protocol in the tree and the node the tests start actually speaks the pointer messages.
Three resilience fixes, all about not trusting one peer. A PUT now stores on a close-group majority rather than stopping at the first success. Nodes do not replicate pointers yet, so the copies that exist are the ones this client wrote — stopping at one meant one node losing the record lost the pointer. A partial store reports a shortfall rather than claiming success. Every acknowledgement is checked against what was sent. A peer answering Success for a different address or state is claiming to hold something this client never submitted; taking that at face value let one peer end the write, after payment, while storing nothing. A read merges the close group's answers under the same merge rule instead of taking the first valid one. Any single peer is entitled to serve a stale record, and taking the first reply let it pin a reader to that record — the fork case the merge rule exists to settle, arriving through the read path instead.
A PUT counted successes from the payment plan's put-targets, which run to twenty peers, while a GET asks only the closest seven. A write could therefore be acknowledged entirely outside the set a read queries — stored, paid for and immediately unreadable. Both now target the strict closest-K set. Both also run concurrently instead of peer by peer. Sequentially, one unreachable peer stalled the whole operation for the store timeout before the next was tried. A read now needs a majority of the close group to answer before it reports a value. With no replication between nodes, one reachable peer holding a stale record is indistinguishable from the whole group agreeing, so a lone answer is reported as a shortfall rather than presented as the value.
A store required CLOSE_GROUP_MAJORITY — a fixed four — while both store and read query a configurable close-group width. At the default seven that is a majority and the two quorums intersect. At twenty it is not: a write could land on four peers and a disjoint four could answer the read, so an acknowledged pointer would read back missing or stale. Both sides now take a majority of the peers actually returned, and a test asserts the two intersect at every width from 1 to 64. A read also returns as soon as that majority has answered, instead of draining every future — one unreachable peer was holding the read open for its whole timeout after the answer was already known — and it uses the read timeout rather than the merkle store timeout.
The pointer tests asserted against copies of the client's logic -- a local fold, a local comparison -- so they could pass while the code they described was wrong. Three pieces are now named functions the tests drive directly: - pointer_group: the one call that picks the peers a write lands on and a read asks. They are the same set by construction now, not by two call sites that happen to agree. - ask_the_group: the shared fan-out. A majority ends the operation and the rest are dropped mid-flight; the tests prove a dead peer cannot stall it by racing four ready answers against three futures that never finish. - read_put_reply / read_get_reply: what one peer's reply means. Fed every reply a dishonest storer can give -- another pointer's record, tampered bytes, an acknowledgement naming a state that was never sent -- and each must be refused. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It waited on the merkle store timeout, 270 seconds. That budget exists because a merkle pool makes the storer run an authoritative network closeness lookup first; a pointer PUT carries a single-node proof -- merkle proofs are refused for pointers outright -- and does no such lookup. It now uses the same STORE_RESPONSE_TIMEOUT a non-merkle chunk PUT uses, so an unreachable peer costs ten seconds rather than four and a half minutes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Everything else about pointers was tested against a function. This runs the client's own create / update / get / resolve over real nodes, real QUIC and real Anvil settlement, which is the only way to show the feature works rather than that its parts do. Six cases: a create is paid for and reads back at the address the client derived offline; each update is paid for at counter + 1 and is what the network then serves; re-submitting a stored state is accepted and moves nothing; a chain of pointers resolves to the chunk at its end and that chunk fetches; an address nobody wrote reads as absent rather than as an error; and a record signed by the owner at a counter that skips ahead is refused, leaving what is held untouched. The harness gives every node a pointer store and service, as a real node has. Without it the nodes refuse every pointer message and the tests would prove nothing about the code a node actually runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit staged these files before the format run, so what landed was the unformatted version and the Format check went red.
A read took the best state any peer offered. An owner can sign a state without paying for it, and one close-group peer willing to serve it was enough: the record verifies, belongs at the address, and wins the merge, so every reader took it while every honest peer answered NotFound. A read now returns a state only when two of the answering peers name it, and a write must reach a majority plus one so that two of them always do -- with |W| + |R| >= K + 2 the write and read sets overlap in at least two peers, checked over every group width from 1 to 64. A pointer costs one more storer than a chunk for the reason a chunk does not need it: one copy of a chunk is the chunk, while a pointer read has to decide which of several signed states is current, and that is not one peer's decision to make. Also: a NotFound reply has to be about the address that was asked for, exactly as an acknowledgement does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every pointer reply names the address it is about, and every other kind was already checked against the address that was sent. A Stale refusal was not: one about somebody else's pointer would have been reported as this write losing a race. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both take their peers from one function, which is what keeps a write from landing outside the set a read asks. But each does its own lookup, so membership that changed in between is not covered -- that is churn, and what covers churn is replication, which is not built. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…grees on The corroboration rule tracked a single running winner and its backer count, so a higher state arriving from one peer replaced the winner and discarded the count behind it -- and every later reply naming the agreed state was dropped as stale. A healthy pointer then read back as uncorroborated: denial of service in place of the forgery the rule was added to stop. An ordinary read taken while an update is in flight looks exactly the same, so this failed without an attacker too. Replies are now tallied per state, and a read answers with the best state that enough peers named -- a stale state the group agrees on beats a newer one only one peer has heard of, because the second is one peer's word and the first is the network's. Tested in every arrival order, with the singleton first, last, and in the middle. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
write_quorum(0) is zero, so a write to an empty close group counted zero acknowledgements as enough: the client settled a payment and reported the record stored on no peers at all. Both paths now refuse an empty group, and the write looks for its storers before it pays rather than after -- there is no sense buying storage when there is nowhere to put it. The doc on pointer_put also stops implying a retry is free. Every call settles a payment for the state it carries; a state the network already holds is answered as stored, but it was paid for either way, because the node cannot tell a retry from a fresh submission before it has been paid to look. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Moving the lookup ahead of the payment stopped the client buying storage with nowhere to put it, but then it wrote to that same set afterwards. Settling on chain takes time, and the peers responsible for an address can change while it does -- the write would go to peers that are no longer responsible and be refused by their closeness gate, while the ones that are responsible were never asked. There are two questions and they now get two lookups: whether there is anywhere to store this, asked before paying, and where, asked after. Refusing an empty group moves into pointer_group itself, so no caller has to notice that a quorum of an empty group is zero. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pointer_resolve hands the caller a target whose kind this build does not know, because the record is signed and the bytes are authentic. The Errors section said it failed instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ADR-0015 was taken by direct browser clients over WebRTC-direct while this branch was open.
main's WebRTC-direct transport landed while this branch was open. The pointer client imports native-only APIs unconditionally, so the Browser WASM job could not compile it -- and compiling it would not have helped: the node admits, sanitizes and classifies message kinds at that boundary by explicit list, and pointer requests are in none of them. A browser API that always fails at the first hop is worse than no API, and half-opening a security boundary during a rebase is not the way to add one. ADR-0016 records what it would take. Also restores main's windows-core edge, which repinning re-resolved by accident, so the lock differs from main only by the two branch pins.
ant-client's main has now taken the same release ant-node did, and it landed exactly the moves this branch had made ahead of it: ant-protocol 3.0.0, saorsa-transport 0.37.0 with webrtc, and no patch table. Those changes are main's now, so what is left here is only what pointers actually need: - [patch.crates-io] with one entry, ant-protocol at the pointer branch's 3.1.0, because the published 3.0.0 has no Pointer record or wire messages. - ant-node pinned by rev to the rebased pointer node, because the published 0.20.0 has no pointer handler or store. The lock is main's with those two entries edited by hand, so no unrelated dependency edge moves.
5df5740 to
aee2111
Compare
Nodes now take any paid pointer state that beats what they hold, whatever its counter. The suite is changed to prove what that buys, against real nodes and real settlement: - every close-group node answers a paid update to a pointer it already holds with Success -- none refuses it as already existing -- and the client's own pointer_update lands on top; - a paid record that skips counters is taken, and the next update carries on from it (this replaces the test that expected skips to be refused); - two nodes that missed two updates take the next one, three counters ahead of what they hold; - a fork at one counter across the close group reads through the merge rule, and one update at the next counter heals it on every node. The fork and catch-up tests pay for a state once, as the client does, and deliver it to chosen close-group nodes through each node's own request handler, so payment verification, the closeness gate and the merge all run as they do for a PUT that came over the network. Run against the previous node, the skip and catch-up tests fail with the node refusing "exactly one increment". Docs follow: pointer_update signs one past what the network serves because that is the smallest counter that beats it, not because a node requires it. Pins ant-protocol to 4412b6a and ant-node to 0d3a3bf.
An adversarial audit found the suite proved the happy paths but not payment enforcement, the read's safety rules, or the client's failure modes over the network, and that one fork assertion depended on reply order. The suite now covers them, 13 tests against real nodes, QUIC and Anvil: - Payment: create, every update and a retry each lower the token balance. Over QUIC, a node refuses an unpaid record, a forged one (for its signature), and a proof paid for a different state, and stores only the genuine record paid for itself. - Reads: a state named by one peer never wins; fewer than four answers never settle a read; four answers with exactly two agreeing do. - Writes: two competing writes that both complete read as the merge winner deterministically. A public pointer_put that reaches three of seven fails with the exact count of replicas it left, and the fork it leaves is healed by the next counter on every node. - Resolve: a self-cycle, a broken chain, and the depth limit on both sides of the boundary. The harness gains PointerSilence: a node ignores pointer GETs or PUTs while staying in the network, decided as each message arrives, so a test can make a read or write reach only part of the group through the client's own path. Tests silence every node except the ones meant to answer, so the result does not depend on which seven peers the client's lookup settles on. Each rule is shown to be load-bearing: with the corroboration bar at 1 or 3, the read quorum at 3, or the write quorum at 3, the corresponding test fails.
- The pointer API is transport-neutral and compiled for the browser: getPointer, resolvePointer, createPointer, updatePointer and storePaidPointer on BrowserNetworkClient, and pointerAddress(seed). They run the native read quorum, corroboration and write quorum; only the wallet is the page's. A paid write is PUT on the data lane and held exclusively like a chunk PUT, only to nodes that advertise pointer_protocol and share the page's payment network. - The write is split into prepare_pointer_payment and pointer_put_paid, for external signers and for the browser. onPaid hands the page the paid record and proof so storePaidPointer can finish without paying again. - A write that falls short is retried with the proof already paid; a peer's refusal (a newer state, a refused payment) ends it instead. - Creating a pointer that exists is refused before anything is paid. - `ant pointer keygen|address|create|update|get|resolve`, with the owner key kept as a 32-byte seed in an 0600 file that is never overwritten. - Tests: mock-transport JS tests for every browser path, a real-Chromium run against local nodes with on-chain payment, and a client E2E for the refused duplicate create. - Repins ant-node to 3c70728.
Testnet evidence 1/2 — smoke run on 100 nodes (DEV-01, registry id 614, 2026-09-26)Deployment-side evidence for this PR set. A dedicated pointer client tier (saorsa-deploy, V2-1339) drove Build: Workload: per client, create a new pointer (fresh ML-DSA-65 key) every 60 s, update a random owned pointer every 30 s, read back after every write, sweep every owned pointer plus the other client's published pointers at least every 600 s; 10% of updates re-point at another owned pointer and are followed with Result — every criterion passed, all failure classes zero:
Full run (3.1h): 303/303 writes, 1,658/1,658 reads, still all zeros. Decay with age (the ADR-0016 question): sweep reads bucketed by pointer age at read, own and foreign, 30-minute buckets out to 150–180 min — 0.00% stale and 0 not_found in every bucket. No decay. Node side (Elasticsearch, full run): One observation worth a look: write latency p50 ~21 s / p95 ~35 s means a paid 5-ack write frequently outlasts the 30 s update interval, so the serial client produced ~100 writes/hour rather than the nominal 180. Nothing failed; throughput is write-latency-bound. Whether p95 ~35–38 s is inherent to a single-node-quote paid write is the open question. |
Testnet evidence 2/2 — staging scale, 990 nodes (DEV-02, registry id 615, 2026-09-26/27)Same branches, same SHAs and same pointer workload as the smoke run, on the full staging shape: 7 bootstraps + genesis, 990 nodes across DigitalOcean, Vultr, OVH and OVH 3-AZ (66 VMs × 15 services), NAT 10%, 10 native uploaders (20–1000 MB) + 2 downloaders, 3 WASM (browser-path) uploaders + 1 WASM downloader also built from this ant-client branch, 2 pointer clients (DO Result — every criterion passed, all failure classes zero, same as at 100 nodes:
Full run (~5h): 371/371 writes, 2,545/2,545 reads, all zeros. Decay with age: own and foreign sweep reads in 30-minute buckets out to 270–300 min — 0.00% stale, 0 not_found in every bucket. A 10x fleet and a 2x longer window changed nothing about correctness. Scale cost, 100 → 990 nodes: write p50 +25% (21.2 → 26.5 s), p95 +9%; read p50 +48% (5.1 → 7.6 s); cost per write +8.8% ANT / +2.0% ETH. Latency only. Node side (Elasticsearch, full run, 124M forwarded node log lines): Alongside (not pointer-related): native transfers 4,517/4,517 uploads and 796/796 downloads; node CPU median 31.8%, per-service RSS p95 459 MB across 997 services. WASM downloads 568/569; WASM uploads failed almost only at 900 MB (8 ok / 145 fail, V2-1305 quote-collection livelock) while the 300 MB and 20 MB WASM uploaders ran at 93% and 99%. Carried forward: the write-latency observation from the smoke run is confirmed and slightly worse at scale — at p50 26.5 s against a 30 s update interval the serial client produced ~74 writes/hour, not 180. If a target write count matters for a staging train it has to be derived from measured write latency at fleet size. |
Resolves 4 conflicts after merging origin/main into the pointer branch: - harness.js: keep PR #206's BrowserNodeClient removal, add pointerAddress import and runPointerIntegration - shared.rs: merge PR #206's diagnostics tracing with pointer lane/capability/payment/exclusive changes - test_utils.rs: merge PR #206's node_session/hello_payment/eviction tests with pointer mock handlers - mock-webrtc.mjs: merge PR #206's deferred close() with pointer store-copying and put_pointer handling
The previous pin (b0263b32, Sep 23) predates pointer support in ant-node. Bump to main (c092f225) which includes the pointer capability from PR #231, so the devnet nodes advertise POINTER_PROTOCOL_CAPABILITY and the Chromium pointer test passes.
mickvandijke
left a comment
There was a problem hiding this comment.
Nice work on the client side. Checking every acknowledgement against the address and state_id that was sent, requiring two peers to agree on a read, and the write/read quorum overlap are all well thought through. It's great to see it running end to end in a real browser too. Thanks Anselme!
Linear issue
Closes V2-1277
Risk tier
Client-side only in code, but it sends a new wire message and pays on a new identity, so it is judged on what it touches rather than where it lives.
Compatibility
pointer_protocol.Client::pointer_{create,update,put,get,resolve},pointer_sign_{create,update},prepare_pointer_paymentandpointer_put_paid(for a proof paid outside this client),ChunkPaymentPlan::proof, andMAX_POINTER_RESOLVE_DEPTH.getPointer,resolvePointer,createPointer,updatePointerandstorePaidPointeronBrowserNetworkClient, andpointerAddress(seed).ant pointer keygen|address|create|update|get|resolve.pay_for_storageandget_store_quote_plankeep their exact signatures and behaviour.Semver impact
Test evidence
cargo test --lib --all: 741 passed.cargo test -p ant-cli: 25 passed.cargo test -p ant-core --features test-utils --test e2e_pointer: 14 passed, against ant-node at the pinned revision.--test e2e_chunk: 10 passed.wasm-pack … --features browser-wasm,test-utils, then the node test runner): 166 passed, 9 of them pointer tests over the mock transport with the real quorum and payment code.ant-core/browser-tests) against seven real nodes and Anvil: the existing upload test and a new pointer test pass. The pointer test creates, updates, re-targets to a second pointer, reads back from a fresh client and resolves the chain.cargo clippy --all-targets --all-features -- -D warnings, the wasm32browser-wasmclippy,cargo fmt --all -- --checkandRUSTDOCFLAGS="-D warnings" cargo doc --all-features --no-deps: clean.What the client is responsible for, each tested against the production function rather than a restatement of it:
onPaidis stored again without payingNew dependency
None new to the lockfile:
ant-clitakesrand0.8 and, for tests,tempfile, both already in the tree.ADR
https://github.com/WithAutonomi/ant-node/blob/feat/pointers-immutable-owner/docs/adr/ADR-0016-pointers-immutable-owner.md —
ant-nodeADR-0016, pointers with an immutable owner. Lands in WithAutonomi/ant-node#231.Mitigation / rollback
Purely additive surface: nothing existing changes behaviour, and the refactored payment helpers keep their signatures, with chunks taking the path they took before. Reverting removes the pointer API and leaves the chunk path untouched. Depends on the protocol and node PRs landing first; until then the branch pins them by revision.