A walk that reaches a position only a leaf may occupy marks the map Invalid and abandons the
descent. Reaching that position is otherwise fatal: SHAMapNodeID::getChildNodeID() throws
std::logic_error past kLeafDepth, and there is no try/catch around getMissingNodes() in
InboundLedger::trigger(), around job.doJob() in JobQueue, or in Workers::Worker::run(), so the
exception leaves the thread function and reaches std::terminate(). It is reachable without any node
passing through addKnownNode(): InboundLedgers::gotStaleData() stores any parseable node from an
unsolicited liAS_NODE reply into the fetch pack keyed by its own hash, with no relatedness check, and
getFetchPack() re-verifies only that hash, so such a node canonicalizes into the tree and the walk
descends onto it. Steering which hash a node acquires needs it to have no trusted validations -
starting up, or an empty or misconfigured UNL - with the attacker holding its peer slots, which makes
this a conditioned remote denial of service rather than a single-packet one.
As in addKnownNode(), the depth test precedes the full-below cache lookup, since that cache is keyed
by node hash and shared across maps and a hash does not cover depth, so a hit would carry the whole
branch past the guard. Four further tests of isValid() bound what the walk does once the verdict
lands: it short-circuits on entry rather than re-deriving a verdict already reached; it breaks out of
the descent but falls through to the deferred-read drain, since posted reads hold a reference to the
MissingNodes block on this frame; it discards whatever was collected, which belongs to a tree that
cannot exist; and it re-tests before clearSynching(), since another thread's addKnownNode() can write
the verdict after the loop's own test. Callers therefore have to re-check isValid() before reading an
empty result as nothing left to fetch, which getMissingNodes()'s docstring states. Three of those
four guards are defensive and no test drives them; only the entry short-circuit and the verdict itself
are pinned.
DeepChain gains withDecoys(): the same chain, but with a second and unresolvable child at every
level, so a backed map's descendAsync() posts a real asynchronous read at every level. That is what
leaves reads in flight when a walk reaches kLeafDepth, which is what the sanitizer case below needs.
Seven cases cover this. Five drive the walk through a ChainFilter, which stands in for a fetch pack by
serving nodes by hash and never structurally: the walk reaches the verdict itself on an unbacked map;
it does so on a backed map with the offending node already marked full below; it drains the reads a
decoy child at every level leaves in flight, which only a sanitizer can see; it refuses an
already-invalid map; and it leaves a walk that stops one level short alone, reporting the genuinely
missing child. The sixth pins that addRootNode() cannot clear the synching flag on an invalid map,
since that call site needs a leaf root and so a zero root hash. The seventh races a walk against
setImmutable() under ThreadSanitizer, and asserts only what trySetState() offers: the verdict stands,
whatever the interleaving. It is skipped at run time rather than compiled out, so every build parses
it.
An immutable map is treated as persistable, so SHAMap::setImmutable() returns [[nodiscard]] bool and
refuses a map that has been proven impossible. Every state change goes through trySetState(),
whose compare-exchange refuses to leave Invalid however it interleaves with another thread's, so
setSynching() and clearSynching() cannot launder an abandoned map back into a state that passes
isValid() either. clearSynching() reports its refusal and carries on, since peer data produces that
verdict and an abort there would be one a peer could ask for, while setSynching() keeps an UNREACHABLE
and says why it is out of reach: it only ever runs on a map that has just been constructed. The
snapshot constructor carries Invalid over rather than promoting it, since a snapshot shares the
source's root, and reads the source's state once into a local so a concurrent walk cannot have it
report two different things.
Ledger::setImmutable() and Ledger::setAccepted() do the same one level up. Both return [[nodiscard]]
bool, and setImmutable() asks mapsValid() before writing anything, so a refusal leaves the header
exactly as it was rather than relabelled on its way to failing. Past that check it settles both maps
through setMapsImmutable(), which is deliberately not short-circuited: each map becomes Immutable or
stays Invalid on its own, and neither is left mid-sync because the other refused. immutable_ is set
last, so isImmutable() never reports a ledger whose maps are not both immutable. The guard is
best-effort by nature, which the comments say: setInvalid() outranks Immutable, so a walk that reaches
the verdict after both maps are settled narrows the window rather than closing it.
All fourteen call sites branch on the result, and the rule that a ledger built or loaded locally
cannot have an invalid map is stated once, in Ledger::setImmutable()'s docstring, with each such site
pointing there. The tiers differ by what the caller can do: the two genesis paths and buildLedgerImpl()
call logicError(), the last of those explaining why it takes the harsher tier on the consensus hot
path; loadLedgerFromFile(), getLastFullLedger() and finishLoadByIndexOrHash() return instead, and the
last of those clears the pointer, since nothing gates usability on the full flag and a caller that
took the ledger would abort further on. InboundLedger and TransactionAcquire recover, since for them a
refusal is an outcome a peer can produce: each withdraws complete_ alongside the failure, so a guard
that reads that flag before failed_ cannot go on treating the result as delivered.
Six cases cover it. Five are gtest: an invalid map refuses repeatedly and is not synching either; a
refusal leaves both header map hashes and the ledger hash untouched; an invalid transaction map and an
invalid state map each block the enclosing ledger, covering both operands of the test in
setImmutable(); and a snapshot of an invalid map is invalid and unpersistable in both flavors. The
boost case drives InboundLedger::done() with a map invalidated after the have-flags were set, and
checks the acquisition reports neither complete nor delivered and remembers the hash as a failure.
SHAMap::addKnownNode() reports invalid() for the two kinds of node it does not hook into the map: an
inner node arriving at kLeafDepth, a depth only a leaf may occupy, and a node whose ID does not match
the position the descent stopped at. That is the verdict callers already handle as bad data, so
neither counts as forward progress. Each emits one warning naming the node and where the descent
stopped, matching the sibling branches beside them, and carries a SOMETIMES() hint for the fuzzer.
The depth rule is spelled once, as a file-local isLeafDepth() that hasLeafNode() reads too. The
verdict on a map-invalidating node belongs to the root hash that was asked for rather than to this
copy of the tree, since every node from the root down hash-verified to get there: no peer can satisfy
such a hash, retrying is futile, and it cannot arise by accident. A charge for it is a deterrent
rather than a control even so, which the comment says, because the same node can reach a map through a
fetch pack or an unsolicited object reply and neither passes through here.
The depth test precedes the full-below cache lookup on the way down. That cache is keyed by node hash
and shared by every map of a family, and a hash covers a node's children but not its depth, so the
same subtree hash can be cached as complete at one depth and reached at kLeafDepth here. Testing the
depth first is what keeps the verdict independent of whatever an unrelated map cached, which is the
determinism the acquisition paths need from it. Skipping the shortcut at the deepest level only
forgoes an optimization, and the depth it guards cannot occur in a real tree.
Three tests drive the map through DeepChain's fill() and addOffendingNode(), so a case names the
position no valid tree can occupy without spelling out the descent. They cover the map-invalidating
node; the three ways a node cannot be hooked anywhere while leaving the map sound; and the same
offending node with a full-below entry already in place, so the depth test is what has to reach the
verdict. A file-local tallyIs() reads each verdict as counts, which leaves get()'s wording pinned in
one place rather than at every site with a verdict to check.
Background ledger acquisition reads and writes SHAMap::state_, SHAMap::full_, SHAMap::ledgerSeq_ and
SHAMapInnerNode::fullBelowGen_ concurrently with the thread driving it, so all four are std::atomic.
state_ is read through state() and written through setInvalid() and the existing setters, and
SHAMapState carries an explicit std::uint8_t underlying type. finishFetch() withdraws full_ with an
exchange behind a relaxed load, so exactly one of the reader threads that miss reports the gap, and a
map that is already not full stays off the exclusive-write path: full_ shares a cache line with
state_ and ledgerSeq_, and a walk posts up to 512 reads per pass. ledgerSeq_ is read through
ledgerSeq() and relaxed both ways, since it only serves as a lookup hint for a nodestore keyed by
hash.
Ledger::setFull() sets each map's ledger sequence before its full flag. A release store publishes
only what is sequenced before it, and the finishFetch() thread that wins the exchange on the flag
reads the sequence after that, so this is the order that makes the sequence visible to the once-only
gap report.
Static assertions pin the three SHAMap members lock-free, and pin fullBelowGen_'s size and alignment
to those of a plain std::uint32_t, so SHAMapInnerNode's packed layout stays byte-identical and
isFullBelow() takes no mutex once per node of every walk. Its accessors are relaxed, since a
generation is only ever compared for equality and the children it vouches for are published through
the node's own child lock. Three tests cover this: sixteen unresolvable branches posted at a backed
map with four nodestore reader threads, so finishFetch() runs concurrently for one map and the single
gap report is observable; that Ledger::setFull() publishes the sequence that report names; and that
every node the sync path hands to a filter carries it.
SHAMapAddNode gains getBad() and getDuplicate() beside getGood(), so a verdict can be read as
counts. Every accessor, mutator and factory is documented, and reset(), get() and operator+= move
after the static factories, so the counters and the ways to read or combine them stay grouped. get()
is a log format, and src/tests/libxrpl/shamap/SHAMapAddNode.cpp is the one place that depends on its
wording: it pins that format, the counts the three accessors report, and the way one tally
accumulates into another.
* upstream/release/3.3.x: (41 commits)
chore: Bump version to 3.3.0
chore: Bump version to 3.3.0-rc7
fix: Increase manifest protocol message size cap and fix manifests relay
fix: Cap untrusted manifests per message and drop oversized ones
chore: Bump version to 3.2.1
chore: Bump version to 3.2.1-rc1
fix: Cap untrusted manifests per message and drop oversized ones
fix: Reject oversized validator manifest before decoding
fix: Reduce untrusted manifest cache cap to 100
fix: Bound untrusted manifest cache
chore: Bump version to 3.3.0-rc6
feat: Package validator-keys inside rippled
chore: Bump version to 3.3.0-rc5
fix: Switch SponsorshipSet to use a delta for sfFeeAmount
fix: Re-revert "fix: Set request size limits and differential pricing for get-object-by-hash calls"
chore: Bump version to 3.3.0-rc4
fix: Revert "fix: Set request size limits and differential pricing for get-object-by-hash calls"
chore: Bump version to 3.3.0-rc3
fix: Reduce untrusted manifest cache cap to 100
fix: Revert "fix: Reject oversized SHAMap nodes in gotStaleData and fetch-pack path"
...