Files
rippled/docker/telemetry
Pratik Mankawde c3e4c4244a docs(telemetry): explain why the overlay dial metrics report one fewer series
overlay_connect_total and the three overlay_dial_latency_ms series come back
with 4 series per run while sibling families such as dns_resolve_* come back
with 5, one per node. That was unexplained, so anyone reading the group had to
choose between suspecting the exporter and re-deriving the cause. It is a
topology artefact and nothing is wrong.

OverlayImpl::connect asks peerFinder().newOutboundSlot for a slot and returns
early when it gets a null one (OverlayImpl.cpp:464-470), before it constructs
the ConnectAttempt that emits both signals (:472). Every node is seeded to dial
every other node -- run-full-validation.sh:324-331 builds IPS_FIXED from all
NUM_NODES-1 peers and :373-374 writes it into [ips] -- so all five nodes do try.
In a full mesh each pair is dialled from both ends, and the node whose peer got
there first is refused an outbound slot for an address it already holds inbound:
no ConnectAttempt, so neither the counter nor the histogram. dns_resolve_*
reports 5 because reportDnsResolve fires inside the resolver handler
(OverlayImpl.cpp:603), which runs before any slot allocation.

The note also records why the four entries assert series presence rather than a
count: hard-coding 4 would bake today's mesh into the contract and break on any
cluster-size change, while gaining nothing -- and it says that fewer than 4
would be worth investigating, since that means a node did not dial at all.

Verified in code: the early return and its position relative to the
ConnectAttempt, the [ips] construction, and the resolver call site. Not verified
against a run: which node is missing on any given run, because dial ordering is
not controlled and the identity is not expected to be stable. No assertion
changed -- this commit adds documentation only.
2026-08-26 19:45:15 +01:00
..