Completes the join left open when the acquire and consensus work landed in
parallel. The acquire span now derives its trace id from the ledger hash it
already records, so a fetch shares one trace with that ledger's validation,
acceptance and store, and its three phase children inherit the same id.
Reading one trace now answers the whole question for a slow ledger: whether
the data was slow to arrive, slow to be accepted, or slow to persist.
The acquire span stays an optional member of the join group, for the reason
its own entry already gives: a healthy cluster agreeing from genesis rarely
back-fills, so the span need not appear on every run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A slow fresh-sync ledger produced spans scattered across threads with no
way to relate them. They now share a trace id derived from the ledger's own
hash, the one value every participating site already holds, so nothing new
is plumbed across threads. This is the pattern the transaction pipeline
already uses for its tx id.
Joined: ledger.validate, ledger.store, and a new
consensus.validation.accept recorded when a trusted validation arrives. In
Tempo, searching one ledger hash returns them together, so an operator can
tell whether the ledger was slow to arrive, slow to be accepted, or slow to
be stored. They are siblings rather than a chain because the accept gate is
entered from three different threads, so no fixed parent order exists.
consensus.validation.accept also records why an arriving validation did or
did not advance the gate, which makes "validations arrive but are all
rejected" visible for the first time.
consensus_round_duration_ms turns the existing round-time span attribute
into a histogram, so a fleet trend needs a metric query rather than raw
trace inspection. An explicit bucket view is required, not optional: the
SDK default tops out at ten seconds while consensus abandons a round at two
minutes, so slow rounds would all fall in one bucket and every quantile
would read exactly ten seconds. Cost is one record per round.
Record layer: the histogram is native and needs no collector change. The
two new bounded attributes are added as span-metric dimensions to both
collector configs. The ledger hash stays out of them, since a per-ledger
dimension mints a series per ledger; it is indexed in Tempo as the join key.
The ledger.acquire span is not joined yet, because that file was being
changed concurrently. It is registered as an optional member of the join
group so nothing fails, and switching it is a one-line follow-up.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four blind spots in the sync exchange, each now a span:
- txset.acquire: transaction-set acquisition had no span at all, though it
is the sibling of ledger.acquire and runs every consensus round. A round
that falls behind because its tx set never arrived was indistinguishable
from one that deliberated slowly.
- ledger.acquire.{header,astree,txtree}: the acquire span was flat, so the
account-state tree, which dominates a fresh sync, could not be separated
from the transaction tree. These are children, closed before the parent.
- peer.dial: the outbound dial already had outcome counters; the span adds
the per-attempt timeline, so a slow stage is visible rather than only its
terminal reason.
- ledger.serve: serving a peer's ledger request was uninstrumented, so this
node's contribution to someone else's sync was invisible.
Every span finalizes exactly once. Outcomes come from shared compile-time
rules rather than a literal per branch, so no exit can mislabel itself and
an exit added later cannot omit one. Destructor paths are noexcept.
One rule needed care: the timeout path also sets the failure flag, because
that is how the timeout loop stops, so precedence puts timeout ahead of
failure or a timed-out acquire would read as a data fault.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The quorum and publish gauges were emitted but never surfaced: no panel, no
harness assertion, no reference entry. Completes those layers.
- Four panels: trusted validations against the quorum target on one axis so
a tally climbing toward quorum is visually distinct from one flat below
it; publish lag; pre-accept shortfall rate; and time to first validated
ledger.
- Both signals are asserted by the workload validator. The shortfall
counter does fire on a healthy cluster, because this node validates and
then immediately re-enters the accept gate before its peers' validations
arrive, so the first evaluation of every round tallies short. The panel
and note say so, and give the fault signature instead: the shortfall rate
outpacing the ledger-close rate while the tally stays flat and nothing
ever reaches first-validated.
- The quorum target is deliberately drawn as its own line rather than as a
headroom stat, so the disabled-quorum sentinel reads as an unreachable
target instead of an unreadable negative number.
Also removes three reference rows that were appended twice when two agents
each documented the same back-fill signals.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Registers the new gauges, renders them, asserts them and documents them, so
each signal reaches an operator rather than stopping at the emit site:
- MetricsRegistry: gauge registration for ledger_quorum_publish,
nodestore_latency, peer_ledger_supply, peerfinder_slot_census and
amendment_block, each guarded by the detached-callbacks check and
tolerant of services that are not ready yet.
- Ledger Sync Health dashboard: panels for the new signals, filtered by
the node template variable like every other board.
- Workload validation: the new series are asserted, so a signal that
regresses to absent fails CI. Signals the local cluster structurally
cannot produce, such as a replay fallback or an amendment block, are
noted rather than asserted, which would fail red on a healthy run.
- Reference, runbook and glossary entries, including the diagnosis order
for a node that has peers and validators but never validates.
- Regenerated levelization baseline: three new one-way edges from the
telemetry and test modules, no new cycles.
Also drops an unused cstddef include from the macro tests, which the
include checker rejects.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The span only recorded an outcome on the normal completion path. An
acquire that stalled and was later swept ended with no outcome at all, and
its duration stretched to the sweep interval rather than the real fetch
time. So the one case these signals exist to catch, a fetch that never
finishes, was the one case that could not be traced, and aggregate outcome
and timeout rates read low exactly when nodes are stuck.
- Adds an abandoned outcome value for the swept-while-fetching case.
- Routes every exit through one idempotent finalizer, so a span is
finalized exactly once whether it completes, fails, short-circuits on
local data, or is destroyed mid-fetch. The destructor path cannot throw.
- Adds the ledger hash to the span and backfills the sequence once known,
since by-hash acquires start without one and could not otherwise be tied
to a specific ledger.
- Record layer: outcome stays a span-metrics dimension in both collector
configs, which drift apart if only one is edited. The ledger hash is
indexed in Tempo for trace search instead, because a per-ledger value as
a metric dimension would mint a new series every ledger.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Quorum and publish (A5):
- ledger_quorum_publish gauge: the trusted-validation tally against the
quorum the candidate ledger must reach, plus the gap between the
validated and published sequences. The published sequence was never
exported, so a publish pipeline falling behind healthy validation was
invisible.
- ledger_quorum_shortfall_total: counts the pre-accept early return where
a node has peers and validators yet still declines to declare a ledger
validated. That path was log-only, and it is the difference between
accumulating toward quorum and never reaching it.
Back-fill and persistence (A6):
- nodestore_latency: read and write service time. storeDurationUs_ was
declared but never written and had no accessor, so there was no write
latency signal at all. This is the direct fingerprint of a node with an
existing database syncing slower than a fresh one, where the node cache
is cold and every tree step reaches disk. Distinct from the existing
NuDB read panels, which show volume and hit ratio rather than service
time.
- ledger_replay_fallback_total and ledger_replay_outcome_total: the replay
path silently falls back to plain acquisition on timeout or failure, so
a defeated optimisation looked like ordinary slow back-fill.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Whether the network can even serve this node, and whether the node is
about to be shut out of validation, were both invisible:
- peer_ledger_supply: how many connected peers advertise a range covering
the sequence being fetched. Peers each track a range from status
changes, but nothing aggregated them, so "nobody has what I need" looked
identical to "peers are slow".
- peerfinder_slot_census: outbound active against capacity, connection
attempts, inbound, fixed configured against active, and the bootcache
and livecache sizes. All were computed already; only two were exported,
read at unrelated instants, so they could not be compared.
- peer_disconnect_total{reason,direction} and peer_accept_total{outcome}:
every disconnect previously collapsed into one number, so our own
backpressure could not be told from topology or network faults. Reasons
are a fixed set of literals recorded on the peer and emitted once at
close, never data supplied by the remote end.
- serve_refused_total{request,reason}: the other half of the sync
exchange, when this node declines to serve a peer.
- amendment_block: whether an unsupported amendment is expected and how
long until it activates. Amendment-blocked is terminal for validation,
so the countdown is the only leading indicator. The amendment id is
deliberately not a label, since the network can vote an id this build
has never heard of; it is already logged.
- ledger_jump_total: repeated last-closed-ledger switches, which mean the
node is thrashing between chains.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sync-critical job types run at very low concurrency limits (ledgerRequest
and ledgerData allow 3 each), so a node can stall simply because those
jobs are held back behind other work. Nothing exposed that until now:
the existing job metrics are rates and quantiles of jobs that already
moved, or a single queue-wide depth.
- jobq_backlog{metric,job_type}: instantaneous waiting, running and
deferred counts per job type. Deferred is the starvation signal and had
no exposure anywhere; it is set when a type is at its concurrency limit.
- jobq_saturation{metric}: running tasks, worker-thread count and total
waiting, so a slowdown spanning several subsystems can be attributed to
worker-pool exhaustion instead of being diagnosed once per victim.
Both read through two new const accessors on JobQueue that take the
existing mutex once and copy integers, so a single reading is internally
consistent and no per-job cost is added. The job_type label reuses the
same JobTypes name helper the existing job counters use, so the two label
sets join.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signals that separate a sync that is merely slow from one that will never
finish:
- sync_acquire{missing_state_nodes_max, missing_tx_nodes_max, in_flight,
received_data_depth}: how many SHAMap nodes each in-flight acquire is
still waiting for. getMissingNodes already computed this and the callers
discarded it after a trace log. A count that stays flat means the
acquire is wedged; a shrinking count means it is progressing. Recorded
once per sweep, never inside the per-node walk, and reset when a tree
completes so a finished acquire does not read as stuck forever.
- shamap_cache_hit_rate{treenode}: hit rate of the in-memory tree-node
cache, which sits above the node store, so it is distinct from the
existing NuDB ratio. A cold cache on a fresh node sends every traversal
step to disk.
- sync_acquire_no_progress_total: timer ticks where an acquire made no
progress, previously only logged.
- sync_addnode_total{good,duplicate,invalid}: whether arriving nodes are
useful, duplicated or rejected, so wasted fetch work is visible.
- sync_acquire_source_total{local,network}: whether a ledger was served
from the local store or had to be fetched.
Adds getBad()/getDuplicate() to SHAMapAddNode and an acquireProgress()
accessor on InboundLedgers so the xrpld gauge can read these without
libxrpl depending on telemetry.
ledger_seq is deliberately not a metric label: it is unbounded. Per-ledger
identity stays on the ledger.acquire span; the metrics expose bounded
aggregates instead.
The full-below cache hit rate is not exported: KeyCache updates different
counters than getHitRate() reads, so it would always report zero. That
libxrpl bug is documented rather than papered over.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five signals that explain why a node is not advancing toward full, none of
which were observable before:
- state_changes_total now carries {from,to} mode labels, emitted at
setMode using the existing strOperatingMode helper. A bare count could
not distinguish a healthy climb from a node flapping between tracking
and connected. Removes the now-unused incrementStateChanges wrapper.
- sync_state{initial_full_duration_us}: time to first reach full, which
StateAccounting already computed but exposed only in server_info.
- sync_state{network_ledger_gate}: whether the node is still refusing to
build ledgers because it has no network ledger.
- sync_state{server_stall_seconds} and server_stall_events_total: how
long the main thread has been unresponsive. LoadManager computed this
and only logged it, so a stall was invisible until the fatal threshold.
The episode rule is a pure function so it can be tested without adding
a test-only mutator to LoadManager.
- sync_state{ledgers_behind}: how far our validated sequence trails the
best sequence any peer advertises, read from already-cached peer ranges
so no extra network traffic is added.
Also fixes the naming checker: it derived only the first label of a
multi-label instrument, so a dashboard querying the second label was
wrongly rejected.
Note: the clang-tidy hook cannot run in this worktree (no build
directory); the remaining pre-commit hooks, the naming check, dashboard
schema and harness syntax all pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three defects found reviewing the WP-A1 commit:
- Handshake.cpp moved a std::string into std::runtime_error, which has no
rvalue constructor. The move never happened and clang-tidy rejects it
under performance-move-const-arg, so CI would fail even though the
local hook only runs clang-tidy with TIDY=1. Takes the message by const
reference instead, and drops the <utility> include that existed only
for that move.
- ValidatorList disables quorum by returning SIZE_MAX. Casting that to
int64_t wrapped it to -1, so the headroom panel computed
0 - (-1) = +1 and coloured yellow on a node that can never validate:
the sign inverted in exactly the bootstrap failure these signals exist
to catch. Reports the disabled state as int64 max so headroom goes
strongly negative instead.
- The runbook claimed an expired list loads no keys. Expired counts as
accepted, so its keys are loaded and then dropped by the expiry sweep,
which calls for a different fix than replacing validators.txt. Pending
is likewise a future-dated refresh, not a rejection. Documents both,
plus how the quorum-disabled state now reads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A freshly started node most often stalls before it ever peers or reaches
quorum, and that whole chain had no telemetry. Adds the six signals that
make it observable:
- dns_resolve_total / dns_resolve_latency_ms: configured-peer hostname
resolution, emitted from OverlayImpl so libxrpl stays independent.
- overlay_connect_total / overlay_dial_latency_ms: outbound dial outcome
by terminal reason, plus dial duration.
- handshake_negotiation_fail_total: protocol and network-id negotiation
rejections, labelled by reason, so a misconfigured network is no longer
indistinguishable from unreachable peers.
- unl_fetch_total and the unl_quorum gauge: validator-list fetch outcome
per site and trusted key count against the required quorum. Without
these a bad validators.txt leaves the node syncing forever with no
signal.
- clock_close_offset_seconds: network close-time offset, which server_info
hides below 60s but which stalls consensus participation.
Panels land in the Bootstrap row of the Ledger Sync Health dashboard, the
metrics are asserted by the workload validator, and both the reference and
the runbook flow describe them.
Levelization baseline regenerated: overlay now includes MetricMacros.h, so
the overlay/telemetry pair is reported one-way instead of bidirectional.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Match the panel height the operator settled on in Grafana (18 rows) so
all 20 ranked bars render without an inner scrollbar, and reflow the
Sync Diagnostics row and the panels below it accordingly.
Adds the anchors the sync-diagnostics signals attach to, with no signals
emitted yet:
- New "Ledger Sync Health" dashboard (uid ledger-sync-health) with the
standard template-variable block copied from an existing board, plus
empty "Bootstrap (Domain 0)" and "Sync pipeline" rows.
- Signal index section in the data-collection reference, an operator-flow
stub in the telemetry runbook, and a glossary anchor.
- A sync_diagnostics group in expected_metrics.json and a matching
assertion helper in validate_telemetry.py so CI fails when a signal
regresses to absent.
Also registers the new dashboard uid with the harness so the board is
covered by validation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bar labels on 'Overlay Traffic Heatmap (All Categories, Bytes In)'
were raw metric names carrying a redundant '_bytes_in' suffix, and the
half-width 8-row panel truncated both the category and the node
identity.
- Strip the '_bytes_in' suffix from the derived series label; the panel
title already states the metric is inbound bytes.
- Widen the panel to full width and grow it to 12 rows so 20 bars render
with their full category and node labels.
- Shift the Sync Diagnostics row and the panels below it down by 4 to
keep the layout gap-free.
- Note the label derivation in the panel description.
The panel plotted only Full and Tracking, so the states a node actually
passes through while catching up were invisible. On the dev box's cold
sync the node spent ~19 minutes in Connected and ~3 seconds in Syncing,
while Tracking totalled 2 microseconds -- the one non-Full line the
panel did draw was the least informative of the five.
Add Syncing, Connected and Disconnected series (state-ladder order,
matching the Operating Mode Transitions panel) and rename the panel to
"State Duration Rate (All States)". Because the node is always in
exactly one state, the five rates sum to ~1.0, so the panel now reads as
"which state is time going into right now", and a handover between two
lines marks a state change with its width showing the dwell time.
Description expanded to the full section set (keywords, computation
boundary, references) used by the other panels on this dashboard.
Verified live: rate(state_accounting_connected_duration)/1e6 reaches
1.0 then 0.387 across the catch-up window and syncing reaches 0.0103 --
both previously undrawable.
The dashboard showed 54 individual signals but no single answer to
"is this node healthy and doing its job?". A reader had to correlate
server state, ledger age and peer counts by eye.
Add a stat panel at the top that reduces the two conditions that define
a healthy observer node to one green/red verdict:
server_state == 4 (Full) AND validated_ledger_age < 30s
Both terms use `== bool` so each yields 0/1, joined with
`* on(service_instance_id)` so the verdict is per node. Value mappings
render 1 as "Healthy" (green) and 0 as "Not Healthy" (red) with
background colouring, so the state is readable at a glance.
Placed at y=0 per the dashboard convention that gauges/stats lead;
existing panels shift down by 4 rows with no other change.
Verified live against Grafana Cloud: returns 1 for aws-dev-xrpl-1
(Full, validated seq 105,824,596).
The four aggregated Sync Diagnostics panels grouped by xrpl_work_item,
which no layer in this repo emits (it is injected by the perf-iac
harness). That tripped Rule D of the telemetry naming check:
D ledger-data-sync.json xrpl_work_item
must exist in L1, a metric label, or be a builtin
Align with the convention used by every other aggregated panel in the
dashboard set: group by (service_instance_id, xrpl_branch,
xrpl_node_role) and leave the xrpl_ident label_join untouched. The
legend is unaffected -- label_join over a label dropped by the
aggregation contributes an empty segment, which the trailing
label_replace already strips.
Verified live: both the NuDB ratio and the job-queue p95 queries still
return per-node series rendering as "[aws-dev-xrpl-1]".
The 0-4 values now render as named states via value mappings, so the numeric
axis label is misleading. Rename to "Server State".
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add value mappings (0 Disconnected, 1 Connected, 2 Syncing, 3 Tracking,
4 Full) and matching red->green thresholds to the Sync State panel, so the
line, tooltip, and legend render state names and colors instead of bare 0-4.
Keeps the minimal custom-key convention; only the Sync State panel changes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Rework the 10 Sync Diagnostics panels' legends to match the format used by the
existing panels: wrap each query in label_join + label_replace to build the
xrpl_ident label ([service_instance_id, xrpl_branch, xrpl_work_item] with empty
values stripped) and set displayName to "${series} ${xrpl_ident}". Verified
against a live node: renders as "Series [aws-dev-xrpl-1]" with no empty-label
gaps, identical mechanism to the original 15 panels. Only the new panels change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a 10-panel "Sync Diagnostics" row that shows ledger-sync slowdown and its
causes, laid out top-to-bottom as symptom -> latency -> root cause:
- Red flags: Sync State, Validated Ledger Age, Ledger Close Rate
- Latency: Job Queue Wait p95 by type, NuDB Read Latency, I/O Scheduler p95
- Root cause: NuDB Cache Hit Ratio, NuDB Read Pressure, Job Queue Depth,
Load Factor & Peers
All panels use native beast::insight metrics introduced on this branch
(nodestore_state, ledgermaster_*, jobq_*_q, ios_latency, load_factor_metrics,
peer_finder_*) plus the consensus.mode_change span metric. Queries filter by
the dashboard's $node and tier variables and were verified against a live
node. Nodestore ratio panels use sum-by(identity) matching to bridge the
metric= sub-label.
Format matches the dashboard convention: minimal custom keys, ${DS_PROMETHEUS}
datasource uid, and the standard What/How/Reading/Healthy/Watch/Source/Function
description structure. Pure-append; existing panels untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>