Commit Graph

11245 Commits

Author SHA1 Message Date
Pratik Mankawde
da35290f27 test(telemetry): compile the span-name and stall-rule tests in every build
Both files gated their whole contents on XRPL_ENABLE_TELEMETRY, so 54 tests
were skipped whenever telemetry was compiled out. The stated reason was that
only a telemetry build puts `src/` on this target's include path, but that
include path is unconditional, so the tests were reachable all along.

Nothing in either file needs the OpenTelemetry SDK. The span-name and outcome
headers hold constants and constexpr functions with no telemetry guards,
LoadManager::evaluateStall is a static constexpr member, and the handful of
guard assertions construct a default SpanGuard, which is inactive in either
configuration. SpanNames.h documents this contract for its own constants.
2026-08-27 11:10:04 +01:00
Pratik Mankawde
fa2a09c758 perf(telemetry): increment the copy-forward total only when it is read
copyForwardTotal_ is a second atomic increment beside copyForwardCount_ on the
same event, kept only so a metric never goes backwards: rotate() zeroes the
per-rotation tally for its log line, which leaves that counter unusable as a
rate. Its only reader is MetricsRegistry.cpp:1153, through copyForwardTotal().

Guard the increment. During a rotation window every archive-served
non-duplicate read pays for it, and with telemetry compiled out there is
nothing to read it back.

copyForwardCount_ is untouched: rotate() exchanges it for the "copied forward N
archive-served reads" warning, which is real logging, not instrumentation. The
virtual and the member stay declared unconditionally, so the nodestore
interface has the same shape in every configuration.
2026-08-27 11:06:48 +01:00
Pratik Mankawde
0ad3587462 perf(telemetry): skip the dial clock and span when nothing records them
dialStart_, outcomeReported_ and dialSpan_ are telemetry-only. dialStart_ is
read only by the two elapsed-time computations in reportOutcome();
outcomeReported_ is written and read only there; dialSpan_ is opened in run(),
ended in reportOutcome() and reset in the destructor, and read nowhere else.

outcomeReported_ is not load-bearing for anything but telemetry. It is a
first-call-wins latch over the histogram, the counter and the span attributes.
Every terminal path calls close() or fail() itself, beside its reportOutcome()
call rather than inside it, so suppressing a second report cannot suppress any
teardown.

Guard the three sites: the clock and span setup in run(), the whole body of
reportOutcome(), and the destructor's reset(). Per outbound dial that removes a
steady_clock reading, an optional emplace and reset of a span handle, and the
latch write. The members stay declared in every configuration so the class has
one shape; only the writes are compiled out.

SpanGuard.h, SpanNames.h, MetricMacros.h and the cstdint header move behind the
guard with the code that names them. ConnectAttempt.h still includes SpanGuard.h
for the member.
2026-08-27 11:06:38 +01:00
Pratik Mankawde
d3c1fc67ce perf(telemetry): skip the per-second stall bookkeeping nobody reads
updateStallState() is wholly telemetry: it applies evaluateStall(), stores the
result in currentStallSeconds_ and bumps stallEventCount_. Those two members
have exactly one reader each, MetricsRegistry.cpp:1986 and :2023, both inside
the registry's own XRPL_ENABLE_TELEMETRY region, reached through
getCurrentStallSeconds() and getStallEventCount(), which nothing else calls.
The monitor thread ran it once per second for the life of the process.

Guard the body, not the members or the accessors: a member set that differs
between build configurations is the hazard that once made a test mock abstract.
evaluateStall() stays where it is, being a public constexpr rule with its own
GTest coverage in SyncStateSignals.cpp.

The atomic header moves behind the same guard, as the relaxed memory orders are
named only in the guarded body; LoadManager.h includes it for the members.
2026-08-27 11:06:30 +01:00
Pratik Mankawde
27dc3236d1 perf(telemetry): build the UNL fetch site label only when recorded
reportFetchOutcome() exists only to label unl_fetch_total. It reads the parsed
URI parts, copies the domain, erases any userinfo with an rfind, joins scheme,
host and optional port, then appends a substr of the path -- two std::string
allocations and several copies -- and that label has no other reader. It ran on
every validator-list fetch, so about once per site every five minutes, whether
or not anything could record the counter.

Guard the whole body with XRPL_ENABLE_TELEMETRY rather than change the
signature: the two failure call sites pass a compile-time constant, so an empty
body is all they need. The success call site is guarded too, because its
to_string(bestDisposition()) builds a std::string that only the label consumes.
bestDisposition() itself keeps running, since lastRefreshStatus stores it.

MetricMacros.h moves behind the same guard, as the macro is now named only
inside the guarded body. MetricNames.h stays unconditional, because the fetch
handlers name the outcome constants either way.
2026-08-27 11:06:20 +01:00
Pratik Mankawde
fef1443a65 docs(telemetry): correct what the span-liveness guard actually skips
The comment claimed the guard skips work for a span that is "not being
recorded", which reads as sampling awareness. It has none: operator bool() is
impl_ != nullptr, and the span factories return an empty guard only when
telemetry is absent, disabled at runtime, or the trace category is off. A span
that exists but was sampled out still pays.

There is no isRecording() in the telemetry API, so the guard is still the
strongest available; only the justification was overstated.
2026-08-26 19:58:16 +01:00
Pratik Mankawde
06792a508d perf(telemetry): read the acquire peer count only when the span records it
finalizeAcquireSpan() took the peer count as an argument, so all three real exit
paths called getPeerCount() before entering it. That walks the acquire's peer set
calling findPeerByShortID for each one, taking the Overlay lock every time, and
the value is used only to set one span attribute -- so an acquire whose span was
never recorded paid for the whole walk.

Pass whether the lookup is safe instead of its result, and make the call at its
point of use, inside the span-active branch. The destructor keeps passing false
for the reason it always had: it can run under the InboundLedgers collection
lock, where taking the Overlay lock underneath would be unsafe.

The two getPeerCount() calls that drive peer recruitment are untouched; they are
real logic, not instrumentation.
2026-08-26 19:13:41 +01:00
Pratik Mankawde
08026b46b9 fix(telemetry): make this branch's files compile clean with telemetry off
The metric macros discard their arguments when telemetry is compiled out, so
anything named only as a macro argument disappears in that build. That produced
fifteen errors across these files.

- guard MetricNames.h in the nine files whose only uses of it are macro
  arguments; the files that pass those constants as ordinary function
  arguments still need it unconditionally
- drop the prevMode local in setMode, reading the mode being left inline in the
  macro argument so nothing is computed when telemetry is off
- compile out the emit helper in recordBatchOutcome and its three calls, which
  exist only to report per-outcome counters
- drop two includes the telemetry-off test block never used
- suppress the static and const suggestions on four methods whose bodies only
  record metrics; each reads the app_ member when telemetry is enabled
2026-08-26 18:08:09 +01:00
Pratik Mankawde
33956ec240 fix(test): alias the second namespace the merged test file needs
1e341d5413 fixed half of this. The slice that assembled the union dropped
`using namespace xrpl::telemetry;`, which supplied TWO names: `consensus` and
`seg`. Aliasing only the first left ConsensusSpanNames.cpp:156 --
`seg::consensus` -- unresolved, and clang, gcc and MSVC all failed there. It was
invisible in the previous round only because clang stops after 20 errors and the
50 `consensus` sites filled that budget, so the same defect had been present since
the merge rather than being introduced by the partial fix.

Why it was missed: the prefix scan behind the first fix sampled a line range that
did not contain line 156. This time every leading namespace qualifier in the file
was enumerated from comment- and string-stripped source and checked against the
names the file makes available, with the line each becomes available: attr, val,
op, part and span from the using-directive, consensus and seg from the aliases,
AvalancheState from a function-local using-declaration inside the test that uses
it, and std/xrpl needing nothing. Every one is declared before its first use, and
nothing else is qualified anywhere in the file.

The predicted hazard did not materialise. No `reference to 'attr' is ambiguous`
error appeared in any of the three compilers, so keeping xrpl::telemetry out of
scope and naming the two members explicitly was the right shape.

This also clears the misc-include-cleaner error that came with it. clang-tidy
reported SpanNames.h as not used directly and its exported fix deleted the
include; that fix was wrong. `seg` is declared at SpanNames.h:103, so the alias
makes the include genuinely used and the diagnostic goes away rather than needing
the include removed.

Verification: compiled. `c++ -fsyntax-only` with this file's real flags from its
compile_commands.json entry exits 0. Proven non-vacuous by removing the alias
again and reproducing CI's exact message -- "'seg' was not declared in this
scope; did you mean 'xrpl::telemetry::seg'?" -- then restoring it and returning to
exit 0. Braces balance 27/27 on stripped source, 17 TEST cases intact, pre-commit
clean including clang-format and clang-tidy. The tests still have not RUN: this is
a syntax-only check of one translation unit, so nothing linked and no assertion
executed.
2026-08-26 13:24:57 +01:00
Pratik Mankawde
1e341d5413 fix(test): repair the merged ConsensusSpanNames test file
The add/add resolution in 493475a9d4 did not compile. Two defects, one cause.

Both branches wrote this file independently and the union was assembled by
slicing each side's tests out of its own copy. That slice was asymmetric: it
started at the first TEST( line, which skipped the pre-merge side's
`using namespace xrpl::telemetry;` and its `namespace {` opener, but ended at the
`#endif`, which kept that side's `}  // namespace` closer. So the file lost the
scope its own tests depend on and gained a brace with nothing to close: 27
openers against 28 closers.

Clang reported 50 "use of undeclared identifier 'consensus'" sites from line 144
and gave up at 20; GCC got far enough to also report "expected declaration
before '}' token". Neither parent has this combination -- it exists only in the
merge.

The nine validation-accept tests spell their constants out from `consensus`,
which the surviving directive does not make visible: `using namespace
xrpl::telemetry::consensus::span` imports the MEMBERS of consensus::span, not the
enclosing namespace name. Fixed with a namespace alias rather than by restoring
the blanket `using namespace xrpl::telemetry;`, because that would pull
xrpl::telemetry::attr (SpanNames.h:117) into scope alongside
consensus::span::attr and make all eleven bare `attr::` references in the
phase-span tests ambiguous -- trading a hard error for a subtler one.

The orphan closer is removed rather than matched with a new opener. The nine
tests it used to wrap declare no helpers, so the anonymous namespace bought no
internal linkage, and the eight tests from the other side were never inside one.

Verification: braces balance 27/27 counted on comment- and string-stripped
source; 17 TEST cases intact; the alias at line 59 precedes the first bare
`consensus::` in CODE at 155 and the directive at 48 precedes the first bare
`attr::` at 64, both measured after stripping comments, because an earlier check
matched its own explanatory prose and reported a false ordering violation; no
blanket using-directive in code; pre-commit clean including clang-format and
clang-tidy. NOT compiled -- there is no approval to build here, so this fix is
structurally verified only, and CI is the first compiler to see it.
2026-08-26 12:31:12 +01:00
Pratik Mankawde
493475a9d4 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to 3836078a78 (74 commits). Seven conflicts, resolved
per-hunk; no side was taken wholesale.

validate_telemetry.py, four hunks. The module docstring keeps both category
lists, renumbered. _log_prometheus_metric_names takes phase-10's version: it
returns the family list that the new reverse-coverage check consumes, and
_log_name_list prints every family sorted one per line, which supersedes the
hand-maintained prefix filter this branch had been extending -- that filter
existed only to keep the log readable and listed strictly less. validate_metrics
takes phase-10's _metric_check_targets call. assert_sync_diagnostics_metrics and
phase-10's _check_metric_label both landed at the same place; both are kept.

That third hunk carried the hazard this branch had flagged in advance.
_metric_check_targets selects every group satisfying isinstance(dict) and had no
equivalent of SKIPPED_METRIC_GROUPS, because on phase-10 there was no group that
needed excluding. Merged as-is it would have walked sync_diagnostics while
assert_sync_diagnostics_metrics also walks it, polling and reporting all 61
metrics twice. The exclusion is reinstated inside that function, and its
docstring claim that the isinstance test "selects exactly the same groups the
previous name-based exclusion list did" is corrected -- true on phase-10, false
here, and the reason is ownership, which no structural test can express.

expected_metrics.json: both sides appended to metrics_excluded, so both sets are
kept, 29 entries. expected_spans.json: phase-10's fuller pathfind.compute
skip_reason replaces this branch's, and this branch's ledger.acquire ->
ledger.acquire.astree relationship is kept.

Both docs carried stale counts, and the two sides disagreed with each other --
15 dashboards against 16, and both claiming 41 span types when the contract holds
48. Rather than pick a stale side, every figure is recomputed from the resolved
contract: 48 span types as 28 required and 20 optional, 145 metric checks across
26 asserting categories as 140 names plus 5 required_labels, of which 61 are the
sync_diagnostics names, and 16 dashboards, which matches both the uid list and
the files on disk. The runbook keeps phase-10's table, which adds the reverse-
coverage row.

ConsensusSpanNames.h: both sides added different constants to namespace val;
both kept. Confirmed no identifier is redefined -- the merged file's duplicate
set is identical to this branch's, and those duplicates are distinct namespaces
(op::round against the enclosing span::round), not redefinitions.

ConsensusSpanNames.cpp was an add/add: both branches wrote this file
independently, 9 tests here and 8 on phase-10, with no name in common. All 17
are kept. The guard is dropped rather than applied to the union: SpanNames.h
documents that its constants are deliberately NOT guarded by
XRPL_ENABLE_TELEMETRY, ConsensusSpanNames.h has no guard, and the tests this
branch contributed reference ValStatus only in comments while calling
validationStatusValue with plain ints. So they compile without telemetry, and
unguarding them gains coverage in a -Dtelemetry=OFF build rather than losing it.

One defect belongs to the merge itself, appearing on neither parent. phase-10
added a job_queue_per_type_gauges group holding six jobq_<type>_running/_waiting
names; this branch declared jobq_saturation in MetricNames.h. Rule K checks a
name only when its family is owned, so declaring that constant made jobq_ owned
and turned phase-10's six entries into violations. They are beast::insight gauges
created per job type by JobTypeData's constructor, so no constant can exist for
them -- one triple per job type, minted at runtime. The group joins
NON_OTEL_METRIC_GROUPS alongside statsd_gauges for the same reason.

Verification: no conflict markers repo-wide and no unmerged index entries; both
JSON contracts parse; validate_telemetry.py compiles; every count written into
the docs re-derived from the resolved files and matching; span counters still 48
and 74; check_otel_naming.py exits 0, and Rule K proven still able to fail by
injecting a bogus name in an owned family; levelization baseline clean after
regeneration, with 17 incoming include changes; doxygen style clean across all
tracked C++ at CI scope; pre-commit --all-files clean except cargo-fmt, which
reports "Executable `cargo` not found" and touches none of the 0 Rust files here.
NOT compiled -- no approval to build, so the incoming C++ is unverified by a
compiler on this branch.
2026-08-26 12:00:25 +01:00
Pratik Mankawde
919e490d1f fix(build): add the includes clang-tidy include-cleaner requires
CI's clang-tidy job failed on this branch. clang-tidy itself succeeded; the
step that failed was the gate that fails the job when findings exist, and the
uploaded diff named exactly two missing includes.

Event.h calls notify(std::uint64_t) but never included <cstdint>, relying on
it arriving transitively. This is pre-existing on the branch rather than new,
surfaced now because clang-tidy only inspects changed files and this one has
not been touched since.

InboundTransactions.cpp uses uint256 throughout and had no direct include for
it either. The round-request work added further direct uses, which is what
brought the file into clang-tidy's changed-file set.

Both fixes are the ones clang-tidy generated itself, applied with the repo's
angle-bracket include style; the include-style and clang-format hooks accept
the placement unchanged.

Note on why this was not caught before pushing: TIDY=1 clang-tidy was run
locally and passed, but against the compile database under
~/sourceCode/.clangdcache/<worktree>/, which was stale -- generated before
these edits. A stale database makes a local clang-tidy pass meaningless, as
project-map.md warns. CI holds the only current one.
2026-08-24 21:01:28 +01:00
Pratik Mankawde
dfda1b2eda Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill 2026-08-24 20:45:37 +01:00
Pratik Mankawde
113a7a9a5a Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation 2026-08-24 20:45:37 +01:00
Pratik Mankawde
f8845b0e00 Merge branch 'pratik/otel-phase6-statsd' into pratik/otel-phase7-native-metrics 2026-08-24 20:45:37 +01:00
Pratik Mankawde
b5e3414f63 Merge branch 'pratik/otel-phase5-docs-deployment' into pratik/otel-phase6-statsd 2026-08-24 20:45:37 +01:00
Pratik Mankawde
162351cfd4 Merge branch 'pratik/otel-phase4-consensus-tracing' into pratik/otel-phase5-docs-deployment 2026-08-24 20:45:37 +01:00
Pratik Mankawde
258adf491e refactor(telemetry): split consensus span labels into their own header
Review of the preceding commits found a clang-tidy failure and a convention
break, both rooted in the same place: the enum-to-label helpers were put in
ConsensusSpanNames.h, which pulled two domain headers into it.

misc-include-cleaner rejected the new test: it used xrpl::LedgerCloseReason
without directly including ConsensusTypes.h, relying on the transitive
include. misc-* is enabled and this path is not in IgnoreHeaders, so it would
have failed CI.

ConsensusSpanNames.h had also become the only one of the eight *SpanNames.h
headers to include anything beyond SpanNames.h. That cost is paid by every
consumer: PeerImp.cpp, ConsensusReceiveTracing.h and RCLConsensus.cpp want
only name and key constants, but were newly compiling ConsensusTypes.h and
DisputedTx.h through it.

Move both helpers to a new ConsensusSpanLabels.h, which owns the domain
includes. ConsensusSpanNames.h is dependency-free again like its siblings, and
the labels reach their only production caller, Consensus.h, directly.

Also from the review:

- phaseOpen() had grown to 81 lines, over the 80-line limit. Extract
  annotateOpenStart() and annotateOpenClose(), which also removes the repeated
  span guards. phaseOpen is 72 lines; startRoundInternal drops 103 to 93,
  still over the limit but it was 99 before this work began.
- Note at the CLOG why the log text keeps the shouldCloseLedger name: existing
  consumers match on it.
- whyCloseLedger's doc claimed "both log identically", implying the wrapper
  logs too. It delegates, so the logging happens once either way.
- Cross-reference proposers_validated and proposers_finished, which sit eight
  lines apart and count different things: validators of the previous ledger
  versus those already past it.
- The two static_asserts no longer sit inside TEST bodies with SUCCEED(); they
  fire at compile time regardless. Also "consteval-safe" was wrong; they are
  constexpr.
- SpanGuardFactory.cpp claimed a libxrpl test cannot include the consensus
  span-name header. The new test in the same directory does exactly that, so
  the claim is corrected to name the real constraint: the rpc_* constants it
  needs live in an xrpld-level header.
2026-08-24 20:45:24 +01:00
Pratik Mankawde
9801b1e297 Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill 2026-08-24 20:42:01 +01:00
Pratik Mankawde
1b62ab096f Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation 2026-08-24 20:42:01 +01:00
Pratik Mankawde
1515d7fe46 Merge branch 'pratik/otel-phase6-statsd' into pratik/otel-phase7-native-metrics
# Conflicts:
#	src/libxrpl/telemetry/TelemetryConfig.cpp
2026-08-24 20:41:43 +01:00
Pratik Mankawde
ccd209b388 Merge branch 'pratik/otel-phase5-docs-deployment' into pratik/otel-phase6-statsd 2026-08-24 20:40:18 +01:00
Pratik Mankawde
26ba0c3698 Merge branch 'pratik/otel-phase4-consensus-tracing' into pratik/otel-phase5-docs-deployment 2026-08-24 20:40:18 +01:00
Pratik Mankawde
f4718cee08 fix(telemetry): correct mistimed and ambiguous consensus phase attributes
Three defects found in review of the two preceding commits.

Drop disputes_count_initial. It claimed to be the dispute count carried in
from the positions held at close, but startEstablishTracing() runs a full
timer tick after closeLedger(): timerEntry() dispatches
`if (phase_ == Open) phaseOpen(); else if (phase_ == Establish)
phaseEstablish();`, and phase_ was Open on the closing tick, so the else-if
cannot run. With ledgerGRANULARITY at 1s the value absorbed up to a second of
dispute growth from peer proposals and arriving tx sets. Making it honest
needs either a member captured at close or moving span creation into
closeLedger(), so it is removed rather than shipped mislabelled.

Record close_time_avalanche_state on recovered rounds. startRoundInternal()
reset establishSpan_ inline, discarding the span before the attribute was
written, so the value was present only on rounds that reached Accepted --
survivor bias in exactly the rounds worth investigating. It now calls
endEstablishTracing(). The comment claiming this avoided "reporting a stale
regime" was wrong: closeTimeAvalancheState_ is not reset until 39 lines
later, so the value was still that span's terminal regime.

Rename avalanche_state to close_time_avalanche_state. DisputedTx carries a
second, per-transaction avalanche tracker; the bare name invited reading a
close-time-only value as the transaction one, which is the tracker that
actually escalates in a stuck round.

Also: both label helpers now fall through to "unknown" instead of a
plausible-looking regime, matching to_string(ConsensusPhase); and the header
now records that the end-of-open attributes are absent on recovered and
simulated rounds, and that tx_sets_acquired can skew either way because
handleWrongLedger clears currPeerPositions_ but not acquired_.

Tests: the minimum-open-time assertion used prevRoundTime=10s, where
openTime=1s trips the too-fast branch as well, so deleting the ledgerMinClose
check entirely left it green. Replaced with prevRoundTime=2s, which isolates
the branch. Added the others-closed boundary, which is strict and was
untested in either direction, its integer truncation for odd prevProposers,
and its precedence over the no-transactions and minimum-open branches.
2026-08-24 20:40:00 +01:00
Pratik Mankawde
c5829df68f feat(telemetry): record which consensus rounds requested a tx-set fetch
A tx-set fetch carried no key tying it to the consensus round that needed
the set, so attributing a stalled fetch to a round meant guessing from
timestamps. One fetch is wanted by many rounds -- it is keyed by set hash,
survives the round sweep, and the round never blocks on it -- so a single
parent, link or attribute cannot describe the relationship.

Instead the fetch span records one timestamped event per requesting round,
carrying the round's parent-ledger hash and the ledger it is building. Both
attribute keys already existed in the shared telemetry namespace with
exactly this meaning, and the existing addEvent API is used as-is, so no
new telemetry surface is added and the whole feature compiles out with
telemetry disabled.

The event fires once per round rather than once per peer proposal, keyed on
the round's parent-ledger hash: that distinguishes rounds started on
different forks at the same height, which a ledger-height compare cannot.
A mid-round wrong-ledger recovery re-enters consensus without re-caching the
round identity, so a fetch begun after that switch is attributed to the
pre-switch round; the limitation is documented where the values are cached.

Also fixes the fetch span's end time, which depended on when the C++ object
was destroyed. Three of the four exits that stop pursuing a fetch -- the set
arriving from elsewhere, the round sweep, and shutdown -- ended the span
only via the destructor, so the recorded duration included however long any
reference happened to be held. Each exit now ends the span itself, plus
cancel() and container teardown, and the destructor asserts the span is
already closed rather than closing it: a fallback that can never legitimately
fire should fail loudly instead of hiding a missed exit. An abandoned fetch
also no longer asks peers for a set nobody wants, which previously led to
charging those peers for answering our own request.

Verification: pre-commit and TIDY=1 clang-tidy pass; levelization is
unchanged. NOT compiled -- the branch is blocked by a gcc-15 internal
compiler error in the unrelated xrpl.libxrpl.rdb unity translation unit.
Runtime behaviour is unasserted: xrpl_tests links only xrpl.libxrpl, so
TransactionAcquire is unreachable from GTest; the added tests cover the new
span-name and attribute constants only.
2026-08-24 20:15:11 +01:00
Pratik Mankawde
6eaf7316e5 feat(telemetry): record the ledger close reason on consensus.phase.open
The open phase ended for one of four distinct reasons, but
shouldCloseLedger() collapsed them into a bool, so a trace could say when a
phase ended and never why. "The network closed without us" and "nothing was
waiting" are the same span today.

Add whyCloseLedger(), which holds the decision and returns
LedgerCloseReason. shouldCloseLedger() keeps its exact signature and becomes
a one-line delegation, so its callers and unit tests are untouched and the
branch logic is not duplicated. phaseOpen() calls whyCloseLedger() directly;
both emit the same journal and CLOG output, so only one is called.

New attributes on consensus.phase.open, both set once on the closing tick:

  close_reason         anomaly | others_closed | idle | normal
  proposers_validated  trusted peers that had already validated the prior
                       ledger, reusing the value the decision was made on

Absent on the simulate() close path, which bypasses the decision rather than
having a reason invented for it.

Skipped has_open_transactions: hasOpenTransactions() is
!getOpenLedger().empty(), which is false on a quiet network for most of a
round, and close_reason=idle already implies it. The sibling
consensus.ledger_close span carries tx_count_open, which is the same fact
with a count instead of a boolean.

shouldCloseLedger() now has no production caller; it stays exported so the
public API and its tests are unchanged.

Tests pin every input vector from should_close_ledger to its literal reason,
including that the anomaly check outranks others-closed, and cover the
inclusive idle boundary either side by one millisecond.
2026-08-24 20:13:39 +01:00
Pratik Mankawde
31619219cc feat(telemetry): add set-once state attrs to consensus phase spans
consensus.phase.open and consensus.establish carried almost no state of
their own. Span attributes are not inherited, so the ledger context on the
parent consensus.round span does not describe either child, and the few
attributes the establish span did carry are rewritten on every iteration
and therefore only ever report the final value.

Add seven attributes that are read from state already in scope, are
written exactly once, and are not duplicates of the parent round span:

  consensus.phase.open (start)
    start_reason            initial, or recovered on a handleWrongLedger
                            re-entry, which emplaces a SECOND phase.open
                            span under the same round
    previous_close_agree    feeds the sinceClose branch in phaseOpen()
    peer_positions_at_open  positions in hand after playbackProposals(),
                            the head start the round began with
    early_close_triggered   the round skipped the timer because enough
                            peers had already closed

  consensus.phase.open (end)
    tx_sets_acquired        candidate tx sets held at close, read before
                            our own position is added; a low count against
                            a high peer_positions_at_close means tx-set
                            fetches did not land, not disagreement

  consensus.establish (start)
    disputes_count_initial  disputes carried in from the positions held at
                            close, as opposed to disputes_count, which is
                            overwritten each iteration

  consensus.establish (end)
    avalanche_state         terminal close-time convergence regime; the
                            derived avalanche_threshold is a weight and
                            cannot be inverted back to the state

The avalanche label is mapped by a new constexpr avalancheStateLabel() in
ConsensusSpanNames.h rather than an inline switch, so the four labels stay
under the naming check's L1 ownership and are unit-testable.

Deliberately not added: ledger_seq and consensus_mode, which would only
copy the parent round span's values down; tx set size and position hash,
which the TxSet concept does not expose portably across RCLTxSet and the
csf simulator; and the peer-unchanged and dead-node counters, whose
underlying state is reset mid-round and so would report a misleading value.

Behaviour is unchanged. The early-close condition is hoisted into a named
local so the annotation happens before timerEntry(), which can reach
closeLedger() and end the open-phase span.

Tests pin the wire strings for every new key and value and cover all four
enumerators of the avalanche mapping. They need no telemetry runtime: the
csf simulator returns an invalid round span context and a null Telemetry,
so consensus spans there are null guards and attribute writes are no-ops.
2026-08-24 19:54:15 +01:00
Pratik Mankawde
7c70e142e9 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
# Conflicts:
#	src/xrpld/telemetry/MetricsRegistry.cpp
2026-08-21 13:09:09 +01:00
Pratik Mankawde
6e2b2da772 fix(telemetry): resolve microsecond latencies below 100us
The microsecond ladder's first edge was 100us, which sat ABOVE the mass of
every instrument using it. Measured on devnet: 99.3% of job_queued_us
samples, 92.5% of job_running_us and 90.4% of getobject_lookup_us fell in
that first bucket. histogram_quantile then interpolated inside bucket 0 and
returned `quantile / fraction_in_bucket_0 x first_edge` -- p75/p95/p99 of
job_queued_us read 75.52/95.66/99.69us against a prediction of
75.53/95.67/99.70. Three-decimal agreement: those panels were reporting
arithmetic on the bucket edge, not latency.

The fix was already half-written. kSubMillisecondBoundaries had been parked
in MetricsRegistry.cpp as [[maybe_unused]] with a comment noting exactly this
problem for nodestore reads. Its edges are now folded into kMicrosecondBuckets
rather than deleted, so the parked intent is carried forward: 1..1000us
resolution where the mass is, upper edges unchanged so multi-second stalls
stay measurable.

Also moves the GetObject count and charge ladders into HistogramBuckets.h, so
all five ladders have one owner and one set of invariant tests (29 now).

Adds check_bucket_parity.py, wired into the existing OTel naming workflow.
The C++ millisecond ladder and the collector's spanmetrics ladder are
specified to agree over their shared range; they were identical when shipped,
then the collector side alone was extended and nothing noticed for eleven
phases. The check asserts containment rather than equality, because jobs
outlive spans -- jobq_updatepaths averages ~60s, which no span approaches, so
demanding equality would force a ceiling that censors it. Verified it rejects
a missing collector edge, a bogus in-range edge, and a return to the 5s
ceiling.

ledger-data-sync's "Job Queue Wait p95 By Type" moves off the beast
jobq_*_q_milliseconds pair onto job_queued_us filtered by job_type. Those
beast metrics are ms-quantised at the source (Event rounds up to a whole
millisecond), so 94-100% of their samples sat in the first bucket and no
ladder change could fix them. Note the label values are camelCase
(job_type="ledgerData"), not the lowercase metric-name fragments.

Both histogram-fed alert thresholds re-validated and left unchanged, with the
measured basis recorded so neither gets tuned against the old artefact: only
0.0022% of job_queued_us samples exceed the 1s threshold, and every edge
bracketing the 1000ms ios_latency threshold survived the ladder change.

Docs: the rpc_size "known issue -- tracked separately" notes in the runbook
and 09-data-collection-reference are now resolved notes, the stale 10-edge
span_duration bucket list is corrected to the collector's real 20, and the
runbook gains a "Reading A Histogram Percentile" section covering both
saturation traps and the expected discontinuity after a ladder change.
2026-08-21 12:46:56 +01:00
Pratik Mankawde
7735d725fb Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill
# Conflicts:
#	docker/telemetry/grafana/dashboards/rpc-pathfinding.json
#	src/libxrpl/telemetry/Telemetry.cpp
2026-08-21 12:33:13 +01:00
Pratik Mankawde
8169894ada Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation 2026-08-21 12:30:50 +01:00
Pratik Mankawde
24094e427b fix(telemetry): give each histogram unit its own bucket ladder
This is the change that actually lifts the 5 s ceiling. Until now the
millisecond ladder and the Unit type existed but nothing consumed them.

Telemetry.cpp registered ONE histogram view: instrument name pattern "*",
unit exactly "ms", boundaries {1, 5, ..., 1000, 5000}. Verified against the
installed SDK, "*" matches every name and "ms" matches exactly, so that view
governed every beast::insight Event -- all 54 of them, whatever they measure.
Measured on devnet: 24.9% of rpc_size samples and 100% of jobq_updatepaths
samples fell above 5000. A quantile landing in the `+Inf` bucket reads back
as the second-highest edge, so those p95s reported a flat 5000 rather than a
measurement, and the 1 s to 5 s span was a single four-second-wide bucket
that any quantile inside it had to interpolate across.

Replaces it with one view per unit, keyed on the unit an instrument declares:

- `ms` gets kMillisecondBuckets: every representable edge of the collector's
  spanmetrics ladder, plus 60 s and 120 s. The extensions are deliberate --
  jobq_updatepaths was measured averaging 59,956 ms, which no span
  approaches, so parity alone would still censor it.
- `By` gets kByteBuckets, placed from the measured response distribution
  (mean 2131 B, half under 1 kB, tail mean bounded at 7538 B).

OTelEventImpl now derives its declared unit AND its description from unit()
instead of hardcoding "Duration in ms"/"ms", so rpc_size exports as
rpc_size_bytes on the byte ladder. rpc-pathfinding's "RPC Response Size"
panel follows the rename; its unit was already decbytes and is now truthful.

Also corrects Phase7_taskList.md, which still specified the 5000 ladder as
"matching SpanMetrics". That was true when written and became false when the
collector ladder was extended on its own -- implementing the plan as written
reproduced the bug, so the spec is where the defect had come to live. The
edges now have exactly one owner and the plan points at it.
2026-08-21 12:30:38 +01:00
Pratik Mankawde
251cd181a7 fix: Reject unreadable [telemetry] TLS certificate paths at startup
With telemetry enabled and use_tls=1, makeTelemetrySetup now reads each
non-empty tls_ca_cert / tls_client_cert / tls_client_key path and refuses to
start when the file is missing or cannot be read. The message names the config
key, the path and the OS error, instead of leaving the problem to surface much
later as an opaque TLS handshake failure inside the exporter.

Reading the file with getFileContents, as the gRPC server already does for its
own ssl_cert and ssl_key pair, proves the file is both present and readable; an
existence test alone would miss a permissions problem. The contents are
discarded.

Both gates are deliberate. The check is skipped when enabled is 0, so a stale
cert line still cannot stop a node from booting, and when use_tls is 0, where
the exporter never opens the files. An empty path stays valid; for tls_ca_cert
it selects the system CA store.

Six GTest cases cover the three keys that can fail, the all-readable case, and
each gate on its own.
2026-08-21 12:29:02 +01:00
Pratik Mankawde
63dc5cce65 Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill 2026-08-21 12:13:29 +01:00
Pratik Mankawde
0136291aa0 Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation 2026-08-21 12:13:19 +01:00
Pratik Mankawde
76c9051203 feat(insight): let an Event declare what it measures
beast::insight::Event documents itself as carrying "a millisecond time, or
other integral value", but both backends assumed the first case: the OTel
bridge declared every instrument with unit `ms` and StatsD tagged every
sample `|ms`. One Event does not measure time -- ServerHandler's "size"
records the serialized RPC response length -- so it exported as
rpc_size_milliseconds and inherited the millisecond bucket ladder. A quarter
of its samples landed above that ladder's top edge, and since Prometheus
returns the second-highest edge for a quantile in the `+Inf` bucket, its p95
panel showed a flat 5.00 kB rather than a measurement.

Adds beast::insight::Unit (Millis, Bytes) plus otelUnitCode(), carried on
EventImpl and selectable at makeEvent(). Naming the unit at creation is what
lets a backend pick the export unit and, through it, the bucket ladder.

- Collector gains a virtual makeEvent(name, Unit) whose default delegates to
  the millisecond overload, so a collector that cannot act on a unit keeps
  working unchanged. NullCollector and the Groups wrapper override it.
- The Groups override matters most: call sites reach a collector through a
  Group, so forwarding only the prefixed name would silently drop the unit.
  A test covers that hop specifically.
- Event gains notify(std::uint64_t) for non-duration samples, replacing
  ServerHandler's `Event::value_type{response.size()}` -- wrapping a byte
  count in a std::chrono::milliseconds compiles but reads as a duration to
  everything downstream.
- EventImpl::value_type stays std::chrono::milliseconds. Widening it would
  change the wire value of every existing StatsD timer, and metrics needing
  finer resolution use the OTel-native microsecond instruments.

The StatsD collector deliberately keeps emitting `|ms`: that path is retired
here (its UDP port is commented out of the compose file and the integration
test fails if anything listens on 8125), so changing its wire format would
alter a legacy contract with no consumer and no way to verify it.

The exported name does not change yet -- OTelEventImpl still hardcodes its
unit. That follows with the unit-keyed histogram views.
2026-08-21 12:11:32 +01:00
Pratik Mankawde
cbfbea67f2 feat(telemetry): own every histogram ladder in one tested header
The bucket edges for the OTel histograms lived as file-local `namespace {}`
constants, unreachable from any test, and they drifted from the collector's
spanmetrics ladder they were specified to match. The millisecond ladder
stayed capped at 5 s after the collector side was extended to 30 s, so any
quantile above 5 s read back as a flat 5000 -- Prometheus returns the
second-highest edge for a quantile in the `+Inf` bucket, which looks like a
measurement rather than an error.

Adds include/xrpl/telemetry/HistogramBuckets.h as the single owner of the
ladders, with a constexpr validator plus static_asserts so a descending or
duplicated edge cannot compile, and gtest coverage that pins the floor and
ceiling against the measured distributions:

- kMillisecondBuckets carries every representable collector edge and extends
  to 120 s, because the updatepaths job type averages ~60 s and a 30 s
  ceiling would censor it exactly as 5 s does today. Sub-millisecond
  collector edges are omitted: beast::insight::Event rounds durations up to
  whole milliseconds, so they would collect nothing.
- kByteBuckets is new, for Events whose samples are sizes rather than
  durations. Edges follow the measured RPC response distribution (mean
  2131 B, half under 1 kB, tail mean bounded at 7538 B) rather than a guess,
  so the resolution sits between 512 B and 64 kB.

No behaviour change yet -- nothing consumes the header until the views are
rewired.
2026-08-21 11:51:22 +01:00
Pratik Mankawde
3a458f2873 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics 2026-08-20 19:12:45 +01:00
Pratik Mankawde
3ed5baf485 Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill 2026-08-20 19:11:28 +01:00
Pratik Mankawde
74d0bb4ae7 Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation 2026-08-20 18:19:34 +01:00
Pratik Mankawde
d8035a35aa fix(telemetry): include the headers the new code depends on
clang-tidy runs misc-include-cleaner with WarningsAsErrors, so a symbol
reached only transitively fails CI. Add the direct includes for JLOG,
beast::Journal, StartUpType, TokenType, toBase58, std::exception and
std::size_t.
2026-08-20 18:19:22 +01:00
Pratik Mankawde
71d44e7a7b Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics 2026-08-20 16:54:19 +01:00
Pratik Mankawde
308e77d003 Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill 2026-08-20 16:54:11 +01:00
Pratik Mankawde
dbef6a705e Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation 2026-08-20 16:54:11 +01:00
Pratik Mankawde
c6491eef27 fix(telemetry): log without JLOG in the beast insight collector
JLOG is defined in xrpl/basics/Log.h, and libxrpl.beast cannot include
xrpl.basics -- basics depends on beast, not the reverse. Use the journal
stream idiom the rest of the file already uses.
2026-08-20 16:54:00 +01:00
Pratik Mankawde
d450b16b71 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics 2026-08-20 16:52:45 +01:00
Pratik Mankawde
f572aedeec Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill
# Conflicts:
#	OpenTelemetryPlan/05-configuration-reference.md
#	docs/telemetry-runbook.md
2026-08-20 16:50:46 +01:00
Pratik Mankawde
b2ba81a56c Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation
# Conflicts:
#	docs/telemetry-runbook.md
2026-08-20 16:46:56 +01:00
Pratik Mankawde
466660564f Merge branch 'pratik/otel-phase6-statsd' into pratik/otel-phase7-native-metrics
# Conflicts:
#	src/xrpld/app/main/Main.cpp
2026-08-20 16:45:37 +01:00
Pratik Mankawde
597b0c2bab Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill
# Conflicts:
#	docker/telemetry/xrpld-telemetry.cfg
#	src/libxrpl/beast/insight/OTelCollector.cpp
#	src/libxrpl/telemetry/Telemetry.cpp
#	src/xrpld/app/main/Application.cpp
2026-08-20 16:43:32 +01:00