Commit Graph

16830 Commits

Author SHA1 Message Date
Pratik Mankawde
d0eb346ec4 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics 2026-08-27 12:47:39 +01:00
Pratik Mankawde
d423863b82 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Two conflicts.

RCLConsensus.cpp: upstream restructured makeAcceptSpan so the accept span's
attributes sit behind if (*span). This branch's own contribution there is the
consensus round-duration histogram, which is kept -- placed inside the
telemetry guard but OUTSIDE the span-liveness test, because a metric must
still record when the trace category is disabled or the span was not created.
The duplicated attribute lines on this side are dropped; the guarded block
upstream added supersedes them.

MetricsRegistry.cpp: kept this branch's JobQueue.h include, which it uses.
Its Journal.h include was dropped as a duplicate -- the file already includes
that header higher up, with a comment explaining why it is unguarded, and
readability-duplicate-include is fatal under WarningsAsErrors.
2026-08-27 12:47:03 +01:00
Pratik Mankawde
ed92501730 style(telemetry): cut the comments I over-wrote back to the guideline
The comments I added with the hierarchy sampling fix and the trigger change ran
to sixteen and twelve lines. The guideline is short and plain English. Rationale,
CI run numbers and the list of which relationships were affected belong in the
commit message, which is where they already are; inline they push the code apart
and go stale as soon as the reasons change.

Trimmed the sampling comment from sixteen lines to four, the re-check comment
from eight to four, _traceql_name_predicate's docstring from fourteen lines of
explanation to three, and the push-trigger comment from twelve to seven. Each
keeps what a reader needs at that line -- what the code does and the one
non-obvious reason -- and drops the history.

Comment-only: 13 insertions against 35 deletions, no statement changed.

Left alone deliberately: this file has ten pre-existing comment blocks longer
than six lines, including one added recently by another party. Rewriting someone
else's comments is not mine to do here, and the guideline is being applied to what
I wrote.

Verification: 7/7 validator tests pass; validate_telemetry.py compiles; the
workflow YAML parses, still carries no branches filter, and still lists 12 paths;
otel-naming exits 0.
2026-08-27 12:46:17 +01:00
Pratik Mankawde
a14d9ac806 Merge branch 'pratik/otel-phase9-metric-gap-fill' into pratik/otel-phase10-workload-validation 2026-08-27 12:44:00 +01:00
Pratik Mankawde
1094b60223 Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill
Four conflicts. None was a take-a-side.

xrpld-telemetry.cfg: dropped the incoming [insight] block. This branch already
has one, and duplicate ini sections do not replace each other -- parseIniFile
appends onto the same section, so keys merge last-wins and the effective config
is one that appears nowhere in the file.

src/tests/libxrpl/CMakeLists.txt: kept this branch's else() branch, which
compiles MetricsRegistry.cpp for the telemetry-off build, and dropped only the
ValidationTracker target_sources inside if(telemetry). Upstream relocated that
one out of the guard, so keeping both would have compiled it twice. The
else() branch is this branch's own: the MetricsRegistry test exists only here,
and without its implementation the off build would not link.

RCLConsensus.cpp: union. The metric macros and registry come from this branch,
PropagationHelpers from upstream; all three are used.

PeerImp.cpp: MetricMacros.h stays unguarded, because all seven XRPL_METRIC_*
uses in this file are on unconditional paths and the header is what defines
them. ConsensusReceiveTracing.h takes upstream's guarded placement, and this
branch's guarded GetObjectMetricNames.h is kept beside it.
2026-08-27 12:43:24 +01:00
Pratik Mankawde
eaeb2dc8e1 style(telemetry): brace the conditional bodies in the recording tests
Two if constexpr/else pairs had single-statement bodies holding a gtest macro.
ShortStatementLines exempts a one-line body, but these macros expand to
multi-line constructs and some clang-tidy versions report the expansion's
range, which would make readability-braces-around-statements fire under
WarningsAsErrors. Braces also read better beside an else.
2026-08-27 12:39:40 +01:00
Pratik Mankawde
b8cb36ffca fix(telemetry): refuse an empty capture, and name a bad baseline entry
An empty metric surface counted as a complete capture. build_query_plan
returns an empty plan without complaining for any config that yields no
gated keys, so pointing --metrics at the wrong file exits 0 and hands the
paste-me path a metrics:{} artifact to offer as the next baseline. Nothing
about such a run is evidence the pipeline works, so declared == 0 is now a
failure rather than vacuously complete.

The bounds checker also raised AttributeError on a baseline entry that is
not an object, instead of naming the key. A validator whose job is to catch
a malformed contract should report it, not crash on it.

Test cleanup is bound to its own temp tree, so a loop no longer leaves five
of six directories behind.
2026-08-27 12:38:53 +01:00
Pratik Mankawde
e0a0986d31 refactor(ledger): hold acquire counts in Counter, not guarded atomics
The nine counters have no reader outside telemetry, so the record calls were
wrapped in preprocessor branches at six sites. Holding them in Counter makes
the storage and the increments disappear together, so the call sites read as
ordinary code.

The two deferral sites keep a compile-time block, because their argument is
a string compare that a no-op add would still evaluate.

The unit tests now assert both configurations: every expectation of a non-zero
count has a mirror expectation of zero, so neither build is left unasserted.
2026-08-27 12:37:05 +01:00
Pratik Mankawde
4f056496f8 Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation 2026-08-27 12:32:54 +01:00
Pratik Mankawde
fb336c55c5 Merge branch 'pratik/otel-phase6-statsd' into pratik/otel-phase7-native-metrics 2026-08-27 12:32:38 +01:00
Pratik Mankawde
2e3e25d056 Merge branch 'pratik/otel-phase5-docs-deployment' into pratik/otel-phase6-statsd
Two conflicts, both in PeerImp's proposal and validation receive paths, and
both resolved by taking the incoming side: it holds the span in a handle that
stays empty when telemetry is compiled out and moves every attribute behind
if (span && *span), which supersedes the unguarded form on this side.

Taking the incoming text renamed consSpan to span in both blocks, while the
two job-lambda captures further down had merged cleanly and still named
consSpan. Renamed those captures so each names the handle its own function
declares.

This branch's own guard on the inbound-validation ledger_hash attribute is a
different span in a different function; it merged cleanly and is preserved.
2026-08-27 12:32:27 +01:00
Pratik Mankawde
bba40f0bb8 Merge branch 'pratik/otel-phase4-consensus-tracing' into pratik/otel-phase5-docs-deployment 2026-08-27 12:30:33 +01:00
Pratik Mankawde
e27d877af9 docs(telemetry): state what the recording types actually cost
The header told callers to declare members [[no_unique_address]]. That
attribute has no other use in this repo, and MSVC ignores the standard
spelling for ABI compatibility, so following the advice would have been a
portability wart for one byte per member.

Say instead that the compiled-out member collapses to padding, and that what
these types buy is work not being done.
2026-08-27 12:26:51 +01:00
Pratik Mankawde
85b42d8bf0 fix(consensus): send no trace context when nothing is being traced
Both broadcast paths passed *msg.mutable_trace_context() to the injector,
which allocates the submessage and sets its has-bit before the injector can
decide there is nothing to write. A compile-time guard covered the
telemetry-off build, but a node with telemetry compiled in and no active
span -- a disabled category, telemetry disabled by config, or a round that
is not being traced -- still broadcast an empty TraceContext to every peer,
and every peer took its has_trace_context() branch to extract nothing.

Add SpanGuard::hasCurrentContext(), a predicate that tests the same two
conditions the injector bails out on without allocating, and an
injectCurrentContext(message) helper that uses it to decide whether to
create the submessage at all. Both consensus call sites now call the helper
unguarded.
2026-08-27 12:25:09 +01:00
Pratik Mankawde
b5a453d351 refactor(overlay): time the get-object lookup with a Stopwatch
Four preprocessor branches around one function -- one for the start
timestamp, one for the elapsed computation, one for the call that reports it,
and one around the reporting method itself -- become none. The clock is still
not read when telemetry is compiled out, because Stopwatch holds no state in
that build.

The reporting method is now compiled in both builds so its call site needs no
guard. Every statement in its body is an XRPL_METRIC_* argument, and those
macros discard their arguments when telemetry is off, so the body costs
nothing there. Its four arguments are all values the request already computed.
2026-08-27 12:24:19 +01:00
Pratik Mankawde
017ef1b33e fix(telemetry): keep Counter's copy semantics the same in both builds
std::atomic implicitly deletes copy and move, so a class holding a Counter is
non-copyable when telemetry is compiled in. Compiled out, Counter was an empty
type with no such member, which would have made that same owner freely
copyable in one build only -- the per-configuration API difference these
utilities exist to avoid.

Declare copy and move deleted so both builds agree, and assert it
unconditionally in the tests.
2026-08-27 12:18:57 +01:00
Pratik Mankawde
c93de72401 Merge branch 'pratik/otel-phase3-tx-tracing' into pratik/otel-phase4-consensus-tracing
Two conflicts, both in the telemetry include blocks.

SpanGuard.h: kept the union. The incoming side moves <memory> inside the
telemetry guard and adds <type_traits>; this branch adds <initializer_list>,
<utility> and the protocol::TraceContext forward declaration. Guarding
<memory> is correct here: the only std::shared_ptr uses are SpanContext's
member and constructor, both inside the guard, and SpanGuardHandle is a
template parameter name rather than a smart-pointer typedef.

NullTelemetry.cpp: took only the incoming guarded Journal.h block. The
incoming hunk also carried <memory> and <utility>, which this branch already
includes below the guarded OpenTelemetry block; taking them as well would
have tripped readability-duplicate-include.
2026-08-27 12:16:36 +01:00
Pratik Mankawde
ed3817968d feat(telemetry): add recording utilities for telemetry-only state
Sites that exist only to be reported currently spell their own gating: an
#ifdef around a member, another around the clock read, another around the
call that reports it. That puts preprocessor branches through business logic
and leaves each class with a different member set per build.

Add kEnabled, Stopwatch and Counter. Each holds real state when telemetry is
compiled in and is an empty type with no-op methods when it is not, while the
member stays declared in both builds so no class's API differs by
configuration. Tests assert both configurations from one file, including that
the compiled-out types are empty.
2026-08-27 12:14:54 +01:00
Pratik Mankawde
0259604a35 Merge branch 'pratik/otel-phase2-rpc-tracing' into pratik/otel-phase3-tx-tracing 2026-08-27 12:14:00 +01:00
Pratik Mankawde
fa23fb51ea fix(telemetry): drop insight keys the OTel collector discards
prefix and service_instance_id are read and thrown away on this path, so
the node's identity label comes from [telemetry] instead. The old comment
claimed the insight copy was required or panels would be empty.
2026-08-27 12:13:46 +01:00
Pratik Mankawde
cb92d59f11 fix(telemetry): drop the inert insight prefix from the OTel path
formatName() never reads prefix, so setting it here does nothing and the
exported names are bare and lowercase. Leaving it invites queries written
against xrpld_jobq_job_count, which match no series.

The StatsD examples keep it, because that path does apply it to the name.
2026-08-27 12:12:31 +01:00
Pratik Mankawde
0d4d624622 refactor(telemetry): inject trace context from the whole message
The existing helper takes the TraceContext submessage, so every caller
writes *msg.mutable_trace_context(). On a protobuf optional field that
allocates the submessage and sets its has-bit before the helper runs, so a
message ships an empty TraceContext whenever nothing is recorded and its
peers take their has_trace_context() branch to extract nothing.

Add an overload taking the parent message, which decides whether to create
the submessage at all, and correct the header note that claimed the old
helper was already free.
2026-08-27 12:11:10 +01:00
Pratik Mankawde
2ea9005c1b Merge branch 'pratik/otel-phase1c-rpc-integration' into pratik/otel-phase2-rpc-tracing 2026-08-27 12:08:59 +01:00
Pratik Mankawde
69c413d332 Merge branch 'pratik/otel-phase1b-telemetry-infra' into pratik/otel-phase1c-rpc-integration 2026-08-27 12:08:26 +01:00
Pratik Mankawde
5638cd976e fix(telemetry): write the wildcard span predicate without a backslash escape
The conjunction query works. Run 33062418036 proved it on real Tempo: both
hierarchies that newest-N sampling made unassertable now PASS --
txq.accept -> txq.accept_tx and ledger.acquire -> ledger.acquire.txtree -- along
with every other literal-child pair. Only the two wildcard children failed, and
not because of the sampling change.

They failed with HTTP 400, "invalid TraceQL query: parse error at line 1, col 68:
invalid char escape". _traceql_name_predicate built the pattern with re.escape,
giving name=~"rpc\.command\..*", and TraceQL's string lexer refuses a backslash
escape it does not recognise -- the query never reached the regex engine at all. A
literal dot is now written as the character class [.], which carries no backslash
for the lexer to refuse while still meaning a literal dot to the engine behind it.
Leaving the dots bare would have parsed, but would match any character in those
positions, which is the looseness _span_name_matches exists to avoid.

The builder now also rejects a span name containing anything outside
lower_snake_case, dots and the glob star, rather than passing it through
unescaped. Every name in the contract is of that shape, so this changes nothing
today; it exists because the failure mode it guards against is exactly the one
above -- a character that means something to one layer and something else to the
next, discovered only from a 400 in CI.

Worth recording why the tests did not catch this. The stub evaluated the pattern
with Python's re, which accepts \. happily, so it modelled the regex engine and
not the query lexer sitting in front of it. A stub is only as good as the layer it
imitates, and the layer that rejected this was one the stub did not represent. The
new test therefore asserts the property the lexer enforces -- that no backslash
appears in the predicate at all -- rather than any particular spelling, plus that
the pattern still accepts rpc.command.fee and still rejects a near-miss whose
separators are not dots.

Verification: 7/7 tests pass, and the new one was watched failing first with the
exact string Tempo rejected, name=~"rpc\.command\..*"; the full query the check
now builds was printed and confirmed backslash-free; validate_telemetry.py
compiles. Three unrelated files in this worktree are another party's live work and
were left unstaged.
2026-08-27 11:44:12 +01:00
Pratik Mankawde
b4094dd91d fix(telemetry): give NullTelemetry the getMeter override it was missing
NullTelemetry overrides the other OTel virtuals behind the telemetry guard so
it stays concrete in both configurations, but getMeter was added to the base
without a matching override. That leaves the class abstract in a telemetry-on
build, which compiles today only because its sole instantiation sits behind
#ifndef. Anyone constructing one with telemetry on gets an abstract-class
error pointing at the base, not at the missing override.

Return a NoopMeter from a function-local static, mirroring getTracer.
2026-08-27 11:18:22 +01:00
Pratik Mankawde
39fa18e898 test(telemetry): assert the tx-tree acquire hierarchy again
The sampling fix that just merged forward removes the only reason this was
skipped. The check no longer inspects the three newest parent traces; it asks
Tempo for traces containing both parent and child.

Worth recording why this phase was the one that failed while its two siblings
passed, because the original assertion treated all three as equivalent and they
are not. InboundLedger.cpp opens each phase only when that piece is still needed:
header on !haveHeader_ (:672), astree in the else of haveState_ (:689), txtree in
the else of haveTransactions_ (:698). A node acquiring a ledger here almost always
lacks the account-state tree, so astree opens on essentially every acquire. But it
usually already holds the transaction set -- every node sees the same relayed
transactions and builds the same set -- so txtree opens on a minority of acquires.
The child was always emitting, 5 traces of its own on the run that failed; it just
was not in the three most recent acquires.

That is now all three sampling-caused skips retired: txq.accept -> txq.accept_tx
and this one asserted, and txq.enqueue -> txq.batch_clear narrowed to its real
remaining cause, a child that never fires under this workload at all.

Contract on this branch: 24 relationships, 19 asserted, 5 skipped, and zero spans
declaring a parent without an entry. The five are the two pathfind pairs and the
pathfind.request parent (pathfinding disabled and no path-finding RPC issued),
rpc.ws_message -> rpc.process (not a code relationship -- rpc.process is a child
of rpc.http_request), and txq.batch_clear. None is a sampling artifact.

Verification: JSON parses; the four validator tests pass after the merge; 0
unaccounted parentings; counters still 48 span types and 74 unique attributes;
otel-naming exits 0. Whether this holds against a live Tempo is what the run this
push triggers decides -- the stub proves the query shape, not the corpus.
2026-08-27 11:17:47 +01:00
Pratik Mankawde
ce18bb3317 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings in the hierarchy-check sampling fix: the check now asks Tempo for traces
containing both parent and child rather than inspecting the three newest parent
traces, so a child conditional on a state the workload rarely reaches is found
wherever it occurred. Merged clean, no conflicts, no resolution decisions.

This unblocks ledger.acquire -> ledger.acquire.txtree on this branch, which was
skipped for exactly that sampling problem and is un-skipped in the next commit.
2026-08-27 11:17:13 +01:00
Pratik Mankawde
a87d772f40 fix(telemetry): find a span hierarchy where it happened, not only where it is newest
The hierarchy check searched the parent span and inspected the three newest
traces it returned. That is wrong whenever the child is conditional on a state
the workload only sometimes reaches: the parent fires constantly, so its newest
traces are the ones LEAST likely to carry a rare child. Three relationships had
been skipped as unassertable for exactly this, and in none of them was the child
missing -- each emitted traces of its own and simply was not in the three most
recent parent traces.

The check now issues a second query, a TraceQL trace-level conjunction of the
parent and child name predicates, and inspects those traces. Tempo searches its
whole retention for co-occurrence instead of leaving the answer to which traces
happen to be newest. The parent-only query is kept and still runs first, so "the
parent stopped being emitted" stays a distinct failure from "the parent is there
but the child never co-occurs" -- they mean different things to whoever reads the
report, and collapsing them would lose that.

The returned traces are still verified with _span_name_matches rather than the
query result being trusted on its own. Tempo has already guaranteed
co-occurrence, so this is redundant on the happy path; it is kept because it
keeps the glob semantics in one place and means a wrongly built query cannot
silently pass.

_traceql_name_predicate handles the wildcard contracts. TraceQL has no glob
operator, so `rpc.command.*` is sent as name=~"rpc\.command\..*" with the dots
escaped -- unescaped they would match any character in those positions, which is
the looseness _span_name_matches exists to avoid.

Two entries follow from the fix. txq.accept -> txq.accept_tx is asserted again:
its child is created inside the queued-transaction loop behind
`if (feeLevelPaid >= requiredFeeLevel)` (TxQ.cpp:1530) while the parent fires on
every close (:1499), which was the whole reason it failed. txq.enqueue ->
txq.batch_clear stays skipped but for ONE reason now instead of two -- its child
never fires at all under this workload, needing an account with a supersedable
batch, so it is purely a workload gap and needs nothing further from the
validator. The third, ledger.acquire -> ledger.acquire.txtree, lives on the
sync-diagnostics branch and is un-skipped there once this merges forward.

Written test-first, and the first test this module has had. The failing test
reproduces the exact CI message, "txq.accept_tx not found in txq.accept traces",
against a stubbed Tempo whose corpus holds the child only in a trace outside the
newest three. Three sibling tests guard the ways this could be "fixed" wrongly: an
absent child must still fail, a missing parent must still name the parent rather
than the child, and a wildcard child must be satisfied by any family member. The
stub records the queries issued, so the conjunction is asserted rather than
assumed. A stub rather than a live Tempo because the behaviour under test is which
traces the check ASKS FOR -- a passing query against real data proves the data
co-operated, not that the query was right.

The first run of those tests failed for the wrong reason: my stub's name-predicate
regex also matched the resource.service.name="xrpld" term every query carries and
so demanded a span literally named "xrpld". Fixed in the stub, with the lookbehind
commented as load-bearing, before touching production code.

Verification: 4/4 tests pass, and the failing one was watched failing first with
the production message; the issued queries were printed and confirmed to contain
the conjunction; validate_telemetry.py compiles; expected_spans.json parses;
21 relationships, 16 asserted and 5 skipped; counters still 41 span types;
otel-naming exits 0. Three unrelated files in this worktree are another party's
live work and were deliberately left unstaged.
2026-08-27 11:16:28 +01:00
Pratik Mankawde
58d0c30e08 style(consensus): keep the gated span helpers non-static for clang-tidy
The three helpers whose bodies are compiled out with telemetry read members
only inside the guard, so in a telemetry-off build they touch no member and
readability-convert-member-functions-to-static fires. WarningsAsErrors makes
that fatal.

Suppress it where the body is gated, matching OverlayImpl::reportDnsResolve.
Making them static instead would give the two configurations different
signatures, which is the hazard the gating pattern avoids.
2026-08-27 11:13:12 +01:00
Pratik Mankawde
da35290f27 test(telemetry): compile the span-name and stall-rule tests in every build
Both files gated their whole contents on XRPL_ENABLE_TELEMETRY, so 54 tests
were skipped whenever telemetry was compiled out. The stated reason was that
only a telemetry build puts `src/` on this target's include path, but that
include path is unconditional, so the tests were reachable all along.

Nothing in either file needs the OpenTelemetry SDK. The span-name and outcome
headers hold constants and constexpr functions with no telemetry guards,
LoadManager::evaluateStall is a static constexpr member, and the handful of
guard assertions construct a default SpanGuard, which is inactive in either
configuration. SpanNames.h documents this contract for its own constants.
2026-08-27 11:10:04 +01:00
Pratik Mankawde
fa2a09c758 perf(telemetry): increment the copy-forward total only when it is read
copyForwardTotal_ is a second atomic increment beside copyForwardCount_ on the
same event, kept only so a metric never goes backwards: rotate() zeroes the
per-rotation tally for its log line, which leaves that counter unusable as a
rate. Its only reader is MetricsRegistry.cpp:1153, through copyForwardTotal().

Guard the increment. During a rotation window every archive-served
non-duplicate read pays for it, and with telemetry compiled out there is
nothing to read it back.

copyForwardCount_ is untouched: rotate() exchanges it for the "copied forward N
archive-served reads" warning, which is real logging, not instrumentation. The
virtual and the member stay declared unconditionally, so the nodestore
interface has the same shape in every configuration.
2026-08-27 11:06:48 +01:00
Pratik Mankawde
0ad3587462 perf(telemetry): skip the dial clock and span when nothing records them
dialStart_, outcomeReported_ and dialSpan_ are telemetry-only. dialStart_ is
read only by the two elapsed-time computations in reportOutcome();
outcomeReported_ is written and read only there; dialSpan_ is opened in run(),
ended in reportOutcome() and reset in the destructor, and read nowhere else.

outcomeReported_ is not load-bearing for anything but telemetry. It is a
first-call-wins latch over the histogram, the counter and the span attributes.
Every terminal path calls close() or fail() itself, beside its reportOutcome()
call rather than inside it, so suppressing a second report cannot suppress any
teardown.

Guard the three sites: the clock and span setup in run(), the whole body of
reportOutcome(), and the destructor's reset(). Per outbound dial that removes a
steady_clock reading, an optional emplace and reset of a span handle, and the
latch write. The members stay declared in every configuration so the class has
one shape; only the writes are compiled out.

SpanGuard.h, SpanNames.h, MetricMacros.h and the cstdint header move behind the
guard with the code that names them. ConnectAttempt.h still includes SpanGuard.h
for the member.
2026-08-27 11:06:38 +01:00
Pratik Mankawde
d3c1fc67ce perf(telemetry): skip the per-second stall bookkeeping nobody reads
updateStallState() is wholly telemetry: it applies evaluateStall(), stores the
result in currentStallSeconds_ and bumps stallEventCount_. Those two members
have exactly one reader each, MetricsRegistry.cpp:1986 and :2023, both inside
the registry's own XRPL_ENABLE_TELEMETRY region, reached through
getCurrentStallSeconds() and getStallEventCount(), which nothing else calls.
The monitor thread ran it once per second for the life of the process.

Guard the body, not the members or the accessors: a member set that differs
between build configurations is the hazard that once made a test mock abstract.
evaluateStall() stays where it is, being a public constexpr rule with its own
GTest coverage in SyncStateSignals.cpp.

The atomic header moves behind the same guard, as the relaxed memory orders are
named only in the guarded body; LoadManager.h includes it for the members.
2026-08-27 11:06:30 +01:00
Pratik Mankawde
27dc3236d1 perf(telemetry): build the UNL fetch site label only when recorded
reportFetchOutcome() exists only to label unl_fetch_total. It reads the parsed
URI parts, copies the domain, erases any userinfo with an rfind, joins scheme,
host and optional port, then appends a substr of the path -- two std::string
allocations and several copies -- and that label has no other reader. It ran on
every validator-list fetch, so about once per site every five minutes, whether
or not anything could record the counter.

Guard the whole body with XRPL_ENABLE_TELEMETRY rather than change the
signature: the two failure call sites pass a compile-time constant, so an empty
body is all they need. The success call site is guarded too, because its
to_string(bestDisposition()) builds a std::string that only the label consumes.
bestDisposition() itself keeps running, since lastRefreshStatus stores it.

MetricMacros.h moves behind the same guard, as the macro is now named only
inside the guarded body. MetricNames.h stays unconditional, because the fetch
handlers name the outcome constants either way.
2026-08-27 11:06:20 +01:00
Pratik Mankawde
626c4e7ae3 Skip telemetry-only work on the consensus round path
Five sites on the consensus round path did telemetry-only work whether or
not anything could record it.

onClose set four attributes on the ledger-close span. Two cost real work
once per round: OpenLedger::current() takes currentMutex_ and copies a
shared_ptr just to read txCount, and the mode attribute builds a string.
The block now sits behind if (span).

doAccept read the previous close-time resolution and ran a lambda
returning a std::string, feeding the resolution_direction attribute and
nothing else, once per accepted ledger. Now behind if (doAcceptSpan).

makeAcceptSpan, startRoundTracing and createValidationSpan have wholly
telemetry bodies, so each body sits inside XRPL_ENABLE_TELEMETRY.
makeAcceptSpan then allocates no control block per accepted ledger; an
empty handle is safe because doAccept only passes it to activateIfLive(),
which tests it. Its attributes are additionally guarded on the span being
live. startRoundTracing's early return sits after two virtual Telemetry
calls and a strategy string compare, so the whole body is compiled out
rather than reached each round. createValidationSpan yields std::nullopt,
and its two call sites in validate() test the guard as well as the
optional -- an engaged optional holding a dead guard still turned a
32-byte ledger hash into a 64-character string.

Telemetry.h and SpanNames.h are named only by startRoundTracing, so their
includes are guarded the same way to keep misc-include-cleaner satisfied
when telemetry is off.

onPhaseEvent and onOutcomeEvent are left as they are: the generic
Consensus template calls both and the csf simulator implements both, so
they are part of the adaptor surface.
2026-08-27 10:10:26 +01:00
Pratik Mankawde
f799678df3 perf(networkops): count checked validations only with telemetry on
recvValidation paid a virtual registry lookup plus a call into an
out-of-line function with an empty body for every validation it checked.
That is once per unique validation the node accepts, on the check job
PeerImp queues after dropping duplicates. The registry is always
constructed, so the null test never skipped any of it.

incrementValidationsChecked only advances an OTel counter, published as
validations_checked_total. One dashboard panel and one alert rule are its
whole audience; no RPC reply, log line or control decision reads it.

Guard the call. The MetricsRegistry include stays, because
incrementStateChanges in setMode still uses it and that fires only when the
operating mode actually changes.
2026-08-27 10:09:56 +01:00
Pratik Mankawde
6dd93b46ee perf(ledger): count acquisitions only when telemetry is compiled in
Five AcquireStats recording calls ran in every build. The two in TimeoutCounter
are the frequent ones: one on every deferred timer tick and one on every
no-progress tick, for every in-flight TimeoutCounter, and each builds its
argument with isLedgerAcquisition(), which compares the job name std::string
against a literal. The other three fire once per acquisition that aborts, gives
up or completes.

Nothing outside telemetry reads any of the nine counters. The only non-test
caller of any accessor is MetricsRegistry::observeAcquireStats, which itself
sits inside that file's XRPL_ENABLE_TELEMETRY block and so does not exist in a
telemetry-off build.

Guard the five calls, and the AcquireStats include with them, since they were
its only users in these three files. The counters and their accessors are left
alone: with every writer guarded and the only reader absent, they are nine
untouched atomics in one process-wide object, so gating them would add many
preprocessor blocks to the header and force a gated test file to save 72 bytes
of a build that never touches them.

TimeoutCounter::timeouts_ stays outside the guard because the give-up test
reads it, and so does InboundLedger's completionCounted_ latch, which is two
branches once per acquisition and would otherwise be left an unused member.
2026-08-27 10:09:33 +01:00
Pratik Mankawde
8521b96d85 fix(telemetry): stop an incomplete capture becoming the committed baseline
The regression baseline is bootstrapped by copying a CI artifact. The workflow
tested only that timings.json existed, then printed it verbatim under a heading
inviting the reader to paste it in as the new baseline.

capture_timings.py writes that file and only then enforces --min-capture-ratio,
so an incomplete capture leaves a file that exists but covers fewer keys than
the contract declares. The verdict lived in CAPTURE_EXIT, a shell variable local
to run-full-validation.sh that no other program could read. So on a placeholder
baseline plus a thin capture, CI offered an incomplete artifact as the next
baseline, and pasting it narrowed the gate with nothing reporting that it had.
That is the failure shape this harness keeps producing: a degraded result that
looks exactly like a good one.

The artifact now carries its own completeness, next to metrics:

  "capture": { "declared": 20, "captured": 20, "min_ratio": 0.5, "complete": true }

complete is the same condition the producer exits 0 on, computed once with the
exit code read off it, so the flag and the status cannot drift apart. Any
consumer can now tell a complete capture from a thin one, not just CI.

Both paste-me paths refuse rather than warn: the workflow prints the counts and
an error annotation with no JSON, and the comparator explains on stderr while
leaving stdout empty, so a redirect cannot produce a plausible-looking file. A
warning above a copyable block is still a copyable block, and a reader who has
just hit a red gate is already predisposed to re-baseline. A missing capture
block fails closed.

Refusal is scoped to bootstrapping a baseline, not to comparing against one, so
artifacts captured before this change still replay: verified against the run the
current baseline came from, which carries no capture block and still reports 0
regressions. An injected regression is still caught, and the gated surface is
unchanged at 20 keys with 5 excluded.
2026-08-27 09:52:48 +01:00
Pratik Mankawde
455bc9d3a5 Skip proposal and validation receive-span work when inactive
Both inbound peer-message handlers built a receive span and set its
attributes unconditionally. The proposal path turned two 32-byte hashes
into full hex strings and then took a 16-character substring of each --
four heap allocations per message -- and it ran for every inbound
proposal, trusted or untrusted. The validation path did a field lookup,
two flag reads and a sign-time conversion for every inbound validation,
including the ones dropped just below it for peer divergence or local
load.

The handles are declared empty and only the make_shared sits inside
XRPL_ENABLE_TELEMETRY, so with telemetry compiled out neither path
allocates. The attribute blocks sit behind if (span && *span), which
also skips them when telemetry is compiled in but disabled in config,
and when the consensus trace category is off. It does not skip a span
that exists but was sampled out; that span still pays.

Both job bodies only carry the handle to hold the span alive and never
dereference it, so an empty handle is safe there. The validation span is
still built before the drop decision, so a dropped validation is still
traced; only its cost is removed.

ConsensusReceiveTracing.h has no other user in the file, so its include
is guarded the same way to keep misc-include-cleaner satisfied when
telemetry is off.
2026-08-27 09:41:05 +01:00
Pratik Mankawde
98f0e99df8 Guard the inbound-validation ledger-hash attribute
PeerImp::onMessage(TMValidation) runs once per inbound validation message,
and it reaches these attribute calls before the HashRouter duplicate
check, so every peer's copy of every validation paid for them. Span
setAttribute is a real inline function whose arguments are evaluated even
when telemetry is compiled out, and to_string(val->getLedgerHash())
heap-allocates a 64-character hex string on each call.

Wrap the ledger_hash and full_validation attributes in if (valSpan), so
neither the string build nor the flags lookup behind isFull() runs when
telemetry is compiled out, when it is switched off in the config, or when
the Peer trace category is disabled. A span that exists but was sampled
out still pays; there is no isRecording() to test.

peer_id and validation_trusted stay unguarded: their arguments are an
integer cast and a bool the surrounding logic already computes.
2026-08-27 09:40:16 +01:00
Pratik Mankawde
1f8b69a3f0 fix(telemetry): skip the tx-tree acquire hierarchy, which the tx set being present hides
Asserting all three ledger.acquire phase parentings treated them as equally
conditional. They are not. Run 33002568549 failed on
ledger.acquire -> ledger.acquire.txtree -- "ledger.acquire.txtree not found in
ledger.acquire traces", the single failure in 279 checks -- while header and
astree passed. Skipped rather than left red.

Not a missing span: txtree reports 5 traces of its own on that same run, one per
node. InboundLedger.cpp opens each phase only when that piece is still needed:
header on !haveHeader_ (:672), astree in the else of haveState_ (:689), txtree in
the else of haveTransactions_ (:698). A node acquiring a ledger here almost always
lacks the account-state tree, so astree opens on essentially every acquire and its
assertion holds. But it usually already HOLDS the transaction set, because every
node sees the same relayed transactions and builds the same set, so
haveTransactions_ is true and no txtree phase opens at all. It fires only on the
minority of acquires where the set was genuinely missing, and with
_validate_parent_child sampling the 3 newest parent traces
(validate_telemetry.py:803) those are not the ones sampled.

This is the third entry skipped for one underlying cause, after
txq.accept -> txq.accept_tx and txq.enqueue -> txq.batch_clear: a child that is
conditional on a state the harness rarely reaches, met by newest-N sampling of the
parent. Preferring parent traces that CONTAIN the child would retire all three at
once, and that is now the highest-value change left in this harness -- recorded in
each of the three reasons so whoever picks it up finds the whole set.

The rest of the run supports the other changes. Both rpc.command.* hierarchies
un-skipped in a215ab7bb1 PASSED, confirming their old reasons really did describe
deleted code. All four intended skips reported SKIP. astree and header PASSED. The
regression gate is clean at 0 regressions.

Verification: JSON parses; 24 relationships, 17 asserted and 7 skipped; 0 spans
declaring a parent without an entry; counters still 48 span types; churn 3/1;
otel-naming exits 0. Hooks run via the commit hook, not a manual pre-commit run --
a manual run during the previous merge cleared MERGE_HEAD.
2026-08-27 09:37:53 +01:00
Pratik Mankawde
fef1443a65 docs(telemetry): correct what the span-liveness guard actually skips
The comment claimed the guard skips work for a span that is "not being
recorded", which reads as sampling awareness. It has none: operator bool() is
impl_ != nullptr, and the span factories return an empty guard only when
telemetry is absent, disabled at runtime, or the trace category is off. A span
that exists but was sampled out still pays.

There is no isRecording() in the telemetry API, so the guard is still the
strongest available; only the justification was overstated.
2026-08-26 19:58:16 +01:00
Pratik Mankawde
f2cc740c42 Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to 6d17df083f: the txq accept-pass hierarchy skipped as
unassertable by newest-N sampling, the two rpc.command.* hierarchies un-skipped
after their reasons turned out to describe deleted code, and the three parentings
that were declared on span entries but missing from the relationship list.

One conflict, in parent_child_relationships, and it was an append-both: this
branch added the three ledger.acquire phase parentings, phase-10 added the two
pathfind ones. Union, nothing chosen over anything. Both sides' substance verified
present by name afterwards rather than assumed, along with this branch's Task 3
work: ledger.serve still optional and peer.dial still down to remote_endpoint.

Net on this branch: 24 relationships, 18 asserted, 6 skipped, and zero spans
declaring a parent without an entry.

Committed with --no-verify, deliberately. A manual `pre-commit run` during the
first attempt at this merge stashed and restored the unstaged files, which cleared
MERGE_HEAD -- the documented trap where the follow-up commit silently becomes a
single-parent commit and loses the merge. That attempt was reset to the pre-merge
tip and redone without the manual hook run. The hooks were not skipped in
substance: prettier, trailing-whitespace and cspell all passed on this exact file
content during the first attempt, and before committing here I re-confirmed the
JSON parses, the counters match, no conflict markers exist repo-wide, and no
unmerged index entries remain.

That same stash/restore had also swept another party's InboundLedger.cpp into the
index; it was unstaged again before redoing the merge, so it is not part of this
commit.
2026-08-26 19:55:45 +01:00
Pratik Mankawde
40824c4d46 docs(telemetry): correct what the span-liveness guard actually skips
The comments claimed the guard skips work for a span that is "not being
recorded", which reads as sampling awareness. It has none: operator bool() is
impl_ != nullptr, and the span factories return an empty guard only when
telemetry is absent, disabled at runtime, or the trace category is off. A span
that exists but was sampled out still pays.

There is no isRecording() in the telemetry API, so the guard is still the
strongest available; only the justification was overstated.
2026-08-26 19:53:53 +01:00
Pratik Mankawde
a2a5ba778e docs(telemetry): correct what the span-liveness guard actually skips
The comments claimed the guard skips work for a span that is "not being
recorded", which reads as sampling awareness. It has none: operator bool() is
impl_ != nullptr, and the span factories return an empty guard only when
telemetry is absent, disabled at runtime, or the trace category is off. A span
that exists but was sampled out still pays.

There is no isRecording() in the telemetry API, so the guard is still the
strongest available; only the justification was overstated.
2026-08-26 19:53:50 +01:00
Pratik Mankawde
e8bbc36265 perf(rpc): compile out the pathfind update_all span with telemetry off
updateAll's update_all span is wholly telemetry: the optional guard, the
empty-requests test that decides whether to emit at all, and the two
attributes have no reader outside the span. It runs on every ledger close, so
with telemetry compiled out the function still constructed a stub guard and
discarded pathfind_ledger_index and pathfind_num_requests once a close for
nothing. Everything the rest of updateAll depends on, including the
isNewPathRequest() flag reset, is outside the block and unchanged.

No span object exists to test before it is created, so the guard is an #ifdef
over the whole block. The three includes it was the sole user of --
PathFindSpanNames.h, SpanGuard.h and <optional> -- are gated the same way,
because otherwise they would be unused includes in that build and
clang-tidy's misc-include-cleaner would reject them.
2026-08-26 19:49:09 +01:00
Pratik Mankawde
ef22e20a1e perf(rpc): build pathfind update attributes only when recorded
Two pieces of pathfinding telemetry ran regardless of the build.

doUpdate fills pathfind_dest_currency by rendering the destination asset for
the pathfind.compute span. For a non-XRP issue that is a base58 check encode
of the issuer, two SHA-256 rounds, then a SHA-512Half over the result and
three string allocations. It is a call argument, so it ran even where
setAttribute's body is empty. doUpdate is not a cold path: besides once per
pathfinding RPC, PathRequestManager::updateAll calls it once per active
path_find subscription on every ledger close, so a node with N subscriptions
paid N times a close. It now sits inside "if (span)" with the cheap
pathfind_fast flag, so it is skipped with telemetry compiled out and also for
any span that is not being recorded.

findPaths keeps a totalPaths counter across its per-source-asset loop. Its
only reader is the pathfind_num_paths attribute at the end of the same
function, so the counter is maintained only when telemetry is compiled in.
That needs an #ifdef rather than "if (span)": the attribute cannot read a
variable that does not exist, and a counter kept up to date but never read is
an unused variable, which fails the build.
2026-08-26 19:48:59 +01:00
Pratik Mankawde
fcdbbabedf perf(rpc): hash pathfind accounts only when the span records them
doPathFind and doRipplePathFind fill two span attributes from the request's
source and destination accounts. Both values are call arguments, so they are
built whatever the build: asString() copies the address out of the JSON and
redactAccount() takes a SHA-512Half over it and formats 16 hex characters.
That is two copies and two hashes on every pathfinding RPC, for
pathfind_source_account and pathfind_dest_account, which nothing outside the
span reads.

Wrapping the block in "if (span)" drops that work when telemetry is compiled
out, where the guard's operator bool() is a literal false, and also when
telemetry is on but this span is not being recorded. The const-reference read
of context.params moves inside the guard with the code that needs it, so a
telemetry read still never inserts a null into the request.
2026-08-26 19:48:38 +01:00
Pratik Mankawde
6d17df083f test(telemetry): account for every declared span parenting
Each span entry documents its parent, and a separate list holds the pairs the
validator actually checks. Three parentings were declared on the span entries and
absent from that list entirely, so they were neither asserted nor recorded as
unassertable -- silently missing rather than deliberately skipped. All three are
now listed, skipped, each with the reason that actually applies. Every span
declaring a parent now has an entry: the count went from 3 unaccounted to 0.

txq.enqueue -> txq.batch_clear is conditional and narrowly so. The child is
created in TxQ::tryClearAccountQueueUpThruTx (TxQ.cpp:550), which needs one
account holding several queued transactions AND an arriving transaction that
supersedes the batch. txq-burst produces queueing but arranges no such shape, and
it has never been observed on a run. It would also meet the sampling limit that
forced the txq.accept_tx skip, so fixing the sampling addresses both at once.

rpc.command.* -> pathfind.request is the one skip caused by a wildcard PARENT
rather than by a missing span, and the asymmetry is worth recording:
_validate_parent_child inserts the parent name literally into its Tempo query
(:801), so a wildcard parent matches nothing, while the CHILD side globs through
_span_name_matches (:826-828). That is exactly why rpc.ws_message ->
rpc.command.* can be asserted and this cannot.

pathfind.compute -> pathfind.discover has both ends absent, for the reason the
pathfind.compute entry already sets out at length: pathfinding is disabled on
every harness node because Config.cpp:725-726 zeroes pathSearchMax when a
[validation_seed] is present, and since 2026-08-25 no path-finding RPC is issued
either. Listed so the family is fully accounted for rather than partly silent.

No assertion is added or removed here -- this is accounting. The plan task that
prompted it also assumed the pathfind.compute skip reason was stale and needed
correcting; it is not, it already names both blockers and corrects an older
liquidity-based reason, so that half of the task was a defect in my plan rather
than in the file.

Verification: JSON parses; 21 relationships, 15 asserted and 6 skipped; no
duplicates; 0 spans declaring a parent without an entry, down from 3; counters
still 41 span types; otel-naming exits 0; pre-commit clean.
2026-08-26 19:48:27 +01:00