Commit Graph

15619 Commits

Author SHA1 Message Date
Pratik Mankawde
23c1a6f5dd fix(docs): spell it "XRPL epoch" in the data-collection reference
The rename script rewrites "Ripple epoch" to "XRPL epoch", so the old
spelling in a tracked .md makes the check-rename job fail on a dirty tree.

The attribute keys are left alone: the script's pattern needs a space, and
those keys are a cross-layer contract.
2026-09-08 17:12:55 +01:00
Pratik Mankawde
30d165da43 fix(telemetry): bound integration-test assertions to the run under test
check_statsd_metric queried rippled_rpc_requests, which no pipeline
produces: the collector's statsd receiver runs with is_monotonic_counter,
so the Prometheus exporter appends _total. A wrong name returns zero
series rather than an error, so the assertion could not be told apart
from a broken pipeline. All eight assertions were re-derived from how
each metric is created in code; this was the only counter.

Tempo searches carried no start/end, and tempo-data is a named volume
that `docker compose down` preserves under a one-hour block retention, so
the 17 span assertions could pass on an earlier local run's traces. Bound
every search to this run, and tear the stack down with -v before starting
so no earlier data is present to match. The service-name check now
matches a whole line, because the tag-values endpoint ignores start/end.

Add a gtest for the StatsD gauge that publishes its initial zero and for
the counter that must publish nothing. Assert two metrics the harness
never checked: a traffic-category gauge no message reaches, and
io_context latency.
2026-09-08 16:43:41 +01:00
Pratik Mankawde
9977810c6d merge: bring the review fixes forward from phase5-docs-deployment
One conflict, in docs/telemetry-runbook.md: both sides had independently
corrected the same consensus_round_id example. This branch kept the pipe form,
which Tempo rejects as a parse error; upstream moved the predicate inside the
braces, which parses and returns data. Upstream's query is kept, with this
branch's note that the value is the previous ledger sequence plus one.
2026-09-08 15:33:02 +01:00
Pratik Mankawde
868edce7e5 merge: bring the review fixes forward from phase4-consensus-tracing
Two conflicts, both additive.

TelemetryConfig.cpp: this branch added requireHttpsEndpoint next to
requireReadableFile; upstream added readConsensusTraceStrategy at the same spot.
Both kept.

05-configuration-reference.md: this branch added the two client-certificate rows
while upstream corrected the consensus strategy value from attribute to random.
Both kept. Also drops the stale "not yet implemented" row for
consensus_trace_strategy, which the merged table now contradicts twice over: the
option is parsed, and its value is no longer spelled attribute.
2026-09-08 15:31:40 +01:00
Pratik Mankawde
0c1babf9fb merge: bring the review fixes forward from otel-phase3-tx-tracing 2026-09-08 15:28:59 +01:00
Pratik Mankawde
7b302f5e77 merge: bring the review fixes forward from otel-phase2-rpc-tracing 2026-09-08 15:28:59 +01:00
Pratik Mankawde
a6dc9edf08 merge: bring the review fixes forward from otel-phase1c-rpc-integration 2026-09-08 15:28:59 +01:00
Pratik Mankawde
e417a4d434 merge: bring the review fixes forward from otel-phase1b-telemetry-infra 2026-09-08 15:28:59 +01:00
Pratik Mankawde
566af67bb8 merge: bring the review fixes forward from otel-phase1a-plan-docs 2026-09-08 15:28:59 +01:00
Pratik Mankawde
2399f5763f docs(telemetry): name the consensus trace strategy value "random"
The plan doc offered `"attribute"` as the alternative to `"deterministic"`
for consensus_trace_strategy. The parser accepts `"random"`; "attribute"
described the correlation mechanism rather than the setting's value. Note
also that the alternative is experimental and not used.
2026-09-08 14:42:32 +01:00
Pratik Mankawde
8da882dccb docs(telemetry): state when retries_remaining is recorded
retries_remaining is stamped on the txq.accept_tx span before the
transaction is applied and before the retry counter is decremented, so a
span with txq_status="retried" always shows a non-zero count and exhaustion
shows up as txq_status="failed" with zero. The attribute comment said only
"retries left before discard", which reads as a post-decrement value and
led to a runbook query that could never match.

Also rename the drifted consensus_trace_strategy value in the plan docs
from "attribute" to "random", the spelling the parser accepts, and note that
it is experimental and not used.
2026-09-08 14:42:05 +01:00
Pratik Mankawde
039c2768ba fix(telemetry): require an https endpoint when a client certificate is set
The OTLP/HTTP exporter selects TLS from the endpoint URL scheme alone
(HttpSslOptions in the pinned SDK matches "https:" exactly), so a client
certificate handed to it alongside an http:// traces_endpoint is loaded and
never presented. The parser checked the cert/key pairing, use_tls and file
readability, but never the scheme, and the default traces_endpoint is plain
HTTP. makeTelemetrySetup() now requires traces_endpoint to start with
"https://" whenever tls_client_cert is set, including when the key is left
at its default.

Nothing asserted the client options reaching the exporter, so a swapped
certificate and key would have passed every test. Move the options mapping
into makeTraceExporterOptions() and assert it at that boundary with
distinct certificate and key paths, plus a one-way-TLS control and a
use_tls=0 control. One case runs the whole path from a [telemetry] section.

Runbook and example-config fixes:

- tx.included is emitted per transaction of the agreed consensus set,
  before buildLCL() applies anything, so it is a superset of the accepted
  ledger rather than proof of inclusion.
- the dispute.resolve query used the descendant operator, but the event is
  on the consensus.update_positions span itself, so it matched nothing.
- the exhausted-retries query asked for txq_status="retried" with
  retries_remaining=0, which cannot occur: the attribute is stamped before
  the attempt and the retried branch only runs while retries are left.
  Exhaustion is txq_status="failed" with a zero count.
- consensus_round_id is an int64, so the two queries comparing it to a
  quoted string matched nothing.
- note that consensus_trace_strategy=random is experimental and not used.
- note that a trailing "| attr = value" is rejected by current Tempo;
  attribute filters belong inside the braces.
2026-09-08 14:41:39 +01:00
Pratik Mankawde
fb827dc0f1 fix(telemetry): make the consensus trace strategy an enum
consensus_trace_strategy was read as a std::string and compared against the
literal "attribute" in startRoundTracing(), while the runbook documented
"deterministic" and "random". The documented value "random" therefore fell
through to the default and did nothing.

Parse the setting once into ConsensusTraceStrategy, so the consensus code
branches on a type. The accepted spellings are now "deterministic" and
"random"; anything else fails at startup instead of silently defaulting.
The behaviour behind the old "attribute" name is unchanged and is now
reached by "random".

Document consensus_trace_strategy in xrpld-example.cfg, stating that
"random" is experimental and not used: it gives each node its own trace id,
so one round arrives as one trace per node.

Also state on the tx.included event that it covers the agreed consensus set
before the ledger is built, so it is a superset of the accepted ledger.
2026-09-08 14:39:05 +01:00
Pratik Mankawde
de61119e5b merge: bring the review fixes forward from otel-phase5-docs-deployment 2026-09-07 15:15:05 +01:00
Pratik Mankawde
de895c6d1e merge: bring the review fixes forward from phase4-consensus-tracing
Two conflicts, both additive on each side.

TelemetryConfig.cpp: include blocks only. This branch added FileUtilities.h for
the certificate readability checks; upstream added <limits> and <optional> for
the bounds parser. Both kept.

The TelemetryConfig test: this branch's mutual-TLS cases and upstream's
batch-bounds cases were added at the same positions, so the file is rebuilt from
both stages and carries all 32 tests. Two shared cases were each edited by one
side only, so the edited side wins in each: upstream asserts the batch defaults
in parse_empty_section, and this branch's parse_full_section writes a real
certificate file, which is now required since the parser opens it.
2026-09-07 15:14:52 +01:00
Pratik Mankawde
3d9ea4b9da merge: bring the review fixes forward from phase3-tx-tracing
One conflict, in 06-implementation-phases.md: this branch had rewritten the
phase-4 task table with a Status column, a descoping note and a Spans Produced
section, while upstream corrected the class name in the old plain table. This
branch's section is kept and the name correction re-applied to its 4.1 row.
2026-09-07 15:04:39 +01:00
Pratik Mankawde
08df97f431 merge: bring the review fixes forward from phase2-rpc-tracing
One conflict, in SpanGuard.h: this branch added struct TraceBytes and upstream
added enum SpanRole at the same position after TraceCategory. Unrelated
declarations, so both are kept.
2026-09-07 15:03:44 +01:00
Pratik Mankawde
161d28b6dd test(telemetry): cover the [telemetry] batch-setting bounds
Guards the validation the parser gained upstream: zero rejected for all three
keys, a non-numeric value raising std::runtime_error rather than leaking
boost::bad_lexical_cast, a negative value rejected instead of wrapping to
4294967295, both bounds accepted exactly, and batch_size held at or below
max_queue_size.

The catch is std::runtime_error, not std::exception, on purpose: if the parser
ever stops wrapping, a bad_cast escapes and the suite fails loudly instead of
swallowing it.
2026-09-07 15:02:48 +01:00
Pratik Mankawde
caa704de10 merge: bring the review fixes forward from phase1c-rpc-integration
Two conflicts, both resolved by composing the sides rather than taking one.

cfg/xrpld-example.cfg: this branch had moved the batch-processor keys under
their own heading while upstream edited them in place, so a merge-both would
have documented them twice. Upstream's range sentences are applied to the
relocated block and the head-sampling note keeps its position.

02-design-decisions.md: the summary table changed on both sides for different
reasons. Upstream renamed ledger_index to current_ledger_seq and ledger_seq;
this branch had corrected the PathFinding row to the keys it actually emits.
Both are kept.
2026-09-07 15:00:53 +01:00
Pratik Mankawde
72372c42db fix(telemetry): mark the internal RPC spans Internal rather than Server
The category mapped every Rpc span to kServer, so one inbound request emitted
several nested server spans. Per the trace spec, SERVER covers server-side
handling of a remote request the client awaits, while INTERNAL is an operation
with a local parent. rpc.process and both rpc.command sites have a local parent,
so they now pass SpanRole::Internal.

The four transport-edge roots keep the category default: rpc.http_request,
rpc.ws_upgrade, rpc.ws_message and the gRPC span each begin a remote call. This
matters to Tempo's service-graph and span-metrics generators, which pair server
spans with client spans and leave a surplus one unpaired.
2026-09-07 14:58:31 +01:00
Pratik Mankawde
a8d678f354 merge: bring the childSpan pseudocode fix forward from phase1b-telemetry-infra 2026-09-07 14:56:32 +01:00
Pratik Mankawde
0f49aecbf0 docs(telemetry): pass a full dotted constant in the childSpan pseudocode
childSpan() takes the span name verbatim, so a bare op:: suffix names the span
"process" rather than "rpc.process". The same defect was corrected in the
SpanGuard and Telemetry examples; this is the last copy.
2026-09-07 14:56:29 +01:00
Pratik Mankawde
a6c24848ba merge: bring the review fixes forward from phase1b-telemetry-infra 2026-09-07 14:55:22 +01:00
Pratik Mankawde
83145db060 merge: bring the review fixes forward from phase1a-plan-docs 2026-09-07 14:55:11 +01:00
Pratik Mankawde
b89a83f8f8 fix(telemetry): stamp the round span's mode when the engine applies it
startRoundTracing runs as an argument to Consensus::startRound, so it creates
consensus.round before startRoundInternal applies the new mode. Reading mode_
there recorded the previous round's value, and a validator switching from
observing to proposing got a round span labelled observing that nothing corrected.

The attribute is now written in onModeChange, from the mode being applied. All
three MonitoredMode::set paths funnel through there, so round start, a wrong-ledger
switch and a bow-out all correct the parent span with one statement. Every path
reaches it under RCLConsensus::mutex_ on the thread that created the span.

The stale write is removed rather than kept alongside: neither Consensus::startRound
nor startRoundInternal has an early return before mode_.set, so every round span is
stamped. If a future path ever skipped it the attribute would be absent, which reads
as a gap, instead of confidently wrong. onClose also sets consensus_mode, from the
engine's own parameter, and is correct as it stands.

Also tests addEvent's attribute overload on a live span, reading the exported event
name and each value back off the in-memory exporter. It was previously only ever
called on a null guard, so a dropped attribute exported nothing and failed nothing.
2026-09-07 14:39:50 +01:00
Pratik Mankawde
b5415233cd fix(telemetry): stop reporting applied_direct for a failed direct apply
tryDirectApply returns an engaged optional whenever the fee bar was cleared,
including when xrpl::apply() failed, so testing the optional labelled failures as
applied. TxQ_test's fail-in-preclaim case hits exactly this: the fee clears the
bar and preclaim then rejects with terINSUF_FEE_B.

The stamp now branches on ApplyResult::applied. A failure reports failed rather
than falling through to the default rejected, because rejected means the
transaction got nowhere, while this one cleared the fee bar and ran through
apply(). ter_code is recorded either way, so the failure is diagnosable. Both
values already existed and are used the same way by the queued-apply path in this
file, so the vocabulary is unchanged.

The value set in the phase-3 task list is updated to match, including a ter_code
row for txq.accept_tx that was already emitted but undocumented.
2026-09-07 13:42:29 +01:00
Pratik Mankawde
18abd100b5 feat(telemetry): let a call site choose a span's role, not just its category
Span kind was derived from TraceCategory alone, so every Rpc-category span was
kServer. A category cannot tell an inbound handler from the internal work under
it, and trace backends pair kServer with kClient, so internal spans left as
kServer become unpaired edges in a service graph and read as extra inbound
requests.

SpanRole is a new xrpl-owned enum, orthogonal to TraceCategory: the category
names the subsystem and gates the span on config, the role says whether the span
handles a remote call. It is a defaulted fourth parameter on span(), freshRoot()
and the ScopedSpanGuard equivalents, defaulting to SpanRole::FromCategory, so no
existing call site changes. resolveSpanKind() applies an explicit role and falls
back to the category map, which keeps its single responsibility. The
telemetry-disabled stubs mirror all four signatures.

No call site passes a role yet. The two that need it are on a later branch.

Also fixes a ScopedSpanGuard example that passed a bare op:: suffix to
childSpan(), which takes the name verbatim. Naming the child rpc.command made it
a child of rpc.command.<cmd>, inverting the hierarchy, so the example's parent is
now rpc.process and the command attribute moved onto the command span.
2026-09-07 13:39:57 +01:00
Pratik Mankawde
f8e0a19b9f fix(telemetry): correct the childSpan doc examples and make Rule D tests real
Review feedback on the RPC integration PR.

The childSpan examples could not work as written. childSpan() takes its parent
from the ambient context and uses impl_ only as a liveness gate, so an unscoped
SpanGuard parent produced two siblings rather than a parent and child. The parent
is now a ScopedSpanGuard, the child no longer reuses the parent's name, and the
examples pass a full dotted constant because childSpan() takes the name verbatim.

Five of the ten Rule D tests could not fail. Four passed an empty L1 key set,
which makes the rule skip validation altogether; the fifth asserted an empty
result against an escaped-quote selector that extracted no labels at all. Each
now passes a nonempty L1 set and carries a known-bad label in the same
expression, so it asserts both that the intended labels are accepted and that
Rule D ran. Verified by disabling the rule: the old tests stay green, the new
ones all fail.

Span kind is not fixed here. categoryToSpanKind and the span factories belong to
the telemetry library, so the role parameter is routed to that branch, and the
two call sites here follow once it exists.
2026-09-07 13:25:22 +01:00
Pratik Mankawde
4d2841ccda fix(telemetry): reject invalid [telemetry] batch settings and make isValid() honest
Two review findings on the telemetry library.

SpanContext::isValid() returned impl_ != nullptr, so it answered true for a
context holding no span. threadLocalContext() wraps whatever GetCurrent()
returns, and that is an empty Context on a thread with no active span, which
contradicted the documented "invalid context if none is active". It now asks the
Context for its span. childSpan(name, ctx) is the one caller whose behaviour
changes: a context with no span used to produce a new root span, and now returns
a null guard as its @return already promised.

The three batch settings went to the OTel BatchSpanProcessor unchecked. Three
ways that failed: zero was accepted for all of them; batch_size could exceed
max_queue_size, which the SDK documents as a precondition and does not enforce;
and a mistyped value let boost::bad_lexical_cast escape, which derives from
std::bad_cast rather than std::runtime_error, so the operator saw a bare "bad
cast" naming no key. Reading unsigned also turned "-1" into 4294967295 instead
of failing, so the value is parsed signed and negatives are rejected.
xrpld-example.cfg now states the ranges.
2026-09-07 13:16:59 +01:00
Pratik Mankawde
9aebdc292c docs(telemetry): fix stale symbols, attribute keys and TraceQL in the plan docs
Review feedback on the plan documents. Four kinds of error:

- Symbols that do not exist: ConsensusProposal::prevLedger_ (it is
  previousLedger_), RCLConsensusAdaptor (it is RCLConsensus::Adaptor, and
  startRound() is on RCLConsensus itself), and RPCHandler::doCommand (a free
  function, xrpl::rpc::doCommand).
- Attribute keys: the tables used ledger_index, which no telemetry code emits.
  Same concept as ledger_seq but a different referent, so the code disambiguates
  by prefix: current_ledger_seq for the open ledger a transaction targeted,
  ledger_seq for a closed or validated one. A note now states which is which.
- TraceQL that does not parse: span-field predicates need braces, status.code
  is not an intrinsic (status = error), and avg(duration) does not take a by
  clause (avg_over_time does). All five re-tested against Tempo.
- The StatsD comparison omitted the Histogram instrument, which aggregates at
  the point of measure, and the when-to-use table had no row for a metric that
  spans cannot afford to carry.
2026-09-07 13:16:32 +01:00
Pratik Mankawde
981323f071 merge: bring the close-time attr doc fixes forward from phase5-docs-deployment 2026-09-04 12:40:21 +01:00
Pratik Mankawde
652c03f5cd merge: bring the close-time attr doc fixes forward from phase4-consensus-tracing 2026-09-04 12:40:21 +01:00
Pratik Mankawde
68364f9b8a merge: bring the close-time attr doc fixes forward from phase1c-rpc-integration 2026-09-04 12:40:21 +01:00
Pratik Mankawde
f2ef748b6b merge: bring the close-time attr doc fixes forward from phase3-tx-tracing 2026-09-04 12:40:21 +01:00
Pratik Mankawde
f2b3e5c53c merge: bring the close-time attr doc fixes forward from phase2-rpc-tracing 2026-09-04 12:40:21 +01:00
Pratik Mankawde
b91ab6c5b9 merge: bring the close-time attr doc fixes forward from phase1b-telemetry-infra 2026-09-04 12:40:21 +01:00
Pratik Mankawde
ff3f41eeaf merge: bring the close-time attr doc fixes forward from phase1a-plan-docs 2026-09-04 12:40:20 +01:00
Pratik Mankawde
ca22f57919 docs(telemetry): follow the close-time attr rename in the data reference
The emitted keys carry the unit and epoch suffix. Update the consensus
and ledger attribute tables and the ledger.build span row to match.

The Close Time Drift panel row is left alone: phase-7 removes that whole
table, so editing it here would only conflict on the way forward.
2026-09-04 12:37:19 +01:00
Pratik Mankawde
1afa35d54d docs(telemetry): follow the close-time attr rename in the runbook
The consensus.accept.apply row listed close_time, parent_close_time and
close_time_self. The emitted keys carry the unit and epoch suffix, so
update the row to match.
2026-09-04 12:35:59 +01:00
Pratik Mankawde
372fba129f docs(telemetry): follow the close-time attr rename in the phase-4 docs
The emitted keys are close_time_ripple_epoch_s,
parent_close_time_ripple_epoch_s and close_time_self_ripple_epoch_s.
Update the consensus.accept.apply attribute tables and the close-time
attribute descriptions to match.
2026-09-04 12:35:25 +01:00
Pratik Mankawde
8e15a81f96 docs(telemetry): follow the close-time attr rename in the ledger attr table
The emitted key is close_time_ripple_epoch_s, which names its unit and
epoch. Update the ledger attribute table to match.

ledger_index and ledger_tx_count in the same table belong to a separate
rename and are left as they are.
2026-09-04 12:34:07 +01:00
Pratik Mankawde
911d7610ef merge: bring telemetry config and doc fixes forward from phase5-docs-deployment 2026-09-03 20:29:29 +01:00
Pratik Mankawde
5381bf722d merge: bring telemetry config and doc fixes forward from phase4-consensus-tracing 2026-09-03 20:29:28 +01:00
Pratik Mankawde
7ba73bf819 merge: bring telemetry config and doc fixes forward from phase3-tx-tracing 2026-09-03 20:29:28 +01:00
Pratik Mankawde
1f97b772b3 merge: bring telemetry config and doc fixes forward from phase2-rpc-tracing 2026-09-03 20:29:13 +01:00
Pratik Mankawde
ea9f778638 merge: bring telemetry config and doc fixes forward from phase1c-rpc-integration
Both sides documented the same six [telemetry] keys, so the automatic merge
duplicated all of them. Resolved by keeping this branch's structure - which
already covers all 14 keys and groups them under TLS and batch-processor
headings - and folding in the corrections from the upstream side:

- endpoint is renamed to traces_endpoint, which is what the parser reads, and
  described as used verbatim including its signal path.
- use_tls no longer claims to enable TLS. The exporter's URL scheme selects
  TLS; this key only decides whether tls_ca_cert reaches it as a CA bundle.
- tls_ca_cert records that the path is not opened while the config is parsed,
  so an unreadable file shows up as an export failure rather than at startup.
- service_instance_id explains that it is normally left unset and filled in
  from the node public key during startup.

Section::value_or in 05-configuration-reference.md becomes Section::valueOr;
that member does not exist under the other spelling.
2026-09-03 20:26:40 +01:00
Pratik Mankawde
c018113c12 merge: bring telemetry config and doc fixes forward from phase1b-telemetry-infra 2026-09-03 20:21:25 +01:00
Pratik Mankawde
9f55acef38 docs(telemetry): name the label the consensus dashboard actually filters on
The Consensus Health template-variable table documented $node as resolving via
exported_instance. That dashboard defines $node as
label_values(target_info, service_instance_id) and its panels filter on
service_instance_id; exported_instance appears in it zero times.

exported_instance is a real label, but it belongs to the StatsD boards shipped
alongside, where Prometheus renames a scraped instance label that collides with
the target's own. Documenting it against an OTel dashboard pointed readers at
the wrong pipeline's label.
2026-09-03 20:19:27 +01:00
Pratik Mankawde
6aedee785a docs(telemetry): document every [telemetry] key and correct stale parser names
The commented [telemetry] block in cfg/xrpld-example.cfg documented 8 of the 14
keys the parser accepts. Add the six that were missing - service_instance_id,
use_tls, tls_ca_cert, batch_size, batch_delay_ms and max_queue_size - each with
the unit and default read from the parser, and rename the documented endpoint
key to traces_endpoint so it matches what makeTelemetrySetup reads.

use_tls is documented for what it does rather than what its name suggests: it
gates whether tls_ca_cert reaches the exporter as a CA bundle, while the scheme
of traces_endpoint is what selects TLS. The path is not opened during parsing,
so an unreadable file surfaces as an export failure at runtime.

05-configuration-reference.md named three symbols that do not exist:
setup_Telemetry, make_Telemetry and Section::value_or. Correct them to
makeTelemetrySetup, makeTelemetry and Section::valueOr.
2026-09-03 20:19:11 +01:00
Pratik Mankawde
61fc616bf8 merge: bring develop forward from phase5-docs-deployment 2026-09-03 17:33:52 +01:00