Commit Graph

16118 Commits

Author SHA1 Message Date
Pratik Mankawde
547cc8bac5 merge: bring the review fixes forward from otel-phase7-native-metrics 2026-09-08 15:46:05 +01:00
Pratik Mankawde
7d679ce596 merge: bring the review fixes forward from otel-phase6-statsd 2026-09-08 15:45:22 +01:00
Pratik Mankawde
9977810c6d merge: bring the review fixes forward from phase5-docs-deployment
One conflict, in docs/telemetry-runbook.md: both sides had independently
corrected the same consensus_round_id example. This branch kept the pipe form,
which Tempo rejects as a parse error; upstream moved the predicate inside the
braces, which parses and returns data. Upstream's query is kept, with this
branch's note that the value is the previous ledger sequence plus one.
2026-09-08 15:33:02 +01:00
Pratik Mankawde
868edce7e5 merge: bring the review fixes forward from phase4-consensus-tracing
Two conflicts, both additive.

TelemetryConfig.cpp: this branch added requireHttpsEndpoint next to
requireReadableFile; upstream added readConsensusTraceStrategy at the same spot.
Both kept.

05-configuration-reference.md: this branch added the two client-certificate rows
while upstream corrected the consensus strategy value from attribute to random.
Both kept. Also drops the stale "not yet implemented" row for
consensus_trace_strategy, which the merged table now contradicts twice over: the
option is parsed, and its value is no longer spelled attribute.
2026-09-08 15:31:40 +01:00
Pratik Mankawde
0c1babf9fb merge: bring the review fixes forward from otel-phase3-tx-tracing 2026-09-08 15:28:59 +01:00
Pratik Mankawde
7b302f5e77 merge: bring the review fixes forward from otel-phase2-rpc-tracing 2026-09-08 15:28:59 +01:00
Pratik Mankawde
a6dc9edf08 merge: bring the review fixes forward from otel-phase1c-rpc-integration 2026-09-08 15:28:59 +01:00
Pratik Mankawde
e417a4d434 merge: bring the review fixes forward from otel-phase1b-telemetry-infra 2026-09-08 15:28:59 +01:00
Pratik Mankawde
566af67bb8 merge: bring the review fixes forward from otel-phase1a-plan-docs 2026-09-08 15:28:59 +01:00
Pratik Mankawde
2399f5763f docs(telemetry): name the consensus trace strategy value "random"
The plan doc offered `"attribute"` as the alternative to `"deterministic"`
for consensus_trace_strategy. The parser accepts `"random"`; "attribute"
described the correlation mechanism rather than the setting's value. Note
also that the alternative is experimental and not used.
2026-09-08 14:42:32 +01:00
Pratik Mankawde
8da882dccb docs(telemetry): state when retries_remaining is recorded
retries_remaining is stamped on the txq.accept_tx span before the
transaction is applied and before the retry counter is decremented, so a
span with txq_status="retried" always shows a non-zero count and exhaustion
shows up as txq_status="failed" with zero. The attribute comment said only
"retries left before discard", which reads as a post-decrement value and
led to a runbook query that could never match.

Also rename the drifted consensus_trace_strategy value in the plan docs
from "attribute" to "random", the spelling the parser accepts, and note that
it is experimental and not used.
2026-09-08 14:42:05 +01:00
Pratik Mankawde
039c2768ba fix(telemetry): require an https endpoint when a client certificate is set
The OTLP/HTTP exporter selects TLS from the endpoint URL scheme alone
(HttpSslOptions in the pinned SDK matches "https:" exactly), so a client
certificate handed to it alongside an http:// traces_endpoint is loaded and
never presented. The parser checked the cert/key pairing, use_tls and file
readability, but never the scheme, and the default traces_endpoint is plain
HTTP. makeTelemetrySetup() now requires traces_endpoint to start with
"https://" whenever tls_client_cert is set, including when the key is left
at its default.

Nothing asserted the client options reaching the exporter, so a swapped
certificate and key would have passed every test. Move the options mapping
into makeTraceExporterOptions() and assert it at that boundary with
distinct certificate and key paths, plus a one-way-TLS control and a
use_tls=0 control. One case runs the whole path from a [telemetry] section.

Runbook and example-config fixes:

- tx.included is emitted per transaction of the agreed consensus set,
  before buildLCL() applies anything, so it is a superset of the accepted
  ledger rather than proof of inclusion.
- the dispute.resolve query used the descendant operator, but the event is
  on the consensus.update_positions span itself, so it matched nothing.
- the exhausted-retries query asked for txq_status="retried" with
  retries_remaining=0, which cannot occur: the attribute is stamped before
  the attempt and the retried branch only runs while retries are left.
  Exhaustion is txq_status="failed" with a zero count.
- consensus_round_id is an int64, so the two queries comparing it to a
  quoted string matched nothing.
- note that consensus_trace_strategy=random is experimental and not used.
- note that a trailing "| attr = value" is rejected by current Tempo;
  attribute filters belong inside the braces.
2026-09-08 14:41:39 +01:00
Pratik Mankawde
fb827dc0f1 fix(telemetry): make the consensus trace strategy an enum
consensus_trace_strategy was read as a std::string and compared against the
literal "attribute" in startRoundTracing(), while the runbook documented
"deterministic" and "random". The documented value "random" therefore fell
through to the default and did nothing.

Parse the setting once into ConsensusTraceStrategy, so the consensus code
branches on a type. The accepted spellings are now "deterministic" and
"random"; anything else fails at startup instead of silently defaulting.
The behaviour behind the old "attribute" name is unchanged and is now
reached by "random".

Document consensus_trace_strategy in xrpld-example.cfg, stating that
"random" is experimental and not used: it gives each node its own trace id,
so one round arrives as one trace per node.

Also state on the tx.included event that it covers the agreed consensus set
before the ledger is built, so it is a superset of the accepted ledger.
2026-09-08 14:39:05 +01:00
Pratik Mankawde
efa6535427 merge: bring the review fixes forward from otel-phase7-native-metrics 2026-09-07 15:24:42 +01:00
Pratik Mankawde
4ad84cae14 merge: bring the review fixes forward from otel-phase6-statsd 2026-09-07 15:23:22 +01:00
Pratik Mankawde
de61119e5b merge: bring the review fixes forward from otel-phase5-docs-deployment 2026-09-07 15:15:05 +01:00
Pratik Mankawde
de895c6d1e merge: bring the review fixes forward from phase4-consensus-tracing
Two conflicts, both additive on each side.

TelemetryConfig.cpp: include blocks only. This branch added FileUtilities.h for
the certificate readability checks; upstream added <limits> and <optional> for
the bounds parser. Both kept.

The TelemetryConfig test: this branch's mutual-TLS cases and upstream's
batch-bounds cases were added at the same positions, so the file is rebuilt from
both stages and carries all 32 tests. Two shared cases were each edited by one
side only, so the edited side wins in each: upstream asserts the batch defaults
in parse_empty_section, and this branch's parse_full_section writes a real
certificate file, which is now required since the parser opens it.
2026-09-07 15:14:52 +01:00
Pratik Mankawde
3d9ea4b9da merge: bring the review fixes forward from phase3-tx-tracing
One conflict, in 06-implementation-phases.md: this branch had rewritten the
phase-4 task table with a Status column, a descoping note and a Spans Produced
section, while upstream corrected the class name in the old plain table. This
branch's section is kept and the name correction re-applied to its 4.1 row.
2026-09-07 15:04:39 +01:00
Pratik Mankawde
08df97f431 merge: bring the review fixes forward from phase2-rpc-tracing
One conflict, in SpanGuard.h: this branch added struct TraceBytes and upstream
added enum SpanRole at the same position after TraceCategory. Unrelated
declarations, so both are kept.
2026-09-07 15:03:44 +01:00
Pratik Mankawde
161d28b6dd test(telemetry): cover the [telemetry] batch-setting bounds
Guards the validation the parser gained upstream: zero rejected for all three
keys, a non-numeric value raising std::runtime_error rather than leaking
boost::bad_lexical_cast, a negative value rejected instead of wrapping to
4294967295, both bounds accepted exactly, and batch_size held at or below
max_queue_size.

The catch is std::runtime_error, not std::exception, on purpose: if the parser
ever stops wrapping, a bad_cast escapes and the suite fails loudly instead of
swallowing it.
2026-09-07 15:02:48 +01:00
Pratik Mankawde
caa704de10 merge: bring the review fixes forward from phase1c-rpc-integration
Two conflicts, both resolved by composing the sides rather than taking one.

cfg/xrpld-example.cfg: this branch had moved the batch-processor keys under
their own heading while upstream edited them in place, so a merge-both would
have documented them twice. Upstream's range sentences are applied to the
relocated block and the head-sampling note keeps its position.

02-design-decisions.md: the summary table changed on both sides for different
reasons. Upstream renamed ledger_index to current_ledger_seq and ledger_seq;
this branch had corrected the PathFinding row to the keys it actually emits.
Both are kept.
2026-09-07 15:00:53 +01:00
Pratik Mankawde
72372c42db fix(telemetry): mark the internal RPC spans Internal rather than Server
The category mapped every Rpc span to kServer, so one inbound request emitted
several nested server spans. Per the trace spec, SERVER covers server-side
handling of a remote request the client awaits, while INTERNAL is an operation
with a local parent. rpc.process and both rpc.command sites have a local parent,
so they now pass SpanRole::Internal.

The four transport-edge roots keep the category default: rpc.http_request,
rpc.ws_upgrade, rpc.ws_message and the gRPC span each begin a remote call. This
matters to Tempo's service-graph and span-metrics generators, which pair server
spans with client spans and leave a surplus one unpaired.
2026-09-07 14:58:31 +01:00
Pratik Mankawde
a8d678f354 merge: bring the childSpan pseudocode fix forward from phase1b-telemetry-infra 2026-09-07 14:56:32 +01:00
Pratik Mankawde
0f49aecbf0 docs(telemetry): pass a full dotted constant in the childSpan pseudocode
childSpan() takes the span name verbatim, so a bare op:: suffix names the span
"process" rather than "rpc.process". The same defect was corrected in the
SpanGuard and Telemetry examples; this is the last copy.
2026-09-07 14:56:29 +01:00
Pratik Mankawde
a6c24848ba merge: bring the review fixes forward from phase1b-telemetry-infra 2026-09-07 14:55:22 +01:00
Pratik Mankawde
83145db060 merge: bring the review fixes forward from phase1a-plan-docs 2026-09-07 14:55:11 +01:00
Pratik Mankawde
0128feb70f fix(telemetry): count each validated ledger once in ValidationTracker
Both record entry points reached their record with pending_[ledgerHash], and
operator[] default-constructs on a missing key. An all-zero recordTime was the
only signal for "first time seeing this ledger", so it could not tell a new
ledger apart from one already counted and evicted.

Reconciled entries are removed on purpose once they leave the late-repair
window, and reconcile() adds a ledger to the totals only on its first
reconcile. A validation arriving after its entry was counted and removed
therefore rebuilt a blank record, looked new, and was counted a second time -
in the lifetime totals and in all three rolling windows, so the agreement
percentages skewed too. A re-arrival carrying only one side manufactured a miss
for a ledger first counted as an agreement.

Both entry points now go through pendingEvent(), which looks in pending_
first so a ledger still awaiting late repair keeps updating its record, then
checks tallied_ and records nothing for a ledger already counted. That retires
the all-zero sentinel, which was the defect itself.

tallied_ holds hashes only, since the record is gone by then, and is bounded at
kMaxTalliedEvents with oldest-first eviction to match evictOldPending. A
validation for a ledger counted more than that many ledgers ago is counted
again; the comment says so.

The message counters keep incrementing on every call, including for an ignored
ledger. They count messages rather than ledgers.

kMaxPendingEvents becomes public so the new test sizes its fill from the
production constant instead of copying the number. Its value and meaning are
unchanged.
2026-09-07 14:50:41 +01:00
Pratik Mankawde
b89a83f8f8 fix(telemetry): stamp the round span's mode when the engine applies it
startRoundTracing runs as an argument to Consensus::startRound, so it creates
consensus.round before startRoundInternal applies the new mode. Reading mode_
there recorded the previous round's value, and a validator switching from
observing to proposing got a round span labelled observing that nothing corrected.

The attribute is now written in onModeChange, from the mode being applied. All
three MonitoredMode::set paths funnel through there, so round start, a wrong-ledger
switch and a bow-out all correct the parent span with one statement. Every path
reaches it under RCLConsensus::mutex_ on the thread that created the span.

The stale write is removed rather than kept alongside: neither Consensus::startRound
nor startRoundInternal has an early return before mode_.set, so every round span is
stamped. If a future path ever skipped it the attribute would be absent, which reads
as a gap, instead of confidently wrong. onClose also sets consensus_mode, from the
engine's own parameter, and is correct as it stands.

Also tests addEvent's attribute overload on a live span, reading the exported event
name and each value back off the in-memory exporter. It was previously only ever
called on a null guard, so a dropped attribute exported nothing and failed nothing.
2026-09-07 14:39:50 +01:00
Pratik Mankawde
b5415233cd fix(telemetry): stop reporting applied_direct for a failed direct apply
tryDirectApply returns an engaged optional whenever the fee bar was cleared,
including when xrpl::apply() failed, so testing the optional labelled failures as
applied. TxQ_test's fail-in-preclaim case hits exactly this: the fee clears the
bar and preclaim then rejects with terINSUF_FEE_B.

The stamp now branches on ApplyResult::applied. A failure reports failed rather
than falling through to the default rejected, because rejected means the
transaction got nowhere, while this one cleared the fee bar and ran through
apply(). ter_code is recorded either way, so the failure is diagnosable. Both
values already existed and are used the same way by the queued-apply path in this
file, so the vocabulary is unchanged.

The value set in the phase-3 task list is updated to match, including a ter_code
row for txq.accept_tx that was already emitted but undocumented.
2026-09-07 13:42:29 +01:00
Pratik Mankawde
18abd100b5 feat(telemetry): let a call site choose a span's role, not just its category
Span kind was derived from TraceCategory alone, so every Rpc-category span was
kServer. A category cannot tell an inbound handler from the internal work under
it, and trace backends pair kServer with kClient, so internal spans left as
kServer become unpaired edges in a service graph and read as extra inbound
requests.

SpanRole is a new xrpl-owned enum, orthogonal to TraceCategory: the category
names the subsystem and gates the span on config, the role says whether the span
handles a remote call. It is a defaulted fourth parameter on span(), freshRoot()
and the ScopedSpanGuard equivalents, defaulting to SpanRole::FromCategory, so no
existing call site changes. resolveSpanKind() applies an explicit role and falls
back to the category map, which keeps its single responsibility. The
telemetry-disabled stubs mirror all four signatures.

No call site passes a role yet. The two that need it are on a later branch.

Also fixes a ScopedSpanGuard example that passed a bare op:: suffix to
childSpan(), which takes the name verbatim. Naming the child rpc.command made it
a child of rpc.command.<cmd>, inverting the hierarchy, so the example's parent is
now rpc.process and the command attribute moved onto the command span.
2026-09-07 13:39:57 +01:00
Pratik Mankawde
f8e0a19b9f fix(telemetry): correct the childSpan doc examples and make Rule D tests real
Review feedback on the RPC integration PR.

The childSpan examples could not work as written. childSpan() takes its parent
from the ambient context and uses impl_ only as a liveness gate, so an unscoped
SpanGuard parent produced two siblings rather than a parent and child. The parent
is now a ScopedSpanGuard, the child no longer reuses the parent's name, and the
examples pass a full dotted constant because childSpan() takes the name verbatim.

Five of the ten Rule D tests could not fail. Four passed an empty L1 key set,
which makes the rule skip validation altogether; the fifth asserted an empty
result against an escaped-quote selector that extracted no labels at all. Each
now passes a nonempty L1 set and carries a known-bad label in the same
expression, so it asserts both that the intended labels are accepted and that
Rule D ran. Verified by disabling the rule: the old tests stay green, the new
ones all fail.

Span kind is not fixed here. categoryToSpanKind and the span factories belong to
the telemetry library, so the role parameter is routed to that branch, and the
two call sites here follow once it exists.
2026-09-07 13:25:22 +01:00
Pratik Mankawde
4d2841ccda fix(telemetry): reject invalid [telemetry] batch settings and make isValid() honest
Two review findings on the telemetry library.

SpanContext::isValid() returned impl_ != nullptr, so it answered true for a
context holding no span. threadLocalContext() wraps whatever GetCurrent()
returns, and that is an empty Context on a thread with no active span, which
contradicted the documented "invalid context if none is active". It now asks the
Context for its span. childSpan(name, ctx) is the one caller whose behaviour
changes: a context with no span used to produce a new root span, and now returns
a null guard as its @return already promised.

The three batch settings went to the OTel BatchSpanProcessor unchecked. Three
ways that failed: zero was accepted for all of them; batch_size could exceed
max_queue_size, which the SDK documents as a precondition and does not enforce;
and a mistyped value let boost::bad_lexical_cast escape, which derives from
std::bad_cast rather than std::runtime_error, so the operator saw a bare "bad
cast" naming no key. Reading unsigned also turned "-1" into 4294967295 instead
of failing, so the value is parsed signed and negatives are rejected.
xrpld-example.cfg now states the ranges.
2026-09-07 13:16:59 +01:00
Pratik Mankawde
9aebdc292c docs(telemetry): fix stale symbols, attribute keys and TraceQL in the plan docs
Review feedback on the plan documents. Four kinds of error:

- Symbols that do not exist: ConsensusProposal::prevLedger_ (it is
  previousLedger_), RCLConsensusAdaptor (it is RCLConsensus::Adaptor, and
  startRound() is on RCLConsensus itself), and RPCHandler::doCommand (a free
  function, xrpl::rpc::doCommand).
- Attribute keys: the tables used ledger_index, which no telemetry code emits.
  Same concept as ledger_seq but a different referent, so the code disambiguates
  by prefix: current_ledger_seq for the open ledger a transaction targeted,
  ledger_seq for a closed or validated one. A note now states which is which.
- TraceQL that does not parse: span-field predicates need braces, status.code
  is not an intrinsic (status = error), and avg(duration) does not take a by
  clause (avg_over_time does). All five re-tested against Tempo.
- The StatsD comparison omitted the Histogram instrument, which aggregates at
  the point of measure, and the when-to-use table had no row for a metric that
  spans cannot afford to carry.
2026-09-07 13:16:32 +01:00
Pratik Mankawde
1bee0b1c63 merge: bring the close-time attr doc fixes forward from phase7-native-metrics 2026-09-04 12:42:22 +01:00
Pratik Mankawde
87270ef642 merge: bring the close-time attr doc fixes forward from phase6-statsd
The consensus and ledger attribute tables conflicted: this branch had
already rewritten both, adding the open-phase and avalanche attributes and
correcting tx_count/tx_failed to sit on tx.apply alone. Keep this branch's
tables and apply the close-time rename to their rows, rather than taking
either side whole.
2026-09-04 12:42:08 +01:00
Pratik Mankawde
981323f071 merge: bring the close-time attr doc fixes forward from phase5-docs-deployment 2026-09-04 12:40:21 +01:00
Pratik Mankawde
652c03f5cd merge: bring the close-time attr doc fixes forward from phase4-consensus-tracing 2026-09-04 12:40:21 +01:00
Pratik Mankawde
68364f9b8a merge: bring the close-time attr doc fixes forward from phase1c-rpc-integration 2026-09-04 12:40:21 +01:00
Pratik Mankawde
f2ef748b6b merge: bring the close-time attr doc fixes forward from phase3-tx-tracing 2026-09-04 12:40:21 +01:00
Pratik Mankawde
f2b3e5c53c merge: bring the close-time attr doc fixes forward from phase2-rpc-tracing 2026-09-04 12:40:21 +01:00
Pratik Mankawde
b91ab6c5b9 merge: bring the close-time attr doc fixes forward from phase1b-telemetry-infra 2026-09-04 12:40:21 +01:00
Pratik Mankawde
ff3f41eeaf merge: bring the close-time attr doc fixes forward from phase1a-plan-docs 2026-09-04 12:40:20 +01:00
Pratik Mankawde
ca22f57919 docs(telemetry): follow the close-time attr rename in the data reference
The emitted keys carry the unit and epoch suffix. Update the consensus
and ledger attribute tables and the ledger.build span row to match.

The Close Time Drift panel row is left alone: phase-7 removes that whole
table, so editing it here would only conflict on the way forward.
2026-09-04 12:37:19 +01:00
Pratik Mankawde
1afa35d54d docs(telemetry): follow the close-time attr rename in the runbook
The consensus.accept.apply row listed close_time, parent_close_time and
close_time_self. The emitted keys carry the unit and epoch suffix, so
update the row to match.
2026-09-04 12:35:59 +01:00
Pratik Mankawde
372fba129f docs(telemetry): follow the close-time attr rename in the phase-4 docs
The emitted keys are close_time_ripple_epoch_s,
parent_close_time_ripple_epoch_s and close_time_self_ripple_epoch_s.
Update the consensus.accept.apply attribute tables and the close-time
attribute descriptions to match.
2026-09-04 12:35:25 +01:00
Pratik Mankawde
8e15a81f96 docs(telemetry): follow the close-time attr rename in the ledger attr table
The emitted key is close_time_ripple_epoch_s, which names its unit and
epoch. Update the ledger attribute table to match.

ledger_index and ledger_tx_count in the same table belong to a separate
rename and are left as they are.
2026-09-04 12:34:07 +01:00
Pratik Mankawde
771ea277a5 merge: bring telemetry config and doc fixes forward from phase7-native-metrics 2026-09-03 20:30:33 +01:00
Pratik Mankawde
a27d3e85d5 merge: bring telemetry config and doc fixes forward from phase6-statsd
phase-6 corrected the Consensus Health template-variable table to name
service_instance_id, the label that dashboard actually filters on. This branch
removes that table entirely - the section is restructured around a
Prometheus-label reference and a pointer to the runbook - so the corrected row
has nothing to land in. Resolved by keeping the restructured section; phase-6's
fix remains correct for phase-6, where the table still exists.

The [telemetry] cfg block merged without conflict: the composed 14-key block
from upstream and this branch's metrics_endpoint, metric_export_interval_ms and
metric_export_timeout_ms entries coexist, 17 keys with one entry each.
2026-09-03 20:30:19 +01:00
Pratik Mankawde
911d7610ef merge: bring telemetry config and doc fixes forward from phase5-docs-deployment 2026-09-03 20:29:29 +01:00
Pratik Mankawde
5381bf722d merge: bring telemetry config and doc fixes forward from phase4-consensus-tracing 2026-09-03 20:29:28 +01:00