Compare commits

..

16 Commits

Author SHA1 Message Date
Nik Bougalis
55b265a8c2 Improve BaseUInt:
This commit reworks the BaseUInt class, which was originally taken from
Bitcoin and has since been heavily modified.

The internal limb type now widens from 32 to 64 bits when the requested
width is a multiple of 64. This halves the iteration count when working
on limbs; coupled with the use of add-with-carry intrinsics, the result
is better code generation and optimized use of processor resources.

The constructors have been reshuffled, constraining how an instance can
be initialized. As a result, many common initialization errors will now
fail at compile time instead of run time.

A `BaseUInt` can now be initialized with:

 - A byte sequence of fixed size, i.e an `std::array`, a C array, or a
   fixed-extent `std::span` of `unsigned char` or `std::byte`.
 - A string literal, formatted as a hexadecimal string with no leading
   0x. Malformed or improperly sized inputs are a compile-time error.
 - An unsigned integer. This constructor is mostly used for tests, and
   must execute at compile time.
 - A byte sequence whose length is not known at compile time, via the
   fromRaw() function which returns a seated `std::optional` only if
   the input is byte sequence is properly sized.

The runtime raw-pointer and `std::uint64_t` constructors are gone, as
are the container assignment operator and the the `fromVoid` and
`fromVoidChecked` functions.

Other changes:

 - The SFINAE-based container trait is replaced by the ByteCopySource
   and FixedByteRange concepts.
 - operator== and operator<=> are now hidden friends. The workaround
   in the old operator<=> is gone.
 - Efficient tag conversion is now possible.
 - `BaseUInt` is now `noexcept` and (almost) fully `constexpr`: every
   operation except byte access, hashing and text output is usable in
   constant evaluation.

Cleanups:

 - Most invocations of `stringIsUInt256Sized` in PeerImp are now gone
   and leverage `fromRaw`.
2026-10-05 19:42:52 -07:00
Nik Bougalis
02e40e7a00 Restructure the spinlock code:
- Add function-based spinlock API
- Improve comments
- Reduce bouncing with contested locks
2026-10-05 19:42:50 -07:00
Nik Bougalis
f14395427e Modernize CountedObject infrastructure:
- Replace the lazy singleton with a constinit registry of per-type
  static counters. The `getInstance` method is removed.
- A concurrency bug that would corrupt the list of counters during
  insertion has been fixed.
- Debug asserts can detect incorrect accounting that can result in
  counter under- or overflow.
- `CountedObject` can only be used as a CRTP base for its own type
  parameter.
- Track the current and maximum values of counter. This results in
  a change in the output of `get_counts`, with individual counters
  now reported as sub-objects of a `counters` object.
2026-10-05 19:42:50 -07:00
Nik Bougalis
b707bac74f Improve AccountID base58 conversion cache:
The previous cache guarded a configurably-sized vector with 64
packed spinlocks. Two aspects of that setup were not worth the
cost:

- Callers were forced to perform atomic RMW operations on a
  single word which was shared by all 64 locks; this caused
  the lock line to bounce between cores on every lookup.

- The configurable size and runtime initialization resulted
  in complexity that did not yield a meaningful improvement
  in performance.

This commit ditches the 64 packed spinlocks, replacing them
with per-entry sequence locks, allowing readers on the fast
path to avoid performing any stores at all.

Cache entries are now carefully sized and aligned to fit into
a typical cache line, which helps avoid false sharing and any
unnecessary coherence traffic.

The cache is now a fixed array of 65,536 entries. The size was
chosen to balance capacity against memory overhead and to keep
the indexing operation fast and simple: since AccountID values
are uniformly distributed, two bytes can serve directly as the
index without requiring additional hashing.

Note: The reader path will perform unsynchronized reads against
      concurrent writes; the sequence lock protocol detects and
      discards any data observed mid-update, so these races are
      benign. A sanitizer like TSAN can still flag them and the
      warnings are expected. A compile-time knob to disable the
      cache is available, if necessary.
2026-10-05 19:42:50 -07:00
Nik Bougalis
da2b5d5eaf Remove beast::typeName
Replace Beast's demangling wrapper (originally by @HowardHinnant)
with calls to `boost::core::type_name`, eliminating the low-level
dependency and manual memory management from the codebase.
2026-10-05 19:42:50 -07:00
Nik Bougalis
4de4902200 Rework and clean up the safe_cast framework:
The SafeToCast concept transposed its signedness clause: it was written in
<Dest, Src> order but it declared <Src, Dest>. This introduced what can be
best described as a polarity bug. As a result:

* Signed-to-wider-unsigned casts were wrongly deemed safe.
* Unsigned-to-wider-signed casts were wrongly deemed unsafe.

The root cause was drift caused by the safety condition being repeated in
the concept and again as a static_assert in safeCast. The checking is now
done only in the concept and is expressed in terms of range coverage, and
not via a sizeof/signedness proxy. This change also improves the handling
of bool (which previously compiled despite truncating) and eliminates the
platform-dependent accept/reject behavior for same-size types.

Additional fixes:

* The pointer form of safeDowncast now performs the same static_cast
  in all builds. The dynamic_cast check is now entirely contained in
  the XRPL_ASSERT.
* Extended integer types wider than intmax_t (e.g. __int128) are now
  excluded from SafeToCast instead of silently misevaluating through
  wrapped bounds; they remain expressible via unsafeCast.

Cleanups:

* Remove the single-parameter enum overload of safeCast: it had an
  explicitly-specified template argument binding to Src, where all
  sibling overload bound to Dest, silently yielding the underlying
  type from expressions that read as casts *to* an enum.
* Add checkedCast for conversions guarded by runtime bounds checks
  where safety varies per instantiation.
* Constrain safeDowncast to polymorphic sources and genuine public
  unambiguous downcasts, rejecting upcasts and unrelated types.
* Route enum conversions through std::to_underlying; clarify the
  unsafeCast/checkedCast documentation split.
2026-10-05 19:42:49 -07:00
Nik Bougalis
2ea1f09dcb Modernize beast::Zero for C++20
- Replace the twelve comparison operators with two
  `constexpr`-capable replacements: `==` and `<=>`
  and allow the compiler to synthesize the rest.
- Constrain `==` and `<=>` using concepts and mark
  them as conditionally noexcept).
- Remove the detail::zero_helper indirection which
  did not do what its comment claimed.
2026-10-05 19:30:38 -07:00
Nik Bougalis
68b4b894de Make ApplyFlags a scoped enumeration:
Convert ApplyFlags to a scoped enum, drop the now-redundant `Tap`
prefix from its enumerators and replace the hand-written bitwise
operators by opting in to enum_bitops.

No functional change.
2026-10-05 19:30:38 -07:00
Nik Bougalis
defd83e5a6 Make STPathElement::Type a scoped enumeration:
`STPathElement` stored its type as an `unsigned int`, returned it
as a `std::uint32_t`, while the flags themselves were an unscoped
`enum` (i.e. an `int`). The type field has always been serialized
as a single byte.

This commit turns `STPathElement::Type` into a scoped `enum` with
a fixed sized that matches the serialized width, and opts it into
the bitwise operators from `enum_bitops.h`:

- Store the type as `Type`, and use `Type` for the constructor
  parameter and the return type of `getNodeType()`.
- Keep the existing spellings such as `STPathElement::TypeAccount`
  valid through `using enum`.
- Replace raw mask tests in the serializer, JSON output, strand
  construction and pathfinder with the existing predicates.
- Validate the type byte once when deserializing, before any field
  is read, and restructure the loop around that.
- Make `operator==` a hidden friend.
- Mark the trivial accessors `noexcept`.

No functional changes are intended. The one observable difference
is the rejection error generated for a malformed path element: if
the currency and MPT bits are both set and the input is truncated
within the account field, the invalid bit combination is detected
first; previously the short read would be detected. Such elements
are rejected either way.
2026-10-05 19:30:38 -07:00
Nik Bougalis
452e31a05f Introduce opt-in bitwise operators for scoped enumerations:
Scoped enumerations do not support bitwise operators, so using an
`enum class` as a set of flags requires extensive casting at call
sites or a hand-written set of operators for each type.

Add `enum_bitops.h`, which provides `&`, `|`, `^`, `~` along with
the compound assignment forms for any scoped enumeration that has
an unsigned underlying type and is opted in:

    enum class Flags : std::uint8_t { a = 1, b = 2 };

    // Opt into bitwise operations
    template <>
    struct enum_bitops::OptIn<Flags> : std::true_type {};

All synthesized operators are constexpr and noexcept, and preserve
the enumeration type.
2026-10-05 19:30:37 -07:00
Mayukha Vadari
1798262d7a refactor: Use LedgerHashesEntry everywhere (#8369) 2026-10-05 18:15:02 +00:00
Mayukha Vadari
cd633b9ac9 refactor: Use EscrowEntry everywhere (#8354) 2026-10-05 17:51:18 +00:00
Ayaz Salikhov
5c5f7d315a chore: Update nix flake file (#8479) 2026-10-05 16:58:41 +00:00
Mayukha Vadari
0f7493ce50 refactor: Use FeeSettingsEntry everywhere (#8370) 2026-10-05 16:56:41 +00:00
Harshit Gupta
40f61f828a fix: Add string type validation for channel_id and signature (#7582) 2026-10-05 16:55:14 +00:00
Mayukha Vadari
108277f4b7 test: Migrate three beast suites from beast::unit_test to gtest (#7993) 2026-10-05 15:28:05 +00:00
321 changed files with 3848 additions and 27745 deletions

View File

@@ -58,17 +58,3 @@ ignore:
- "src/tests/"
- "include/xrpl/beast/test/"
- "include/xrpl/beast/unit_test/"
# Telemetry modules. Telemetry compiles in by default (conanfile.py,
# CMakeLists.txt). SpanGuardScope.cpp sends spans to an in-memory exporter,
# but no unit test sends spans through the real OTLP export path, so these
# files are left out of the coverage report. This block belongs on the
# earliest branch that adds telemetry code — codecov config only flows
# child-ward, so an ignore added on a later branch can never cover the
# branches before it.
- "src/xrpld/telemetry/"
- "src/libxrpl/telemetry/"
- "include/xrpl/telemetry/"
# Per-module span-name and span-label constant headers: compile-time
# constants only, colocated with their subsystem rather than under telemetry/.
- "**/*SpanNames.h"
- "**/*SpanLabels.h"

View File

@@ -69,7 +69,6 @@ words:
- Btrfs
- Buildx
- canonicality
- CGNAT
- canonicalised
- cctools
- changespq
@@ -122,7 +121,6 @@ words:
- envrc
- exceptioned
- EXPECT_STREQ
- exfiltration
- Falco
- fcontext
- finalizers
@@ -131,8 +129,6 @@ words:
- fsanitize
- funclets
- Gamal
- gantt
- Gantt
- gcov
- gcovr
- ghead
@@ -142,8 +138,6 @@ words:
- gpgkey
- Hinnant
- hotwallet
- hicpp
- htpasswd
- hwaddress
- hwrap
- ifndef
@@ -185,10 +179,11 @@ words:
- mathbunnyru
- mcmodel
- MEMORYSTATUSEX
- MPTAMM
- MPTDEX
- Merkle
- misprediction
- missingok
- MPTAMM
- mptbalance
- MPTDEX
- mptflags
@@ -223,7 +218,6 @@ words:
- nonxrp
- noreplace
- noripple
- nostd
- nostdinc
- notifempty
- nudb
@@ -232,7 +226,6 @@ words:
- Nyffenegger
- onlatest
- ostr
- otelc
- otool
- oxalica
- pargs
@@ -243,7 +236,6 @@ words:
- permdex
- perminute
- permissioned
- pimpl
- pointee
- populator
- preauth
@@ -264,7 +256,6 @@ words:
- Raphson
- rcflags
- reencrypted
- reparent
- replayer
- repodata
- repomd
@@ -337,7 +328,6 @@ words:
- TMEndpointv2
- toolchain
- tparam
- traceql
- trixie
- tx
- txid
@@ -345,7 +335,6 @@ words:
- txjson
- txn
- txns
- txqueue
- txs
- ubsan
- UBSAN
@@ -360,7 +349,6 @@ words:
- unfindable
- unflatten
- unfund
- ungated
- unimpair
- unroutable
- unscalable
@@ -399,8 +387,6 @@ words:
- xrplf
- xxhash
- xxhasher
- xychart
- zpages
- zstdio
- pratik
- dedup
- CGNAT
- ungated

View File

@@ -53,10 +53,6 @@ libxrpl.shamap > xrpl.basics
libxrpl.shamap > xrpl.nodestore
libxrpl.shamap > xrpl.protocol
libxrpl.shamap > xrpl.shamap
libxrpl.telemetry > xrpl.basics
libxrpl.telemetry > xrpl.config
libxrpl.telemetry > xrpl.protocol
libxrpl.telemetry > xrpl.telemetry
libxrpl.tx > xrpl.basics
libxrpl.tx > xrpl.conditions
libxrpl.tx > xrpl.core
@@ -64,7 +60,6 @@ libxrpl.tx > xrpl.json
libxrpl.tx > xrpl.ledger
libxrpl.tx > xrpl.protocol
libxrpl.tx > xrpl.server
libxrpl.tx > xrpl.telemetry
libxrpl.tx > xrpl.tx
test.app > test.jtx
test.app > test.unit_test
@@ -153,7 +148,6 @@ test.overlay > xrpl.protocol
test.overlay > xrpl.resource
test.overlay > xrpl.server
test.overlay > xrpl.shamap
test.overlay > xrpl.telemetry
test.protocol > test.jtx
test.protocol > test.unit_test
test.protocol > xrpl.basics
@@ -199,14 +193,10 @@ tests.libxrpl > xrpl.protocol_autogen
tests.libxrpl > xrpl.resource
tests.libxrpl > xrpl.server
tests.libxrpl > xrpl.shamap
tests.libxrpl > xrpl.telemetry
tests.libxrpl > xrpl.tx
tests.xrpld > xrpl.basics
tests.xrpld > xrpld.rpc
tests.xrpld > xrpld.telemetry
tests.xrpld > xrpl.json
tests.xrpld > xrpl.protocol
tests.xrpld > xrpl.telemetry
xrpl.conditions > xrpl.basics
xrpl.conditions > xrpl.protocol
xrpl.config > xrpl.basics
@@ -214,7 +204,6 @@ xrpl.consensus > xrpl.basics
xrpl.consensus > xrpl.json
xrpl.consensus > xrpl.ledger
xrpl.consensus > xrpl.protocol
xrpl.consensus > xrpl.telemetry
xrpl.core > xrpl.basics
xrpl.core > xrpl.json
xrpl.core > xrpl.protocol
@@ -250,20 +239,16 @@ xrpl.server > xrpl.resource
xrpl.shamap > xrpl.basics
xrpl.shamap > xrpl.nodestore
xrpl.shamap > xrpl.protocol
xrpl.telemetry > xrpl.basics
xrpl.telemetry > xrpl.config
xrpl.tx > xrpl.basics
xrpl.tx > xrpl.core
xrpl.tx > xrpl.ledger
xrpl.tx > xrpl.protocol
xrpl.tx > xrpl.telemetry
xrpld.app > test.unit_test
xrpld.app > xrpl.basics
xrpld.app > xrpl.config
xrpld.app > xrpl.consensus
xrpld.app > xrpl.core
xrpld.app > xrpld.core
xrpld.app > xrpld.telemetry
xrpld.app > xrpl.json
xrpld.app > xrpl.ledger
xrpld.app > xrpl.net
@@ -274,7 +259,6 @@ xrpld.app > xrpl.rdb
xrpld.app > xrpl.resource
xrpld.app > xrpl.server
xrpld.app > xrpl.shamap
xrpld.app > xrpl.telemetry
xrpld.app > xrpl.tx
xrpld.core > xrpl.basics
xrpld.core > xrpl.config
@@ -288,7 +272,6 @@ xrpld.overlay > xrpl.consensus
xrpld.overlay > xrpl.core
xrpld.overlay > xrpld.core
xrpld.overlay > xrpld.peerfinder
xrpld.overlay > xrpld.telemetry
xrpld.overlay > xrpl.json
xrpld.overlay > xrpl.ledger
xrpld.overlay > xrpl.peerfinder
@@ -296,7 +279,6 @@ xrpld.overlay > xrpl.protocol
xrpld.overlay > xrpl.resource
xrpld.overlay > xrpl.server
xrpld.overlay > xrpl.shamap
xrpld.overlay > xrpl.telemetry
xrpld.overlay > xrpl.tx
xrpld.peerfinder > xrpl.basics
xrpld.peerfinder > xrpld.app
@@ -324,13 +306,9 @@ xrpld.rpc > xrpl.rdb
xrpld.rpc > xrpl.resource
xrpld.rpc > xrpl.server
xrpld.rpc > xrpl.shamap
xrpld.rpc > xrpl.telemetry
xrpld.rpc > xrpl.tx
xrpld.shamap > xrpl.basics
xrpld.shamap > xrpld.core
xrpld.shamap > xrpl.nodestore
xrpld.shamap > xrpl.protocol
xrpld.shamap > xrpl.shamap
xrpld.telemetry > xrpl.basics
xrpld.telemetry > xrpl.consensus
xrpld.telemetry > xrpl.telemetry

View File

@@ -1,70 +0,0 @@
# OTel naming-consistency check
`check_otel_naming.py` enforces the OpenTelemetry span-attribute naming
convention documented in
[CONTRIBUTING.md](../../../CONTRIBUTING.md#telemetry-span-attribute-naming)
across every layer of the telemetry pipeline. The `*SpanNames.h` constants are
the single source of truth (L1); every other layer must agree with them.
## Running locally
```
python .github/scripts/otel-naming/check_otel_naming.py
```
It takes no arguments, can be run from any directory inside the repo, and uses
only the Python standard library (no `pip install`, matching the levelization
check). A non-zero exit code means a violation was found; the output lists each
violation as `RULE | location | token | expected`.
## What it checks
The valid key set is **derived dynamically from the OTel code** — there is no
hardcoded allowlist:
- **L1 keys** come from the `namespace attr { ... }` blocks of every
`*SpanNames.h`, resolving the `makeStr("x")` / `join(seg::a, seg::b)` DSL
(cross-file, so `join(seg::rpc, ...)` resolves `seg::rpc` from the base
`SpanNames.h`). Each constant is resolved against **its own** header, so two
headers that define a same-named constant (e.g. a base `attr::ledgerHash` and
a domain `attr::ledgerHash`) each contribute their real wire key — a later
header cannot clobber an earlier one's value in a flat table.
- **Legitimate dotted keys** = ONLY the keys the code actually sets as resource
attributes, i.e. the entries inside `Telemetry.cpp`'s `Resource::Create({...})`
call: the `semconv::service::*` keys (`service.*`) plus any `attr::<name>`
constants passed there (`xrpl.network.*`). A dotted key that is _declared_ in a
header but never set as a resource attr is a span attribute in resource
clothing — a Rule-A violation, even if it lives in the base `SpanNames.h`.
### Rules (each fails the build, when its inputs are present)
| Rule | Check |
| ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| A | No stray dotted span-attribute key (only the derived resource keys may be dotted). |
| G | Attribute keys are `lower_snake_case` (`^[a-z][a-z0-9_]*$` per dot-segment) — no camelCase, UPPERCASE, or spaces. |
| F | No string literals as attribute keys or span-name arguments in `setAttribute`/`addEvent`/`span`/`rootSpan`/`childSpan` (`rootSpan` shares `span`'s `(cat, prefix, name)` signature). Attribute _values_ are exempt (runtime data); `*SpanNames.h` definitions and test files are exempt. |
| B | Every collector `spanmetrics.dimensions` name exists in the L1 key set. |
| C | Every Tempo span-filter tag exists in the L1 key set. |
| D | Every dashboard label resolves to an L1 span attribute, a native-metric label (L6, emitted by MetricsRegistry), or a Prometheus/Grafana builtin. TraceQL scope prefixes (`span.`/`resource.`/…) are stripped before the L1 lookup. |
| E | No dotted `xrpl.<domain>.<field>` attribute key in the runbook (only the L1 resource attrs `xrpl.network.*` may be dotted). Span names, filenames, OTel-standard keys, and metric labels are not flagged. |
Rule F runs **unconditionally** (it is a purely syntactic check on the
call-sites and needs no `*SpanNames.h`), so a code path that calls
`SpanGuard::span`/`setAttribute` directly without ever defining a header is
still caught.
### Warnings (printed, never fail the build)
| Rule | Check |
| ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| H | A namespace-qualified constant (e.g. `foo::bar::myKey`) used at a telemetry call-site is not defined in any `*SpanNames.h`. The constant should live in the proper header; defining it in-place bypasses rules A/G/F. Warns rather than fails — the argument may be a legitimately dynamic value, and the header may live on a later branch. Bare locals and `std::` names are not warned. |
## Presence-gated
Every rule runs **only when the source files it needs are present** in the tree
and is otherwise skipped (printed as `SKIP: <rule> — <reason>`), never failed.
This keeps the check correct no matter how telemetry work is split across PRs —
a stacked chain, one large PR, or independent per-stage PRs where (for example)
the collector config lands before the dashboards. The collector/Tempo/dashboard/
runbook layers are introduced in later phases; on a branch without them, only
the L1-intrinsic rules (A, G, F) run.

View File

@@ -1,929 +0,0 @@
#!/usr/bin/env python3
"""
Usage: check_otel_naming.py
This script takes no parameters and can be called from any directory inside the
repository (it locates the repo root via `git rev-parse`).
Enforces the OpenTelemetry span-attribute naming convention documented in
CONTRIBUTING.md ("Telemetry span attribute naming") across every layer of the
telemetry pipeline. The `*SpanNames.h` constants are the single source of truth
(L1); every other layer must agree with them.
Design principles
-----------------
1. No hardcoded allowlist. The set of valid attribute keys — including which
dotted keys are legitimate resource attributes — is derived dynamically by
parsing the repository's own OTel code:
* `*SpanNames.h` `namespace attr { ... }` blocks (the underscore/bare keys
and the `join(seg::..., ...)` dotted resource compositions), and
* the keys the code passes to `Resource::Create({ ... })` in Telemetry.cpp
(the standard `semconv::service::*` keys -> service.name/version/...).
The one narrow, explicit exception is EXTERNAL_INFRA_LABELS (Rule D):
identity labels stamped by infrastructure outside this repo's OTel code
(the perf-iac harness), which by definition have no source in-tree to
derive from. Kept separate from the generic Prometheus/Grafana builtins
set so the exception stays visible rather than blending into "things
every OTel setup has".
2. Presence-gated enforcement. Every rule runs ONLY when the source files it
needs are present in the tree, and is otherwise skipped (never failed). This
keeps the check correct no matter how work is split across PRs: a stacked
chain, one large PR, or independent per-stage PRs where (for example) the
collector config lands in a different PR than the dashboards. The check never
assumes a file from another phase/PR exists.
Layers
------
L1 code : src/**/*SpanNames.h, include/**/*SpanNames.h (ground truth)
L1 resource : src/libxrpl/telemetry/Telemetry.cpp (dotted allowlist)
L1 callsites : setAttribute/addEvent/span/rootSpan/childSpan in src/**,
include/**
L2 collector : docker/telemetry/otel-collector-config.yaml (spanmetrics dims)
L3 tempo : docker/telemetry/tempo.yaml (span filter tags)
L4 dashboards: docker/telemetry/grafana/dashboards/*.json (PromQL labels)
L5 runbook : docs/telemetry-runbook.md (attr tables)
L6 metrics : MetricsRegistry.cpp instrument labels (native-metric
label keys, a valid dashboard-label source besides L1)
Rules (each FAILS the build, when its inputs are present)
---------------------------------------------------------
A No stray dotted span-attribute key. A dotted `<a>.<b>` used as a span
attribute that is not in the derived resource-key set is a violation.
G Attribute keys must be lower_snake_case (^[a-z][a-z0-9_]*$ per segment).
Flags camelCase, UPPERCASE, spaces, and other stray characters.
F No string literals as attribute keys or span-name arguments. The
setAttribute/addEvent key and the span/rootSpan/childSpan prefix/name args
must reference a *SpanNames.h constant, never a "literal". Attribute VALUES
are exempt (runtime data). Definitions inside *SpanNames.h are exempt, and
test files are exempt (they pass arbitrary literals to exercise the API).
B Every collector spanmetrics dimension exists in the L1 key set.
C Every tempo span-filter tag exists in the L1 key set.
D Every dashboard label resolves to an L1 span attribute, an L6
native-metric label, or a builtin. TraceQL `span.`/`resource.` scope
prefixes are stripped before the L1 lookup.
E No dotted `xrpl.<domain>.<field>` attribute key in the runbook (only the
L1 resource attrs xrpl.network.* may be dotted). Span names, filenames,
OTel-standard keys, and metric labels are not flagged.
Warnings (printed, but do NOT fail the build)
----------------------------------------------
H A constant referenced at a telemetry call-site is not defined in any
*SpanNames.h. Span constants should live in the corresponding
*SpanNames.h (single source of truth); defining one in-place bypasses the
naming rules. A warning (not a failure) because the argument may instead
be a legitimately dynamic local (e.g. a computed span-name leaf).
Exit code is non-zero if any present-and-enforced rule finds a violation.
Warnings never change the exit code.
"""
import re
import subprocess
import sys
from pathlib import Path
from typing import Dict, List, Optional, Set, Tuple
# ---------------------------------------------------------------------------
# Repo location
# ---------------------------------------------------------------------------
def repo_root() -> Path:
"""Return the repository root, so the script works from any CWD.
Exits with a readable message (not a traceback) if git is unavailable or the
CWD is outside a repository."""
try:
out = subprocess.run(
["git", "rev-parse", "--show-toplevel"],
capture_output=True,
text=True,
check=True,
)
except (subprocess.CalledProcessError, FileNotFoundError):
print(
"error: check_otel_naming.py must be run inside the git repository.",
file=sys.stderr,
)
sys.exit(2)
return Path(out.stdout.strip())
def read_source(path: Path) -> str:
"""Read a file as UTF-8, tolerating stray non-UTF-8 bytes rather than
crashing the whole check on one bad byte."""
return path.read_text(encoding="utf-8", errors="ignore")
# ---------------------------------------------------------------------------
# Regexes (compiled once)
# ---------------------------------------------------------------------------
# A segment/string constant definition: `inline constexpr auto NAME = <expr>;`
CONST_DEF = re.compile(r"inline\s+constexpr\s+auto\s+(\w+)\s*=\s*(.+?);", re.DOTALL)
MAKESTR = re.compile(r'makeStr\(\s*"([^"]*)"\s*\)')
# A `namespace <name> {` opener, to track which namespace a constant lives in.
NS_OPEN = re.compile(r"namespace\s+([\w:]+)\s*\{")
# A `using ::a::b::field;` re-export inside an attr block; captures the leaf.
USING_DECL = re.compile(r"using\s+(?:::)?[\w:]*::(\w+)\s*;")
# Telemetry call-sites whose string arguments must be constants, not literals.
# Require a receiver so we match real SpanGuard calls, not std::span / a math
# `span(...)` / a bare method declaration:
# - `SpanGuard::span(` / `SpanGuard::rootSpan(` / `SpanGuard::childSpan(`
# (static factories)
# - `<obj>.span(` / `<obj>->setAttribute(` etc. (member call)
# `span`/`rootSpan`/`childSpan` additionally require the `SpanGuard`/`.`/`->`
# receiver; `setAttribute`/`addEvent` only ever exist on a guard, so a `.`/`->`
# suffices. `rootSpan` shares `span`'s (cat, prefix, name) signature.
CALLSITE = re.compile(
r"(?:SpanGuard::|\.|->)\s*(setAttribute|addEvent|span|rootSpan|childSpan)\s*\("
)
# A C++ string literal (used to flag literals inside call-site argument lists).
STRING_LITERAL = re.compile(r'"((?:[^"\\]|\\.)*)"')
# A C++ line comment (`//` ... end of line) and a block comment (`/* ... */`).
LINE_COMMENT = re.compile(r"//[^\n]*")
BLOCK_COMMENT = re.compile(r"/\*.*?\*/", re.DOTALL)
# A TraceQL scope prefix on a label (`span.`, `resource.`, `event.`, etc.).
# Dashboards reference span attributes in TraceQL as `span.<attr>`; the bare
# attribute is what must exist in L1, so strip the scope before validating.
TRACEQL_SCOPE = re.compile(r"^(?:span|resource|event|link|instrumentation_scope)\.")
# An OTel metric label key as emitted in C++: `Add(.., {{"label", ...}})` /
# `{{"label", value}}` instrument calls in MetricsRegistry.
METRIC_LABEL = re.compile(r'\{\{\s*"([a-z_][a-z0-9_]*)"\s*,')
def strip_comments(text: str) -> str:
"""Remove C/C++ `//` line comments and `/* ... */` block comments.
Used only for L1 attribute-key extraction so that a commented-out or
illustrative `makeStr("...")` inside a `namespace attr` block does not leak
into the authoritative key set. Rule F deliberately does NOT strip comments
— it must still see `@code` doc-comment examples so their call-site
arguments are held to the constant-only convention.
String literals are not specially handled; a `//` or `/*` appearing inside a
string is vanishingly rare in the *SpanNames.h headers and would at worst
drop a constant from L1 (a conservative direction).
"""
text = BLOCK_COMMENT.sub("", text)
text = LINE_COMMENT.sub("", text)
return text
# ---------------------------------------------------------------------------
# L1: parse *SpanNames.h into the authoritative key set
# ---------------------------------------------------------------------------
def find_spanname_headers(root: Path) -> List[Path]:
return sorted(
p
for p in list((root / "src").rglob("*SpanNames.h"))
+ list((root / "include").rglob("*SpanNames.h"))
if p.is_file()
)
def resolve_constants(
text: str, symbols: Optional[Dict[str, str]] = None
) -> Dict[str, str]:
"""Resolve `inline constexpr auto NAME = <makeStr/join expr>` to strings.
Supports the small constexpr DSL used by SpanNames.h:
makeStr("x") -> "x"
join(a, b) -> resolve(a) + "." + resolve(b)
seg::xrpl / attr::foo -> looked up in the symbol table
The optional `symbols` argument seeds (and is updated in place with) the
table, so a global pass over ALL *SpanNames.h headers can resolve
cross-file references such as `join(seg::rpc, ...)` where `seg::rpc` is
defined in the base SpanNames.h. Keys are stored by their bare name
(last `::` component), so `seg::rpc` and `rpc` both resolve.
"""
if symbols is None:
symbols = {}
def resolve_expr(expr: str) -> Optional[str]:
expr = expr.strip()
m = MAKESTR.fullmatch(expr)
if m:
return m.group(1)
if expr.startswith("join(") and expr.endswith(")"):
args = split_top_level_args(expr[len("join(") : -1])
parts = [resolve_expr(a) for a in args]
if any(p is None for p in parts):
return None
return ".".join(p for p in parts if p is not None)
# Bare or qualified symbol reference, e.g. `seg::xrpl` or `networkId`.
key = expr.split("::")[-1]
return symbols.get(key, symbols.get(expr))
# Iterate definitions in source order so earlier symbols are available.
for m in CONST_DEF.finditer(text):
name, expr = m.group(1), m.group(2)
val = resolve_expr(expr)
if val is not None:
symbols[name] = val
return symbols
def build_global_symbols(headers: List[Path]) -> Dict[str, str]:
"""Resolve constants across ALL headers so cross-file `seg::`/`join`
references (e.g. `join(seg::rpc, ...)` in RpcSpanNames.h, where `seg::rpc`
lives in the base SpanNames.h) resolve. Base SpanNames.h is processed
first so its `seg::` segments seed the table."""
symbols: Dict[str, str] = {}
ordered = sorted(headers, key=lambda p: (p.name != "SpanNames.h", str(p)))
# Two passes: the first seeds segments, the second resolves dependents.
# Comments are stripped so a commented-out constant cannot seed the table.
for _ in range(2):
for h in ordered:
resolve_constants(strip_comments(read_source(h)), symbols)
return symbols
def split_top_level_args(s: str) -> List[str]:
"""Split a comma-separated arg list, respecting nested parentheses and
ignoring parens/commas that appear inside a "string literal" (so a value
like `setAttribute(k, ",")` does not get mis-split)."""
args, depth, cur = [], 0, ""
in_str = False
escaped = False
for ch in s:
if in_str:
cur += ch
if escaped:
escaped = False
elif ch == "\\":
escaped = True
elif ch == '"':
in_str = False
continue
if ch == '"':
in_str = True
cur += ch
elif ch == "(":
depth += 1
cur += ch
elif ch == ")":
depth -= 1
cur += ch
elif ch == "," and depth == 0:
args.append(cur)
cur = ""
else:
cur += ch
if cur.strip():
args.append(cur)
return args
def attr_namespace_spans(text: str) -> List[str]:
"""Return the source text of each `namespace attr { ... }` block in `text`.
Brace-matched over the whole (comment-stripped) text, so a definition that
wraps across several physical lines is contained in one span. Nested braces
inside the block are balanced correctly."""
spans: List[str] = []
for opener in NS_OPEN.finditer(text):
if opener.group(1).split("::")[-1] != "attr":
continue
# Walk from the opening brace, balancing nesting to the matching close.
i = opener.end() # one char past the namespace's `{`
depth = 1
start = i
while i < len(text) and depth > 0:
c = text[i]
if c == "{":
depth += 1
elif c == "}":
depth -= 1
i += 1
spans.append(text[start : i - 1])
return spans
def attr_keys_from_header(path: Path, symbols: Dict[str, str]) -> Set[str]:
"""Return the set of attribute-key strings declared in a header's
`namespace attr { ... }` block(s). `symbols` is the global cross-file
table, used ONLY to seed `seg::`/segment references for `join(...)`
resolution — never to look up an attr constant's value.
A constant DEFINED in this header is resolved against this header's OWN
text, so two headers that each define a same-named constant (e.g. the base
`attr::ledgerHash = xrpl.ledger.hash` and consensus
`attr::ledgerHash = ledger_hash`) each report their real wire key. The
global table is keyed by bare name and would otherwise let a later header
clobber an earlier one, erasing the real key from L1 (a Rule-A blind spot).
A `using`-re-export, by contrast, imports a constant defined elsewhere, so
it is resolved against the global table.
Comments are stripped first (a commented constant must not enter L1), and
each attr block is brace-matched over the whole text so multi-line
`inline constexpr auto NAME = join(\\n ...);` definitions are captured."""
text = strip_comments(read_source(path))
# Local table: the global segments/symbols seed cross-file `join` parts,
# then this header's own definitions overwrite any same-named global entry
# so a locally-defined attr resolves to ITS value, not another header's.
local = dict(symbols)
resolve_constants(text, local)
keys: Set[str] = set()
for block in attr_namespace_spans(text):
for md in CONST_DEF.finditer(block):
# Resolve a locally-defined constant against the LOCAL table; this
# captures makeStr("x") and join(seg::y, ...) with the header's own
# value, immune to cross-header bare-name collisions.
val = local.get(md.group(1))
if val is not None:
keys.add(val)
# `using ::ns::attr::field;` re-exports a constant defined in ANOTHER
# header (e.g. PeerSpanNames imports the base ledgerHash). Resolve the
# imported name against the global table.
for um in USING_DECL.finditer(block):
val = symbols.get(um.group(1))
if val is not None:
keys.add(val)
return keys
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
class Report:
def __init__(self) -> None:
self.violations: List[Tuple[str, str, str, str]] = []
self.warnings: List[Tuple[str, str, str, str]] = []
self.skips: List[str] = []
self.checked: List[str] = []
def violation(self, rule: str, loc: str, token: str, expected: str) -> None:
self.violations.append((rule, loc, token, expected))
def warning(self, rule: str, loc: str, token: str, note: str) -> None:
"""A non-fatal finding: printed, but does not fail the build. Used where
the script cannot be certain a finding is wrong (e.g. a constant used at
a call-site that is not defined in any *SpanNames.h — it might be a
misplaced constant, or a legitimately dynamic value)."""
self.warnings.append((rule, loc, token, note))
def skip(self, rule: str, reason: str) -> None:
self.skips.append(f"SKIP: {rule} — {reason}")
def ok(self, msg: str) -> None:
self.checked.append(f"OK: {msg}")
def render_and_exit(self) -> None:
for line in self.skips:
print(line)
for line in self.checked:
print(line)
if self.warnings:
print("\nNaming-convention warnings (non-fatal):\n")
print(f" {'RULE':<5} {'LOCATION':<48} {'TOKEN':<28} NOTE")
print(f" {'-' * 5} {'-' * 48} {'-' * 28} {'-' * 30}")
for rule, loc, token, note in self.warnings:
print(f" {rule:<5} {loc:<48} {token:<28} {note}")
if self.violations:
print("\nNaming-convention violations:\n")
print(f" {'RULE':<5} {'LOCATION':<48} {'TOKEN':<28} EXPECTED")
print(f" {'-' * 5} {'-' * 48} {'-' * 28} {'-' * 30}")
for rule, loc, token, expected in self.violations:
print(f" {rule:<5} {loc:<48} {token:<28} {expected}")
print(
"\nSee CONTRIBUTING.md -> 'Telemetry span attribute naming'. "
"The *SpanNames.h constants are the single source of truth."
)
sys.exit(1)
print("\nAll present telemetry naming layers are consistent.")
sys.exit(0)
def main() -> None:
root = repo_root()
report = Report()
# --- Build the L1 ground-truth key set (presence-gated) ----------------
headers = find_spanname_headers(root)
l1_keys: Set[str] = set()
if headers:
symbols = build_global_symbols(headers)
# Map each key to the header(s) that declare it, so Rule A can tell a
# legitimate resource attr (declared in the base SpanNames.h) from a
# stray dotted key declared in a domain header.
keys_by_header: Dict[Path, Set[str]] = {}
for h in headers:
hk = attr_keys_from_header(h, symbols)
keys_by_header[h] = hk
l1_keys |= hk
report.ok(
f"L1: {len(l1_keys)} attribute keys from {len(headers)} "
f"*SpanNames.h header(s)"
)
else:
report.skip("L1", "no *SpanNames.h present (not a naming-relevant tree)")
keys_by_header = {}
# --- Derive the legitimate dotted (resource) keys dynamically ----------
# ONLY the keys actually passed to Resource::Create() in Telemetry.cpp
# (semconv service.* + the attr:: constants set there, e.g. xrpl.network.*).
# A dotted key declared in a header but NOT set as a resource attr is a
# Rule-A violation, not an allowlist entry.
resource_symbols = symbols if headers else {}
dotted_allow = derive_dotted_resource_keys(root, resource_symbols, report)
# --- Rule A: no stray dotted span-attribute keys -----------------------
if l1_keys:
run_rule_a(keys_by_header, dotted_allow, report)
# --- Rule G: keys must be lower_snake_case -----------------------------
if l1_keys:
run_rule_g(keys_by_header, report)
# --- Rule F (+ Rule H): scan telemetry call-sites ----------------------
# Runs UNCONDITIONALLY: Rule F is a purely syntactic check (is this argument
# a literal?) and does not need the L1 key set, so a code path that uses
# SpanGuard::span/setAttribute directly without ever defining a *SpanNames.h
# is still caught. Rule H (warning) additionally flags constant references
# not defined in any *SpanNames.h.
header_symbols = spanname_symbol_names(headers)
run_rule_f(root, report, header_symbols)
# --- Cross-layer rules B/C/D/E (each presence-gated) -------------------
# L6 native-metric labels: span attributes are not the only valid dashboard
# labels — the MetricsRegistry emits OTel metrics whose label keys are an
# additional source of truth. Derive them dynamically (same principle as L1)
# so dashboards may reference them without tripping Rule D.
metric_labels = metric_label_names(root)
run_rule_b_collector(root, l1_keys, report)
run_rule_c_tempo(root, l1_keys, report)
run_rule_d_dashboards(root, l1_keys, metric_labels, report)
run_rule_e_runbook(root, l1_keys, report)
report.render_and_exit()
def resource_create_block(text: str) -> str:
"""Return the text inside the first `Resource::Create({ ... })` argument
list, brace-matched so nested `{key, value}` initializers are contained.
Empty string if the call is absent."""
m = re.search(r"Resource::Create\(\s*\{", text)
if not m:
return ""
i = m.end() # one char past the opening `{`
depth, start = 1, i
while i < len(text) and depth > 0:
c = text[i]
if c == "{":
depth += 1
elif c == "}":
depth -= 1
i += 1
return text[start : i - 1]
def derive_dotted_resource_keys(
root: Path, symbols: Dict[str, str], report: Report
) -> Set[str]:
"""Legitimate dotted keys = ONLY the keys the code actually sets as RESOURCE
attributes, i.e. the entries inside Telemetry.cpp's `Resource::Create({...})`
call: the standard semconv keys (`service.*`) plus any `attr::<name>`
constants passed there (resolved to their wire key via the global symbol
table, e.g. `attr::networkId` -> `xrpl.network.id`).
A dotted key DECLARED in a `*SpanNames.h` header but NOT passed to
Resource::Create() is a span attribute wearing the resource form — a Rule-A
violation, never allowlisted. Deriving the allowlist from the actual
resource call (not from "any dotted key in the base header") is what lets
Rule A catch a stray dotted span attr such as `xrpl.ledger.hash`."""
allow: Set[str] = set()
tele = root / "src" / "libxrpl" / "telemetry" / "Telemetry.cpp"
if not tele.is_file():
report.skip("resource-derive", "Telemetry.cpp not present")
return allow
block = resource_create_block(read_source(tele))
# semconv::<group>::k<CamelKey> -> the dotted OTel-standard key. The
# CamelKey already embeds the group, e.g. service::kServiceInstanceId
# -> service.instance.id. Split the CamelCase name into dotted lowercase
# segments; if it does not lead with the group, prepend the group.
for m in re.finditer(r"semconv::(\w+)::k(\w+)", block):
group, camel = m.group(1), m.group(2)
segments = camel_to_dotsegments(camel)
if segments and segments[0] == group:
allow.add(".".join(segments))
else:
allow.add(group + "." + ".".join(segments))
# attr::<name> constants set as resource attrs (e.g. networkId/networkType);
# resolve each to its wire key and allowlist only the dotted ones.
for m in re.finditer(r"attr::(\w+)", block):
val = symbols.get(m.group(1))
if val is not None and "." in val:
allow.add(val)
report.ok(f"resource dotted-key allowlist derived: {sorted(allow)}")
return allow
def camel_to_dotsegments(s: str) -> List[str]:
"""Split a CamelCase identifier into lowercase dot-segment parts, e.g.
`ServiceInstanceId` -> ['service', 'instance', 'id']."""
return [w.lower() for w in re.findall(r"[A-Z][a-z0-9]*", s)]
def run_rule_a(
keys_by_header: Dict[Path, Set[str]], dotted_allow: Set[str], report: Report
) -> None:
"""Any dotted attribute key that is not an allowed resource key is a
violation, reported against the header that declares it."""
found = False
for h in sorted(keys_by_header):
for key in sorted(keys_by_header[h]):
if "." in key and key not in dotted_allow:
found = True
report.violation("A", h.name, key, "underscore form, not dotted")
if not found:
report.ok("A: no stray dotted span-attribute keys")
# A lower_snake_case identifier segment: starts lowercase, then lowercase /
# digits / underscores. No uppercase, no spaces, no camelCase.
SNAKE_SEGMENT = re.compile(r"^[a-z][a-z0-9_]*$")
def run_rule_g(keys_by_header: Dict[Path, Set[str]], report: Report) -> None:
"""Every attribute key must be lower_snake_case. Bare/underscore keys must
match ^[a-z][a-z0-9_]*$; dotted resource keys must be lowercase
dot-separated segments (each segment lower_snake_case). Flags camelCase,
UPPERCASE, spaces, and other stray characters."""
found = False
for h in sorted(keys_by_header):
for key in sorted(keys_by_header[h]):
segments = key.split(".")
if all(SNAKE_SEGMENT.match(seg) for seg in segments):
continue
found = True
report.violation("G", h.name, key, "must be lower_snake_case")
if not found:
report.ok("G: all attribute keys are lower_snake_case")
# Which argument positions of each call must be a constant (0-based). The
# attribute VALUE position is intentionally absent: values are runtime data
# (command names, hashes, counts), not naming-convention surface.
# setAttribute(key, value) -> check arg 0 (key); value (arg 1) exempt
# addEvent(name[, attrs]) -> check arg 0 (event name)
# span(category, prefix, name) -> check args 1,2 (prefix + span-name leaf)
# rootSpan(category, prefix, name)-> check args 1,2 (same signature as span)
# childSpan(name[, parentCtx]) -> check arg 0 (span-name leaf)
CONSTANT_ARG_POSITIONS: Dict[str, Set[int]] = {
"setAttribute": {0},
"addEvent": {0},
"span": {1, 2},
"rootSpan": {1, 2}, # same signature as span(cat, prefix, name)
"childSpan": {0},
}
def is_test_path(path: Path, root: Path) -> bool:
"""True if the path is test code. Tests legitimately pass arbitrary literal
keys/names to exercise the API mechanics, so Rule F does not apply to them.
Matches a `test`/`tests` directory anywhere in the path below `root` (e.g.
src/test/, src/tests/, .../detail/tests/). Only the part below `root` is
read, so a checkout inside a folder named `test` is not all test code."""
return any(part in ("test", "tests") for part in path.relative_to(root).parts)
# A constant reference passed at a call-site, e.g. `rpc_span::attr::command`
# or a bare `myKey`. We capture the leaf identifier (after the last `::`).
IDENTIFIER_ARG = re.compile(r"^[\s&*]*([A-Za-z_][\w:]*)\s*$")
def spanname_symbol_names(headers: List[Path]) -> Set[str]:
"""Every `inline constexpr auto NAME = ...;` symbol defined across the
*SpanNames.h headers, by bare name. Used by Rule H to tell whether a
constant referenced at a call-site actually lives in a SpanNames header."""
names: Set[str] = set()
for h in headers:
for m in CONST_DEF.finditer(strip_comments(read_source(h))):
names.add(m.group(1))
return names
def run_rule_f(root: Path, report: Report, header_symbols: Set[str]) -> None:
"""Walk every telemetry call-site (non-test, non-*SpanNames.h) and check the
constant-only argument positions of
setAttribute/addEvent/span/rootSpan/childSpan:
Rule F (FAIL): a string literal in a key / span-name position. Attribute
VALUES are exempt (runtime data).
Rule H (WARN): a constant reference whose name is not defined in any
*SpanNames.h. The constant should live in the corresponding
*SpanNames.h (single source of truth); defining it in-place bypasses
the naming rules. Warn rather than fail — the argument may instead be a
legitimately dynamic local (e.g. a computed span-name leaf)."""
found_f = False
sources = [
p
for base in ("src", "include")
for ext in ("*.h", "*.cpp")
for p in (root / base).rglob(ext)
if p.is_file()
]
for path in sorted(sources):
if path.name.endswith("SpanNames.h") or is_test_path(path, root):
continue
text = read_source(path)
rel = path.relative_to(root)
for call, arglist, lineno in iter_calls(text):
positions = CONSTANT_ARG_POSITIONS.get(call, set())
args = split_top_level_args(arglist)
for idx in positions:
if idx >= len(args):
continue
arg = args[idx]
lit = STRING_LITERAL.search(arg)
if lit:
found_f = True
report.violation(
"F",
f"{rel}:{lineno}",
f'{call} arg{idx} "{lit.group(1)}"',
"use a *SpanNames.h constant",
)
continue
# Not a literal: Rule H warns when a NAMESPACE-QUALIFIED constant
# reference (e.g. `consensus::span::accept`) is not defined in
# any *SpanNames.h — i.e. the constant was defined in-place
# instead of in the proper header. We only consider qualified
# refs (containing `::`): a bare lowercase identifier is almost
# always a legitimately dynamic local (a computed span-name leaf
# or attribute value), not a misplaced constant, so warning on it
# would be noise. Standard-library types (std::...) are skipped.
ident = IDENTIFIER_ARG.match(arg)
if not (ident and header_symbols):
continue
ref = ident.group(1)
if "::" not in ref or ref.startswith("std::"):
continue
leaf = ref.split("::")[-1]
if leaf not in header_symbols:
report.warning(
"H",
f"{rel}:{lineno}",
f"{call} arg{idx} {ref}",
"not defined in any *SpanNames.h",
)
if not found_f:
report.ok("F: no string-literal keys/names at telemetry call-sites")
def iter_calls(text: str):
"""Yield (call_name, raw_arglist, lineno) for each setAttribute/addEvent/
span/rootSpan/childSpan invocation, spanning multiple physical lines if
needed."""
for m in CALLSITE.finditer(text):
name = m.group(1)
# Walk from the opening paren, balancing nesting to find the close.
# Parens inside a "string literal" are ignored so a value such as
# `setAttribute(k, ")")` does not close the call early.
i = m.end() # one char past the '('
depth = 1
in_str = False
escaped = False
while i < len(text) and depth > 0:
c = text[i]
if in_str:
if escaped:
escaped = False
elif c == "\\":
escaped = True
elif c == '"':
in_str = False
elif c == '"':
in_str = True
elif c == "(":
depth += 1
elif c == ")":
depth -= 1
i += 1
arglist = text[m.end() : i - 1]
lineno = text.count("\n", 0, m.start()) + 1
yield name, arglist, lineno
def run_rule_b_collector(root: Path, l1_keys: Set[str], report: Report) -> None:
path = root / "docker" / "telemetry" / "otel-collector-config.yaml"
if not path.is_file():
report.skip("B", "collector config not present")
return
text = read_source(path)
if "spanmetrics" not in text:
report.skip("B", "no spanmetrics block in collector config")
return
dims = extract_spanmetrics_dimensions(text)
if not l1_keys:
report.skip("B", "no L1 key set to validate against")
return
miss = [d for d in dims if d not in l1_keys]
for d in miss:
report.violation("B", str(path.relative_to(root)), d, "must exist in L1")
if not miss:
report.ok(f"B: {len(dims)} collector dimension(s) all in L1")
def extract_spanmetrics_dimensions(text: str) -> List[str]:
"""The `- name: <dim>` entries listed under each `dimensions:` key.
Blank and comment lines are skipped, so a comment between entries never ends
the list, even when it holds a colon. The list ends at the first other line
indented no deeper than its `dimensions:` key. A `-` entry at the key's own
indentation still belongs to the list, as YAML allows."""
dims: List[str] = []
key_col: Optional[int] = None
for line in text.splitlines():
content = line.strip()
if not content or content.startswith("#"):
continue
indent = len(line) - len(line.lstrip())
if key_col is not None:
is_entry = re.match(r"-(\s|$)", content) is not None
if indent > key_col or (indent == key_col and is_entry):
m = re.search(r"-\s*name\s*:\s*([A-Za-z0-9_.]+)", line)
if m:
dims.append(m.group(1))
continue
key_col = None
key = re.match(r"(\s*(?:-\s+)?)dimensions\s*:", line)
if key:
key_col = len(key.group(1))
return dims
def run_rule_c_tempo(root: Path, l1_keys: Set[str], report: Report) -> None:
# The trace-search filter tags live in the Grafana Tempo DATASOURCE
# provisioning file (search.filters[].{tag,scope}); the Tempo server
# tempo.yaml has no such tags. Prefer the datasource file; fall back to the
# server file so the rule still does something if the layout changes.
candidates = [
root / "docker/telemetry/grafana/provisioning/datasources/tempo.yaml",
root / "docker/telemetry/tempo.yaml",
]
path = next((p for p in candidates if p.is_file()), None)
if path is None:
report.skip("C", "tempo datasource provisioning not present")
return
if not l1_keys:
report.skip("C", "no L1 key set to validate against")
return
# Pair each filter's `tag:` with its `scope:` (a few lines below it) and
# validate only span-scope tags — resource/intrinsic tags (service.*, name,
# status, duration) are not span attributes. Strip a TraceQL span. prefix.
lines = read_source(path).splitlines()
span_tags: List[str] = []
for i, line in enumerate(lines):
m = re.search(r"^\s*tag:\s*(\S+)", line)
if not m:
continue
scope = next(
(
sm.group(1)
for j in range(i, min(i + 4, len(lines)))
for sm in [re.search(r"scope:\s*(\S+)", lines[j])]
if sm
),
"",
)
if scope == "span":
span_tags.append(TRACEQL_SCOPE.sub("", m.group(1)))
if not span_tags:
report.skip("C", "no span-scope filter tags in tempo datasource")
return
miss = [t for t in span_tags if t not in l1_keys]
for t in sorted(set(miss)):
report.violation("C", str(path.relative_to(root)), t, "must exist in L1")
if not miss:
report.ok(f"C: {len(span_tags)} tempo span-filter tag(s) all in L1")
def metric_label_names(root: Path) -> Set[str]:
"""L6: OTel native-metric label keys emitted by the telemetry code, e.g.
`counter->Add(1, {{"job_type", value}})` in MetricsRegistry.cpp. These are
a valid source of dashboard labels distinct from span attributes (L1)."""
labels: Set[str] = set()
for base in ("src", "include"):
for p in (root / base).rglob("*.cpp"):
if not p.is_file():
continue
text = read_source(p)
if "MetricsRegistry" not in p.name and "metric" not in text.lower():
continue
labels |= set(METRIC_LABEL.findall(text))
return labels
# Identity labels stamped by EXTERNAL infrastructure the OTel pipeline in this
# repo does not own: the perf-iac harness attaches these to every metric it
# scrapes so dashboards can filter by which build/role produced a series. They
# have no L1 (*SpanNames.h), L2 (collector config), or L6 (MetricsRegistry.cpp)
# source to derive from, so — unlike every other dashboard label — they cannot
# be validated dynamically. This is a deliberate, narrow exception to the "no
# hardcoded allowlist" design principle, kept separate from the generic
# Prometheus/Grafana `builtins` set below so it stays visible and auditable.
# Add a label here ONLY if it is genuinely injected by infra outside this
# repo's OTel code (never as a workaround for a dashboard querying a label
# that nothing actually emits — that is a real Rule D violation).
EXTERNAL_INFRA_LABELS = {
"xrpl_branch", # perf-iac: git ref of the xrpld build under test
"xrpl_node_role", # perf-iac: validator/peer role in the perf cluster
}
def run_rule_d_dashboards(
root: Path, l1_keys: Set[str], metric_labels: Set[str], report: Report
) -> None:
dash_dir = root / "docker" / "telemetry" / "grafana" / "dashboards"
files = sorted(dash_dir.glob("*.json")) if dash_dir.is_dir() else []
if not files:
report.skip("D", "no dashboard JSON present")
return
if not l1_keys:
report.skip("D", "no L1 key set to validate against")
return
builtins = {
"__name__", # Prometheus reserved label for the metric name itself
"le",
"exported_instance",
"span_name",
"status_code",
"service_name",
"service_version",
"service_instance_id",
"job",
"instance",
}
# A dashboard label is valid if it is a span attribute (L1), a native-metric
# label (L6), a Prometheus/Grafana builtin, or an external-infra identity
# label (EXTERNAL_INFRA_LABELS).
valid = l1_keys | metric_labels | builtins | EXTERNAL_INFRA_LABELS
found = False
for f in files:
try:
text = read_source(f)
except OSError:
continue
# PromQL `sum by (a, b)` and `{label="..."}` references.
labels: Set[str] = set()
for m in re.finditer(r"by\s*\(([^)]*)\)", text):
labels |= {x.strip() for x in m.group(1).split(",") if x.strip()}
for m in re.finditer(r"\b([a-z_][a-z0-9_.]*)\s*(?:=~|!~|!=|=)\s*\"", text):
labels.add(m.group(1))
for lbl in sorted(labels):
# Strip a TraceQL scope prefix (span./resource./...) — the bare
# attribute is what must resolve against L1.
bare = TRACEQL_SCOPE.sub("", lbl)
if bare in valid:
continue
found = True
report.violation(
"D",
str(f.relative_to(root)),
lbl,
"must exist in L1, a metric label, or be a builtin",
)
if not found:
report.ok(f"D: dashboard PromQL labels all resolve ({len(files)} file(s))")
def run_rule_e_runbook(root: Path, l1_keys: Set[str], report: Report) -> None:
path = root / "docs" / "telemetry-runbook.md"
if not path.is_file():
report.skip("E", "runbook not present")
return
if not l1_keys:
report.skip("E", "no L1 key set to validate against")
return
text = read_source(path)
found = False
# Only the dotted `xrpl.<domain>.<field>` attribute form is a violation. The
# `xrpl.`-with-trailing-dot anchor is the discriminator: it matches the old
# dotted attribute convention being migrated away from, while everything
# else legitimately dotted in the runbook does NOT match it —
# * span names (`consensus.round`, `tx.process`) no `xrpl.` prefix
# * filenames (`xrpld.cfg`, `RCLConsensus.cpp`) `xrpld.`/`.cpp`, not `xrpl.`
# * OTel-standard (`service.name`, `http.method`) no `xrpl.` prefix
# * metric labels (`xrpl_rpc_command`) underscore, no dot
# Legitimate dotted resource attrs (`xrpl.network.id`/`.type`) are in L1 and
# are skipped. A dotted `xrpl.` token absent from L1 is a genuine doc/code
# mismatch (e.g. `xrpl.tx.hash` where the code emits `tx_hash`).
for m in re.finditer(r"`(xrpl\.[a-z][a-z0-9_.]*)`", text):
token = m.group(1)
if token in l1_keys: # legitimate dotted resource attr (xrpl.network.*)
continue
found = True
report.violation(
"E", str(path.relative_to(root)), token, "underscore, not dotted"
)
if not found:
report.ok("E: runbook attribute references consistent with L1")
if __name__ == "__main__":
main()

View File

@@ -1,961 +0,0 @@
#!/usr/bin/env python3
"""Unit tests for check_otel_naming.py.
Stdlib-only (unittest), matching the dependency-free policy of the check itself.
Run from anywhere:
python .github/scripts/otel-naming/test_check_otel_naming.py
Each rule is exercised in isolation against a synthetic tree / synthetic L1 key
set, covering positive (must flag), negative (must not flag), and boundary
cases. Rule E (runbook dotted-attribute detection) has the densest coverage
because its discriminator — the `xrpl.<domain>.` prefix vs span names,
filenames, OTel-standard keys, and metric labels — is the subtlest.
"""
import contextlib
import importlib.util
import io
import shutil
import tempfile
import unittest
from pathlib import Path
# Load the check module by path (it is not an importable package).
_spec = importlib.util.spec_from_file_location(
"check_otel_naming", str(Path(__file__).with_name("check_otel_naming.py"))
)
chk = importlib.util.module_from_spec(_spec)
_spec.loader.exec_module(chk)
# A controlled L1 set used across tests: the two legitimate dotted resource
# attrs plus a handful of underscore span-attribute keys.
L1 = {
"xrpl.network.id",
"xrpl.network.type",
"tx_hash",
"peer_id",
"consensus_mode",
"command",
"rpc_status",
"ledger_seq",
}
def _run_rule_e(runbook_text: str):
"""Run Rule E against a synthetic runbook; return the flagged tokens."""
d = Path(tempfile.mkdtemp())
try:
(d / "docs").mkdir()
(d / "docs" / "telemetry-runbook.md").write_text(runbook_text)
report = chk.Report()
chk.run_rule_e_runbook(d, set(L1), report)
return sorted(v[2] for v in report.violations)
finally:
shutil.rmtree(d)
class RuleERunbook(unittest.TestCase):
"""Rule E: only dotted `xrpl.<domain>.<field>` attribute keys are flagged."""
# ----- positive: genuine dotted attribute-key violations -----
def test_single_dotted_attr(self):
self.assertEqual(_run_rule_e("`xrpl.tx.hash`"), ["xrpl.tx.hash"])
def test_multiple_dotted_attrs(self):
self.assertEqual(
_run_rule_e("`xrpl.tx.hash` and `xrpl.consensus.mode`"),
["xrpl.consensus.mode", "xrpl.tx.hash"],
)
def test_deep_dotted_three_segments(self):
self.assertEqual(
_run_rule_e("`xrpl.consensus.ledger.seq`"), ["xrpl.consensus.ledger.seq"]
)
def test_dotted_attr_with_underscore_field(self):
self.assertEqual(
_run_rule_e("`xrpl.consensus.round_id`"), ["xrpl.consensus.round_id"]
)
def test_repeated_token_reported_each_occurrence(self):
self.assertEqual(
_run_rule_e("`xrpl.tx.hash` ... `xrpl.tx.hash`"),
["xrpl.tx.hash", "xrpl.tx.hash"],
)
def test_resource_attr_not_in_l1_is_flagged(self):
self.assertEqual(
_run_rule_e("`xrpl.network.unknown`"), ["xrpl.network.unknown"]
)
# ----- negative: legitimately-dotted tokens that must NOT be flagged -----
def test_span_name_single(self):
self.assertEqual(_run_rule_e("`consensus.round`"), [])
def test_span_name_multi_segment(self):
self.assertEqual(
_run_rule_e("`consensus.phase.open` `rpc.command.server_info`"), []
)
def test_filename_cfg(self):
self.assertEqual(_run_rule_e("`xrpld.cfg`"), [])
def test_filename_cpp(self):
self.assertEqual(_run_rule_e("`RCLConsensus.cpp`"), [])
def test_otel_standard_service_name(self):
self.assertEqual(_run_rule_e("`service.name`"), [])
def test_otel_standard_http_method(self):
self.assertEqual(_run_rule_e("`http.method`"), [])
def test_metric_label_underscore(self):
self.assertEqual(_run_rule_e("`xrpl_rpc_command`"), [])
def test_bare_underscore_attrs(self):
self.assertEqual(_run_rule_e("`tx_hash` `consensus_mode`"), [])
def test_legit_dotted_resource_attrs_in_l1(self):
self.assertEqual(_run_rule_e("`xrpl.network.id` `xrpl.network.type`"), [])
def test_prose_word(self):
self.assertEqual(_run_rule_e("the `command` attribute"), [])
def test_plain_prose_no_backticks(self):
self.assertEqual(_run_rule_e("xrpl.tx.hash without backticks is prose"), [])
# ----- boundary -----
def test_empty_runbook(self):
self.assertEqual(_run_rule_e(""), [])
def test_lookalike_prefix_xrpld(self):
# `xrpld.` is NOT `xrpl.` — must not match.
self.assertEqual(_run_rule_e("`xrpld.foo`"), [])
def test_lookalike_prefix_underscore(self):
# `xrpl_rpc.command` starts with `xrpl_`, not `xrpl.`.
self.assertEqual(_run_rule_e("`xrpl_rpc.command`"), [])
def test_uppercase_segment_not_matched(self):
# The pattern requires a lowercase char after `xrpl.`; uppercase keys are
# caught by Rule G at the L1 layer, not by the runbook text scan.
self.assertEqual(_run_rule_e("`xrpl.TX.hash`"), [])
def test_token_touching_table_pipes(self):
self.assertEqual(_run_rule_e("| `xrpl.tx.hash` | desc |"), ["xrpl.tx.hash"])
def test_mixed_line_only_xrpl_dotted_flagged(self):
self.assertEqual(
_run_rule_e("`consensus.round` uses `xrpl.tx.hash` and `service.name`"),
["xrpl.tx.hash"],
)
def test_skips_when_runbook_absent(self):
d = Path(tempfile.mkdtemp())
try:
report = chk.Report()
chk.run_rule_e_runbook(d, set(L1), report)
self.assertEqual(report.violations, [])
self.assertTrue(any("SKIP: E" in s for s in report.skips))
finally:
shutil.rmtree(d)
def test_skips_when_l1_empty(self):
d = Path(tempfile.mkdtemp())
try:
(d / "docs").mkdir()
(d / "docs" / "telemetry-runbook.md").write_text("`xrpl.tx.hash`")
report = chk.Report()
chk.run_rule_e_runbook(d, set(), report)
self.assertEqual(report.violations, [])
self.assertTrue(any("SKIP: E" in s for s in report.skips))
finally:
shutil.rmtree(d)
class DslParser(unittest.TestCase):
"""The makeStr/join/seg:: constexpr DSL resolver — the foundation of the
L1 key set. Covers flat, nested, cross-file, alias, and multi-line forms."""
def test_flat_join(self):
syms = chk.resolve_constants(
'inline constexpr auto a = makeStr("xrpl");\n'
'inline constexpr auto b = makeStr("network");\n'
"inline constexpr auto c = join(a, b);\n"
)
self.assertEqual(syms["c"], "xrpl.network")
def test_nested_join_three_segments(self):
syms = chk.resolve_constants(
'inline constexpr auto xrpl = makeStr("xrpl");\n'
'inline constexpr auto network = makeStr("network");\n'
"inline constexpr auto networkId = "
'join(join(xrpl, network), makeStr("id"));\n'
)
self.assertEqual(syms["networkId"], "xrpl.network.id")
def test_qualified_seg_reference(self):
# `seg::rpc` resolves by its bare leaf `rpc`.
syms = chk.resolve_constants('inline constexpr auto rpc = makeStr("rpc");\n')
syms2 = chk.resolve_constants(
'inline constexpr auto command = join(seg::rpc, makeStr("command"));\n',
syms,
)
self.assertEqual(syms2["command"], "rpc.command")
def test_alias_reference(self):
syms = chk.resolve_constants('inline constexpr auto rpc = makeStr("rpc");\n')
chk.resolve_constants("inline constexpr auto alias = seg::rpc;\n", syms)
self.assertEqual(syms["alias"], "rpc")
def test_unresolvable_expr_omitted(self):
syms = chk.resolve_constants("inline constexpr auto x = join(unknown, y);\n")
self.assertNotIn("x", syms)
def test_split_top_level_args_respects_nesting(self):
self.assertEqual(
chk.split_top_level_args("join(seg::a, b), c"),
["join(seg::a, b)", " c"],
)
def test_split_top_level_args_ignores_comma_in_string(self):
self.assertEqual(
chk.split_top_level_args('key, ","'),
["key", ' ","'],
)
def test_strip_comments_removes_line_and_block(self):
self.assertEqual(
chk.strip_comments("a // line\nb /* blk */ c").split(),
["a", "b", "c"],
)
def _write(path: Path, text: str) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(text)
def _header(ns_attr_body: str, prefix_seg: str = "") -> str:
"""A minimal *SpanNames.h body: optional seg defs + a namespace attr block."""
return (
"#pragma once\n"
+ prefix_seg
+ "namespace xrpl::telemetry::demo::span {\n"
+ "namespace attr {\n"
+ ns_attr_body
+ "} // namespace attr\n"
+ "}\n"
)
class AttrKeyExtraction(unittest.TestCase):
"""attr_keys_from_header: comment-stripping + multi-line + using re-export."""
def _l1(self, header_text):
d = Path(tempfile.mkdtemp())
try:
h = d / "src" / "DemoSpanNames.h"
_write(h, header_text)
syms = chk.build_global_symbols([h])
return chk.attr_keys_from_header(h, syms)
finally:
shutil.rmtree(d)
def test_single_line_makestr(self):
keys = self._l1(_header('inline constexpr auto k = makeStr("tx_hash");\n'))
self.assertIn("tx_hash", keys)
def test_multiline_constexpr_captured(self):
keys = self._l1(
_header("inline constexpr auto k =\n" ' makeStr("round_time_ms");\n')
)
self.assertIn("round_time_ms", keys)
def test_commented_makestr_not_leaked(self):
keys = self._l1(
_header(
'inline constexpr auto k = makeStr("good");\n'
'// inline constexpr auto bad = makeStr("old.dotted");\n'
)
)
self.assertIn("good", keys)
self.assertNotIn("old.dotted", keys)
def test_block_commented_makestr_not_leaked(self):
keys = self._l1(
_header(
'inline constexpr auto k = makeStr("good");\n'
'/* makeStr("blockbad") */\n'
)
)
self.assertNotIn("blockbad", keys)
class CamelToDotSegments(unittest.TestCase):
"""semconv CamelCase -> dotted OTel-standard key derivation."""
def test_service_instance_id(self):
self.assertEqual(
chk.camel_to_dotsegments("ServiceInstanceId"),
["service", "instance", "id"],
)
def test_service_name(self):
self.assertEqual(chk.camel_to_dotsegments("ServiceName"), ["service", "name"])
def test_derive_keys_from_telemetry_cpp(self):
d = Path(tempfile.mkdtemp())
try:
tele = d / "src" / "libxrpl" / "telemetry" / "Telemetry.cpp"
_write(
tele,
"resource::Resource::Create({\n"
" {semconv::service::kServiceName, x},\n"
" {semconv::service::kServiceInstanceId, y},\n"
"});\n",
)
report = chk.Report()
allow = chk.derive_dotted_resource_keys(d, {}, report)
self.assertIn("service.name", allow)
self.assertIn("service.instance.id", allow)
finally:
shutil.rmtree(d)
class SymbolCollision(unittest.TestCase):
"""attr_keys_from_header must resolve a constant against ITS OWN header, so
two headers defining a same-named constant each report their real wire key.
Regression for the flat-symbol-table collision that let a later header
clobber an earlier one and erased a dotted key from L1 (a Rule-A blind
spot)."""
def _build(self, files):
d = Path(tempfile.mkdtemp())
paths = {}
for rel, text in files.items():
p = d / rel
_write(p, text)
paths[rel] = p
return d, paths
def test_same_named_const_not_clobbered_across_headers(self):
base = (
"#pragma once\n"
"namespace xrpl::telemetry {\n"
'namespace seg { inline constexpr auto xrpl = makeStr("xrpl");\n'
'inline constexpr auto ledger = makeStr("ledger"); }\n'
"namespace attr {\n"
"inline constexpr auto ledgerHash = "
'join(join(seg::xrpl, seg::ledger), makeStr("hash"));\n'
"}\n}\n"
)
cons = (
"#pragma once\n"
"namespace xrpl::telemetry::consensus::span {\n"
"namespace attr { inline constexpr auto ledgerHash = "
'makeStr("ledger_hash"); }\n}\n'
)
d, paths = self._build(
{
"include/xrpl/telemetry/SpanNames.h": base,
"src/xrpld/consensus/ConsensusSpanNames.h": cons,
}
)
try:
headers = chk.find_spanname_headers(d)
syms = chk.build_global_symbols(headers)
by_name = {p.name: chk.attr_keys_from_header(p, syms) for p in headers}
# The base header keeps its dotted key; consensus keeps the bare one.
self.assertIn("xrpl.ledger.hash", by_name["SpanNames.h"])
self.assertEqual(by_name["ConsensusSpanNames.h"], {"ledger_hash"})
finally:
shutil.rmtree(d)
def test_using_reexport_still_resolves_globally(self):
# A `using`-re-export imports a constant defined elsewhere; it must
# resolve against the global table, not the local header.
base = (
"#pragma once\n"
"namespace xrpl::telemetry {\n"
"namespace attr { inline constexpr auto txHash = "
'makeStr("tx_hash"); }\n}\n'
)
dom = (
"#pragma once\n"
"namespace xrpl::telemetry::tx::span {\n"
"namespace attr { using ::xrpl::telemetry::attr::txHash; }\n}\n"
)
d, paths = self._build(
{
"include/xrpl/telemetry/SpanNames.h": base,
"src/xrpld/app/misc/TxSpanNames.h": dom,
}
)
try:
headers = chk.find_spanname_headers(d)
syms = chk.build_global_symbols(headers)
keys = chk.attr_keys_from_header(
paths["src/xrpld/app/misc/TxSpanNames.h"], syms
)
self.assertEqual(keys, {"tx_hash"})
finally:
shutil.rmtree(d)
class ResourceAllowlistScope(unittest.TestCase):
"""derive_dotted_resource_keys must allowlist ONLY the dotted keys actually
passed to Resource::Create() — not every dotted key in the base header. A
dotted attr declared in a header but not set as a resource attr is a Rule-A
violation."""
def _derive(self, tele_text, span_text):
d = Path(tempfile.mkdtemp())
try:
_write(d / "src" / "libxrpl" / "telemetry" / "Telemetry.cpp", tele_text)
_write(d / "include" / "xrpl" / "telemetry" / "SpanNames.h", span_text)
headers = chk.find_spanname_headers(d)
syms = chk.build_global_symbols(headers)
allow = chk.derive_dotted_resource_keys(d, syms, chk.Report())
return allow, syms, headers, d
except Exception:
shutil.rmtree(d)
raise
def test_dotted_span_attr_not_allowlisted_and_flagged(self):
span = (
"#pragma once\n"
"namespace xrpl::telemetry {\n"
'namespace seg { inline constexpr auto xrpl = makeStr("xrpl");\n'
'inline constexpr auto ledger = makeStr("ledger");\n'
'inline constexpr auto network = makeStr("network"); }\n'
"namespace attr {\n"
"inline constexpr auto networkId = "
'join(join(seg::xrpl, seg::network), makeStr("id"));\n'
"inline constexpr auto ledgerHash = "
'join(join(seg::xrpl, seg::ledger), makeStr("hash"));\n'
"}\n}\n"
)
tele = (
"auto r = resource::Resource::Create({\n"
" {semconv::service::kServiceName, x},\n"
" {std::string(attr::networkId), n},\n"
"});\n"
)
allow, syms, headers, d = self._derive(tele, span)
try:
# networkId IS a resource attr; ledgerHash is NOT, despite living in
# the base header.
self.assertIn("xrpl.network.id", allow)
self.assertNotIn("xrpl.ledger.hash", allow)
kbh = {h: chk.attr_keys_from_header(h, syms) for h in headers}
report = chk.Report()
chk.run_rule_a(kbh, allow, report)
self.assertEqual([v[2] for v in report.violations], ["xrpl.ledger.hash"])
finally:
shutil.rmtree(d)
def test_resource_block_brace_matched(self):
# A nested {key,value} initializer must not truncate the block scan.
tele = (
"auto r = resource::Resource::Create({\n"
" {semconv::service::kServiceName, x},\n"
" {std::string(attr::networkType), t},\n"
"});\n"
)
span = (
"#pragma once\n"
"namespace xrpl::telemetry {\n"
'namespace seg { inline constexpr auto xrpl = makeStr("xrpl");\n'
'inline constexpr auto network = makeStr("network"); }\n'
"namespace attr { inline constexpr auto networkType = "
'join(join(seg::xrpl, seg::network), makeStr("type")); }\n}\n'
)
allow, _syms, _headers, d = self._derive(tele, span)
try:
self.assertIn("xrpl.network.type", allow)
self.assertIn("service.name", allow)
finally:
shutil.rmtree(d)
def _run_rule_a(keys_by_header, allow):
report = chk.Report()
chk.run_rule_a(keys_by_header, allow, report)
return sorted(v[2] for v in report.violations)
class RuleADotted(unittest.TestCase):
def test_dotted_attr_not_in_allow_flagged(self):
kbh = {Path("src/RpcSpanNames.h"): {"xrpl.tx.hash", "command"}}
self.assertEqual(_run_rule_a(kbh, {"xrpl.network.id"}), ["xrpl.tx.hash"])
def test_resource_attr_in_allow_passes(self):
kbh = {Path("src/SpanNames.h"): {"xrpl.network.id"}}
self.assertEqual(_run_rule_a(kbh, {"xrpl.network.id"}), [])
def test_bare_key_never_flagged(self):
kbh = {Path("src/TxSpanNames.h"): {"tx_hash", "command"}}
self.assertEqual(_run_rule_a(kbh, set()), [])
def _run_rule_g(keys_by_header):
report = chk.Report()
chk.run_rule_g(keys_by_header, report)
return sorted(v[2] for v in report.violations)
class RuleGSnakeCase(unittest.TestCase):
def test_camelcase_flagged(self):
self.assertEqual(_run_rule_g({Path("h"): {"txHash"}}), ["txHash"])
def test_uppercase_flagged(self):
self.assertEqual(_run_rule_g({Path("h"): {"TX_HASH"}}), ["TX_HASH"])
def test_space_flagged(self):
self.assertEqual(_run_rule_g({Path("h"): {"bad key"}}), ["bad key"])
def test_snake_case_passes(self):
self.assertEqual(_run_rule_g({Path("h"): {"tx_hash", "rpc_status"}}), [])
def test_dotted_resource_segments_pass(self):
self.assertEqual(_run_rule_g({Path("h"): {"xrpl.network.id"}}), [])
def test_dotted_with_bad_segment_flagged(self):
self.assertEqual(
_run_rule_g({Path("h"): {"xrpl.Network.id"}}), ["xrpl.Network.id"]
)
class RuleFAndH(unittest.TestCase):
"""run_rule_f: literal keys/span-names flagged; values & tests exempt.
Rule H: qualified constant not in any header warns (non-fatal)."""
def _run(self, rel_path, source, header_symbols=frozenset()):
d = Path(tempfile.mkdtemp())
try:
_write(d / rel_path, source)
report = chk.Report()
chk.run_rule_f(d, report, set(header_symbols))
return (
sorted(v[2] for v in report.violations),
sorted(w[2] for w in report.warnings),
)
finally:
shutil.rmtree(d)
def test_literal_key_flagged(self):
v, _ = self._run("src/Foo.cpp", 'g.setAttribute("lit_key", v);\n')
self.assertEqual(v, ['setAttribute arg0 "lit_key"'])
def test_literal_value_exempt(self):
v, _ = self._run("src/Foo.cpp", 'g.setAttribute(attr::command, "submit");\n')
self.assertEqual(v, [])
def test_span_name_args_flagged(self):
v, _ = self._run("src/Foo.cpp", 'SpanGuard::span(cat, "rpc", "command");\n')
self.assertEqual(v, ['span arg1 "rpc"', 'span arg2 "command"'])
def test_rootspan_literal_flagged_by_rule_f(self):
# rootSpan(cat, prefix, name) shares span()'s signature, so a string
# literal in the prefix/name position must FAIL rule F exactly as it
# does for span() — otherwise a call switched to rootSpan silently
# escapes span-name validation.
v, _ = self._run(
"src/Foo.cpp",
'SpanGuard::rootSpan(cat, "peer", "validation.receive");\n',
)
self.assertEqual(
v, ['rootSpan arg1 "peer"', 'rootSpan arg2 "validation.receive"']
)
def test_rootspan_constant_args_accepted(self):
# Constant references in the prefix/name position are accepted (no
# rule F), mirroring span()'s constant-arg handling.
v, _ = self._run(
"src/Foo.cpp",
"SpanGuard::rootSpan(TraceCategory::Peer, seg::peer, "
"peer_span::op::validationReceive);\n",
)
self.assertEqual(v, [])
def test_test_path_exempt(self):
v, _ = self._run("src/test/Foo.cpp", 'g.setAttribute("lit_key", v);\n')
self.assertEqual(v, [])
def test_spannames_header_exempt(self):
v, _ = self._run("src/DemoSpanNames.h", 'g.setAttribute("lit_key", v);\n')
self.assertEqual(v, [])
def test_bare_span_call_not_matched(self):
# No SpanGuard/./-> receiver -> not a telemetry call-site.
v, _ = self._run("src/Foo.cpp", 'auto s = span("not", "telemetry");\n')
self.assertEqual(v, [])
def test_multiline_call_reports_first_line(self):
v, _ = self._run("src/Foo.cpp", 'g.setAttribute(\n "k",\n v);\n')
self.assertEqual(v, ['setAttribute arg0 "k"'])
def test_paren_in_string_value_does_not_break_parsing(self):
# The ")" inside the value must not end the call early; key still seen.
v, _ = self._run("src/Foo.cpp", 'g.setAttribute("k", ")");\n')
self.assertEqual(v, ['setAttribute arg0 "k"'])
def test_rule_h_qualified_constant_warns(self):
v, w = self._run(
"src/Foo.cpp",
"g.setAttribute(consensus::span::accept, v);\n",
header_symbols={"command"},
)
self.assertEqual(v, [])
self.assertEqual(w, ["setAttribute arg0 consensus::span::accept"])
def test_rule_h_known_constant_no_warning(self):
_, w = self._run(
"src/Foo.cpp",
"g.setAttribute(rpc_span::attr::command, v);\n",
header_symbols={"command"},
)
self.assertEqual(w, [])
def test_rule_h_bare_local_no_warning(self):
_, w = self._run(
"src/Foo.cpp", "g.setAttribute(myLeaf, v);\n", header_symbols={"command"}
)
self.assertEqual(w, [])
class RuleBCollector(unittest.TestCase):
def _run(self, yaml_text, l1):
d = Path(tempfile.mkdtemp())
try:
_write(d / "docker" / "telemetry" / "otel-collector-config.yaml", yaml_text)
report = chk.Report()
chk.run_rule_b_collector(d, set(l1), report)
return sorted(v[2] for v in report.violations), report.skips
finally:
shutil.rmtree(d)
def test_dimension_not_in_l1_flagged(self):
y = "spanmetrics:\n dimensions:\n - name: bogus_dim\n - name: command\n"
v, _ = self._run(y, {"command"})
self.assertEqual(v, ["bogus_dim"])
def test_all_dimensions_in_l1_pass(self):
y = "spanmetrics:\n dimensions:\n - name: command\n - name: rpc_status\n"
v, _ = self._run(y, {"command", "rpc_status"})
self.assertEqual(v, [])
def test_comment_does_not_end_list_but_dedent_does(self):
# The entries sit at the key's own indentation, which YAML allows. A
# comment holding a colon sits between them, so the entry after it must
# still be read. The sibling key `other:` ends the list, so the
# `- name:` under it is not a dimension.
y = (
"connectors:\n"
" spanmetrics:\n"
" dimensions:\n"
" - name: command\n"
" # Note: a comment with a colon.\n"
" - name: bogus_after_comment\n"
" other:\n"
" - name: bogus_outside_list\n"
)
v, _ = self._run(y, {"command"})
self.assertEqual(v, ["bogus_after_comment"])
def test_skip_when_no_spanmetrics_block(self):
v, skips = self._run("receivers:\n otlp:\n", {"command"})
self.assertEqual(v, [])
self.assertTrue(any("SKIP: B" in s for s in skips))
class RuleCTempo(unittest.TestCase):
"""Rule C reads the Grafana Tempo DATASOURCE file's search.filters and
validates only span-scope tags against L1."""
DS = "docker/telemetry/grafana/provisioning/datasources/tempo.yaml"
def _run(self, yaml_text, l1):
d = Path(tempfile.mkdtemp())
try:
_write(d / self.DS, yaml_text)
report = chk.Report()
chk.run_rule_c_tempo(d, set(l1), report)
return sorted(v[2] for v in report.violations), report.skips
finally:
shutil.rmtree(d)
def _filter(self, fid, tag, scope):
return (
f" - id: {fid}\n"
f" tag: {tag}\n"
f' operator: "="\n'
f" scope: {scope}\n"
f" type: static\n"
)
def test_span_tag_not_in_l1_flagged(self):
y = "search:\n filters:\n" + self._filter("f1", "bogus_tag", "span")
v, _ = self._run(y, {"command"})
self.assertEqual(v, ["bogus_tag"])
def test_span_tags_in_l1_pass(self):
y = (
"search:\n filters:\n"
+ self._filter("f1", "command", "span")
+ self._filter("f2", "tx_hash", "span")
)
v, _ = self._run(y, {"command", "tx_hash"})
self.assertEqual(v, [])
def test_resource_and_intrinsic_tags_ignored(self):
# service.* (resource) and name/status/duration (intrinsic) are not
# span attributes — they must not be validated against L1.
y = (
"search:\n filters:\n"
+ self._filter("f1", "service.instance.id", "resource")
+ self._filter("f2", "name", "intrinsic")
+ self._filter("f3", "duration", "intrinsic")
)
v, skips = self._run(y, {"command"})
self.assertEqual(v, [])
self.assertTrue(any("SKIP: C" in s for s in skips))
def test_skip_when_datasource_absent(self):
d = Path(tempfile.mkdtemp())
try:
report = chk.Report()
chk.run_rule_c_tempo(d, {"command"}, report)
self.assertEqual(report.violations, [])
self.assertTrue(any("SKIP: C" in s for s in report.skips))
finally:
shutil.rmtree(d)
class RuleDDashboards(unittest.TestCase):
def _run(self, json_text, l1, metric_labels=frozenset()):
d = Path(tempfile.mkdtemp())
try:
_write(
d / "docker" / "telemetry" / "grafana" / "dashboards" / "x.json",
json_text,
)
report = chk.Report()
chk.run_rule_d_dashboards(d, set(l1), set(metric_labels), report)
return sorted(v[2] for v in report.violations)
finally:
shutil.rmtree(d)
def test_unknown_promql_label_flagged(self):
self.assertEqual(
self._run('"expr": "sum by (bogus_label) (x)"', {"command"}),
["bogus_label"],
)
# Rule D returns early on an empty L1 set, so a test that passes one asserts
# nothing. Each test below passes a nonempty L1 set and puts `bogus_label` in
# the same expression as the labels under test: the assertion then pins both
# halves at once — the accepted labels are absent from the result, and Rule D
# demonstrably ran because it flagged the bad one.
def test_builtin_labels_not_flagged(self):
self.assertEqual(
self._run(
'"expr": "sum by (le, span_name, exported_instance, bogus_label) (x)"',
{"command"},
),
["bogus_label"],
)
def test_external_infra_labels_not_flagged(self):
# EXTERNAL_INFRA_LABELS (perf-iac identity labels with no in-tree
# source) must be recognized as valid, distinct from `builtins`. The
# names are spelled out rather than joined from chk.EXTERNAL_INFRA_LABELS
# because building the query from the set that validates it passes for
# whatever that set happens to hold — including an empty one.
self.assertEqual(
self._run(
'"expr": "sum by (xrpl_branch, xrpl_node_role, bogus_label) (x)"',
{"command"},
),
["bogus_label"],
)
def test_prometheus_name_label_not_flagged(self):
# `__name__` is the Prometheus reserved metric-name label; the renamed
# system-*.json dashboards use `sum by (le, __name__)`.
self.assertEqual(
self._run(
'"expr": "sum by (le, __name__, bogus_label) (rate(x[5m]))"',
{"command"},
),
["bogus_label"],
)
def test_l1_label_passes(self):
# `by (...)` form, not a `{command="x"}` selector: a dashboard stores the
# query inside a JSON string, so its quotes are escaped on disk and the
# selector branch extracts nothing from them here.
self.assertEqual(
self._run('"expr": "sum by (command, bogus_label) (x)"', {"command"}),
["bogus_label"],
)
def test_traceql_span_prefix_stripped(self):
# `span.establish_count` must validate against the bare L1 key.
self.assertEqual(
self._run(
'"expr": "count_over_time(x) by (span.establish_count)"',
{"establish_count"},
),
[],
)
def test_traceql_resource_prefix_stripped(self):
# `resource.service_name` must validate against the bare builtin, same as
# the `span.` case above.
self.assertEqual(
self._run(
'"expr": "count_over_time(x) by (resource.service_name, bogus_label)"',
{"command"},
),
["bogus_label"],
)
def test_native_metric_label_passes(self):
# `job_type` / `reason` are emitted by MetricsRegistry, not span attrs.
self.assertEqual(
self._run(
'"expr": "sum by (job_type, reason) (x)"',
{"command"},
metric_labels={"job_type", "reason"},
),
[],
)
def test_unknown_label_still_flagged_with_metric_labels(self):
# A label that is neither L1, metric label, nor builtin still fails.
self.assertEqual(
self._run(
'"expr": "sum by (bogus) (x)"',
{"command"},
metric_labels={"job_type"},
),
["bogus"],
)
def test_span_prefixed_unknown_still_flagged(self):
# `span.not_a_key` whose bare form is unknown is still a violation.
self.assertEqual(
self._run('"expr": "x by (span.not_a_key)"', {"command"}),
["span.not_a_key"],
)
def test_not_equal_selector_label_flagged(self):
# A label used only with `!=` must still be extracted and validated.
# The operator class `[=!]~?` matched `=`, `=~` and `!~`, but not
# `!=`, so a label that appears only in a `!=` selector slipped past
# Rule D -- a false negative, where a bad label name used with `!=`
# was never caught. The quote is unescaped, which both the phase-1c
# (`\s*\"`) and phase-9 (`\s*\\?\"`) selector forms accept.
#
# Mutation caught: revert the operator class to `[=!]~?` and this
# test fails, because `bogus_label` is never extracted and the
# result is empty instead of ["bogus_label"].
self.assertEqual(
self._run('"expr": "node_metric{bogus_label!="x"}"', {"command"}),
["bogus_label"],
)
class MetricLabelExtraction(unittest.TestCase):
"""L6: native-metric label keys parsed from C++ instrument calls."""
def test_extracts_add_label(self):
d = Path(tempfile.mkdtemp())
try:
_write(
d / "src" / "xrpld" / "telemetry" / "MetricsRegistry.cpp",
'counter->Add(1, {{"job_type", std::string(jobType)}});\n'
'c2->Add(1, {{"reason", std::string(r)}});\n',
)
self.assertEqual(chk.metric_label_names(d), {"job_type", "reason"})
finally:
shutil.rmtree(d)
def test_no_metrics_file_empty(self):
d = Path(tempfile.mkdtemp())
try:
(d / "src").mkdir()
self.assertEqual(chk.metric_label_names(d), set())
finally:
shutil.rmtree(d)
class ReportExitContract(unittest.TestCase):
@staticmethod
def _exit_code(report):
"""Call render_and_exit (which prints + raises SystemExit), swallowing
its stdout, and return the exit code."""
with contextlib.redirect_stdout(io.StringIO()):
try:
report.render_and_exit()
except SystemExit as e:
return e.code
return None # pragma: no cover - render_and_exit always exits
def test_violation_exits_nonzero(self):
r = chk.Report()
r.violation("A", "f", "tok", "exp")
self.assertEqual(self._exit_code(r), 1)
def test_clean_exits_zero(self):
r = chk.Report()
r.ok("all good")
self.assertEqual(self._exit_code(r), 0)
def test_warning_only_exits_zero(self):
r = chk.Report()
r.warning("H", "f", "tok", "note")
self.assertEqual(self._exit_code(r), 0)
class RuleEReportTuple(unittest.TestCase):
"""Assert Rule E records the full (rule, expected) tuple, not just token."""
def test_violation_tuple_fields(self):
d = Path(tempfile.mkdtemp())
try:
(d / "docs").mkdir()
(d / "docs" / "telemetry-runbook.md").write_text("`xrpl.tx.hash`")
report = chk.Report()
chk.run_rule_e_runbook(d, {"xrpl.network.id"}, report)
self.assertEqual(len(report.violations), 1)
rule, _loc, token, expected = report.violations[0]
self.assertEqual(rule, "E")
self.assertEqual(token, "xrpl.tx.hash")
self.assertEqual(expected, "underscore, not dotted")
finally:
shutil.rmtree(d)
def test_clean_runbook_records_ok(self):
d = Path(tempfile.mkdtemp())
try:
(d / "docs").mkdir()
(d / "docs" / "telemetry-runbook.md").write_text(
"`tx_hash` `consensus.round`"
)
report = chk.Report()
chk.run_rule_e_runbook(d, {"tx_hash"}, report)
self.assertEqual(report.violations, [])
self.assertTrue(any("E:" in c for c in report.checked))
finally:
shutil.rmtree(d)
if __name__ == "__main__":
unittest.main(verbosity=2)

View File

@@ -70,10 +70,8 @@ jobs:
files: |
# These paths are unique to `on-pr.yml`.
.github/scripts/levelization/**
.github/scripts/otel-naming/**
.github/scripts/rename/**
.github/workflows/reusable-check-levelization.yml
.github/workflows/reusable-check-otel-naming.yml
.github/workflows/reusable-check-rename.yml
.github/workflows/on-pr.yml
@@ -147,11 +145,6 @@ jobs:
if: ${{ needs.should-run.outputs.go == 'true' }}
uses: ./.github/workflows/reusable-check-levelization.yml
check-otel-naming:
needs: should-run
if: ${{ needs.should-run.outputs.go == 'true' }}
uses: ./.github/workflows/reusable-check-otel-naming.yml
check-rename:
needs: should-run
if: ${{ needs.should-run.outputs.go == 'true' }}
@@ -278,7 +271,6 @@ jobs:
needs:
- check-autogen
- check-levelization
- check-otel-naming
- check-rename
- check-release-commits
- clang-tidy

View File

@@ -1,28 +0,0 @@
# This workflow checks that OpenTelemetry span-attribute names stay consistent
# across the code (*SpanNames.h), collector, Tempo, dashboards, and docs.
# See .github/scripts/otel-naming/check_otel_naming.py and the
# "Telemetry span attribute naming" section in CONTRIBUTING.md.
name: Check OTel naming
# This workflow can only be triggered by other workflows.
on: workflow_call
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}-otel-naming
cancel-in-progress: true
defaults:
run:
shell: bash
jobs:
otel-naming:
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
- name: Check OTel naming
# The script is stdlib-only and reads only files already in the tree;
# it enforces each rule only when the layer it needs is present, so it
# works whether telemetry changes land in one PR or several.
run: python .github/scripts/otel-naming/check_otel_naming.py

3
.gitignore vendored
View File

@@ -100,6 +100,3 @@ target/
# Rust build directory
crates/target
# Env. file carrying environmental setup data for local or cloud runs.
.env.*

View File

@@ -30,6 +30,11 @@ Version 3.5.0 is not yet released.
- `subscribe`, `unsubscribe`: Added an optional `mpt_issuances` request field, an array of MPT issuance IDs (hex strings). Subscribers receive a message with `type` `mptTransaction` for each validated transaction whose metadata affects a subscribed issuance; the message has the same fields as the `transactions` stream. MPT issuance subscriptions count toward the per-connection subscription limit. An empty array, a non-array value, or an invalid ID returns `invalidParams`. ([#5671](https://github.com/XRPLF/rippled/pull/5671))
### Bugfixes in 3.5.0
- `channel_authorize`: The `channel_id` field now returns an `invalidParams` error if the value is not a string. [#7582](https://github.com/XRPLF/rippled/pull/7582)
- `channel_verify`: The `channel_id` and `signature` fields now return an `invalidParams` error if the value is not a string. [#7582](https://github.com/XRPLF/rippled/pull/7582)
## XRP Ledger server version 3.4.0
Version 3.4.0 is not yet released. These changes are available in the 3.4.0 beta releases.

View File

@@ -140,24 +140,6 @@ if(rocksdb)
target_link_libraries(xrpl_libs INTERFACE RocksDB::rocksdb)
endif()
# OpenTelemetry distributed tracing (optional).
# When on, links against opentelemetry-cpp and defines XRPL_ENABLE_TELEMETRY
# so that SpanGuard factory methods produce real OTel spans.
# When off, all tracing code compiles to no-ops with zero overhead.
#
# There is no CMake option. The one switch is `conan install -o telemetry=`,
# which decides whether opentelemetry-cpp is fetched and sets the variable read
# below through the generated toolchain.
#
# The define is set for every target in this project, and also PUBLIC on
# xrpl.libxrpl.telemetry (cmake/XrplCore.cmake), so a parent project that
# adds this one with add_subdirectory() and links that module gets it too.
if(telemetry)
find_package(opentelemetry-cpp CONFIG REQUIRED)
add_compile_definitions(XRPL_ENABLE_TELEMETRY)
message(STATUS "OpenTelemetry tracing enabled")
endif()
# Work around changes to Conan recipe for now.
if(TARGET nudb::core)
set(nudb nudb::core)

View File

@@ -377,66 +377,6 @@ run-clang-tidy -p build -quiet -fix -format -allow-no-checks src tests
`-format` reformats the fixed code with [`.clang-format`](./.clang-format); without it the fixes are inserted in LLVM style and the `clang-format` hook rewrites them afterwards.
## Telemetry span attribute naming
OpenTelemetry span attribute keys follow these rules so they stay consistent
across the code, the OTel collector, Tempo, Grafana dashboards, and docs. The
constants in the `*SpanNames.h` headers are the single source of truth; every
other layer must match them. A CI check enforces this end to end.
1. Per-span unique attribute: bare field name — allowed when the field is
recorded by a single span/workflow, so the span name already supplies the
domain (e.g. `command`, `local`, `version` on `rpc.command` / `tx.process`).
2. Shared attribute (same concept on more than one span): ONE key, reused
verbatim on every span that records it — the span name tells the occurrences
apart, so no per-emitter prefix is added. Pick the name by the field's
meaning: a property of a domain object keeps that object's bare field name
(`ledger_hash`, `ledger_seq`, `tx_hash`, `peer_id`, `full_validation`); a
field already qualified by a sub-kind keeps that qualifier on every emitter
(`proposal_trusted` on both `consensus.proposal.receive` and
`peer.proposal.receive`; `validation_trusted` likewise). Define it once in
the base `SpanNames.h` `namespace attr` block and re-export (`using`) it from
each domain header, so all emitters share the exact string.
3. Collision qualifier: `<domain>_<field>` — only when a bare name would collide
with a DIFFERENT concept in the shared spanmetrics label space, or with the
OTel-reserved `status` key (e.g. `rpc_status`, `grpc_status`,
`consensus_phase`, `consensus_round`). This disambiguates distinct concepts
that share a word; it is NOT used to tag the same concept with the workflow
that emitted it — that is rule 2 (one shared name).
4. Resource attribute: dotted `xrpl.<subsystem>.<field>` — reserved ONLY for
process/network identity set once at startup (`xrpl.network.id`,
`xrpl.network.type`). Never use the dotted `xrpl.` form for span attributes.
5. Span names use `<subsystem>[.<component>]` (dotted). Only attribute _keys_
follow rules 1–4.
All attribute keys are `lower_snake_case` (lowercase letters, digits, and
underscores; each dot-separated segment of a resource key likewise). No
camelCase, uppercase, or spaces.
Standard OpenTelemetry semantic-convention keys keep their canonical dotted
form (e.g. `service.*` resource attributes, `http.*` span attributes); the
"no dotted form" rule above applies to xrpl-custom keys, not to OTel-standard
conventions.
Always reference the `*SpanNames.h` constants for attribute keys and span
names — never pass a string literal as a key or as a `span`/`childSpan` name
argument. (Attribute _values_ may be runtime data.)
These rules are enforced by `.github/scripts/otel-naming/check_otel_naming.py`,
run in CI on every pull request. The check derives the set of valid keys
directly from the `*SpanNames.h` constants and the resource attributes the code
registers, so there is no separate list to keep in sync. It cross-validates the
collector, Tempo, dashboards, and docs against those keys, and each rule runs
only when the file it needs is present — so it works whether telemetry changes
land in one pull request or several. Run it locally with:
```
python .github/scripts/otel-naming/check_otel_naming.py
```
See [.github/scripts/otel-naming/README.md](.github/scripts/otel-naming/README.md)
for the full rule list.
## Contracts and instrumentation
We are using [Antithesis](https://antithesis.com/) for continuous fuzzing,

View File

@@ -1,570 +0,0 @@
# Distributed Tracing Fundamentals
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Next**: [Architecture Analysis](./01-architecture-analysis.md)
---
## What is Distributed Tracing?
Distributed tracing is a method for tracking data objects as they flow through distributed systems. In a network like XRP Ledger, a single transaction touches multiple independent nodes—each with no shared memory or logging. Distributed tracing connects these dots.
**Without tracing:** You see isolated logs on each node with no way to correlate them.
**With tracing:** You see the complete journey of a transaction or an event across all nodes it touched.
---
## Actors and Actions at a Glance
### Actors
| Who (Plain English) | Technical Term |
| ---------------------------------------------- | --------------- |
| A single unit of work being tracked | Span |
| The complete journey of a request | Trace |
| Data that links spans across services | Trace Context |
| Code that creates spans and propagates context | Instrumentation |
| Service that receives and processes traces | Collector |
| Storage and visualization system | Backend (Tempo) |
| Decision logic for which traces to keep | Sampler |
### Actions
| What Happens (Plain English) | Technical Term |
| --------------------------------------- | ----------------------- |
| Start tracking a new operation | Create a Span |
| Connect a child operation to its parent | Set `parent_span_id` |
| Group all related operations together | Share a `trace_id` |
| Pass tracking data between services | Context Propagation |
| Decide whether to record a trace | Sampling (Head or Tail) |
| Send completed traces to storage | Export (OTLP) |
---
## Core Concepts
### 1. Trace
A **trace** represents the entire journey of a request through the system. It has a unique `trace_id` that stays constant across all nodes.
```
Trace ID: abc123
├── Node A: received transaction
├── Node B: relayed transaction
├── Node C: included in consensus
└── Node D: applied to ledger
```
### 2. Span
A **span** represents a single unit of work within a trace. Each span has:
| Attribute | Description | Example |
| ---------------- | -------------------------------- | -------------------------- |
| `trace_id` | Identifies the trace | `event123` |
| `span_id` | Unique identifier | `span456` |
| `parent_span_id` | Parent span (if any) | `p_span123` |
| `name` | Operation name | `rpc.submit` |
| `start_time` | When work began (local time) | `2024-01-15T10:30:00Z` |
| `end_time` | When work completed (local time) | `2024-01-15T10:30:00.050Z` |
| `attributes` | Key-value metadata | `tx_hash=ABC...` |
| `status` | OK, ERROR MSG | `OK` |
### 3. Trace Context
**Trace context** is the data that propagates between services to link spans together. It contains:
- `trace_id` - The trace this span belongs to
- `span_id` - The current span (becomes parent for child spans)
- `trace_flags` - Sampling decisions
---
## How Spans Form a Trace
Spans have parent-child relationships forming a tree structure:
```mermaid
flowchart TB
subgraph trace["Trace: abc123"]
A["tx.submit<br/>span_id: 001<br/>50ms"] --> B["tx.validate<br/>span_id: 002<br/>5ms"]
A --> C["tx.relay<br/>span_id: 003<br/>10ms"]
A --> D["tx.apply<br/>span_id: 004<br/>30ms"]
D --> E["ledger.update<br/>span_id: 005<br/>20ms"]
end
style A fill:#0d47a1,stroke:#082f6a,color:#ffffff
style B fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style C fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style D fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style E fill:#bf360c,stroke:#8c2809,color:#ffffff
```
**Reading the diagram:**
- **tx.submit (blue, root)**: The top-level span representing the entire transaction submission; all other spans are its descendants.
- **tx.validate, tx.relay, tx.apply (green)**: Direct children of tx.submit, representing the three main stages -- validation, relay to peers, and application to the ledger.
- **ledger.update (red)**: A grandchild span nested under tx.apply, representing the actual ledger state mutation triggered by applying the transaction.
- **Arrows (parent to child)**: Each arrow indicates a parent-child span relationship where the parent's completion depends on the child finishing.
The same trace visualized as a **timeline (Gantt chart)**:
```
Time → 0ms 10ms 20ms 30ms 40ms 50ms
├───────────────────────────────────────────┤
tx.submit│▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓│
├─────┤
tx.valid │▓▓▓▓▓│
│ ├──────────┤
tx.relay │ │▓▓▓▓▓▓▓▓▓▓│
│ ├────────────────────────────┤
tx.apply │ │▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓│
│ ├──────────────────┤
ledger │ │▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓│
```
---
## Span Relationships
Spans don't always form simple parent-child trees. Distributed tracing defines several relationship types to capture different causal patterns:
### 1. Parent-Child (ChildOf)
The default relationship. The parent span **depends on** or **contains** the child span. The child runs within the scope of the parent.
```
tx.submit (parent)
├── tx.validate (child) ← parent waits for this
├── tx.relay (child) ← parent waits for this
└── tx.apply (child) ← parent waits for this
```
**When to use:** Synchronous calls, nested operations, any case where the parent's completion depends on the child.
### 2. Follows-From
A causal relationship where the first span **triggers** the second, but does **not wait** for it. The originator fires and moves on.
```
Time →
tx.receive [=======]
↓ triggers (follows-from)
tx.relay [===========] ← runs independently
```
**When to use:** Asynchronous jobs, queued work, fire-and-forget patterns. For example, a node receives a transaction and queues it for relay — the relay span _follows from_ the receive span but the receiver doesn't wait for relaying to complete.
> **OpenTracing** defined `FollowsFrom` as a first-class reference type alongside `ChildOf`.
> **OpenTelemetry** represents this using **Span Links** with descriptive attributes instead (see below).
### 3. Span Links (Cross-Trace and Non-Hierarchical)
Links connect spans that are **causally related but not in a parent-child hierarchy**. Unlike parent-child, links can cross trace boundaries.
```
Trace A Trace B
────── ──────
batch.schedule batch.execute
├─ item.enqueue (span X) ┌──► process.item
├─ item.enqueue (span Y) ───┤ (links to X, Y, Z)
├─ item.enqueue (span Z) └──►
```
**Use cases:**
| Pattern | Description |
| -------------------- | --------------------------------------------------------------------------- |
| **Batch processing** | A batch span links back to all individual spans that contributed to it |
| **Fan-in** | An aggregation span links to the multiple producer spans it merges |
| **Fan-out** | Multiple downstream spans link back to the single span that triggered them |
| **Async handoff** | A deferred job links back to the request that queued it (follows-from) |
| **Cross-trace** | Correlating spans across independent traces (e.g., retries, related events) |
**Link structure:** Each link carries the target span's context plus optional attributes:
```
Link {
trace_id: <target trace>
span_id: <target span>
attributes: { "link.description": "triggered by batch scheduler" }
}
```
### Relationship Summary
```mermaid
flowchart LR
subgraph parent_child["Parent-Child"]
direction TB
P["Parent"] --> C["Child"]
end
subgraph follows_from["Follows-From"]
direction TB
A["Span A"] -.->|triggers| B["Span B"]
end
subgraph links["Span Links"]
direction TB
X["`Span X
(Trace 1)`"] -.-|link| Y["`Span Y
(Trace 2)`"]
end
parent_child ~~~ follows_from ~~~ links
style P fill:#0d47a1,stroke:#082f6a,color:#ffffff
style C fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style A fill:#0d47a1,stroke:#082f6a,color:#ffffff
style B fill:#bf360c,stroke:#8c2809,color:#ffffff
style X fill:#4a148c,stroke:#38006b,color:#ffffff
style Y fill:#4a148c,stroke:#38006b,color:#ffffff
```
| Relationship | Same Trace? | Dependency? | OTel Mechanism |
| ---------------- | ----------- | -------------------------- | ----------------- |
| **Parent-Child** | Yes | Parent depends on child | `parent_span_id` |
| **Follows-From** | Usually | Causal but no dependency | Link + attributes |
| **Span Link** | Either | Correlation, no dependency | Link + attributes |
---
## Trace ID Generation
A `trace_id` is a 128-bit (16-byte) identifier that groups all spans belonging to one logical operation. How it's generated determines how easily you can find and correlate traces later.
### General Approaches
#### 1. Random (W3C Default)
Generate a random 128-bit ID when a trace starts. Standard approach for most services.
```
trace_id = random_128_bits()
```
| Pros | Cons |
| --------------------------- | --------------------------------------------- |
| Simple, standard | No natural correlation to domain events |
| Guaranteed unique per trace | If propagation is lost, trace is broken |
| Works with all OTel tooling | "Find trace for TX abc" requires index lookup |
#### 2. Deterministic (Derived from Domain Data)
Compute the trace_id from a hash of a natural identifier. Every node independently derives the **same** trace_id for the same event.
```
trace_id = SHA-256(domain_identifier)[0:16] // truncate to 128 bits
```
| Pros | Cons |
| --------------------------------------------------- | ---------------------------------------------------------- |
| Propagation-resilient — same ID computed everywhere | Same event processed twice (retry) shares trace_id |
| Natural search — domain ID maps directly to trace | Non-standard (tooling assumes random) |
| No coordination needed between nodes | 256→128 bit truncation (collision risk negligible at ~2⁶⁴) |
#### 3. Hybrid (Deterministic Prefix + Random Suffix)
First 8 bytes derived from domain data, last 8 bytes random.
```
trace_id = SHA-256(domain_identifier)[0:8] || random_64_bits()
```
| Pros | Cons |
| ------------------------------------------- | ---------------------------------------- |
| Prefix search: "find all traces for TX abc" | Must propagate to maintain full trace_id |
| Unique per processing instance | More complex generation logic |
| Retries get distinct trace_ids | Partial correlation only (prefix match) |
### XRPL Workflow Analysis
XRPL has a unique advantage: its core workflows produce **globally unique 256-bit hashes** that are known on every node. This makes deterministic trace_id generation practical in ways most systems can't achieve.
#### Natural Identifiers by Workflow
| Workflow | Natural Identifier | Size | Known at Start? | Same on All Nodes? |
| ------------------- | --------------------------------- | ---------- | ----------------------------- | -------------------------------- |
| **Transaction** | Transaction hash (`tid_`) | 256-bit | Yes — computed before signing | Yes — hash of canonical tx data |
| **Consensus round** | Previous ledger hash + ledger seq | 256+32 bit | Yes — known when round opens | Yes — all validators agree |
| **Validation** | Ledger hash being validated | 256-bit | Yes — from consensus result | Yes — same closed ledger |
| **Ledger catch-up** | Target ledger hash | 256-bit | Yes — we know what to fetch | Yes — identifies ledger globally |
#### Where These Identifiers Live in Code
```
Transaction: STTx::getTransactionID() → uint256 tid_
TMTransaction::rawTransaction → recompute hash from bytes
Consensus: ConsensusProposal::previousLedger_ → uint256 (previous ledger hash)
ConsensusProposal::position_ → uint256 (TxSet hash)
LedgerHeader::seq → uint32_t (ledger sequence)
Validation: STValidation::getLedgerHash() → uint256
STValidation::getNodeID() → NodeID (160-bit)
Ledger fetch: InboundLedger constructor → uint256 hash, uint32_t seq
TMGetLedger::ledgerHash → bytes (uint256)
```
### Recommended Strategy: Workflow-Scoped Deterministic
Each workflow type derives its trace_id from its natural domain identifier:
```
Transaction trace: trace_id = SHA-256("tx" || tx_hash)[0:16]
Consensus trace: trace_id = SHA-256("cons" || prev_ledger_hash || ledger_seq)[0:16]
Ledger catch-up: trace_id = SHA-256("fetch" || target_ledger_hash)[0:16]
```
The string prefix (`"tx"`, `"cons"`, `"fetch"`) prevents collisions between workflows that might share underlying hashes.
**Why this works for XRPL:**
1. **Propagation-resilient** — Even if a P2P message drops trace context, every node independently computes the same trace_id from the same tx_hash or ledger_hash. Spans still correlate.
2. **Zero-cost search** — "Show me the trace for transaction ABC" becomes a direct lookup: compute `SHA-256("tx" || ABC)[0:16]` and query. No secondary index needed.
3. **Cross-workflow linking via Span Links** — A consensus trace links to individual transaction traces. A validation span links to the consensus trace. This connects the full picture without forcing everything into one giant trace.
### Cross-Workflow Correlation
Each workflow gets its own trace. Span Links tie them together:
```mermaid
flowchart TB
subgraph tx_trace["Transaction Trace"]
direction LR
Tn["trace_id = f(tx_hash)"]:::note --> T1["tx.receive"] --> T2["tx.validate"] --> T3["tx.relay"]
end
subgraph cons_trace["Consensus Trace"]
direction LR
Cn["trace_id = f(prev_ledger, seq)"]:::note --> C1["cons.open"] --> C2["cons.propose"] --> C3["cons.accept"]
end
subgraph val_trace["Validation"]
direction LR
Vn["spans within consensus trace"]:::note --> V1["val.create"] --> V2["val.broadcast"]
end
subgraph fetch_trace["Catch-Up Trace"]
direction LR
Fn["trace_id = f(ledger_hash)"]:::note --> F1["fetch.request"] --> F2["fetch.receive"] --> F3["fetch.apply"]
end
C1 -.-|"`span link
(tx traces)`"| T3
C3 --> V1
F1 -.-|"`span link
(target ledger)`"| C3
classDef note fill:none,stroke:#888,stroke-dasharray:5 5,color:#333,font-style:italic
style T1 fill:#0d47a1,stroke:#082f6a,color:#ffffff
style T2 fill:#0d47a1,stroke:#082f6a,color:#ffffff
style T3 fill:#0d47a1,stroke:#082f6a,color:#ffffff
style C1 fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style C2 fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style C3 fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style V1 fill:#bf360c,stroke:#8c2809,color:#ffffff
style V2 fill:#bf360c,stroke:#8c2809,color:#ffffff
style F1 fill:#4a148c,stroke:#38006b,color:#ffffff
style F2 fill:#4a148c,stroke:#38006b,color:#ffffff
style F3 fill:#4a148c,stroke:#38006b,color:#ffffff
```
**Reading the diagram:**
- **Transaction Trace (blue)**: An independent trace whose `trace_id` is deterministically derived from the transaction hash. Contains receive, validate, and relay spans.
- **Consensus Trace (green)**: An independent trace whose `trace_id` is derived from the previous ledger hash and sequence number. Covers the open, propose, and accept phases.
- **Validation (red)**: Validation spans live within the consensus trace (not a separate trace). They are created after the accept phase completes.
- **Catch-Up Trace (purple)**: An independent trace for ledger acquisition, derived from the target ledger hash. Used when a node is behind and fetching missing ledgers.
- **Dotted arrows (span links)**: Cross-trace correlations. Consensus links to transaction traces it included; catch-up links to the consensus trace that produced the target ledger.
- **Solid arrow (C3 to V1)**: A parent-child relationship -- validation spans are direct children of the consensus accept span within the same trace.
**How a query flows:**
```
"Why was TX abc slow?"
1. Compute trace_id = SHA-256("tx" || abc)[0:16]
2. Find transaction trace → see it was included in consensus round N
3. Follow span link → consensus trace for round N
4. See which phase was slow (propose? accept?)
5. If a node was catching up, follow link → catch-up trace
```
### Trade-offs to Consider
| Concern | Mitigation |
| ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| **Retries get same trace_id** | Add `attempt` attribute to root span; spans have unique span_ids and timestamps |
| **256→128 bit truncation** | Birthday-bound collision at ~2⁶⁴ operations — negligible for XRPL's throughput |
| **Non-standard generation** | OTel spec allows any 16-byte non-zero value; tooling works on the hex string |
| **Hash computation cost** | SHA-256 is ~0.3μs per call; XRPL already computes these hashes for other purposes |
| **Late-binding identifiers** | Ledger hash isn't known until after consensus — validation spans use ledger_seq as fallback, then link to the consensus trace |
---
## Distributed Traces Across Nodes
In distributed systems like xrpld, traces span **multiple independent nodes**. The trace context must be propagated in network messages:
```mermaid
sequenceDiagram
participant Client
participant NodeA as Node A
participant NodeB as Node B
participant NodeC as Node C
Client->>NodeA: Submit TX<br/>(no trace context)
Note over NodeA: Creates new trace<br/>trace_id: abc123<br/>span: tx.receive
NodeA->>NodeB: Relay TX<br/>(trace_id: abc123, parent: 001)
Note over NodeB: Creates child span<br/>span: tx.relay<br/>parent_span_id: 001
NodeA->>NodeC: Relay TX<br/>(trace_id: abc123, parent: 001)
Note over NodeC: Creates child span<br/>span: tx.relay<br/>parent_span_id: 001
Note over NodeA,NodeC: All spans share trace_id: abc123<br/>enabling correlation across nodes
```
**Reading the diagram:**
- **Client**: The external entity that submits a transaction. It does not carry trace context -- the trace originates at the first node.
- **Node A**: The entry point that creates a new trace (trace_id: abc123) and the root span `tx.receive`. It relays the transaction to peers with trace context attached.
- **Node B and Node C**: Peer nodes that receive the relayed transaction along with the propagated trace context. Each creates a child span under Node A's span, preserving the same `trace_id`.
- **Arrows with trace context**: The relay messages carry `trace_id` and `parent_span_id`, allowing each downstream node to link its spans back to the originating span on Node A.
---
## Context Propagation
For traces to work across nodes, **trace context must be propagated** in messages.
### What's in the Context (~26 bytes)
| Field | Size | Description |
| ------------- | -------- | ------------------------------------------------------- |
| `trace_id` | 16 bytes | Identifies the entire trace (constant across all nodes) |
| `span_id` | 8 bytes | The sender's current span (becomes parent on receiver) |
| `trace_flags` | 1 byte | Sampling decision (bit 0 = sampled; bits 1-7 reserved) |
| `trace_state` | variable | Optional vendor-specific data (typically omitted) |
### How span_id Changes at Each Hop
Only **one** `span_id` travels in the context - the sender's current span. Each node:
1. Extracts the received `span_id` and uses it as the `parent_span_id`
2. Creates a **new** `span_id` for its own span
3. Sends its own `span_id` as the parent when forwarding
```
Node A Node B Node C
────── ────── ──────
Span AAA Span BBB Span CCC
│ │ │
▼ ▼ ▼
Context out: Context out: Context out:
├─ trace_id: abc123 ├─ trace_id: abc123 ├─ trace_id: abc123
├─ span_id: AAA ──────────► ├─ span_id: BBB ──────────► ├─ span_id: CCC ──────►
└─ flags: 01 └─ flags: 01 └─ flags: 01
│ │
parent = AAA parent = BBB
```
The `trace_id` stays constant, but `span_id` **changes at every hop** to maintain the parent-child chain.
### Propagation Formats
There are two patterns:
### HTTP/RPC Headers (W3C Trace Context)
```
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
│ │ │ │
│ │ │ └── Flags (sampled)
│ │ └── Parent span ID (16 hex)
│ └── Trace ID (32 hex)
└── Version
```
### Protocol Buffers (xrpld P2P messages)
xrpld P2P messages such as `TMTransaction` carry the trace context in two added byte fields alongside the existing payload: `trace_parent` holds the W3C traceparent (`trace_id`, `span_id`, and `trace_flags`), and `trace_state` holds the optional W3C tracestate. Together they propagate the trace across the P2P boundary so a receiving node can attach its spans to the sender's span.
---
## Sampling
Not every trace needs to be recorded. **Sampling** reduces overhead:
### Head Sampling (at trace start)
```
Request arrives → Random N% chance → Record or skip entire trace
```
- ✅ Low overhead
- ❌ May miss interesting traces
> **xrpld note**: xrpld intentionally fixes head sampling at 100% (sample
> everything) and does not expose a configurable ratio. A per-node ratio
> would let different nodes make divergent keep/drop decisions for the same
> distributed trace, producing broken/partial traces. A span with a remote
> parent is decided by the same trace-id ratio sampler, not by the peer's
> sampled flag, so nodes agree on every trace. Volume reduction is
> delegated to collector-side tail sampling.
### Tail Sampling (after trace completes)
```
Trace completes → Collector evaluates:
- Error? → KEEP
- Slow? → KEEP
- Normal? → Sample 10%
```
- ✅ Never loses important traces
- ❌ Higher memory usage at collector
---
## Key Benefits for xrpld
| Challenge | How Tracing Helps |
| ---------------------------------- | ---------------------------------------- |
| "Where is my transaction?" | Follow trace across all nodes it touched |
| "Why was consensus slow?" | See timing breakdown of each phase |
| "Which node is the bottleneck?" | Compare span durations across nodes |
| "What happened during the outage?" | Correlate errors across the network |
---
## Glossary
| Term | Definition |
| -------------------- | ------------------------------------------------------------------- |
| **Trace** | Complete journey of a request, identified by `trace_id` |
| **Span** | Single operation within a trace |
| **Parent-Child** | Span relationship where the parent depends on the child |
| **Follows-From** | Causal relationship where originator doesn't wait for the result |
| **Span Link** | Non-hierarchical connection between spans, possibly across traces |
| **Deterministic ID** | Trace ID derived from domain data (e.g., tx_hash) instead of random |
| **Context** | Data propagated between services (`trace_id`, `span_id`, flags) |
| **Instrumentation** | Code that creates spans and propagates context |
| **Collector** | Service that receives, processes, and exports traces |
| **Backend** | Storage/visualization system (Tempo) |
| **Head Sampling** | Sampling decision at trace start |
| **Tail Sampling** | Sampling decision after trace completes |
---
_Next: [Architecture Analysis](./01-architecture-analysis.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,467 +0,0 @@
# Architecture Analysis
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Related**: [Design Decisions](./02-design-decisions.md) | [Implementation Strategy](./03-implementation-strategy.md)
---
## 1.1 Current xrpld Architecture Overview
> **WS** = WebSocket | **UNL** = Unique Node List | **TxQ** = Transaction Queue | **StatsD** = Statistics Daemon
The xrpld node software consists of several interconnected components that need instrumentation for distributed tracing:
```mermaid
flowchart TB
subgraph xrpld["xrpld Node"]
subgraph services["Core Services"]
RPC["RPC Server<br/>(HTTP/WS/gRPC)"]
Overlay["Overlay<br/>(P2P Network)"]
Consensus["Consensus<br/>(RCLConsensus)"]
ValidatorList["ValidatorList<br/>(UNL Mgmt)"]
end
JobQueue["JobQueue<br/>(Thread Pool)"]
subgraph processing["Processing Layer"]
NetworkOPs["NetworkOPs<br/>(Tx Processing)"]
LedgerMaster["LedgerMaster<br/>(Ledger Mgmt)"]
NodeStore["NodeStore<br/>(Database)"]
InboundLedgers["InboundLedgers<br/>(Ledger Sync)"]
end
subgraph appservices["Application Services"]
PathFind["PathFinding<br/>(Payment Paths)"]
TxQ["TxQ<br/>(Fee Escalation)"]
LoadMgr["LoadManager<br/>(Fee/Load)"]
end
subgraph observability["Existing Observability"]
PerfLog["PerfLog<br/>(JSON)"]
Insight["Insight<br/>(StatsD)"]
Logging["Logging<br/>(Journal)"]
end
services --> JobQueue
JobQueue --> processing
JobQueue --> appservices
end
style xrpld fill:#424242,stroke:#212121,color:#ffffff
style services fill:#1565c0,stroke:#0d47a1,color:#ffffff
style processing fill:#2e7d32,stroke:#1b5e20,color:#ffffff
style appservices fill:#6a1b9a,stroke:#4a148c,color:#ffffff
style observability fill:#e65100,stroke:#bf360c,color:#ffffff
```
**Reading the diagram:**
- **Core Services (blue)**: The entry points into xrpld -- RPC Server handles client requests, Overlay manages peer-to-peer networking, Consensus drives agreement, and ValidatorList manages trusted validators.
- **JobQueue (center)**: The asynchronous thread pool that decouples Core Services from the Processing and Application layers. All work flows through it.
- **Processing Layer (green)**: Core business logic -- NetworkOPs processes transactions, LedgerMaster manages ledger state, NodeStore handles persistence, and InboundLedgers synchronizes missing data.
- **Application Services (purple)**: Higher-level features -- PathFinding computes payment routes, TxQ manages fee-based queuing, and LoadManager tracks server load.
- **Existing Observability (orange)**: The current monitoring stack (PerfLog, Insight, Journal logging) that OpenTelemetry will complement, not replace.
- **Arrows (Services to JobQueue to layers)**: Work originates at Core Services, is enqueued onto the JobQueue, and dispatched to Processing or Application layers for execution.
---
## 1.1.1 Actors and Actions
### Actors
| Who (Plain English) | Technical Term |
| ----------------------------------------- | -------------------------- |
| Network node running XRPL software | xrpld node |
| External client submitting requests | RPC Client |
| Network neighbor sharing data | Peer (PeerImp) |
| Request handler for client queries | RPC Server (ServerHandler) |
| Command executor for specific RPC methods | RPCHandler |
| Agreement process between nodes | Consensus (RCLConsensus) |
| Transaction processing coordinator | NetworkOPs |
| Background task scheduler | JobQueue |
| Ledger state manager | LedgerMaster |
| Payment route calculator | PathFinding (Pathfinder) |
| Transaction waiting room | TxQ (Transaction Queue) |
| Fee adjustment system | LoadManager |
| Trusted validator list manager | ValidatorList |
| Protocol upgrade tracker | AmendmentTable |
| Ledger state hash tree | SHAMap |
| Persistent key-value storage | NodeStore |
### Actions
| What Happens (Plain English) | Technical Term |
| ---------------------------------------------- | ---------------------- |
| Client sends a request to a node | `rpc.request` |
| Node executes a specific RPC command | `rpc.command.*` |
| Node receives a transaction from a peer | `tx.receive` |
| Node checks if a transaction is valid | `tx.validate` |
| Node forwards a transaction to neighbors | `tx.relay` |
| Nodes agree on which transactions to include | `consensus.round` |
| Consensus progresses through phases | `consensus.phase.*` |
| Node builds a new confirmed ledger | `ledger.build` |
| Node fetches missing ledger data from peers | `ledger.acquire` |
| Node computes payment routes | `pathfind.compute` |
| Node queues a transaction for later processing | `txq.enqueue` |
| Node increases fees due to high load | `fee.escalate` |
| Node fetches the latest trusted validator list | `validator.list.fetch` |
| Node votes on a protocol amendment | `amendment.vote` |
| Node synchronizes state tree data | `shamap.sync` |
---
## 1.2 Key Components for Instrumentation
> **TxQ** = Transaction Queue | **UNL** = Unique Node List
| Component | Location | Purpose | Trace Value |
| ------------------ | ------------------------------------------ | ------------------------ | -------------------------------- |
| **Overlay** | `src/xrpld/overlay/` | P2P communication | Message propagation timing |
| **PeerImp** | `src/xrpld/overlay/detail/PeerImp.cpp` | Individual peer handling | Per-peer latency |
| **RCLConsensus** | `src/xrpld/app/consensus/RCLConsensus.cpp` | Consensus algorithm | Round timing, phase analysis |
| **NetworkOPs** | `src/xrpld/app/misc/NetworkOPs.cpp` | Transaction processing | Tx lifecycle tracking |
| **ServerHandler** | `src/xrpld/rpc/detail/ServerHandler.cpp` | RPC entry point | Request latency |
| **RPCHandler** | `src/xrpld/rpc/detail/RPCHandler.cpp` | Command execution | Per-command timing |
| **JobQueue** | `src/xrpl/core/JobQueue.h` | Async task execution | Queue wait times |
| **PathFinding** | `src/xrpld/app/paths/` | Payment path computation | Path latency, cache hits |
| **TxQ** | `src/xrpld/app/misc/TxQ.cpp` | Transaction queue/fees | Queue depth, eviction rates |
| **LoadManager** | `src/xrpld/app/main/LoadManager.cpp` | Fee escalation/load | Fee levels, load factors |
| **InboundLedgers** | `src/xrpld/app/ledger/InboundLedgers.cpp` | Ledger acquisition | Sync time, peer reliability |
| **ValidatorList** | `src/xrpld/app/misc/ValidatorList.cpp` | UNL management | List freshness, fetch failures |
| **AmendmentTable** | `src/xrpld/app/misc/AmendmentTable.cpp` | Protocol amendments | Voting status, activation events |
| **SHAMap** | `src/xrpld/shamap/` | State hash tree | Sync speed, missing nodes |
---
## 1.3 Transaction Flow Diagram
Transaction flow spans multiple nodes in the network. Each node creates linked spans to form a distributed trace:
```mermaid
sequenceDiagram
participant Client
participant PeerA as Peer A (Receive)
participant PeerB as Peer B (Relay)
participant PeerC as Peer C (Validate)
Client->>PeerA: 1. Submit TX
rect rgb(230, 245, 255)
Note over PeerA: tx.receive SPAN START
PeerA->>PeerA: HashRouter Deduplication
PeerA->>PeerA: tx.validate (child span)
end
PeerA->>PeerB: 2. Relay TX (with trace ctx)
rect rgb(230, 245, 255)
Note over PeerB: tx.receive (linked span)
end
PeerB->>PeerC: 3. Relay TX
rect rgb(230, 245, 255)
Note over PeerC: tx.receive (linked span)
PeerC->>PeerC: tx.process
end
Note over Client,PeerC: DISTRIBUTED TRACE (same trace_id: abc123)
```
**Reading the diagram:**
- **Client**: The external entity that submits a transaction to Peer A. It has no trace context -- the trace starts at the first node.
- **Peer A (Receive)**: The entry node that creates the root span `tx.receive`, runs HashRouter deduplication to avoid processing duplicates, and creates a child `tx.validate` span.
- **Peer A to Peer B arrow**: The relay message carries trace context (trace_id + parent span_id), enabling Peer B to create a linked span under the same trace.
- **Peer B (Relay)**: Receives the transaction and trace context, creates a `tx.receive` span linked to Peer A's trace, then relays onward.
- **Peer C (Validate)**: Final hop in this example. Creates a linked `tx.receive` span and runs `tx.process` to fully process the transaction.
- **Blue rectangles**: Highlight the span boundaries on each node, showing where instrumentation creates and closes spans.
### Trace Structure
```
trace_id: abc123
├── span: tx.receive (Peer A)
│ ├── span: tx.validate
│ └── span: tx.relay
├── span: tx.receive (Peer B) [parent: Peer A]
│ └── span: tx.relay
└── span: tx.receive (Peer C) [parent: Peer B]
└── span: tx.process
```
---
## 1.4 Consensus Round Flow
Consensus rounds are multi-phase operations that benefit significantly from tracing:
```mermaid
flowchart TB
subgraph round["consensus.round (root span)"]
attrs["Attributes:<br/>ledger_seq = 12345678<br/>consensus_mode = proposing<br/>proposers = 35"]
subgraph open["consensus.phase.open"]
open_desc["Duration: ~3s<br/>Waiting for transactions"]
end
subgraph establish["consensus.phase.establish"]
est_attrs["proposals_received = 28<br/>disputes_resolved = 3"]
est_children["├── consensus.proposal.receive (×28)<br/>├── consensus.proposal.send (×1)<br/>└── consensus.dispute.resolve (×3)"]
end
subgraph accept["consensus.phase.accept"]
acc_attrs["transactions_applied = 150<br/>ledger_hash = DEF456..."]
acc_children["├── ledger.build<br/>└── ledger.validate"]
end
attrs --> open
open --> establish
establish --> accept
end
style round fill:#f57f17,stroke:#e65100,color:#ffffff
style open fill:#1565c0,stroke:#0d47a1,color:#ffffff
style establish fill:#2e7d32,stroke:#1b5e20,color:#ffffff
style accept fill:#c2185b,stroke:#880e4f,color:#ffffff
```
**Reading the diagram:**
- **consensus.round (orange, root span)**: The top-level span encompassing the entire consensus round, with attributes like ledger sequence, mode, and proposer count.
- **consensus.phase.open (blue)**: The first phase where the node waits (~3s) to collect incoming transactions before proposing.
- **consensus.phase.establish (green)**: The negotiation phase where validators exchange proposals, resolve disputes, and converge on a transaction set. Child spans track each proposal received/sent and each dispute resolved.
- **consensus.phase.accept (pink)**: The final phase where the agreed transaction set is applied, a new ledger is built, and the ledger is validated. Child spans cover `ledger.build` and `ledger.validate`.
- **Arrows (open to establish to accept)**: The sequential flow through the three consensus phases. Each phase must complete before the next begins.
---
## 1.5 RPC Request Flow
> **WS** = WebSocket
RPC requests support W3C Trace Context headers for distributed tracing across services:
```mermaid
flowchart TB
subgraph request["rpc.request (root span)"]
http["HTTP Request — POST /<br/>traceparent:<br/>00-abc123...-def456...-01"]
attrs["Attributes:<br/>http.method = POST<br/>net.peer.ip = 192.168.1.100<br/>command = submit"]
subgraph enqueue["jobqueue.enqueue"]
job_attr["job_type = jtCLIENT_RPC"]
end
subgraph command["rpc.command.submit"]
cmd_attrs["version = 2<br/>rpc_role = user"]
cmd_children["├── tx.deserialize<br/>├── tx.validate_local<br/>└── tx.submit_to_network"]
end
response["Response: 200 OK<br/>Duration: 45ms"]
http --> attrs
attrs --> enqueue
enqueue --> command
command --> response
end
style request fill:#2e7d32,stroke:#1b5e20,color:#ffffff
style enqueue fill:#1565c0,stroke:#0d47a1,color:#ffffff
style command fill:#e65100,stroke:#bf360c,color:#ffffff
```
**Reading the diagram:**
- **rpc.request (green, root span)**: The outermost span representing the full RPC request lifecycle, from HTTP receipt to response. Carries the W3C `traceparent` header for distributed tracing.
- **HTTP Request node**: Shows the incoming POST request with its `traceparent` header and extracted attributes (method, peer IP, command name).
- **jobqueue.enqueue (blue)**: The span covering the asynchronous handoff from the RPC thread to the JobQueue worker thread. The trace context is preserved across this async boundary.
- **rpc.command.submit (orange)**: The span for the actual command execution, with child spans for deserialization, local validation, and network submission.
- **Response node**: The final output with HTTP status and total duration, marking the end of the root span.
- **Arrows (top to bottom)**: The sequential processing pipeline -- receive request, extract attributes, enqueue job, execute command, return response.
---
## 1.6 Key Trace Points
> **TxQ** = Transaction Queue
The following table identifies priority instrumentation points across the codebase:
| Category | Span Name | File | Method | Priority |
| --------------- | ---------------------- | ---------------------- | ----------------------- | -------- |
| **Transaction** | `tx.receive` | `PeerImp.cpp` | `handleTransaction()` | High |
| **Transaction** | `tx.validate` | `NetworkOPs.cpp` | `processTransaction()` | High |
| **Transaction** | `tx.process` | `NetworkOPs.cpp` | `doTransactionSync()` | High |
| **Transaction** | `tx.relay` | `OverlayImpl.cpp` | `relay()` | Medium |
| **Consensus** | `consensus.round` | `RCLConsensus.cpp` | `startRound()` | High |
| **Consensus** | `consensus.phase.*` | `Consensus.h` | `timerEntry()` | High |
| **Consensus** | `consensus.proposal.*` | `RCLConsensus.cpp` | `peerProposal()` | Medium |
| **RPC** | `rpc.request` | `ServerHandler.cpp` | `onRequest()` | High |
| **RPC** | `rpc.command.*` | `RPCHandler.cpp` | `doCommand()` | High |
| **Peer** | `peer.connect` | `OverlayImpl.cpp` | `onHandoff()` | Low |
| **Peer** | `peer.message.*` | `PeerImp.cpp` | `onMessage()` | Low |
| **Ledger** | `ledger.acquire` | `InboundLedgers.cpp` | `acquire()` | Medium |
| **Ledger** | `ledger.build` | `RCLConsensus.cpp` | `buildLCL()` | High |
| **PathFinding** | `pathfind.request` | `PathRequest.cpp` | `doUpdate()` | High |
| **PathFinding** | `pathfind.compute` | `Pathfinder.cpp` | `findPaths()` | High |
| **TxQ** | `txq.enqueue` | `TxQ.cpp` | `apply()` | High |
| **TxQ** | `txq.apply` | `TxQ.cpp` | `processClosedLedger()` | High |
| **Fee** | `fee.escalate` | `LoadManager.cpp` | `raiseLocalFee()` | Medium |
| **Ledger** | `ledger.replay` | `LedgerReplayer.h` | `replay()` | Medium |
| **Ledger** | `ledger.delta` | `LedgerDeltaAcquire.h` | `processData()` | Medium |
| **Validator** | `validator.list.fetch` | `ValidatorList.cpp` | `verify()` | Medium |
| **Validator** | `validator.manifest` | `Manifest.cpp` | `applyManifest()` | Low |
| **Amendment** | `amendment.vote` | `AmendmentTable.cpp` | `doVoting()` | Low |
| **SHAMap** | `shamap.sync` | `SHAMap.cpp` | `fetchRoot()` | Medium |
---
## 1.7 Instrumentation Priority
> **TxQ** = Transaction Queue
```mermaid
quadrantChart
title Instrumentation Priority Matrix
x-axis Low Complexity --> High Complexity
y-axis Low Value --> High Value
quadrant-1 Implement First
quadrant-2 Plan Carefully
quadrant-3 Quick Wins
quadrant-4 Consider Later
RPC Tracing: [0.2, 0.92]
Transaction Tracing: [0.55, 0.88]
Consensus Tracing: [0.78, 0.82]
PathFinding: [0.38, 0.75]
TxQ and Fees: [0.25, 0.65]
Ledger Sync: [0.62, 0.58]
Peer Message Tracing: [0.35, 0.25]
JobQueue Tracing: [0.2, 0.48]
Validator Mgmt: [0.48, 0.42]
Amendment Tracking: [0.15, 0.32]
SHAMap Operations: [0.72, 0.45]
```
---
## 1.8 Observable Outcomes
> **TxQ** = Transaction Queue | **UNL** = Unique Node List
After implementing OpenTelemetry, operators and developers will gain visibility into the following:
### 1.8.1 What You Will See: Traces
| Trace Type | Description | Example Query in Grafana/Tempo |
| -------------------------- | ------------------------------------------------------------------------------------------- | ----------------------------------------------- |
| **Transaction Lifecycle** | Full journey from RPC submission through validation, relay, consensus, and ledger inclusion | `{service.name="xrpld" && tx_hash="ABC123..."}` |
| **Cross-Node Propagation** | Transaction path across multiple xrpld nodes with timing | `{relay_count > 0}` |
| **Consensus Rounds** | Complete round with all phases (open, establish, accept) | `{span.name=~"consensus.round.*"}` |
| **RPC Request Processing** | Individual command execution with timing breakdown | `{command="account_info"}` |
| **Ledger Acquisition** | Peer-to-peer ledger data requests and responses | `{span.name="ledger.acquire"}` |
| **PathFinding Latency** | Path computation time and cache effectiveness for payment RPCs | `{span.name="pathfind.compute"}` |
| **TxQ Behavior** | Queue depth, eviction patterns, fee escalation during congestion | `{span.name=~"txq.*"}` |
| **Ledger Sync** | Full acquisition timeline including delta and transaction fetches | `{span.name=~"ledger.acquire.*"}` |
| **Validator Health** | UNL fetch success, manifest updates, stale list detection | `{span.name=~"validator.*"}` |
### 1.8.2 What You Will See: Metrics (Derived from Traces)
| Metric | Description | Dashboard Panel |
| ----------------------------- | --------------------------------------- | --------------------------- |
| **RPC Latency (p50/p95/p99)** | Response time distribution per command | Heatmap by command |
| **Transaction Throughput** | Transactions processed per second | Time series graph |
| **Consensus Round Duration** | Time to complete consensus phases | Histogram |
| **Cross-Node Latency** | Time for transaction to reach N nodes | Line chart with percentiles |
| **Error Rate** | Failed transactions/RPC calls by type | Stacked bar chart |
| **PathFinding Latency** | Path computation time per currency pair | Heatmap by currency |
| **TxQ Depth** | Queued transactions over time | Time series with thresholds |
| **Fee Escalation Level** | Current fee multiplier | Gauge with alert thresholds |
| **Ledger Sync Duration** | Time to acquire missing ledgers | Histogram |
### 1.8.3 Concrete Dashboard Examples
**Transaction Trace View (Tempo):**
```
┌────────────────────────────────────────────────────────────────────────────────┐
│ Trace: abc123... (Transaction Submission) Duration: 847ms │
├────────────────────────────────────────────────────────────────────────────────┤
│ ├── rpc.request [ServerHandler] ████░░░░░░ 45ms │
│ │ └── rpc.command.submit [RPCHandler] ████░░░░░░ 42ms │
│ │ └── tx.receive [NetworkOPs] ███░░░░░░░ 35ms │
│ │ ├── tx.validate [TxQ] █░░░░░░░░░ 8ms │
│ │ └── tx.relay [Overlay] ██░░░░░░░░ 15ms │
│ │ ├── tx.receive [Node-B] █████░░░░░ 52ms │
│ │ │ └── tx.relay [Node-B] ██░░░░░░░░ 18ms │
│ │ └── tx.receive [Node-C] ██████░░░░ 65ms │
│ └── consensus.round [RCLConsensus] ████████░░ 720ms │
│ ├── consensus.phase.open ██░░░░░░░░ 180ms │
│ ├── consensus.phase.establish █████░░░░░ 480ms │
│ └── consensus.phase.accept █░░░░░░░░░ 60ms │
└────────────────────────────────────────────────────────────────────────────────┘
```
**RPC Performance Dashboard Panel:**
```
┌─────────────────────────────────────────────────────────────┐
│ RPC Command Latency (Last 1 Hour) │
├─────────────────────────────────────────────────────────────┤
│ Command │ p50 │ p95 │ p99 │ Errors │ Rate │
│──────────────────┼────────┼────────┼────────┼────────┼──────│
│ account_info │ 12ms │ 45ms │ 89ms │ 0.1% │ 150/s│
│ submit │ 35ms │ 120ms │ 250ms │ 2.3% │ 45/s│
│ ledger │ 8ms │ 25ms │ 55ms │ 0.0% │ 80/s│
│ tx │ 15ms │ 50ms │ 100ms │ 0.5% │ 60/s│
│ server_info │ 5ms │ 12ms │ 20ms │ 0.0% │ 200/s│
└─────────────────────────────────────────────────────────────┘
```
**Consensus Health Dashboard Panel:**
```mermaid
---
config:
xyChart:
width: 1200
height: 400
plotReservedSpacePercent: 50
chartOrientation: vertical
themeVariables:
xyChart:
plotColorPalette: "#3498db"
---
xychart-beta
title "Consensus Round Duration (Last 24 Hours)"
x-axis "Time of Day (Hours)" [0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24]
y-axis "Duration (seconds)" 1 --> 5
line [2.1, 2.4, 2.8, 3.2, 3.8, 4.3, 4.5, 5.0, 4.7, 4.0, 3.2, 2.6, 2.0]
```
### 1.8.4 Operator Actionable Insights
| Scenario | What You'll See | Action |
| ------------------------- | ---------------------------------------------------------------------------- | ------------------------------------------------ |
| **Slow RPC** | Span showing which phase is slow (parsing, execution, serialization) | Optimize specific code path |
| **Transaction Stuck** | Trace stops at validation; error attribute shows reason | Fix transaction parameters |
| **Consensus Delay** | Phase.establish taking too long; proposer attribute shows missing validators | Investigate network connectivity |
| **Memory Spike** | Large batch of spans correlating with memory increase | Tune batch_size or sampling |
| **Network Partition** | Traces missing cross-node links for specific peer | Check peer connectivity |
| **Path Computation Slow** | pathfind.compute span shows high latency; cache miss rate in attributes | Warm the RippleLineCache, check order book depth |
| **TxQ Full** | txq.enqueue spans show evictions; fee.escalate spans increasing | Monitor fee levels, alert operators |
| **Ledger Sync Stalled** | ledger.acquire spans timing out; peer reliability attributes show issues | Check peer connectivity, add trusted peers |
| **UNL Stale** | validator.list.fetch spans failing; last_update attribute aging | Verify validator site URLs, check DNS |
### 1.8.5 Developer Debugging Workflow
1. **Find Transaction**: Query by `tx_hash` to get full trace
2. **Identify Bottleneck**: Look at span durations to find slowest component
3. **Check Attributes**: Review `validity`, `rpc_status` for errors
4. **Correlate Logs**: Use `trace_id` to find related PerfLog entries
5. **Compare Nodes**: Filter by `service.instance.id` to compare behavior across nodes
---
_Next: [Design Decisions](./02-design-decisions.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,700 +0,0 @@
# Design Decisions
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Related**: [Architecture Analysis](./01-architecture-analysis.md)
---
## 2.1 OpenTelemetry Components
> **OTLP** = OpenTelemetry Protocol
### 2.1.1 SDK Selection
**Primary Choice**: OpenTelemetry C++ SDK (`opentelemetry-cpp`)
| Component | Purpose | Required |
| --------------------------------------- | ---------------------- | ------------------------- |
| `opentelemetry-cpp::api` | Tracing API headers | Yes |
| `opentelemetry-cpp::sdk` | SDK implementation | Yes |
| `opentelemetry-cpp::ext` | Extensions (exporters) | Yes |
| `opentelemetry-cpp::otlp_http_exporter` | OTLP/HTTP export | Yes (shipped in Phase 1b) |
| `opentelemetry-cpp::otlp_grpc_exporter` | OTLP/gRPC export | Future (not yet wired up) |
### 2.1.2 Instrumentation Strategy
**Manual Instrumentation** (recommended):
| Approach | Pros | Cons |
| ---------- | --------------------------------------------------------------- | ------------------------------------------------------- |
| **Manual** | Precise control, optimized placement, xrpld-specific attributes | More development effort |
| **Auto** | Less code, automatic coverage | Less control, potential overhead, limited customization |
---
## 2.2 Exporter Configuration
> **OTLP** = OpenTelemetry Protocol
```mermaid
flowchart TB
subgraph nodes["xrpld Nodes"]
node1["xrpld<br/>Node 1"]
node2["xrpld<br/>Node 2"]
node3["xrpld<br/>Node 3"]
end
collector["OpenTelemetry<br/>Collector<br/>(sidecar or standalone)"]
subgraph backends["Observability Backends"]
tempo["Tempo"]
elastic["Elastic<br/>APM"]
end
node1 -->|"OTLP/HTTP<br/>:4318"| collector
node2 -->|"OTLP/HTTP<br/>:4318"| collector
node3 -->|"OTLP/HTTP<br/>:4318"| collector
collector --> tempo
collector --> elastic
style nodes fill:#0d47a1,stroke:#082f6a,color:#ffffff
style backends fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style collector fill:#bf360c,stroke:#8c2809,color:#ffffff
```
**Reading the diagram:**
- **xrpld Nodes (blue)**: The source of telemetry data. Each xrpld node exports spans via OTLP/HTTP on port 4318 (the only exporter shipped in Phase 1b).
- **OpenTelemetry Collector (red)**: The central aggregation point that receives spans from all nodes. Can run as a sidecar (per-node) or standalone (shared). Handles batching, filtering, and routing.
- **Observability Backends (green)**: The storage and visualization destinations. Tempo is the recommended backend for both development and production, and Elastic APM is an alternative. The Collector routes to one or more backends.
- **Arrows (nodes to collector to backends)**: The data pipeline -- spans flow from nodes to the Collector over HTTP, then the Collector fans out to the configured backends.
### 2.2.1 OTLP/HTTP (Shipped in Phase 1b)
OTLP/HTTP is the only exporter wired up in Phase 1b. It is configured via
`OtlpHttpExporterOptions` with the collector traces endpoint
(`http://localhost:4318/v1/traces` by default) and a JSON content type
(binary protobuf is also available).
### 2.2.2 OTLP/gRPC (Future Work — Planned Upgrade)
OTLP/gRPC is planned as a future upgrade from the HTTP exporter. The gRPC
transport offers lower per-span overhead and tighter back-pressure semantics
than HTTP/JSON, making it attractive for production deployments once the HTTP
path is validated in earlier phases.
Required to land this upgrade:
1. Add `opentelemetry-cpp::otlp_grpc_exporter` to the Conan recipe (the
dependency already exists but is not linked in Phase 1b builds).
2. Extend `TelemetryConfig.cpp` to parse an `exporter` key (`otlp_http`
default, `otlp_grpc` opt-in) and a gRPC endpoint override.
3. In `Telemetry::start()` branch on the parsed exporter type and construct
either `OtlpHttpExporterFactory::Create(httpOpts)` or
`OtlpGrpcExporterFactory::Create(grpcOpts)` accordingly.
4. Update the runbook and dashboards to document the alternate port and TLS
settings.
When wired up, the gRPC path will use `OtlpGrpcExporterOptions` configured with
the collector endpoint (host on port 4317), TLS credentials enabled, and a CA
certificate path.
Until that work lands, `OtlpGrpcExporterOptions` is **not** used by any code
path in Phase 1b through Phase 5.
---
## 2.3 Span Naming Conventions
> **TxQ** = Transaction Queue | **UNL** = Unique Node List | **WS** = WebSocket
### 2.3.1 Naming Schema
```
<component>.<operation>[.<sub-operation>]
```
**Examples**:
- `tx.receive` - Transaction received from peer
- `consensus.phase.establish` - Consensus establish phase
- `rpc.command.server_info` - server_info RPC command
### 2.3.2 Complete Span Catalog
| Span name | Description |
| ------------------------------ | --------------------------------------- |
| `tx.receive` | Transaction received from network |
| `tx.validate` | Transaction signature/format validation |
| `tx.process` | Full transaction processing |
| `tx.relay` | Transaction relay to peers |
| `tx.apply` | Apply transaction to ledger |
| `consensus.round` | Complete consensus round |
| `consensus.phase.open` | Open phase - collecting transactions |
| `consensus.phase.establish` | Establish phase - reaching agreement |
| `consensus.phase.accept` | Accept phase - applying consensus |
| `consensus.proposal.receive` | Receive peer proposal |
| `consensus.proposal.send` | Send our proposal |
| `consensus.validation.receive` | Receive peer validation |
| `consensus.validation.send` | Send our validation |
| `rpc.request` | HTTP/WebSocket request handling |
| `rpc.command.*` | Specific RPC command (dynamic) |
| `peer.connect` | Peer connection establishment |
| `peer.disconnect` | Peer disconnection |
| `peer.message.send` | Send protocol message |
| `peer.message.receive` | Receive protocol message |
| `ledger.acquire` | Ledger acquisition from network |
| `ledger.build` | Build new ledger |
| `ledger.validate` | Ledger validation |
| `ledger.close` | Close ledger |
| `ledger.replay` | Ledger replay executed |
| `ledger.delta` | Delta-based ledger acquired |
| `pathfind.request` | Path request initiated |
| `pathfind.compute` | Path computation executed |
| `txq.enqueue` | Transaction queued |
| `txq.apply` | Queued transaction applied |
| `fee.escalate` | Fee escalation triggered |
| `validator.list.fetch` | UNL list fetched |
| `validator.manifest` | Manifest update processed |
| `amendment.vote` | Amendment voting executed |
| `shamap.sync` | State tree synchronization |
| `job.enqueue` | Job added to queue |
| `job.execute` | Job execution |
### 2.3.3 Attribute Naming Conventions
Span **names** follow §2.3.1 (dotted `<component>.<operation>`). Span
**attribute keys** follow the rules below. The constants in the `*SpanNames.h`
headers are the single source of truth; the collector, Tempo, the Grafana
dashboards, and the runbook all consume these exact keys, so every layer must
agree with the code. A CI check enforces this end to end.
1. **Per-span unique attribute** → bare field name, allowed when the field is
recorded by a single span/workflow so the span name already supplies the
domain (e.g. `command`, `version`, `local` on `rpc.command`).
2. **Shared attribute (same concept on more than one span)** → ONE key, reused
verbatim on every span that records it; the span name tells the occurrences
apart, so no per-emitter prefix is added. Name it by the field's meaning: a
property of a domain object keeps that object's bare field name (`ledger_hash`,
`ledger_seq`, `tx_hash`, `peer_id`, `full_validation`); a field already
qualified by a sub-kind keeps that qualifier on every emitter (`proposal_trusted`
on both `consensus.proposal.receive` and `peer.proposal.receive`;
`validation_trusted` likewise). Defined once in the base `SpanNames.h`
`namespace attr` block and re-exported (`using`) by each domain header.
3. **Collision qualifier** → `<domain>_<field>`, only when a bare name would
collide with a DIFFERENT concept in the shared spanmetrics label space or with
the OTel-reserved `status` key (e.g. `rpc_status`, `grpc_status`,
`consensus_phase`, `consensus_round`, `consensus_mode`). This disambiguates
distinct concepts that share a word; it is NOT used to tag the same concept
with its emitting workflow — that is rule 2 (one shared name).
4. **Resource attribute** → dotted `xrpl.<subsystem>.<field>`, reserved ONLY
for process/network identity set once at startup (`xrpl.network.id`,
`xrpl.network.type`). Span attributes are never dotted in the `xrpl.` form —
it blurs the resource/span scope boundary and parses awkwardly in TraceQL.
5. **Span names** use `<subsystem>[.<component>]` (dotted, per §2.3.1). Only
attribute _keys_ follow rules 1–4.
Standard OpenTelemetry semantic-convention keys keep their canonical dotted
form (e.g. `service.*` resource attributes, `http.*` span attributes); the
"no dotted form" rule applies to xrpl-custom keys only.
The same rules are recorded in `CONTRIBUTING.md` (the permanent home, since
`OpenTelemetryPlan/` is removed once the rollout completes). The attribute
examples in §2.4 below follow these rules.
---
## 2.4 Attribute Schema
> **TxQ** = Transaction Queue | **UNL** = Unique Node List | **OTLP** = OpenTelemetry Protocol
### 2.4.1 Resource Attributes (Set Once at Startup)
Resource attributes identify the process and are set once at startup. They use
the standard OpenTelemetry semantic conventions plus custom dotted `xrpl.*`
keys (the dotted form is reserved for resource scope per §2.3.3).
| Key | Type / value | Description |
| --------------------- | ------------------------------------------------------- | ------------------------------ |
| `service.name` | `"xrpld"` | Standard `SERVICE_NAME` |
| `service.version` | `build_info::getVersionString()` | Standard `SERVICE_VERSION` |
| `service.instance.id` | node public key (base58) | Standard `SERVICE_INSTANCE_ID` |
| `xrpl.network.id` | network id (e.g. 0 for mainnet) | Network identifier |
| `xrpl.network.type` | `"mainnet"` \| `"testnet"` \| `"devnet"` \| `"unknown"` | Network kind |
| `xrpl.node.type` | `"validator"` \| `"stock"` \| `"reporting"` | Node role |
| `xrpl.node.cluster` | cluster name | Cluster name, if clustered |
### 2.4.2 Span Attributes by Category
> Span attribute keys use the underscore form from §2.3.3 (shared/qualified
> keys are `<domain>_<field>`; per-span unique keys are bare). The dotted form
> is reserved for the resource attributes in §2.4.1 above. This catalog lists
> the planned attribute set by category; the exact emitted key for each
> implemented span is defined by the `*SpanNames.h` constants, which are the
> single source of truth where the two differ.
#### Transaction Attributes
| Key | Type | Description |
| -------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `tx_hash` | string | Transaction hash (hex) |
| `tx_type` | string | `"Payment"`, `"OfferCreate"`, etc. |
| `tx_account` | string | Sending account, raw r-address |
| `tx_<field>` | string | One per account-typed top-level field the transaction carries (`tx_destination`, `tx_owner`, `tx_issuer`, ...), raw r-address; keys in `TxAccountSpanNames.h` |
| `tx_sequence` | int64 | Account sequence number |
| `tx_fee` | int64 | Fee in drops |
| `tx_result` | string | `"tesSUCCESS"`, `"tecPATH_DRY"`, etc. |
| `current_ledger_seq` | int64 | Open ledger the transaction targeted |
| `relay_count` | int64 | Peers the transaction was relayed to |
> **Note:** `current_ledger_seq` and `ledger_seq` are the same concept — a ledger's sequence number — but they name different ledgers, so the design keeps two keys rather than one. `current_ledger_seq` is the open or in-flight ledger a transaction's work was applied into; it is named after the RPC field `ledger_current_index`. `ledger_seq` (see [Ledger & Job Attributes](#ledger--job-attributes)) is a closed or validated ledger, set by the ledger and consensus spans. Neither is spelled `ledger_index`: per rule 2 of [Telemetry span attribute naming](../CONTRIBUTING.md#telemetry-span-attribute-naming), one concept gets one key reused verbatim, and a different referent is disambiguated with a prefix rather than a synonym.
#### Consensus Attributes
| Key | Type | Description |
| -------------------- | ------- | ----------------------------------- |
| `consensus_round` | int64 | Round number |
| `consensus_phase` | string | `"open"`, `"establish"`, `"accept"` |
| `consensus_mode` | string | `"proposing"`, `"observing"`, etc. |
| `proposers` | int64 | Number of proposers |
| `prev_ledger_prefix` | string | Previous ledger hash prefix |
| `ledger_seq` | int64 | Ledger sequence |
| `tx_count` | int64 | Transactions in consensus set |
| `round_time_ms` | float64 | Round duration |
Establish-phase gap fill and cross-node correlation attributes (Phase 4a):
| Key | Type | Description |
| --------------------- | ------ | --------------------------------------------------------- |
| `consensus_round_id` | int64 | Consensus round number |
| `consensus_ledger_id` | string | `previousLedger.id()` — shared across nodes |
| `trace_strategy` | string | `"deterministic"` or `"random"` |
| `converge_percent` | int64 | Convergence % (0-100+) |
| `establish_count` | int64 | Number of establish iterations |
| `disputes_count` | int64 | Active disputed transactions |
| `agree_count` | int64 | Peers that agree (haveConsensus) |
| `disagree_count` | int64 | Peers that disagree |
| `threshold_percent` | int64 | Close-time consensus threshold (`avCT_CONSENSUS_PCT`=75%) |
| `consensus_result` | string | `"yes"`, `"no"`, `"moved_on"`, `"expired"` |
| `mode_old` | string | Previous consensus mode |
| `mode_new` | string | New consensus mode |
#### RPC Attributes
| Key | Type | Description |
| ------------- | ------- | ----------------------------------------------------------------------------- |
| `command` | string | Command name (per-span unique on `rpc.command`) |
| `version` | int64 | API version |
| `rpc_role` | string | `"admin"` or `"user"` (qualified — `role` is generic) |
| `params` | string | Sanitized parameters (optional) |
| `rpc_status` | string | Response status: `success` \| `error` (qualified — `status` is OTel-reserved) |
| `duration_ms` | float64 | Request duration in milliseconds |
#### Peer & Message Attributes
| Key | Type | Description |
| -------------------- | ------- | -------------------------- |
| `peer_id` | string | Peer public key (base58) |
| `peer_address` | string | IP:port |
| `peer_latency_ms` | float64 | Measured latency |
| `peer_cluster` | string | Cluster name if clustered |
| `message_type` | string | Protocol message type name |
| `message_size_bytes` | int64 | Message size |
| `message_compressed` | bool | Whether compressed |
#### Ledger & Job Attributes
| Key | Type | Description |
| --------------------------- | ------- | -------------------------------- |
| `ledger_hash` | string | Ledger hash |
| `ledger_seq` | int64 | Closed/validated ledger sequence |
| `close_time_ripple_epoch_s` | int64 | Close time (XRPL epoch seconds) |
| `ledger_tx_count` | int64 | Transaction count |
| `job_type` | string | Job type name |
| `job_queue_ms` | float64 | Time spent in queue |
| `job_worker` | int64 | Worker thread ID |
#### PathFinding Attributes
| Key | Type | Description |
| -------------------------- | ------ | ---------------------------------------------------------------------- |
| `pathfind_source_account` | string | Source r-address, raw |
| `pathfind_dest_account` | string | Destination r-address, raw |
| `pathfind_source_currency` | string | Source currency code |
| `pathfind_dest_currency` | string | Destination asset: `XRP`, `<issuer>/<currency>`, or an MPT issuance id |
| `pathfind_path_count` | int64 | Number of paths found |
| `pathfind_cache_hit` | bool | RippleLineCache hit |
#### TxQ Attributes
| Key | Type | Description |
| --------------------- | ------ | --------------------------- |
| `txq_queue_depth` | int64 | Current queue depth |
| `txq_fee_level` | int64 | Fee level of transaction |
| `txq_eviction_reason` | string | Why transaction was evicted |
#### Fee Attributes
| Key | Type | Description |
| ---------------------- | ----- | ------------------------- |
| `fee_load_factor` | int64 | Current load factor |
| `fee_escalation_level` | int64 | Fee escalation multiplier |
#### Validator Attributes
| Key | Type | Description |
| ------------------------ | ----- | ------------------------- |
| `validator_list_size` | int64 | UNL size |
| `validator_list_age_sec` | int64 | Seconds since last update |
#### Amendment Attributes
| Key | Type | Description |
| ------------------ | ------ | -------------------------------------- |
| `amendment_name` | string | Amendment name |
| `amendment_status` | string | `"enabled"`, `"vetoed"`, `"supported"` |
#### SHAMap Attributes
| Key | Type | Description |
| ---------------------- | ------- | --------------------------------------------- |
| `shamap_type` | string | `"transaction"`, `"state"`, `"account_state"` |
| `shamap_missing_nodes` | int64 | Number of missing nodes during sync |
| `shamap_duration_ms` | float64 | Sync duration |
### 2.4.3 Data Collection Summary
The following table summarizes what data is collected by category:
| Category | Attributes Collected | Purpose |
| --------------- | ---------------------------------------------------------------------------------------------------------------- | ---------------------------- |
| **Transaction** | `tx_hash`, `tx_type`, `tx_result`, `tx_fee`, `current_ledger_seq` | Trace transaction lifecycle |
| **Consensus** | `consensus_round`, `consensus_phase`, `consensus_mode`, `proposers`, `round_time_ms` | Analyze consensus timing |
| **RPC** | `command`, `version`, `rpc_status`, `duration_ms` | Monitor RPC performance |
| **Peer** | `peer_id` (public key), `peer_latency_ms`, `message_type`, `message_size_bytes` | Network topology analysis |
| **Ledger** | `ledger_hash`, `ledger_seq`, `close_time`, `ledger_tx_count` | Ledger progression tracking |
| **Job** | `job_type`, `job_queue_ms`, `job_worker` | JobQueue performance |
| **PathFinding** | `pathfind_fast`, `pathfind_search_level`, `pathfind_num_paths`, `pathfind_ledger_index`, `pathfind_num_requests` | Payment path analysis |
| **TxQ** | `txq_queue_depth`, `txq_fee_level`, `txq_eviction_reason` | Queue depth and fee tracking |
| **Fee** | `fee_load_factor`, `fee_escalation_level` | Fee escalation monitoring |
| **Validator** | `validator_list_size`, `validator_list_age_sec` | UNL health monitoring |
| **Amendment** | `amendment_name`, `amendment_status` | Protocol upgrade tracking |
| **SHAMap** | `shamap_type`, `shamap_missing_nodes`, `shamap_duration_ms` | State tree sync performance |
### 2.4.4 Privacy & Sensitive Data Policy
> **PII** = Personally Identifiable Information
OpenTelemetry instrumentation is designed to collect **operational metadata only**, never sensitive content.
#### Data NOT Collected
The following data is explicitly **excluded** from telemetry collection:
| Excluded Data | Reason |
| ----------------------- | ----------------------------------------- |
| **Private Keys** | Never exposed; not relevant to tracing |
| **Account Balances** | Financial data; privacy sensitive |
| **Transaction Amounts** | Financial data; privacy sensitive |
| **Raw TX Payloads** | May contain sensitive memo/data fields |
| **Personal Data** | No PII collected |
| **IP Addresses** | Configurable; excluded by default in prod |
#### Privacy Protection Mechanisms
| Mechanism | Description |
| ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Account Addresses** | Emitted raw (`pathfind_source_account`, `pathfind_dest_account`). An account address is a public ledger identifier; hashing it protects nothing and breaks the join against explorers, RPC and logs |
| **Collector Tail Sampling** | xrpld head sampling is fixed at 1.0 (every span emitted); the collector retains ~10% of non-error traces, reducing stored data exposure |
| **Sampling** | Only 10% of traces recorded by default, reducing data exposure |
| **Local Control** | Node operators have full control over what gets exported |
| **No Raw Payloads** | Transaction content is never recorded, only metadata (hash, type, result) |
| **Collector-Level Filtering** | Available for a future genuinely sensitive attribute via an `attributes` processor. None is shipped, and none must be added for account addresses |
#### Account Addresses
Account addresses are emitted **raw**, at every layer:
1. **SDK-side** (this node): the path-finding RPC handlers set
`pathfind_source_account` / `pathfind_dest_account` to the request's
r-address, only when it parses as one, and `pathfind_dest_currency` to `to_string(Asset)`,
which carries the IOU issuer's r-address. The rationale sits on the attribute
constants in `PathFindSpanNames.h`.
2. **Collector-side**: no collector configuration in this repository hashes
or deletes these attributes.
Why raw: an r-address is a public, enumerable identifier on the ledger. An
unsalted hash of it is reversible by table lookup, so it protects nothing, and
it breaks the one thing the attribute is for: joining a span to the account as
explorers, RPC responses and logs show it. The helper `redactAccount()`
(`xrpl::telemetry`, `Redaction.h`) remains available for a value that is
genuinely private, but it is applied to no span.
#### Collector-Level Data Protection
No hashing or redaction processor is shipped. If a future span introduces a
genuinely sensitive attribute, an `attributes` processor in the collector is the
place to strip it, and the attribute is added to the §2.4 catalogue with that
note in the same change. Account addresses are not such an attribute.
#### Configuration Options for Privacy
In `xrpld.cfg`, operators control data collection granularity through the
`[telemetry]` section. Besides `enabled`, per-component toggles
(`trace_transactions`, `trace_consensus`, `trace_rpc`, `trace_peer` — the last
often disabled due to high volume) select which spans are emitted. There is no
redaction setting: account addresses are public and are emitted raw.
> **Key Principle**: Telemetry collects **operational metadata** (timing, counts, hashes, public identifiers) — never **sensitive content** (keys, balances, amounts, raw payloads).
---
## 2.5 Context Propagation Design
> **WS** = WebSocket
### 2.5.0 Deterministic Trace ID Strategy
Both transaction and consensus tracing use **deterministic trace IDs** derived from
a globally known hash, so all nodes handling the same workflow independently produce
spans under the same `trace_id`. This is combined with protobuf `span_id` propagation
for parent-child relay ordering when available.
#### Transactions — `trace_id = txHash[0:16]`
Every node that handles a transaction knows its `txID` (the `uint256` transaction
hash). The first 16 bytes of this hash are used as the OTel `trace_id`:
```
uint256 txHash: A1B2C3D4 E5F6A7B8 C9D0E1F2 A3B4C5D6 E7F8A9B0 C1D2E3F4 A5B6C7D8 E9F0A1B2
|---------- trace_id (16 bytes) ---------| (remaining 16 bytes unused)
```
Each node generates a **random 8-byte `span_id`** so its span is unique within the shared trace. When the incoming `TMTransaction` carries a protobuf `TraceContext` whose `trace_id` is `txID[0:16]`, the sender's `span_id` becomes the parent. This keeps the relay chain as a parent-child tree. A parent and its child share one trace. A context that names another trace is not used. When the context is absent (older peers, first hop from client) or not used, the span appears as a root in the same trace. Correlation is preserved. Only the tree structure degrades.
```
Node A (submitter) Node B (relay) Node C (relay)
trace_id: A1B2... trace_id: A1B2... trace_id: A1B2...
span_id: 1234 (random) span_id: 5678 (random) span_id: 9ABC (random)
parent: (none) parent: 1234 (proto) parent: 5678 (proto)
↑ ↑
protobuf propagation protobuf propagation
```
If protobuf propagation fails at Node B (old peer):
```
Node A Node B (old peer) Node C
trace_id: A1B2... trace_id: A1B2... trace_id: A1B2...
span_id: 1234 span_id: 5678 span_id: 9ABC
parent: (none) parent: (none) parent: 5678 (proto)
↑ no parent, but same trace_id — still grouped
```
#### Consensus — `trace_id = prevLedgerHash[0:16]`
All validators in the same consensus round share the same `previousLedger.id()`.
The first 16 bytes are used as trace_id. See [Phase 4a implementation status](./06-implementation-phases.md)
and `createDeterministicContext()` in `RCLConsensus.cpp` for the implementation.
Switchable via `consensus_trace_strategy` config:
`"deterministic"` (default) or `"random"` (random trace_id, correlation via attribute queries).
`"random"` is experimental and not used: it would break cross-node trace correlation.
#### Why Not Random IDs with Propagation Only?
Random trace IDs require **unbroken context propagation** across every hop. In a
mixed-version network (common during upgrades), older peers silently drop the
`trace_context` protobuf field. The trace splits and downstream spans become
impossible to find. Deterministic IDs make correlation **propagation-resilient** — the trace
backend groups all spans for the same transaction/round regardless of whether
propagation succeeded.
#### Why Keep Protobuf Propagation?
Deterministic trace IDs alone provide correlation (all spans grouped) but not
**causality** (which node relayed to which). Protobuf `span_id` propagation adds
parent-child ordering that shows the exact relay path. The two mechanisms complement
each other:
| Mechanism | Provides | Fails when |
| ---------------------------- | --------------------------- | -------------------------------------- |
| Deterministic trace_id | Cross-node correlation | Never (hash is always known) |
| Protobuf span_id propagation | Parent-child relay ordering | Older peer drops `trace_context` field |
#### Implementation Reference
The utility function `createDeterministicTxContext(uint256 const& txHash)` follows
the same pattern as `createDeterministicContext(uint256 const& ledgerId)` in
`RCLConsensus.cpp`. See [Phase 3 Task 3.9](./Phase3_taskList.md) for the full spec.
### 2.5.1 Propagation Boundaries
```mermaid
flowchart TB
subgraph http["HTTP/WebSocket (RPC)"]
w3c["W3C Trace Context Headers:<br/>traceparent:<br/>00-trace_id-span_id-flags<br/>tracestate: xrpld=..."]
end
subgraph protobuf["Protocol Buffers (P2P)"]
proto["message TraceContext {<br/> bytes trace_id = 1; // 16 bytes<br/> bytes span_id = 2; // 8 bytes<br/> uint32 trace_flags = 3;<br/> reserved 4; // trace_state, later<br/>}"]
end
subgraph jobqueue["JobQueue (Internal Async)"]
job["Context captured at job creation,<br/>restored at execution<br/><br/>class Job {<br/> otel::context::Context<br/> traceContext_;<br/>};"]
end
style http fill:#0d47a1,stroke:#082f6a,color:#ffffff
style protobuf fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style jobqueue fill:#bf360c,stroke:#8c2809,color:#ffffff
```
**Reading the diagram:**
- **HTTP/WebSocket - RPC (blue)**: For client-facing RPC requests, trace context is propagated using the W3C `traceparent` header. This is the standard approach and works with any OTel-compatible client.
- **Protocol Buffers - P2P (green)**: For peer-to-peer messages between xrpld nodes, trace context is embedded as a protobuf `TraceContext` message carrying trace_id, span_id, flags, and optional trace_state.
- **JobQueue - Internal Async (red)**: For asynchronous work within a single node, the OTel context is captured when a job is created and restored when the job executes on a worker thread. This bridges the async gap so spans remain linked.
---
## 2.6 Integration with Existing Observability
> **OTLP** = OpenTelemetry Protocol | **WS** = WebSocket
### 2.6.1 Existing Frameworks Comparison
xrpld already has two observability mechanisms. OpenTelemetry complements (not replaces) them:
| Aspect | PerfLog | Beast Insight (StatsD) | OpenTelemetry |
| --------------------- | ----------------------------- | ---------------------------- | ------------------------- |
| **Type** | Logging | Metrics | Distributed Tracing |
| **Data** | JSON log entries | Counters, gauges, histograms | Spans with context |
| **Scope** | Single node | Single node | **Cross-node** |
| **Output** | `perf.log` file | StatsD server | OTLP Collector |
| **Question answered** | "What happened on this node?" | "How many? How fast?" | "What was the journey?" |
| **Correlation** | By timestamp | By metric name | By `trace_id` |
| **Overhead** | Low (file I/O) | Low (UDP packets) | Low-Medium (configurable) |
### 2.6.2 What Each Framework Does Best
#### PerfLog
- **Purpose**: Detailed local event logging for RPC and job execution
- **Strengths**:
- Rich JSON output with timing data
- Already integrated in RPC handlers
- File-based, no external dependencies
- **Limitations**:
- Single-node only (no cross-node correlation)
- No parent-child relationships between events
- Manual log parsing required
A PerfLog entry is a JSON object with fields such as `time`, `method`,
`duration_us`, and `result`.
#### Beast Insight (StatsD)
- **Purpose**: Real-time metrics for monitoring dashboards
- **Strengths**:
- Aggregated metrics (counters, gauges, histograms)
- Low overhead (UDP, fire-and-forget)
- Good for alerting thresholds
- **Limitations**:
- No request-level detail
- No causal relationships
- Single-node perspective
- Aggregation happens on the StatsD server, not in the process
In xrpld, Beast Insight is used through `increment` (counters), `gauge`
(point-in-time values), and `timing` (durations) calls. A `timing` call sends
each measured value as its own raw `|ms` sample, so the histogram a dashboard
reads is built by the StatsD server from that stream of values.
#### OpenTelemetry (NEW)
- **Purpose**: Distributed request tracing across nodes
- **Strengths**:
- **Cross-node correlation** via `trace_id`
- Parent-child span relationships
- Rich attributes per span
- A `Histogram` instrument that aggregates **at the point of measure**
- Industry standard (CNCF)
- **Limitations**:
- Requires collector infrastructure
- Higher complexity than logging
A span is created via `startSpan` (e.g. `"tx.relay"`), annotated with
attributes such as `tx_hash` and `peer_id`, and is automatically linked to its
parent through the active context.
OpenTelemetry is not only spans. The same SDK offers a `Histogram` instrument,
and a `Record()` call folds the value straight into bucket counts inside the
process — no per-event record is shipped and no server-side aggregation step is
needed. That is what makes it affordable in a hot loop where one span per event
would not be, and it is the one thing Beast Insight cannot do, because its
`timing` path ships raw values and aggregates them on the StatsD server.
### 2.6.3 When to Use Each
| Scenario | PerfLog | StatsD | OpenTelemetry |
| --------------------------------------- | ---------- | ------ | ------------- |
| "How many TXs per second?" | ❌ | ✅ | ✅ |
| "What's the p99 RPC latency?" | ❌ | ✅ | ✅ |
| "Why was this specific TX slow?" | ⚠️ partial | ❌ | ✅ |
| "Which node delayed consensus?" | ❌ | ❌ | ✅ |
| "What happened on node X at time T?" | ✅ | ❌ | ✅ |
| "Show me the TX journey across 5 nodes" | ❌ | ❌ | ✅ |
| "p99 NodeStore fetch latency?" | ❌ | ❌ | ✅ |
The last row is the case a span cannot answer. One `TMGetObjectByHash` message
requests up to `tuning::kHardMaxReplyNodes` objects, so a span per NodeStore
fetch is not affordable in that loop. Instead the fetch loop's wall time is
recorded once per message into an OpenTelemetry `Histogram`
(`getobject_lookup_us`), and the quantile is read off its buckets. StatsD is
marked ❌ because that instrument is recorded on the native OpenTelemetry metrics
path, not through Beast Insight.
### 2.6.4 Coexistence Strategy
```mermaid
flowchart TB
subgraph xrpld["xrpld Process"]
perflog["PerfLog<br/>(JSON to file)"]
insight["Beast Insight<br/>(StatsD)"]
otel["OpenTelemetry<br/>(Tracing)"]
end
perflog --> perffile["perf.log"]
insight --> statsd["StatsD Server"]
otel --> collector["OTLP Collector"]
perffile --> grafana["Grafana<br/>(Unified UI)"]
statsd --> grafana
collector --> grafana
style xrpld fill:#212121,stroke:#0a0a0a,color:#ffffff
style grafana fill:#bf360c,stroke:#8c2809,color:#ffffff
```
**Reading the diagram:**
- **xrpld Process (dark gray)**: The single xrpld node running all three observability frameworks side by side. Each framework operates independently with no interference.
- **PerfLog to perf.log**: PerfLog writes JSON-formatted event logs to a local file. Grafana can ingest these via Loki or a file-based datasource.
- **Beast Insight to StatsD Server**: Insight sends aggregated metrics (counters, gauges) over UDP to a StatsD server. Grafana reads from StatsD-compatible backends like Graphite or Prometheus (via StatsD exporter).
- **OpenTelemetry to OTLP Collector**: OTel exports spans over OTLP/HTTP to a Collector, which then forwards to a trace backend (Tempo). (OTLP/gRPC is future work — §2.2.2.)
- **Grafana (red, unified UI)**: All three data streams converge in Grafana, enabling operators to correlate logs, metrics, and traces in a single dashboard.
### 2.6.5 Correlation with PerfLog
Trace IDs can be correlated with existing PerfLog entries for comprehensive
debugging. The design is for `RPCHandler.cpp` to start an `rpc.command.<method>`
span alongside the existing PerfLog `rpcStart`/`rpcFinish`/`rpcError` calls,
extract the span's `trace_id` (when valid), and eventually stamp it onto the
PerfLog entry (a planned `setTraceId` hook) so logs and traces share a key. The
span status is set to OK on success or to error (recording the exception) on
failure.
---
_Previous: [Architecture Analysis](./01-architecture-analysis.md)_ | _Next: [Implementation Strategy](./03-implementation-strategy.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,411 +0,0 @@
# Implementation Strategy
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Related**: [Configuration Reference](./05-configuration-reference.md)
---
## 3.1 Directory Structure
The telemetry implementation follows xrpld's existing code organization pattern:
```
include/xrpl/
├── telemetry/
│ ├── Telemetry.h # Main telemetry interface (global singleton)
│ ├── TelemetryConfig.h # Configuration structures
│ ├── TraceContext.h # Context propagation utilities
│ ├── SpanGuard.h # RAII span management with factory methods + discard()
│ ├── DiscardFlag.h # Thread-local discard flag
│ └── SpanAttributes.h # Attribute helper functions
src/libxrpl/
├── telemetry/
│ ├── Telemetry.cpp # Implementation + FilteringSpanProcessor
│ ├── TelemetryConfig.cpp # Config parsing
│ ├── TraceContext.cpp # Context serialization
│ └── NullTelemetry.cpp # No-op implementation
```
---
## 3.2 Implementation Approach
<div align="center">
```mermaid
%%{init: {'flowchart': {'nodeSpacing': 20, 'rankSpacing': 30}}}%%
flowchart TB
subgraph phase1["Phase 1: Core"]
direction LR
sdk["SDK Integration"] ~~~ interface["Telemetry Interface"] ~~~ config["Configuration"]
end
subgraph phase2["Phase 2: RPC"]
direction LR
http["HTTP Context"] ~~~ rpc["RPC Handlers"]
end
subgraph phase3["Phase 3: P2P"]
direction LR
proto["Protobuf Context"] ~~~ tx["Transaction Relay"]
end
subgraph phase4["Phase 4: Consensus"]
direction LR
consensus["Consensus Rounds"] ~~~ proposals["Proposals"]
end
phase1 --> phase2 --> phase3 --> phase4
style phase1 fill:#1565c0,stroke:#0d47a1,color:#ffffff
style phase2 fill:#2e7d32,stroke:#1b5e20,color:#ffffff
style phase3 fill:#e65100,stroke:#bf360c,color:#ffffff
style phase4 fill:#c2185b,stroke:#880e4f,color:#ffffff
```
</div>
### Key Principles
1. **Minimal Intrusion**: Instrumentation should not alter existing control flow
2. **Zero-Cost When Disabled**: Use compile-time flags and no-op implementations
3. **Backward Compatibility**: Protocol Buffer extensions use high field numbers
4. **Graceful Degradation**: Tracing failures must not affect node operation
---
## 3.3 Performance Overhead Summary
> **OTLP** = OpenTelemetry Protocol
| Metric | Overhead | Notes |
| ------------- | ---------- | ------------------------------------------------ |
| CPU | 1-3% | Of per-transaction CPU cost (~200μs baseline) |
| Memory | ~10 MB | SDK statics + batch buffer + worker thread stack |
| Network | 10-50 KB/s | Compressed OTLP export to collector |
| Latency (p99) | <2% | With proper sampling configuration |
---
## 3.4 Detailed CPU Overhead Analysis
### 3.4.1 Per-Operation Costs
> **Note on hardware assumptions**: The costs below are based on the official OTel C++ SDK CI benchmarks
> (969 runs on GitHub Actions 2-core shared runners). On production server hardware (3+ GHz Xeon),
> expect costs at the **lower end** of each range (~30-50% improvement over CI hardware).
| Operation | Time (ns) | Frequency | Impact |
| --------------------- | --------- | ---------------------- | ---------- |
| Span creation | 500-1000 | Every traced operation | Low |
| Span end | 100-200 | Every traced operation | Low |
| SetAttribute (string) | 80-120 | 3-5 per span | Low |
| SetAttribute (int) | 40-60 | 2-3 per span | Negligible |
| AddEvent | 100-200 | 0-2 per span | Low |
| Context injection | 150-250 | Per outgoing message | Low |
| Context extraction | 100-180 | Per incoming message | Low |
| GetCurrent context | 10-20 | Thread-local access | Negligible |
**Source**: Span creation based on OTel C++ SDK `BM_SpanCreation` benchmark (AlwaysOnSampler +
SimpleSpanProcessor + InMemoryExporter), median ~1,000 ns on CI hardware. AddEvent includes
timestamp read + string copy + vector push + mutex acquisition. Context injection/extraction
confirmed by `BM_SpanCreationWithScope` benchmark delta (~160 ns).
### 3.4.2 Transaction Processing Overhead
<div align="center">
```mermaid
%%{init: {'pie': {'textPosition': 0.75}}}%%
pie showData
"tx.receive (1400ns)" : 1400
"tx.validate (1200ns)" : 1200
"tx.relay (1200ns)" : 1200
"Context inject (200ns)" : 200
```
**Transaction Tracing Overhead (~4.0μs total)**
</div>
**Overhead percentage**: 4.0 μs / 200 μs (avg tx processing) = **~2.0%**
> **Breakdown**: Each span (tx.receive, tx.validate, tx.relay) costs ~1,000 ns for creation plus
> ~200-400 ns for 3-5 attribute sets. Context injection is ~200 ns (confirmed by benchmarks).
> On production hardware, expect ~2.6 μs total (~1.3% overhead) due to faster span creation (~500-600 ns).
### 3.4.3 Consensus Round Overhead
| Operation | Count | Cost (ns) | Total |
| ---------------------- | ----- | --------- | ---------- |
| consensus.round span | 1 | ~1200 | ~1.2 μs |
| consensus.phase spans | 3 | ~1100 | ~3.3 μs |
| proposal.receive spans | ~20 | ~1100 | ~22 μs |
| proposal.send spans | ~3 | ~1100 | ~3.3 μs |
| Context operations | ~30 | ~200 | ~6 μs |
| **TOTAL** | | | **~36 μs** |
> **Why higher**: Each span costs ~1,000 ns creation + ~100-200 ns for 1-2 attributes, totaling ~1,100-1,200 ns.
> Context operations remain ~200 ns (confirmed by benchmarks). On production hardware, expect ~24 μs total.
**Overhead percentage**: 36 μs / 3s (typical round) = **~0.001%** (negligible)
### 3.4.4 RPC Request Overhead
| Operation | Cost (ns) |
| ---------------- | ------------ |
| rpc.request span | ~1200 |
| rpc.command span | ~1100 |
| Context extract | ~250 |
| Context inject | ~200 |
| **TOTAL** | **~2.75 μs** |
> **Why higher**: Each span costs ~1,000 ns creation + ~100-200 ns for attributes (command name,
> version, role). Context extract/inject costs are confirmed by OTel C++ benchmarks.
- Fast RPC (1ms): 2.75 μs / 1ms = **~0.275%**
- Slow RPC (100ms): 2.75 μs / 100ms = **~0.003%**
---
## 3.5 Memory Overhead Analysis
> **OTLP** = OpenTelemetry Protocol
### 3.5.1 Static Memory
| Component | Size | Allocated |
| ------------------------------------ | ----------- | ---------- |
| TracerProvider singleton | ~64 KB | At startup |
| BatchSpanProcessor (circular buffer) | ~16 KB | At startup |
| BatchSpanProcessor (worker thread) | ~8 MB | At startup |
| OTLP/HTTP exporter (client init) | ~64 KB | At startup |
| Propagator registry | ~8 KB | At startup |
| **Total static** | **~8.1 MB** | |
> **Why higher than earlier estimate**: The BatchSpanProcessor's circular buffer itself is only ~16 KB
> (2049 x 8-byte `AtomicUniquePtr` entries), but it spawns a dedicated worker thread whose default
> stack size on Linux is ~8 MB. The OTLP/HTTP exporter allocates a small client and TLS
> initialization buffer. The worker thread stack dominates the static footprint.
### 3.5.2 Dynamic Memory
| Component | Size per unit | Max units | Peak |
| -------------------- | -------------- | ---------- | --------------- |
| Active span | ~500-800 bytes | 1000 | ~500-800 KB |
| Queued span (export) | ~500 bytes | 2048 | ~1 MB |
| Attribute storage | ~80 bytes | 5 per span | Included |
| Context storage | ~64 bytes | Per thread | ~6.4 KB |
| **Total dynamic** | | | **~1.5-1.8 MB** |
> **Why active spans are larger**: An active `Span` object includes the wrapper (~88 bytes: shared_ptr,
> mutex, unique_ptr to Recordable) plus `SpanData` (~250 bytes: SpanContext, timestamps, name, status,
> empty containers) plus attribute storage (~200-500 bytes for 3-5 string attributes in a `std::map`).
> Source: `sdk/src/trace/span.h` and `sdk/include/opentelemetry/sdk/trace/span_data.h`.
> Queued spans release the wrapper, keeping only `SpanData` + attributes (~500 bytes).
### 3.5.3 Memory Growth Characteristics
```mermaid
---
config:
xyChart:
width: 700
height: 400
---
xychart-beta
title "Memory Usage vs Span Rate (bounded by queue limit)"
x-axis "Spans/second" [0, 200, 400, 600, 800, 1000]
y-axis "Memory (MB)" 0 --> 12
line [8.5, 9.2, 9.6, 9.9, 10.0, 10.0]
```
**Notes**:
- Memory increases with span rate but **plateaus at queue capacity** (default 2048 spans)
- Batch export prevents unbounded growth
- At queue limit, oldest spans are dropped (not blocked)
- Maximum memory is bounded: ~8.3 MB static (dominated by worker thread stack) + 2048 queued spans x ~500 bytes (~1 MB) + active spans (~0.8 MB) ≈ **~10 MB ceiling**
- The worker thread stack (~8 MB) is virtual memory; actual RSS depends on stack usage (typically much less)
> **Measured outcome**: A perf-iac comparison (telemetry compiled-in + enabled vs compiled-out,
> 9 nodes — validators and client-handlers — under sustained payment load) recorded **no measurable
> RSS increase over the telemetry-off baseline** (~15 GiB mean / ~18–19 GiB peak on both sides),
> with no OOM, no swap, and no leak across the run. The ~10 MB ceiling above is therefore a
> provisioning safety margin (dominated by virtual thread-stack address space), not an expected
> resident-memory increase. Steady-state cost shows up as throughput (~3–4% at head sampling 1.0),
> not memory.
### 3.5.4 Performance Data Sources
The overhead estimates in Sections 3.3-3.5 are derived from the following sources:
| Source | What it covers | URL |
| ------------------------------------------------ | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| OTel C++ SDK CI benchmarks (969 runs) | Span creation, context activation, sampler overhead | [Benchmark Dashboard](https://open-telemetry.github.io/opentelemetry-cpp/benchmarks/) |
| `api/test/trace/span_benchmark.cc` | API-level span creation (~22 ns no-op) | [Source](https://github.com/open-telemetry/opentelemetry-cpp/blob/main/api/test/trace/span_benchmark.cc) |
| `sdk/test/trace/sampler_benchmark.cc` | SDK span creation with samplers (~1,000 ns AlwaysOn) | [Source](https://github.com/open-telemetry/opentelemetry-cpp/blob/main/sdk/test/trace/sampler_benchmark.cc) |
| `sdk/include/.../span_data.h` | SpanData memory layout (~250 bytes base) | [Source](https://github.com/open-telemetry/opentelemetry-cpp/blob/main/sdk/include/opentelemetry/sdk/trace/span_data.h) |
| `sdk/src/trace/span.h` | Span wrapper memory layout (~88 bytes) | [Source](https://github.com/open-telemetry/opentelemetry-cpp/blob/main/sdk/src/trace/span.h) |
| `sdk/include/.../batch_span_processor_options.h` | Default queue size (2048), batch size (512) | [Source](https://github.com/open-telemetry/opentelemetry-cpp/blob/main/sdk/include/opentelemetry/sdk/trace/batch_span_processor_options.h) |
| `sdk/include/.../circular_buffer.h` | CircularBuffer implementation (AtomicUniquePtr array) | [Source](https://github.com/open-telemetry/opentelemetry-cpp/blob/main/sdk/include/opentelemetry/sdk/common/circular_buffer.h) |
| OTLP proto definition | Serialized span size estimation | [Proto](https://github.com/open-telemetry/opentelemetry-proto/blob/main/opentelemetry/proto/trace/v1/trace.proto) |
---
## 3.6 Network Overhead Analysis
### 3.6.1 Export Bandwidth
> **Bytes per span**: Estimates use ~500 bytes/span (conservative upper bound). OTLP protobuf analysis
> shows a typical span with 3-5 string attributes serializes to ~200-300 bytes raw; with gzip
> compression (~60-70% of raw) and batching (amortized headers), ~350 bytes/span is more realistic.
> The table uses the conservative estimate for capacity planning.
| Sampling Rate | Spans/sec | Bandwidth | Notes |
| ------------- | --------- | --------- | ---------------- |
| 100% | ~500 | ~250 KB/s | Development only |
| 10% | ~50 | ~25 KB/s | Staging |
| 1% | ~5 | ~2.5 KB/s | Production |
| Error-only | ~1 | ~0.5 KB/s | Minimal overhead |
### 3.6.2 Trace Context Propagation
| Message Type | Context Size | Messages/sec | Overhead |
| ---------------------- | ------------ | ------------ | ----------- |
| TMTransaction | 25 bytes | ~100 | ~2.5 KB/s |
| TMProposeSet | 25 bytes | ~10 | ~250 B/s |
| TMValidation | 25 bytes | ~50 | ~1.25 KB/s |
| **Total P2P overhead** | | | **~4 KB/s** |
---
## 3.7 Optimization Strategies
### 3.7.1 Sampling Strategies
#### Tail Sampling
```mermaid
flowchart TD
trace["New Trace"]
trace --> errors{"Is Error?"}
errors -->|Yes| sample["SAMPLE"]
errors -->|No| consensus{"Is Consensus?"}
consensus -->|Yes| sample
consensus -->|No| slow{"Is Slow?"}
slow -->|Yes| sample
slow -->|No| prob{"Random < 10%?"}
prob -->|Yes| sample
prob -->|No| drop["DROP"]
style sample fill:#4caf50,stroke:#388e3c,color:#fff
style drop fill:#f44336,stroke:#c62828,color:#fff
```
### 3.7.2 Batch Tuning Recommendations
| Environment | Batch Size | Batch Delay | Max Queue |
| ------------------ | ---------- | ----------- | --------- |
| Low-latency | 128 | 1000ms | 512 |
| High-throughput | 1024 | 10000ms | 8192 |
| Memory-constrained | 256 | 2000ms | 512 |
### 3.7.3 Conditional Instrumentation
Instrumentation is gated on two levels. A compile-time feature flag (`XRPL_ENABLE_TELEMETRY`) reduces the trace macros to no-ops when telemetry is built out, so disabled builds carry zero cost. At runtime, per-component guards (e.g. `shouldTracePeer()`) skip span creation for components whose tracing is turned off, incurring no overhead beyond a single boolean check.
---
## 3.8 Links to Detailed Documentation
- **[Configuration Reference](./05-configuration-reference.md)**: Configuration options and collector setup
- **[Implementation Phases](./06-implementation-phases.md)**: Detailed timeline and milestones
---
## 3.9 Code Intrusiveness Assessment
> **TxQ** = Transaction Queue
This section provides a detailed assessment of how intrusive the OpenTelemetry integration is to the existing xrpld codebase.
### 3.9.3 Risk Assessment by Component
<div align="center">
**Do First** ↖ ↗ **Plan Carefully**
```mermaid
quadrantChart
title Code Intrusiveness Risk Matrix
x-axis Low Risk --> High Risk
y-axis Low Value --> High Value
RPC Tracing: [0.2, 0.55]
Transaction Relay: [0.55, 0.85]
Consensus Tracing: [0.75, 0.92]
Peer Message Tracing: [0.85, 0.35]
JobQueue Context: [0.3, 0.42]
Ledger Acquisition: [0.48, 0.65]
PathFinding: [0.38, 0.72]
TxQ and Fees: [0.25, 0.62]
Validator Mgmt: [0.15, 0.35]
```
**Optional** ↙ ↘ **Avoid**
</div>
#### Risk Level Definitions
| Risk Level | Definition | Mitigation |
| ---------- | ---------------------------------------------------------------- | ---------------------------------- |
| **Low** | Additive changes only; no modification to existing logic | Standard code review |
| **Medium** | Minor modifications to existing functions; clear boundaries | Comprehensive unit tests |
| **High** | Changes to core logic or data structures; potential side effects | Integration tests + staged rollout |
### 3.9.4 Architectural Impact Assessment
| Aspect | Impact | Justification |
| -------------------- | ------- | -------------------------------------------------------------------------------- |
| **Data Flow** | Minimal | Read-only instrumentation; no modification to consensus or transaction data flow |
| **Threading Model** | Minimal | Context propagation uses thread-local storage (standard OTel pattern) |
| **Memory Model** | Low | Bounded queues prevent unbounded growth; RAII ensures cleanup |
| **Network Protocol** | Low | Optional fields in protobuf (high field numbers); backward compatible |
| **Configuration** | None | New config section; existing configs unaffected |
| **Build System** | Low | Optional CMake flag; builds work without OpenTelemetry |
| **Dependencies** | Low | OpenTelemetry SDK is optional; null implementation when disabled |
### 3.9.5 Backward Compatibility
| Compatibility | Status | Notes |
| --------------- | ------- | ----------------------------------------------------- |
| **Config File** | ✅ Full | New `[telemetry]` section is optional |
| **Protocol** | ✅ Full | Optional protobuf fields with high field numbers |
| **Build** | ✅ Full | `XRPL_ENABLE_TELEMETRY=OFF` produces identical binary |
| **Runtime** | ✅ Full | `enabled=0` produces zero overhead |
| **API** | ✅ Full | No changes to public RPC or P2P APIs |
### 3.9.6 Rollback Strategy
If issues are discovered after deployment:
1. **Immediate**: Set `enabled=0` in config and restart (zero code change)
2. **Quick**: Rebuild with `XRPL_ENABLE_TELEMETRY=OFF`
3. **Complete**: Revert telemetry commits (clean separation makes this easy)
### 3.9.7 Code Change Examples
**Minimal RPC Instrumentation (Low Intrusiveness):** Instrumenting an RPC handler adds roughly 3-4 lines: one macro to start the span and one or two `setAttribute` calls (command name, status). The span ends automatically via RAII, so the existing control flow — process the request, send the result — is untouched.
**Consensus Instrumentation (Medium Intrusiveness):** Consensus is slightly more intrusive because child spans in later phase transitions need the round's context. Beyond the span-start and attribute macros, this requires storing the active context in a new member variable (`currentRoundContext_`) at round start. The existing round logic itself remains unchanged.
---
_Previous: [Design Decisions](./02-design-decisions.md)_ | _Next: [Configuration Reference](./05-configuration-reference.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,267 +0,0 @@
# Configuration Reference
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Related**: [Implementation Phases](./06-implementation-phases.md)
---
## 5.1 xrpld Configuration
> **OTLP** = OpenTelemetry Protocol | **TxQ** = Transaction Queue
### 5.1.1 Configuration File Section
The authoritative `[telemetry]` example lives in `cfg/xrpld-example.cfg`. Telemetry is disabled by default (`enabled=0`); enabling it turns on distributed tracing for transaction flow, consensus, and RPC calls, with traces exported to an OpenTelemetry Collector over OTLP. Head sampling is intentionally fixed at 1.0 (sample everything) and is not configurable — per-node head-sampling would produce broken/partial distributed traces, so volume reduction is delegated to the collector's tail sampling (see Section 7.4.2). The full option reference follows.
### 5.1.2 Configuration Options Summary
| Option | Type | Default | Description |
| -------------------------- | ------ | --------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `enabled` | bool | `false` | Enable/disable telemetry |
| `traces_endpoint` | string | `http://localhost:4318/v1/traces` | Full OTLP/HTTP URL for spans, used verbatim |
| `use_tls` | bool | `false` | Enable TLS for exporter connection |
| `tls_ca_cert` | string | `""` | Path to CA certificate file |
| `batch_size` | uint | `512` | Spans per export batch |
| `batch_delay_ms` | uint | `5000` | Max delay before sending batch (ms) |
| `max_queue_size` | uint | `2048` | Maximum queued spans |
| `trace_transactions` | 0 or 1 | `1` | Enable transaction tracing |
| `trace_consensus` | 0 or 1 | `1` | Enable consensus tracing |
| `trace_rpc` | 0 or 1 | `1` | Enable RPC tracing |
| `trace_peer` | 0 or 1 | `1` | Enable peer message tracing (high volume) |
| `trace_ledger` | 0 or 1 | `1` | Enable ledger tracing |
| `tx_trace_strategy` | string | `"deterministic"` | TX trace ID strategy: `"deterministic"` (trace_id = txHash[0:16]) or `"attribute"` (random) |
| `consensus_trace_strategy` | string | `"deterministic"` | Consensus trace ID strategy: `"deterministic"` (trace_id = prevLedgerHash[0:16]) or `"random"` (experimental, not used) |
| `service_name` | string | `"xrpld"` | Service name (`service.name`) for traces and metrics |
| `service_instance_id` | string | `<node_pubkey>` | Instance identifier |
**Planned (not yet implemented)**: the following options appear in the design
documents but are not parsed by `TelemetryConfig.cpp` in Phase 1b and later
phases. They will be added as the corresponding subsystems are instrumented:
| Option | Planned Phase | Purpose |
| -------------------------- | ------------- | ----------------------------------------------------------------------- |
| `exporter` | Future | Select between OTLP/HTTP and OTLP/gRPC |
| `trace_pathfind` | Phase 2 | Path computation tracing toggle |
| `trace_txq` | Phase 3 | Transaction queue tracing toggle |
| `trace_validator` | Future | Validator list / manifest update tracing |
| `trace_amendment` | Future | Amendment voting tracing |
| `consensus_trace_strategy` | Phase 4 | Trace ID strategy for consensus rounds (`deterministic` \| `attribute`) |
---
## 5.2 Configuration Parser
> **TxQ** = Transaction Queue
The parser `makeTelemetrySetup()` in `src/libxrpl/telemetry/TelemetryConfig.cpp` reads the `[telemetry]` `Section` and populates a `Telemetry::Setup` struct, applying the defaults listed in Section 5.1.2 via `section.valueOr(...)`. It derives `serviceInstanceId` from the node public key when not overridden, selects the exporter endpoint default by exporter type, and leaves the sampling ratio at its fixed 1.0 default (not read from config — see Section 7.4.2).
---
## 5.3 Application Integration
### 5.3.1 ApplicationImp Changes
> **Deferred identity**: The node public key (`nodeIdentity_`) is not
> available during `ApplicationImp`'s member initializer list — it is
> resolved later in `setup()`. The `Telemetry` object is therefore
> constructed with an empty `serviceInstanceId` and patched via
> `setServiceInstanceId()` once `setup()` has called `getNodeIdentity()`.
`ApplicationImp` (in `src/xrpld/app/main/Application.cpp`) owns a `std::unique_ptr<telemetry::Telemetry> telemetry_`. It is built in the member initializer list via `makeTelemetry(makeTelemetrySetup(...))` with an empty `serviceInstanceId`, then patched in `setup()` by calling `setServiceInstanceId()` with the Base58 node public key (unless the user supplied a custom `service_instance_id`). `start()` and `run()` forward to `telemetry_->start()` / `telemetry_->stop()`, and `getTelemetry()` returns the owned instance.
### 5.3.2 ServiceRegistry Interface Addition
`include/xrpl/core/ServiceRegistry.h` gains a pure-virtual `telemetry::Telemetry& getTelemetry()` (with a forward declaration of `telemetry::Telemetry`), giving every component a uniform accessor for the tracing subsystem.
> **Note:** `Application` extends `ServiceRegistry`, so `getTelemetry()` is
> available on both. Components that hold a `ServiceRegistry&` (e.g.
> `NetworkOPsImp`) call `registry_.get().getTelemetry()`. Components that
> still hold an `Application&` (e.g. `ServerHandler`, `PeerImp`,
> `RCLConsensus::Adaptor`) call `app_.getTelemetry()` directly.
---
## 5.4 CMake Integration
> **OTLP** = OpenTelemetry Protocol
### 5.4.1 Find OpenTelemetry Module
A `cmake/FindOpenTelemetry.cmake` module locates the OpenTelemetry C++ SDK. It first tries `find_package(opentelemetry-cpp CONFIG)`, aliasing the imported targets `OpenTelemetry::api`, `OpenTelemetry::sdk`, and `OpenTelemetry::otlp_grpc_exporter`, and falls back to `pkg-config` when no CMake config package is present.
### 5.4.2 CMakeLists.txt Changes
The top-level `CMakeLists.txt` adds an `XRPL_ENABLE_TELEMETRY` option (default `OFF`). When enabled, it runs `find_package(OpenTelemetry REQUIRED)`, defines the `XRPL_ENABLE_TELEMETRY` compile flag, and builds the `xrpl_telemetry` library from the real telemetry sources linked against the OpenTelemetry targets; when disabled, it builds the same target from a no-op `NullTelemetry.cpp` so call sites compile unchanged.
---
## 5.5 OpenTelemetry Collector Configuration
> **OTLP** = OpenTelemetry Protocol | **APM** = Application Performance Monitoring
The authoritative collector config lives in the repo at `docker/telemetry/otel-collector-config.yaml` (with Tempo backend config in `docker/telemetry/tempo.yaml`). The sections below summarize the development and production shapes of that pipeline.
### 5.5.1 Development Configuration
The development collector enables an OTLP receiver on both gRPC (`0.0.0.0:4317`) and HTTP (`0.0.0.0:4318`), a single `batch` processor (1s timeout, batch size 100), and two exporters: a `logging` exporter for console debugging and `otlp_grpc/tempo` (insecure) for trace visualization. The single `traces` pipeline wires receiver → batch → both exporters.
### 5.5.2 Production Configuration
The production collector adds TLS on the OTLP gRPC receiver and a richer processor chain: a `memory_limiter` (OOM guard), `batch` (5s timeout, size 512), `tail_sampling`, and an `attributes` processor that stamps `deployment.environment`. Account addresses are public identifiers and are not hashed at any layer. Tail sampling keeps all `ERROR` traces, slow consensus rounds (>5s) and slow RPC requests (>1s), and probabilistically samples the remainder at 10%. Exporters target Grafana Tempo (TLS) and Elastic APM; `health_check` and `zpages` extensions are enabled for operability.
---
## 5.6 Docker Compose Development Environment
> **OTLP** = OpenTelemetry Protocol
The authoritative development stack lives in the repo at `docker/telemetry/docker-compose.yml`. It brings up four services on a shared `xrpld-telemetry` network: an `otel-collector` (otel/opentelemetry-collector-contrib) exposing OTLP gRPC `4317`, OTLP HTTP `4318`, and health check `13133`; `tempo` for trace storage/visualization; `grafana` with provisioned datasources and dashboards (anonymous admin enabled); and an optional `prometheus` for metric correlation.
---
## 5.7 Configuration Architecture
> **OTLP** = OpenTelemetry Protocol
```mermaid
flowchart TB
subgraph config["Configuration Sources"]
cfgFile["xrpld.cfg<br/>[telemetry] section"]
cmake["CMake<br/>XRPL_ENABLE_TELEMETRY"]
end
subgraph init["Initialization"]
parse["makeTelemetrySetup()"]
factory["makeTelemetry()"]
end
subgraph runtime["Runtime Components"]
tracer["TracerProvider"]
exporter["OTLP Exporter"]
processor["BatchProcessor"]
end
subgraph collector["Collector Pipeline"]
recv["Receivers"]
proc["Processors"]
exp["Exporters"]
end
cfgFile --> parse
cmake -->|"compile flag"| parse
parse --> factory
factory --> tracer
tracer --> processor
processor --> exporter
exporter -->|"OTLP"| recv
recv --> proc
proc --> exp
style config fill:#e3f2fd,stroke:#1976d2
style runtime fill:#e8f5e9,stroke:#388e3c
style collector fill:#fff3e0,stroke:#ff9800
```
**Reading the diagram:**
- **Configuration Sources**: `xrpld.cfg` provides runtime settings (endpoint, per-component trace toggles) while the CMake flag controls whether telemetry is compiled in at all. Head sampling is fixed at 1.0 and is not a config option; volume reduction happens via tail sampling in the collector.
- **Initialization**: `makeTelemetrySetup()` parses config values, then `makeTelemetry()` constructs the provider, processor, and exporter objects.
- **Runtime Components**: The `TracerProvider` creates spans, the `BatchProcessor` buffers them, and the `OTLP Exporter` serializes and sends them over the wire.
- **OTLP arrow to Collector**: Trace data leaves the xrpld process via OTLP/HTTP and enters the external Collector pipeline. (OTLP/gRPC is future work — see design decisions §2.2.2.)
- **Collector Pipeline**: `Receivers` ingest OTLP data, `Processors` apply sampling/filtering/enrichment, and `Exporters` forward traces to storage backends (Tempo, etc.).
---
## 5.8 Grafana Integration
> **APM** = Application Performance Monitoring
Step-by-step instructions for integrating xrpld traces with Grafana.
### 5.8.1 Data Source Configuration
#### Tempo (Recommended)
A Tempo datasource (`grafana/provisioning/datasources/tempo.yaml`, provisioned from `docker/telemetry/grafana/`) points at `http://tempo:3200` and enables `tracesToLogs` (linking to Loki on `service.name`/`tx_hash` and mapping `trace_id` → `traceID`), `serviceMap` against Prometheus, the node graph, and Loki search.
#### Elastic APM
Alternatively, an Elasticsearch datasource (`grafana/provisioning/datasources/elastic-apm.yaml`) of type `elasticsearch` points at `http://elasticsearch:9200` against the `apm-*` index, using `@timestamp` as the time field and mapping the log message/level fields.
### 5.8.2 Dashboard Provisioning
A dashboard provider (`grafana/provisioning/dashboards/dashboards.yaml`) loads the `xrpld` dashboard folder from disk (`/var/lib/grafana/dashboards/rippled`), polling for changes every 30s with deletion disabled.
### 5.8.3 Example Dashboard: RPC Performance
An example `xrpld RPC Performance` dashboard (uid `xrpld-rpc-performance`) sourced from Tempo via TraceQL provides four panels: RPC latency by command (heatmap), RPC error rate by command (timeseries), the top 10 slowest RPC commands by average duration (table), and a recent-traces table.
### 5.8.4 Example Dashboard: Transaction Tracing
An example `xrpld Transaction Tracing` dashboard (uid `xrpld-tx-tracing`) over Tempo provides three panels: transaction throughput (`tx.receive` rate, stat), cross-node relay count (average `span.relay_count` on `tx.relay`, timeseries), and a table of transaction validation errors (`tx.validate` with `status = error`).
### 5.8.5 TraceQL Query Examples
Common queries for xrpld traces:
```
# Find all traces for a specific transaction hash
{resource.service.name="xrpld" && span.tx_hash="ABC123..."}
# Find slow RPC commands (>100ms)
{resource.service.name="xrpld" && name=~"rpc.command.*"} | { duration > 100ms }
# Find consensus rounds taking >5 seconds
{resource.service.name="xrpld" && name="consensus.round"} | { duration > 5s }
# Find failed transactions with error details
{resource.service.name="xrpld" && name="tx.validate" && status = error}
# Find transactions relayed to many peers
{resource.service.name="xrpld" && name="tx.relay"} | { span.relay_count > 10 }
# Compare latency across nodes
{resource.service.name="xrpld" && name="rpc.command.account_info"} | avg_over_time(duration) by (resource.service.instance.id)
```
### 5.8.6 Correlation with PerfLog
To correlate OpenTelemetry traces with existing PerfLog data:
**Step 1: Configure Loki to ingest PerfLog**
Configure a Promtail scrape job (`promtail-config.yaml`) that tails `/var/log/rippled/perf*.log`, parses each JSON line, and promotes `trace_id`, `ledger_seq`, and `tx_hash` to Loki labels.
**Step 2: Add trace_id to PerfLog entries**
Modify PerfLog so its JSON output includes a `trace_id` field whenever a valid span is active: fetch the current span from the OpenTelemetry runtime context, and if its context is valid, render the trace ID as a 32-character lowercase hex string into the log entry.
**Step 3: Configure Grafana trace-to-logs link**
In the Tempo datasource, set the `tracesToLogs` derived field to link to Loki on the `trace_id` and `tx_hash` tags, with `filterByTraceID: true`.
### 5.8.7 Correlation with Insight/StatsD Metrics
To correlate traces with existing Beast Insight metrics:
**Step 1: Export Insight metrics to Prometheus**
Add a Prometheus scrape job (`prometheus.yaml`) named `xrpld-statsd` targeting the StatsD exporter at `statsd-exporter:9102`.
**Step 2: Add exemplars to metrics**
The OpenTelemetry SDK automatically adds exemplars (trace IDs) to metrics when using the Prometheus exporter, linking metric spikes to specific traces.
**Step 3: Configure Grafana metric-to-trace link**
In the Prometheus datasource, set `exemplarTraceIdDestinations` to map the `trace_id` exemplar to the Tempo datasource.
**Step 4: Dashboard panel with exemplars**
Add a timeseries panel over Prometheus (e.g. `histogram_quantile(0.99, rate(xrpld_rpc_duration_seconds_bucket[5m]))`) with `exemplar: true` enabled.
This allows clicking on metric data points to jump directly to the related trace.
---
_Previous: [Implementation Strategy](./03-implementation-strategy.md)_ | _Next: [Implementation Phases](./06-implementation-phases.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,673 +0,0 @@
# Implementation Phases
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Related**: [Configuration Reference](./05-configuration-reference.md) | [Observability Backends](./07-observability-backends.md)
---
## 6.1 Phase Overview
> **TxQ** = Transaction Queue
```mermaid
gantt
title OpenTelemetry Implementation Timeline
dateFormat YYYY-MM-DD
axisFormat Week %W
section Phase 1
Core Infrastructure :p1, 2024-01-01, 2w
SDK Integration :p1a, 2024-01-01, 4d
Telemetry Interface :p1b, after p1a, 3d
Configuration & CMake :p1c, after p1b, 3d
Unit Tests :p1d, after p1c, 2d
Buffer & Integration :p1e, after p1d, 2d
section Phase 2
RPC Tracing :p2, after p1, 2w
HTTP Context Extraction :p2a, after p1, 2d
RPC Handler Instrumentation :p2b, after p2a, 4d
PathFinding Instrumentation :p2f, after p2b, 2d
TxQ Instrumentation :p2g, after p2f, 2d
WebSocket Support :p2c, after p2g, 2d
Integration Tests :p2d, after p2c, 2d
Buffer & Review :p2e, after p2d, 4d
section Phase 3
Transaction Tracing :p3, after p2, 2w
Protocol Buffer Extension :p3a, after p2, 2d
PeerImp Instrumentation :p3b, after p3a, 3d
Fee Escalation Instrumentation :p3f, after p3b, 2d
Relay Context Propagation :p3c, after p3f, 3d
Multi-node Tests :p3d, after p3c, 2d
Buffer & Review :p3e, after p3d, 4d
section Phase 4
Consensus Tracing :p4, after p3, 2w
Consensus Round Spans :p4a, after p3, 3d
Proposal Handling :p4b, after p4a, 3d
Establish Phase (4a) :p4f, after p4b, 3d
Validation Tests :p4c, after p4f, 4d
Buffer & Review :p4e, after p4c, 4d
section Phase 5
Documentation & Deploy :p5, after p4, 1w
```
---
## 6.2 Phase 1: Core Infrastructure (Weeks 1-2)
**Objective**: Establish foundational telemetry infrastructure
### Tasks
| Task | Description |
| ---- | ----------------------------------------------------- |
| 1.1 | Add OpenTelemetry C++ SDK to Conan/CMake |
| 1.2 | Implement `Telemetry` interface and factory |
| 1.3 | Implement `SpanGuard` RAII wrapper |
| 1.4 | Implement configuration parser |
| 1.5 | Integrate into `ApplicationImp` |
| 1.6 | Add conditional compilation (`XRPL_ENABLE_TELEMETRY`) |
| 1.7 | Create `NullTelemetry` no-op implementation |
| 1.8 | Unit tests for core infrastructure |
### Exit Criteria
- [ ] OpenTelemetry SDK compiles and links
- [ ] Telemetry can be enabled/disabled via config
- [ ] Basic span creation works
- [ ] No performance regression when disabled
- [ ] Unit tests passing
---
## 6.3 Phase 2: RPC Tracing (Weeks 3-4)
> **TxQ** = Transaction Queue
**Objective**: Complete tracing for all RPC operations
### Tasks
| Task | Description |
| ---- | -------------------------------------------------------------------------- |
| 2.1 | Implement W3C Trace Context HTTP header extraction |
| 2.2 | Instrument `ServerHandler::onRequest()` |
| 2.3 | Instrument `xrpl::rpc::doCommand()` |
| 2.4 | Add RPC-specific attributes |
| 2.5 | Instrument WebSocket handler |
| 2.6 | PathFinding instrumentation (`pathfind.request`, `pathfind.compute` spans) |
| 2.7 | TxQ instrumentation (`txq.enqueue`, `txq.apply` spans) |
| 2.8 | Integration tests for RPC tracing |
| 2.9 | Performance benchmarks |
| 2.10 | Documentation |
### Exit Criteria
- [ ] All RPC commands traced
- [ ] Trace context propagates from HTTP headers
- [ ] WebSocket and HTTP both instrumented
- [ ] <1ms overhead per RPC call
- [ ] Integration tests passing
---
## 6.4 Phase 3: Transaction Tracing (Weeks 5-6)
**Objective**: Trace transaction lifecycle across network with deterministic cross-node correlation
### Tasks
| Task | Description |
| ---- | -------------------------------------------------------------- |
| 3.1 | Define `TraceContext` Protocol Buffer message |
| 3.2 | Implement protobuf context serialization |
| 3.3 | Instrument `PeerImp::handleTransaction()` |
| 3.4 | Instrument `NetworkOPs::submitTransaction()` |
| 3.5 | Instrument HashRouter integration |
| 3.6 | Fee escalation instrumentation (`fee.escalate` span) |
| 3.7 | Implement relay context propagation |
| 3.8 | Integration tests (multi-node) |
| 3.9 | Deterministic transaction trace ID (`trace_id = txHash[0:16]`) |
| 3.10 | Performance benchmarks |
### Deterministic Trace ID (Task 3.9)
Transaction spans use **deterministic trace IDs** derived from the transaction hash:
`trace_id = txHash[0:16]`. All nodes handling the same transaction independently
produce spans under the same trace_id. Protobuf `span_id` propagation (Task 3.7)
additionally provides parent-child relay ordering when available. See
[02-design-decisions.md §2.5.0](./02-design-decisions.md) for the design rationale
and [Phase3_taskList.md Task 3.9](./Phase3_taskList.md) for the full implementation spec.
### Exit Criteria
- [ ] Transaction traces span across nodes
- [ ] Trace context in Protocol Buffer messages
- [ ] HashRouter deduplication visible in traces
- [ ] Multi-node integration tests passing
- [ ] <5% overhead on transaction throughput
- [ ] Deterministic trace_id: all nodes produce same trace_id for same transaction
- [ ] Protobuf span_id propagation preserves parent-child ordering when available
---
## 6.5 Phase 4: Consensus Tracing (Weeks 7-8)
**Objective**: Full observability into consensus rounds
### Tasks
| Task | Description | Status |
| ---- | ------------------------------------------- | ------------------ |
| 4.1 | Instrument `RCLConsensus::startRound()` | ✅ Done (via 4a.2) |
| 4.2 | Instrument phase transitions | ✅ Done |
| 4.3 | Instrument proposal handling | ✅ Done |
| 4.4 | Instrument validation handling | ✅ Done |
| 4.5 | Add consensus-specific attributes | ✅ Done |
| 4.6 | Correlate with transaction traces | ✅ Done |
| 4.7 | Build verification and testing | ✅ Done |
| 4.8 | Validation span enrichment (ext. dashboard) | ❌ Not done |
**Note**: The original plan doc listed tasks 4.7-4.11 as "Validator list tracing",
"Amendment voting tracing", "SHAMap sync tracing", "Multi-validator integration tests",
and "Performance validation". These were descoped and replaced by the tasklist's 4.7
(build verification) and 4.8 (validation span enrichment). Validator, amendment, and
SHAMap tracing are not implemented.
### Spans Produced
| Span Name | Location | Attributes |
| --------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `consensus.phase.open` | `Consensus.h` | _(none)_ |
| `consensus.proposal.send` | `RCLConsensus.cpp` | `consensus_round` |
| `consensus.ledger_close` | `RCLConsensus.cpp` | `ledger_seq`, `consensus_mode` |
| `consensus.accept` | `RCLConsensus.cpp` | `proposers`, `round_time_ms`, `quorum` |
| `consensus.accept.apply` | `RCLConsensus.cpp` | `close_time_ripple_epoch_s`, `close_time_correct`, `close_resolution_ms`, `consensus_state`, `proposing`, `round_time_ms`, `ledger_seq`, `parent_close_time_ripple_epoch_s`, `close_time_self_ripple_epoch_s`, `close_time_vote_bins`, `resolution_direction` |
| `consensus.validation.send` | `RCLConsensus.cpp` | `ledger_seq`, `proposing` |
### Exit Criteria
- [x] Complete consensus round traces
- [x] Phase transitions visible (open, establish, close, accept)
- [x] Proposals and validations traced — send and receive; relay deferred to Phase 4b
- [x] Close time agreement tracked (per `avCT_CONSENSUS_PCT`)
- [x] No impact on consensus timing
- [ ] Multi-validator test network validated
- [x] Transaction-consensus correlation (Task 4.6) — `tx.included` events in doAccept
- [ ] Validation span enrichment (Task 4.8) — not implemented
### Implementation Status — Phase 4a Complete
Phase 4a (establish-phase gap fill & cross-node correlation) adds:
- **Deterministic trace ID** derived from `previousLedger.id()` so all validators
in the same round share the same `trace_id` (switchable via
`consensus_trace_strategy` config: `"deterministic"`, or `"random"` which is
experimental and not used).
See [Configuration Reference](./05-configuration-reference.md) for full
configuration options.
- **Round lifecycle spans**: `consensus.round` with round-to-round span links.
- **Establish phase**: `consensus.establish`, `consensus.update_positions` (with
`dispute.resolve` events), `consensus.check` (with threshold tracking).
- **Mode changes**: `consensus.mode_change` spans.
- **Validation**: `consensus.validation.send` with span link to round span
(thread-safe cross-thread access via `roundSpanContext_` snapshot).
- **Separation of concerns**: telemetry extracted to private helpers
(`startRoundTracing`, `createValidationSpan`, `startEstablishTracing`,
`updateEstablishTracing`, `endEstablishTracing`).
See [Phase4_taskList.md](./Phase4_taskList.md) for the full spec and implementation notes.
---
## 6.5a Phase 4a: Establish-Phase Gap Fill & Cross-Node Correlation
**Objective**: Fill tracing gaps in the establish phase and establish cross-node
correlation using deterministic trace IDs derived from `previousLedger.id()`.
**Approach**: Direct instrumentation in `Consensus.h` and `RCLConsensus.cpp`.
All spans use `SpanGuard` factory methods (`span()`, `hashSpan()`, `linkedSpan()`)
with `TraceCategory::Consensus` gating. No macros used — all tracing via direct
`SpanGuard` API calls.
### Tasks
| Task | Description | Effort | Risk | Status |
| ---- | ------------------------------------------------ | ------ | ------ | ------------------------ |
| 4a.0 | Prerequisites: extend SpanGuard & Telemetry APIs | 1d | Medium | ✅ Done (no macros) |
| 4a.1 | Adaptor `getTelemetry()` method | 0.5d | Low | ⏭️ Skipped (not needed) |
| 4a.2 | Switchable round span with deterministic traceID | 2d | High | ✅ Done |
| 4a.3 | Span members in `Consensus.h` | 0.5d | Medium | ✅ Done (with deviation) |
| 4a.4 | Instrument `phaseEstablish()` | 1d | Medium | ✅ Done |
| 4a.5 | Instrument `updateOurPositions()` | 1d | Medium | ✅ Done |
| 4a.6 | Instrument `haveConsensus()` (thresholds) | 1d | Medium | ✅ Done |
| 4a.7 | Instrument mode changes | 0.5d | Low | ✅ Done |
| 4a.8 | Reparent existing spans under round | 0.5d | Low | ✅ Done |
| 4a.9 | Build verification and testing | 1d | Low | ✅ Done |
**Total Effort**: 9 days
### Spans Produced
| Span Name | Location | Key Attributes (actually set) |
| ---------------------------- | ------------------ | ----------------------------------------------------------------------------------------------------------------------------- |
| `consensus.round` | `RCLConsensus.cpp` | `consensus_round_id`, `consensus_ledger_id`, `ledger_seq`, `consensus_mode`, `trace_strategy` |
| `consensus.establish` | `Consensus.h` | `converge_percent`, `establish_count`, `proposers` |
| `consensus.update_positions` | `Consensus.h` | `converge_percent`, `proposers`, `have_close_time_consensus`, `close_time_threshold`, `disputes_count`, `avalanche_threshold` |
| `consensus.check` | `Consensus.h` | `agree_count`, `disagree_count`, `converge_percent`, `have_close_time_consensus`, `threshold_percent`, `consensus_result` |
| `consensus.mode_change` | `RCLConsensus.cpp` | `mode_old`, `mode_new` |
### Exit Criteria
- [x] Establish phase internals traced (establish, update_positions, check spans)
- [x] Establish phase fully traced — `disputes_count`, `avalanche_threshold`, dispute `yays`/`nays` all implemented
- [x] Cross-node correlation works via deterministic trace_id
- [x] Strategy switchable via config (`deterministic` / `attribute`)
- [x] Consecutive rounds linked via follows-from spans
- [x] Build passes with telemetry ON and OFF
- [x] No impact on consensus timing
See [Phase4_taskList.md](./Phase4_taskList.md) for full task details.
---
## 6.5b Phase 4b: Cross-Node Propagation (Future)
**Objective**: Wire `TraceContextPropagator` for P2P messages (proposals,
validations) to enable true distributed tracing between nodes.
**Status**: Partially implemented. Send-side injection (proposals and
validations) and receive-side extraction (`consensus.{proposal,validation}.
receive` spans parented on the sender's context) are wired in Phase 4a.
Remaining Phase 4b work: relay spans in `share(RCLCxPeerPos)` and multi-node
validation of the propagation path.
**Prerequisites**: Phase 4a complete and validated.
See [Phase4_taskList.md § Phase 4b](./Phase4_taskList.md) for full design.
---
## 6.6 Phase 5: Documentation & Deployment (Week 9)
**Objective**: Production readiness
### Tasks
| Task | Description |
| ---- | ----------------------------- |
| 5.1 | Operator runbook |
| 5.2 | Grafana dashboards |
| 5.3 | Alert definitions |
| 5.4 | Collector deployment examples |
| 5.5 | Developer documentation |
| 5.6 | Training materials |
| 5.7 | Final integration testing |
---
## 6.7 Risk Assessment
```mermaid
quadrantChart
title Risk Assessment Matrix
x-axis Low Impact --> High Impact
y-axis Low Likelihood --> High Likelihood
quadrant-1 Mitigate Immediately
quadrant-2 Plan Mitigation
quadrant-3 Accept Risk
quadrant-4 Monitor Closely
SDK Compat: [0.2, 0.18]
Protocol Chg: [0.75, 0.72]
Perf Overhead: [0.58, 0.42]
Context Prop: [0.4, 0.55]
Memory Leaks: [0.85, 0.25]
```
### Risk Details
| Risk | Likelihood | Impact | Mitigation |
| ------------------------------------ | ---------- | ------ | --------------------------------------- |
| Protocol changes break compatibility | Medium | High | Use high field numbers, optional fields |
| Performance overhead unacceptable | Medium | Medium | Sampling, conditional compilation |
| Context propagation complexity | Medium | Medium | Phased rollout, extensive testing |
| SDK compatibility issues | Low | Medium | Pin SDK version, fallback to no-op |
| Memory leaks in long-running nodes | Low | High | Memory profiling, bounded queues |
---
## 6.8 Success Metrics
| Metric | Target | Measurement |
| ------------------------ | -------------------------------------------------------------- | --------------------- |
| Trace coverage | >95% of transaction code paths (independent of sampling ratio) | Sampling verification |
| CPU overhead | <3% | Benchmark tests |
| Memory overhead | <10 MB | Memory profiling |
| Latency impact (p99) | <2% | Performance tests |
| Trace completeness | >99% spans with required attrs | Validation script |
| Cross-node trace linkage | >90% of multi-hop transactions | Integration tests |
---
## 6.9 Quick Wins and Crawl-Walk-Run Strategy
> **TxQ** = Transaction Queue
This section outlines a prioritized approach to maximize ROI with minimal initial investment.
### 6.9.1 Crawl-Walk-Run Overview
<div align="center">
```mermaid
flowchart TB
subgraph crawl["🐢 CRAWL (Week 1-2)"]
direction LR
c1[Core SDK Setup] ~~~ c2[RPC Tracing Only] ~~~ c3[PathFinding + TxQ Tracing] ~~~ c4[Single Node]
end
subgraph walk["🚶 WALK (Week 3-5)"]
direction LR
w1[Transaction Tracing] ~~~ w2[Fee Escalation Tracing] ~~~ w3[Cross-Node Context] ~~~ w4[Basic Dashboards]
end
subgraph run["🏃 RUN (Week 6-9)"]
direction LR
r1[Consensus Tracing] ~~~ r2[Establish Phase<br/>& Cross-Node Correlation] ~~~ r3[StatsD Integration] ~~~ r4[Production Deploy]
end
crawl --> walk --> run
style crawl fill:#1b5e20,stroke:#0d3d14,color:#fff
style walk fill:#bf360c,stroke:#8c2809,color:#fff
style run fill:#0d47a1,stroke:#082f6a,color:#fff
style c1 fill:#1b5e20,stroke:#0d3d14,color:#fff
style c2 fill:#1b5e20,stroke:#0d3d14,color:#fff
style c3 fill:#1b5e20,stroke:#0d3d14,color:#fff
style c4 fill:#1b5e20,stroke:#0d3d14,color:#fff
style w1 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style w2 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style w3 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style w4 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style r1 fill:#0d47a1,stroke:#082f6a,color:#fff
style r2 fill:#0d47a1,stroke:#082f6a,color:#fff
style r3 fill:#0d47a1,stroke:#082f6a,color:#fff
style r4 fill:#0d47a1,stroke:#082f6a,color:#fff
```
</div>
**Reading the diagram:**
- **CRAWL (Weeks 1-2)**: Minimal investment -- set up the SDK, instrument RPC and PathFinding/TxQ handlers, and verify on a single node. Delivers immediate latency visibility.
- **WALK (Weeks 3-5)**: Expand to transaction lifecycle tracing, fee escalation, cross-node context propagation, and basic Grafana dashboards. This is where distributed tracing starts working.
- **RUN (Weeks 6-9)**: Full consensus instrumentation, establish-phase gap fill, cross-node correlation, StatsD integration, and production deployment with sampling and alerting.
- **Arrows (crawl → walk → run)**: Each phase builds on the prior one; you cannot skip ahead because later phases depend on infrastructure established earlier.
### 6.9.2 Quick Wins (Immediate Value)
| Quick Win | Value | When to Deploy |
| ------------------------------ | ------ | -------------- |
| **RPC Command Tracing** | High | Week 2 |
| **RPC Latency Histograms** | High | Week 2 |
| **Error Rate Dashboard** | Medium | Week 2 |
| **Transaction Submit Tracing** | High | Week 3 |
| **Consensus Round Duration** | Medium | Week 6 |
### 6.9.3 CRAWL Phase (Weeks 1-2)
**Goal**: Get basic tracing working with minimal code changes.
**What You Get**:
- RPC request/response traces for all commands
- Latency breakdown per RPC command
- PathFinding and TxQ tracing (directly impacts RPC latency)
- Error visibility with stack traces
- Basic Grafana dashboard
**Code Changes**: ~15 lines in `ServerHandler.cpp`, ~40 lines in new telemetry module
**Why Start Here**:
- RPC is the lowest-risk, highest-visibility component
- PathFinding and TxQ are RPC-adjacent and directly affect latency
- Immediate value for debugging client issues
- No cross-node complexity
- Single file modification to existing code
### 6.9.4 WALK Phase (Weeks 3-5)
**Goal**: Add transaction lifecycle tracing across nodes.
**What You Get**:
- End-to-end transaction traces from submit to relay
- Fee escalation tracing within the transaction pipeline
- Cross-node correlation (see transaction path)
- HashRouter deduplication visibility
- Relay latency metrics
**Code Changes**: ~120 lines across 4 files, plus protobuf extension
**Why Do This Second**:
- Builds on RPC tracing (transactions submitted via RPC)
- Fee escalation is integral to the transaction processing pipeline
- Moderate complexity (requires context propagation)
- High value for debugging transaction issues
### 6.9.5 RUN Phase (Weeks 6-9)
**Goal**: Full observability including consensus.
**What You Get**:
- Complete consensus round visibility
- Phase transition timing
- Validator proposal tracking
- ~~Validator list and manifest tracing~~ — descoped
- ~~Amendment voting tracing~~ — descoped
- ~~SHAMap sync tracing~~ — descoped
- Full end-to-end traces (client → RPC → TX → consensus → ledger) — partial (tx-consensus correlation not yet done)
**Code Changes**: ~100 lines across 3 consensus files
**Why Do This Last**:
- Highest complexity (consensus is critical path)
- Validator, amendment, and SHAMap components were descoped (lower priority)
- Requires thorough testing
- Lower relative value (consensus issues are rarer)
### 6.9.6 ROI Prioritization Matrix
```mermaid
quadrantChart
title Implementation ROI Matrix
x-axis Low Effort --> High Effort
y-axis Low Value --> High Value
quadrant-1 Quick Wins - Do First
quadrant-2 Major Projects - Plan Carefully
quadrant-3 Nice to Have - Optional
quadrant-4 Time Sinks - Avoid
RPC Tracing: [0.15, 0.92]
TX Submit Trace: [0.3, 0.78]
TX Relay Trace: [0.5, 0.88]
Consensus Trace: [0.72, 0.72]
Peer Msg Trace: [0.85, 0.3]
Ledger Acquire: [0.55, 0.52]
```
---
## 6.10 Definition of Done
> **TxQ** = Transaction Queue | **HA** = High Availability
Clear, measurable criteria for each phase.
### 6.10.1 Phase 1: Core Infrastructure
| Criterion | Measurement | Target |
| --------------- | ---------------------------------------------------------- | ---------------------------- |
| SDK Integration | `cmake --build` succeeds with `-DXRPL_ENABLE_TELEMETRY=ON` | ✅ Compiles |
| Runtime Toggle | `enabled=0` produces zero overhead | <0.1% CPU difference |
| Span Creation | Unit test creates and exports span | Span appears in Tempo |
| Configuration | All config options parsed correctly | Config validation tests pass |
| Documentation | Developer guide exists | PR approved |
**Definition of Done**: All criteria met, PR merged, no regressions in CI.
### 6.10.2 Phase 2: RPC Tracing
| Criterion | Measurement | Target |
| ------------------ | ---------------------------------- | -------------------------- |
| Coverage | All RPC commands instrumented | 100% of commands |
| Context Extraction | traceparent header propagates | Integration test passes |
| Attributes | Command, status, duration recorded | Validation script confirms |
| Performance | RPC latency overhead | <1ms p99 |
| Dashboard | Grafana dashboard deployed | Screenshot in docs |
**Definition of Done**: RPC traces visible in Tempo for all commands, dashboard shows latency distribution.
### 6.10.3 Phase 3: Transaction Tracing
| Criterion | Measurement | Target |
| --------------------- | ------------------------------------------------- | -------------------------------------------------------- |
| Local Trace | Submit → validate → TxQ traced | Single-node test passes |
| Cross-Node | Context propagates via protobuf | Multi-node test passes |
| Deterministic TraceID | Same trace_id on all nodes for same tx | Multi-node test: query by txHash[0:16] returns all spans |
| Relay Ordering | Protobuf span_id propagation creates parent-child | Tempo trace tree shows relay chain |
| Graceful Degradation | Old peer drops trace_context | Spans still grouped by deterministic trace_id |
| Relay Visibility | relay_count attribute correct | Spot check 100 txs |
| HashRouter | Deduplication visible | Duplicates produce no span |
| Performance | TX throughput overhead | <5% degradation |
**Definition of Done**: Transaction traces span 3+ nodes in test network with deterministic trace_id correlation, parent-child ordering via protobuf propagation, and performance within bounds.
### 6.10.4 Phase 4: Consensus Tracing
| Criterion | Measurement | Target |
| -------------------- | ----------------------------- | ------------------------- |
| Round Tracing | startRound creates root span | Unit test passes |
| Phase Visibility | All phases have child spans | Integration test confirms |
| Proposer Attribution | Proposer ID in attributes | Spot check 50 rounds |
| Timing Accuracy | Phase durations match PerfLog | <5% variance |
| No Consensus Impact | Round timing unchanged | Performance test passes |
**Definition of Done**: Consensus rounds fully traceable, no impact on consensus timing.
### 6.10.5 Phase 5: Production Deployment
| Criterion | Measurement | Target |
| ------------ | ---------------------------- | -------------------------- |
| Collector HA | Multiple collectors deployed | No single point of failure |
| Sampling | Tail sampling configured | 10% base + errors + slow |
| Retention | Data retained per policy | 7 days hot, 30 days warm |
| Alerting | Alerts configured | Error spike, high latency |
| Runbook | Operator documentation | Approved by ops team |
| Training | Team trained | Session completed |
**Definition of Done**: Telemetry running in production, operators trained, alerts active.
### 6.10.6 Success Metrics Summary
| Phase | Primary Metric | Secondary Metric | Deadline |
| ------- | ---------------------- | --------------------------- | ------------- |
| Phase 1 | SDK compiles and runs | Zero overhead when disabled | End of Week 2 |
| Phase 2 | 100% RPC coverage | <1ms latency overhead | End of Week 4 |
| Phase 3 | Cross-node traces work | <5% throughput impact | End of Week 6 |
| Phase 4 | Consensus fully traced | No consensus timing impact | End of Week 8 |
| Phase 5 | Production deployment | Operators trained | End of Week 9 |
---
## 6.11 Recommended Implementation Order
Based on ROI analysis, implement in this exact order:
```mermaid
flowchart TB
subgraph week1["Week 1"]
t1[1. OpenTelemetry SDK<br/>Conan/CMake integration]
t2[2. Telemetry interface<br/>SpanGuard, config]
end
subgraph week2["Week 2"]
t3[3. RPC ServerHandler<br/>instrumentation]
t4[4. Basic Tempo setup<br/>for testing]
end
subgraph week3["Week 3"]
t5[5. Transaction submit<br/>tracing]
t6[6. Grafana dashboard<br/>v1]
end
subgraph week4["Week 4"]
t7[7. Protobuf context<br/>extension]
t8[8. PeerImp tx.relay<br/>instrumentation]
end
subgraph week5["Week 5"]
t9[9. Multi-node<br/>integration tests]
t10[10. Performance<br/>benchmarks]
end
subgraph week6_8["Weeks 6-8"]
t11[11. Consensus<br/>instrumentation]
t12[12. Full integration<br/>testing]
end
subgraph week9["Week 9"]
t13[13. Production<br/>deployment]
t14[14. Documentation<br/>& training]
end
t1 --> t2 --> t3 --> t4
t4 --> t5 --> t6
t6 --> t7 --> t8
t8 --> t9 --> t10
t10 --> t11 --> t12
t12 --> t13 --> t14
style week1 fill:#1b5e20,stroke:#0d3d14,color:#fff
style week2 fill:#1b5e20,stroke:#0d3d14,color:#fff
style week3 fill:#bf360c,stroke:#8c2809,color:#fff
style week4 fill:#bf360c,stroke:#8c2809,color:#fff
style week5 fill:#bf360c,stroke:#8c2809,color:#fff
style week6_8 fill:#0d47a1,stroke:#082f6a,color:#fff
style week9 fill:#4a148c,stroke:#2e0d57,color:#fff
style t1 fill:#1b5e20,stroke:#0d3d14,color:#fff
style t2 fill:#1b5e20,stroke:#0d3d14,color:#fff
style t3 fill:#1b5e20,stroke:#0d3d14,color:#fff
style t4 fill:#1b5e20,stroke:#0d3d14,color:#fff
style t5 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style t6 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style t7 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style t8 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style t9 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style t10 fill:#ffe0b2,stroke:#ffcc80,color:#1e293b
style t11 fill:#0d47a1,stroke:#082f6a,color:#fff
style t12 fill:#0d47a1,stroke:#082f6a,color:#fff
style t13 fill:#4a148c,stroke:#2e0d57,color:#fff
style t14 fill:#4a148c,stroke:#2e0d57,color:#fff
```
**Reading the diagram:**
- **Week 1 (tasks 1-2)**: Foundation work -- integrate the OpenTelemetry SDK via Conan/CMake and build the `Telemetry` interface with `SpanGuard` and config parsing.
- **Week 2 (tasks 3-4)**: First observable output -- instrument `ServerHandler` for RPC tracing and stand up Tempo so developers can see traces immediately.
- **Weeks 3-5 (tasks 5-10)**: Transaction lifecycle -- add submit tracing, build the first Grafana dashboard, extend protobuf for cross-node context, instrument `PeerImp` relay, then validate with multi-node integration tests and performance benchmarks.
- **Weeks 6-8 (tasks 11-12)**: Consensus deep-dive -- instrument consensus rounds and phases, then run full integration testing across all instrumented paths.
- **Week 9 (tasks 13-14)**: Go-live -- deploy to production with sampling/alerting configured, and deliver documentation and operator training.
- **Arrow chain (t1 → ... → t14)**: Strict sequential dependency; each task's output is a prerequisite for the next.
---
_Previous: [Configuration Reference](./05-configuration-reference.md)_ | _Next: [Observability Backends](./07-observability-backends.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,407 +0,0 @@
# Observability Backend Recommendations
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Related**: [Implementation Phases](./06-implementation-phases.md) | [Appendix](./08-appendix.md)
---
## 7.1 Development/Testing Backends
> **OTLP** = OpenTelemetry Protocol
| Backend | Pros | Cons | Use Case |
| ---------- | ----------------------------------- | ---------------------- | ------------------- |
| **Tempo** | Cost-effective, Grafana integration | Requires Grafana stack | Local dev, CI, Prod |
| **Zipkin** | Simple, lightweight | Basic features | Quick prototyping |
### Quick Start with Tempo
```bash
# Start Tempo with OTLP support
docker run -d --name tempo \
-p 3200:3200 \
-p 4317:4317 \
-p 4318:4318 \
grafana/tempo:2.6.1
```
---
## 7.2 Production Backends
> **APM** = Application Performance Monitoring
| Backend | Pros | Cons | Use Case |
| ----------------- | ----------------------------------------- | ---------------------- | --------------------------- |
| **Grafana Tempo** | Cost-effective, Grafana integration | Requires Grafana stack | Most production deployments |
| **Elastic APM** | Full observability stack, log correlation | Resource intensive | Existing Elastic users |
| **Honeycomb** | Excellent query, high cardinality | SaaS cost | Deep debugging needs |
| **Datadog APM** | Full platform, easy setup | SaaS cost | Enterprise with budget |
### Backend Selection Flowchart
```mermaid
flowchart TD
start[Select Backend] --> budget{Budget<br/>Constraints?}
budget -->|Yes| oss[Open Source]
budget -->|No| saas{Prefer<br/>SaaS?}
oss --> existing{Existing<br/>Stack?}
existing -->|Grafana| tempo[Grafana Tempo]
existing -->|Elastic| elastic[Elastic APM]
existing -->|None| tempo
saas -->|Yes| enterprise{Enterprise<br/>Support?}
saas -->|No| oss
enterprise -->|Yes| datadog[Datadog APM]
enterprise -->|No| honeycomb[Honeycomb]
tempo --> final[Configure Collector]
elastic --> final
honeycomb --> final
datadog --> final
style start fill:#0f172a,stroke:#020617,color:#fff
style budget fill:#334155,stroke:#1e293b,color:#fff
style oss fill:#1e293b,stroke:#0f172a,color:#fff
style existing fill:#334155,stroke:#1e293b,color:#fff
style saas fill:#334155,stroke:#1e293b,color:#fff
style enterprise fill:#334155,stroke:#1e293b,color:#fff
style final fill:#0f172a,stroke:#020617,color:#fff
style tempo fill:#1b5e20,stroke:#0d3d14,color:#fff
style elastic fill:#bf360c,stroke:#8c2809,color:#fff
style honeycomb fill:#0d47a1,stroke:#082f6a,color:#fff
style datadog fill:#4a148c,stroke:#2e0d57,color:#fff
```
**Reading the diagram:**
- **Budget Constraints? (Yes)**: Leads to open-source options. If you already run Grafana or Elastic, pick the matching backend; otherwise default to Grafana Tempo.
- **Budget Constraints? (No) → Prefer SaaS?**: If you want a managed service, choose between Datadog (enterprise support) and Honeycomb (developer-focused). If not, fall back to open-source.
- **Terminal nodes (Tempo / Elastic / Honeycomb / Datadog)**: Each represents a concrete backend choice, all of which feed into the same final step.
- **Configure Collector**: Regardless of backend, you always finish by configuring the OTel Collector to export to your chosen destination.
---
## 7.3 Recommended Production Architecture
> **OTLP** = OpenTelemetry Protocol | **APM** = Application Performance Monitoring | **HA** = High Availability
```mermaid
flowchart TB
subgraph validators["Validator Nodes"]
v1[xrpld<br/>Validator 1]
v2[xrpld<br/>Validator 2]
end
subgraph stock["Stock Nodes"]
s1[xrpld<br/>Stock 1]
s2[xrpld<br/>Stock 2]
end
subgraph collector["OTel Collector Cluster"]
c1[Collector<br/>DC1]
c2[Collector<br/>DC2]
end
subgraph backends["Storage Backends"]
tempo[(Grafana<br/>Tempo)]
elastic[(Elastic<br/>APM)]
archive[(S3/GCS<br/>Archive)]
end
subgraph ui["Visualization"]
grafana[Grafana<br/>Dashboards]
end
v1 -->|OTLP| c1
v2 -->|OTLP| c1
s1 -->|OTLP| c2
s2 -->|OTLP| c2
c1 --> tempo
c1 --> elastic
c2 --> tempo
c2 --> archive
tempo --> grafana
elastic --> grafana
%% Note: simplified single-collector-per-DC topology shown for clarity
style validators fill:#b71c1c,stroke:#7f1d1d,color:#ffffff
style stock fill:#0d47a1,stroke:#082f6a,color:#ffffff
style collector fill:#bf360c,stroke:#8c2809,color:#ffffff
style backends fill:#1b5e20,stroke:#0d3d14,color:#ffffff
style ui fill:#4a148c,stroke:#2e0d57,color:#ffffff
```
**Reading the diagram:**
- **Validator / Stock Nodes**: All xrpld nodes emit trace data via OTLP. Validators and stock nodes are grouped separately because they may reside in different network zones.
- **Collector Cluster (DC1, DC2)**: Regional collectors receive OTLP from nodes in their datacenter, apply processing (sampling, enrichment), and fan out to multiple backends. Enrichment includes deployment-tier tagging: each collector stamps `deployment.environment` and (as a fallback) `xrpl.network.type` so one Grafana stack can filter data from many collectors by tier.
- **Storage Backends**: Tempo and Elastic provide queryable trace storage; S3/GCS Archive provides long-term cold storage for compliance or post-incident analysis.
- **Grafana Dashboards**: The single visualization layer that queries both Tempo and Elastic, giving operators a unified view of all traces.
- **Data flow direction**: Nodes → Collectors → Storage → Grafana. Each arrow represents a network hop; minimizing collector-to-backend hops reduces latency.
> **Note**: Production deployments should use multiple collector instances behind a load balancer for high availability. The diagram shows a simplified single-collector topology for clarity.
---
## 7.4 Architecture Considerations
### 7.4.1 Collector Placement
| Strategy | Description | Pros | Cons |
| ------------- | -------------------- | ------------------------ | ----------------------- |
| **Sidecar** | Collector per node | Isolation, simple config | Resource overhead |
| **DaemonSet** | Collector per host | Shared resources | Complexity |
| **Gateway** | Central collector(s) | Centralized processing | Single point of failure |
**Recommendation**: Use **Gateway** pattern with regional collectors for xrpld networks:
- One collector cluster per datacenter/region
- Tail-based sampling at collector level
- Multiple export destinations for redundancy
### 7.4.2 Sampling Strategy
```mermaid
flowchart LR
subgraph head["Head Sampling (Node)"]
hs[Node-level head sampling<br/>fixed at 100%<br/>not configurable]
end
subgraph tail["Tail Sampling (Collector)"]
ts1[Keep all errors]
ts2[Keep slow >5s]
ts3[Keep 10% rest]
end
head --> tail
ts1 --> final[Final Traces]
ts2 --> final
ts3 --> final
style head fill:#0d47a1,stroke:#082f6a,color:#fff
style tail fill:#1b5e20,stroke:#0d3d14,color:#fff
style hs fill:#0d47a1,stroke:#082f6a,color:#fff
style ts1 fill:#1b5e20,stroke:#0d3d14,color:#fff
style ts2 fill:#1b5e20,stroke:#0d3d14,color:#fff
style ts3 fill:#1b5e20,stroke:#0d3d14,color:#fff
style final fill:#bf360c,stroke:#8c2809,color:#fff
```
**Reading the diagram:**
- **Head Sampling (Node)**: xrpld pins head sampling at 100% (sample everything) and does not expose a configurable ratio. This is intentional: a per-node ratio would let different nodes make divergent keep/drop decisions for the same distributed trace, producing broken/partial traces. A span inheriting a remote parent is decided by the same ratio sampler, not by the peer's sampled flag. Volume reduction is delegated to the collector's tail sampling.
- **Tail Sampling (Collector)**: The second filter -- the collector inspects completed traces and applies rules: keep all errors, keep anything slower than 5 seconds, and keep 10% of the remainder.
- **Arrow head → tail**: All head-sampled traces flow to the collector, where tail sampling further reduces volume while preserving the most valuable data.
- **Final Traces**: The output after both sampling stages; this is what gets stored and queried. The two-stage approach balances cost with debuggability.
### 7.4.3 Data Retention
| Environment | Hot Storage | Warm Storage | Cold Archive |
| ----------- | ----------- | ------------ | ------------ |
| Development | 24 hours | N/A | N/A |
| Staging | 7 days | N/A | N/A |
| Production | 7 days | 30 days | many years |
---
## 7.5 Integration Checklist
- [ ] Choose primary backend (Tempo recommended for cost/features)
- [ ] Deploy collector cluster with high availability
- [ ] Configure tail-based sampling for error/latency traces
- [ ] Set up Grafana dashboards for trace visualization
- [ ] Configure alerts for trace anomalies
- [ ] Establish data retention policies
- [ ] Test trace correlation with logs and metrics
---
## 7.6 Grafana Dashboard Examples
Pre-built dashboards for xrpld observability.
### 7.6.1 Consensus Health Dashboard
A Tempo-backed dashboard (uid `xrpld-consensus-health`) with four panels, all driven by TraceQL:
- **Consensus Round Duration** (timeseries, ms): average `consensus.round` span duration per node instance, with yellow/red thresholds at 4s/5s.
- **Phase Duration Breakdown** (barchart): average duration of `consensus.phase.*` spans grouped by span name.
- **Proposers per Round** (stat): average of the `span.proposers` attribute on `consensus.round` spans.
- **Recent Slow Rounds (>5s)** (table): `consensus.round` spans filtered to `duration > 5s`.
Each panel's TraceQL query is described inline in its bullet above.
### 7.6.2 Node Overview Dashboard
A Tempo-backed dashboard (uid `xrpld-node-overview`) with four panels:
- **Active Nodes** (stat): count of distinct `resource.service.instance.id` values seen for the `xrpld` service.
- **Total Transactions (1h)** (stat): count of `tx.receive` spans.
- **Error Rate** (gauge, percent): ratio of `status = error` spans to all spans, with yellow/red thresholds at 1%/5%.
- **Service Map** (nodeGraph): Tempo-generated service dependency graph.
### 7.6.3 Alert Rules
Grafana provisions three TraceQL-based alert rules (group `xrpld-tracing-alerts`, evaluated every 1m) against the Tempo datasource:
- **Consensus Round Slow** (warning, `for: 5m`): fires when average `consensus.round` duration exceeds 5s.
```
{resource.service.name="xrpld" && name="consensus.round"} | avg(duration) > 5s
```
- **RPC Error Rate Spike** (critical, `for: 2m`): fires when the error rate across `rpc.command.*` spans exceeds 5%. Error _rate_ is a ratio, so it must divide the error-span rate by the total-span rate — a single TraceQL `rate()` returns spans/second, not a percentage, and would fire on traffic volume alone. This uses span metrics emitted by the collector's `spanmetrics` connector (Prometheus datasource), not a TraceQL query:
```
sum(rate(traces_spanmetrics_calls_total{service_name="xrpld", span_name=~"rpc.command.*", status_code="STATUS_CODE_ERROR"}[5m]))
/
sum(rate(traces_spanmetrics_calls_total{service_name="xrpld", span_name=~"rpc.command.*"}[5m]))
> 0.05
```
- **Transaction Throughput Drop** (warning, `for: 10m`): fires when the `tx.receive` span rate falls below 10/s.
```
{resource.service.name="xrpld" && name="tx.receive"} | rate() < 10
```
> **Note**: The Consensus Round Slow and Transaction Throughput Drop rules use TraceQL aggregates (`avg(duration)`, `rate()`), which require Tempo 2.3+ with TraceQL metrics enabled. Verify aggregate query support in your Tempo version before provisioning. The RPC Error Rate Spike rule instead queries Prometheus span metrics (collector `spanmetrics` connector), so it needs that connector enabled in the collector pipeline.
---
## 7.7 PerfLog and Insight Correlation
> **OTLP** = OpenTelemetry Protocol
How to correlate OpenTelemetry traces with existing xrpld observability.
### 7.7.1 Correlation Architecture
```mermaid
flowchart TB
subgraph xrpld["xrpld Node"]
otel[OpenTelemetry<br/>Spans]
perflog[PerfLog<br/>JSON Logs]
insight[Beast Insight<br/>StatsD Metrics]
end
subgraph collectors["Data Collection"]
otelc[OTel Collector]
promtail[Promtail/Fluentd]
statsd[StatsD Exporter]
end
subgraph storage["Storage"]
tempo[(Tempo)]
loki[(Loki)]
prom[(Prometheus)]
end
subgraph grafana["Grafana"]
traces[Trace View]
logs[Log View]
metrics[Metrics View]
corr[Correlation<br/>Panel]
end
otel -->|OTLP| otelc --> tempo
perflog -->|JSON| promtail --> loki
insight -->|StatsD| statsd --> prom
tempo --> traces
loki --> logs
prom --> metrics
traces --> corr
logs --> corr
metrics --> corr
style xrpld fill:#0d47a1,stroke:#082f6a,color:#fff
style collectors fill:#bf360c,stroke:#8c2809,color:#fff
style storage fill:#1b5e20,stroke:#0d3d14,color:#fff
style grafana fill:#4a148c,stroke:#2e0d57,color:#fff
style otel fill:#0d47a1,stroke:#082f6a,color:#fff
style perflog fill:#0d47a1,stroke:#082f6a,color:#fff
style insight fill:#0d47a1,stroke:#082f6a,color:#fff
style otelc fill:#bf360c,stroke:#8c2809,color:#fff
style promtail fill:#bf360c,stroke:#8c2809,color:#fff
style statsd fill:#bf360c,stroke:#8c2809,color:#fff
style tempo fill:#1b5e20,stroke:#0d3d14,color:#fff
style loki fill:#1b5e20,stroke:#0d3d14,color:#fff
style prom fill:#1b5e20,stroke:#0d3d14,color:#fff
style traces fill:#4a148c,stroke:#2e0d57,color:#fff
style logs fill:#4a148c,stroke:#2e0d57,color:#fff
style metrics fill:#4a148c,stroke:#2e0d57,color:#fff
style corr fill:#4a148c,stroke:#2e0d57,color:#fff
```
**Reading the diagram:**
- **xrpld Node (three sources)**: A single node emits three independent data streams -- OpenTelemetry spans, PerfLog JSON logs, and Beast Insight StatsD metrics.
- **Data Collection layer**: Each stream has its own collector -- OTel Collector for spans, Promtail/Fluentd for logs, and a StatsD exporter for metrics. They operate independently.
- **Storage layer (Tempo, Loki, Prometheus)**: Each data type lands in a purpose-built store optimized for its query patterns (trace search, log grep, metric aggregation).
- **Grafana Correlation Panel**: The key integration point -- Grafana queries all three stores and links them via shared fields (`trace_id`, `tx_hash`, `ledger_seq`), enabling a single-pane debugging experience.
### 7.7.2 Correlation Fields
| Source | Field | Link To | Purpose |
| ----------- | ------------------- | ------------- | -------------------------- |
| **Trace** | `trace_id` | Logs | Find log entries for trace |
| **Trace** | `tx_hash` | Logs, Metrics | Find TX-related data |
| **Trace** | `ledger_seq` | Logs | Find ledger-related logs |
| **PerfLog** | `trace_id` (new) | Traces | Jump to trace from log |
| **PerfLog** | `ledger_seq` | Traces | Find consensus trace |
| **Insight** | `exemplar.trace_id` | Traces | Jump from metric spike |
### 7.7.3 Example: Debugging a Slow Transaction
**Step 1: Find the trace**
```
# In Grafana Explore with Tempo
{resource.service.name="xrpld" && span.tx_hash="ABC123..."}
```
**Step 2: Get the trace_id from the trace view**
```
Trace ID: 4bf92f3577b34da6a3ce929d0e0e4736
```
**Step 3: Find related PerfLog entries**
```
# In Grafana Explore with Loki
{job="xrpld"} |= "4bf92f3577b34da6a3ce929d0e0e4736"
```
**Step 4: Check Insight metrics for the time window**
```
# In Grafana with Prometheus
rate(xrpld_tx_applied_total[1m])
@ timestamp_from_trace
```
### 7.7.4 Unified Dashboard Example
A single dashboard (uid `xrpld-unified`) that ties traces, metrics, and logs together across the Tempo, Prometheus, and Loki datasources:
- **Transaction Latency (Traces)** (timeseries, Tempo): `histogram_over_time(duration)` of `tx.receive` spans.
- **Transaction Rate (Metrics)** (timeseries, Prometheus): `rate(xrpld_tx_received_total[5m])` per instance, with a data link that opens the matching `tx.receive` traces in Tempo.
- **Recent Logs** (logs, Loki): `{job="xrpld"} | json`.
- **Trace Search** (table, Tempo): all `xrpld` traces, with per-row data links on `traceID` that jump to the trace in Tempo and to the correlated logs in Loki (`{job="xrpld"} |= "<traceID>"`).
The cross-datasource data links are what make this a single-pane debugging view; the correlation fields they rely on are listed in section 7.7.2.
---
_Previous: [Implementation Phases](./06-implementation-phases.md)_ | _Next: [Appendix](./08-appendix.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,196 +0,0 @@
# Appendix
> **Parent Document**: [OpenTelemetryPlan.md](./OpenTelemetryPlan.md)
> **Related**: [Observability Backends](./07-observability-backends.md)
---
## 8.1 Glossary
> **OTLP** = OpenTelemetry Protocol | **TxQ** = Transaction Queue
| Term | Definition |
| --------------------- | ---------------------------------------------------------- |
| **Span** | A unit of work with start/end time, name, and attributes |
| **Trace** | A collection of spans representing a complete request flow |
| **Trace ID** | 128-bit unique identifier for a trace |
| **Span ID** | 64-bit unique identifier for a span within a trace |
| **Context** | Carrier for trace/span IDs across boundaries |
| **Propagator** | Component that injects/extracts context |
| **Sampler** | Decides which traces to record |
| **Exporter** | Sends spans to backend |
| **Collector** | Receives, processes, and forwards telemetry |
| **OTLP** | OpenTelemetry Protocol (wire format) |
| **W3C Trace Context** | Standard HTTP headers for trace propagation |
| **Baggage** | Key-value pairs propagated across service boundaries |
| **Resource** | Entity producing telemetry (service, host, etc.) |
| **Instrumentation** | Code that creates telemetry data |
### xrpld-Specific Terms
| Term | Definition |
| ----------------- | ------------------------------------------------------------- |
| **Overlay** | P2P network layer managing peer connections |
| **Consensus** | XRP Ledger consensus algorithm (RCL) |
| **Proposal** | Validator's suggested transaction set for a ledger |
| **Validation** | Validator's signature on a closed ledger |
| **HashRouter** | Component for transaction deduplication |
| **JobQueue** | Thread pool for asynchronous task execution |
| **PerfLog** | Existing performance logging system in xrpld |
| **Beast Insight** | Existing metrics framework in xrpld |
| **PathFinding** | Payment path computation engine for cross-currency payments |
| **TxQ** | Transaction queue managing fee-based prioritization |
| **LoadManager** | Dynamic fee escalation based on network load |
| **SHAMap** | SHA-256 hash-based map (Merkle trie variant) for ledger state |
---
## 8.2 Span Hierarchy Visualization
> **TxQ** = Transaction Queue
```mermaid
flowchart TB
subgraph trace["Trace: Transaction Lifecycle"]
rpc["rpc.request<br/>(entry point)"]
validate["tx.validate"]
relay["tx.relay<br/>(parent span)"]
subgraph peers["Peer Spans"]
p1["peer.send<br/>Peer A"]
p2["peer.send<br/>Peer B"]
p3["peer.send<br/>Peer C"]
end
subgraph pathfinding["PathFinding Spans"]
pathfind["pathfind.request"]
pathcomp["pathfind.compute"]
end
consensus["consensus.round"]
apply["tx.apply"]
subgraph txqueue["TxQ Spans"]
txq["txq.enqueue"]
txqApply["txq.apply"]
end
feeCalc["fee.escalate"]
end
subgraph validators["Validator Spans"]
valFetch["validator.list.fetch"]
valManifest["validator.manifest"]
end
rpc --> validate
rpc --> pathfind
pathfind --> pathcomp
validate --> relay
relay --> p1
relay --> p2
relay --> p3
p1 -.->|"context propagation"| consensus
consensus --> apply
apply --> txq
txq --> txqApply
txq --> feeCalc
style trace fill:#0f172a,stroke:#020617,color:#fff
style peers fill:#1e3a8a,stroke:#172554,color:#fff
style pathfinding fill:#134e4a,stroke:#0f766e,color:#fff
style txqueue fill:#064e3b,stroke:#047857,color:#fff
style validators fill:#4c1d95,stroke:#6d28d9,color:#fff
style rpc fill:#1d4ed8,stroke:#1e40af,color:#fff
style validate fill:#047857,stroke:#064e3b,color:#fff
style relay fill:#047857,stroke:#064e3b,color:#fff
style p1 fill:#0e7490,stroke:#155e75,color:#fff
style p2 fill:#0e7490,stroke:#155e75,color:#fff
style p3 fill:#0e7490,stroke:#155e75,color:#fff
style consensus fill:#fef3c7,stroke:#fde68a,color:#1e293b
style apply fill:#047857,stroke:#064e3b,color:#fff
style pathfind fill:#0e7490,stroke:#155e75,color:#fff
style pathcomp fill:#0e7490,stroke:#155e75,color:#fff
style txq fill:#047857,stroke:#064e3b,color:#fff
style txqApply fill:#047857,stroke:#064e3b,color:#fff
style feeCalc fill:#047857,stroke:#064e3b,color:#fff
style valFetch fill:#6d28d9,stroke:#4c1d95,color:#fff
style valManifest fill:#6d28d9,stroke:#4c1d95,color:#fff
```
**Reading the diagram:**
- **rpc.request (blue, top)**: The entry point — every traced transaction starts as an RPC call; this root span is the parent of all downstream work.
- **tx.validate and pathfind.request (green/teal, first fork)**: The RPC request fans out into transaction validation and, for cross-currency payments, a PathFinding branch (`pathfind.request` -> `pathfind.compute`).
- **tx.relay -> Peer Spans (teal, middle)**: After validation, the transaction is relayed to peers A, B, and C in parallel; each `peer.send` is a sibling child span showing fan-out across the network.
- **context propagation (dashed arrow)**: The dotted line from `peer.send Peer A` to `consensus.round` represents the trace context crossing a node boundary — the receiving validator picks up the same `trace_id` and continues the trace.
- **consensus.round -> tx.apply -> TxQ Spans (green, lower)**: Once consensus accepts the transaction, it is applied to the ledger; the TxQ spans (`txq.enqueue`, `txq.apply`, `fee.escalate`) capture queue depth and fee escalation behavior.
- **Validator Spans (purple, detached)**: `validator.list.fetch` and `validator.manifest` are independent workflows for UNL management — they run on their own traces and are linked to consensus via Span Links, not parent-child relationships.
---
## 8.3 References
> **OTLP** = OpenTelemetry Protocol
### OpenTelemetry Resources
1. [OpenTelemetry C++ SDK](https://github.com/open-telemetry/opentelemetry-cpp)
2. [OpenTelemetry Specification](https://opentelemetry.io/docs/specs/otel/)
3. [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/)
4. [OTLP Protocol Specification](https://opentelemetry.io/docs/specs/otlp/)
### Standards
5. [W3C Trace Context](https://www.w3.org/TR/trace-context/)
6. [W3C Baggage](https://www.w3.org/TR/baggage/)
7. [Protocol Buffers](https://protobuf.dev/)
### xrpld Resources
8. [xrpld Source Code](https://github.com/XRPLF/rippled)
9. [XRP Ledger Documentation](https://xrpl.org/docs/)
10. [xrpld Overlay README](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/overlay/README.md)
11. [xrpld RPC README](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/rpc/README.md)
12. [xrpld Consensus README](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/app/consensus/README.md)
---
## 8.4 Version History
| Version | Date | Author | Changes |
| ------- | ---------- | ------ | -------------------------------------------------------------- |
| 1.0 | 2026-02-12 | - | Initial implementation plan |
| 1.1 | 2026-02-13 | - | Refactored into modular documents |
| 1.2 | 2026-03-24 | - | Review fixes: accuracy corrections, cross-document consistency |
---
## 8.5 Document Index
### Plan Documents
| Document | Description |
| ---------------------------------------------------------------- | -------------------------------------------- |
| [OpenTelemetryPlan.md](./OpenTelemetryPlan.md) | Master overview and executive summary |
| [00-tracing-fundamentals.md](./00-tracing-fundamentals.md) | Distributed tracing concepts and OTel primer |
| [01-architecture-analysis.md](./01-architecture-analysis.md) | xrpld architecture and trace points |
| [02-design-decisions.md](./02-design-decisions.md) | SDK selection, exporters, span conventions |
| [03-implementation-strategy.md](./03-implementation-strategy.md) | Directory structure, performance analysis |
| [05-configuration-reference.md](./05-configuration-reference.md) | xrpld config, CMake, Collector configs |
| [06-implementation-phases.md](./06-implementation-phases.md) | Timeline, tasks, risks, success metrics |
| [07-observability-backends.md](./07-observability-backends.md) | Backend selection and architecture |
| [08-appendix.md](./08-appendix.md) | Glossary, references, version history |
### Task Lists
| Document | Description |
| ------------------------------------------ | ------------------------------------ |
| [Phase2_taskList.md](./Phase2_taskList.md) | RPC layer trace instrumentation |
| [Phase3_taskList.md](./Phase3_taskList.md) | Peer overlay & consensus tracing |
| [Phase4_taskList.md](./Phase4_taskList.md) | Transaction lifecycle tracing |
| [Phase5_taskList.md](./Phase5_taskList.md) | Ledger processing & advanced tracing |
---
_Previous: [Observability Backends](./07-observability-backends.md)_ | _Back to: [Overview](./OpenTelemetryPlan.md)_

View File

@@ -1,199 +0,0 @@
# [OpenTelemetry](00-tracing-fundamentals.md) Distributed Tracing Implementation Plan for xrpld
## Executive Summary
> **OTLP** = OpenTelemetry Protocol
This document provides a comprehensive implementation plan for integrating OpenTelemetry distributed tracing into the xrpld XRP Ledger node software. The plan addresses the unique challenges of a decentralized peer-to-peer system where trace context must propagate across network boundaries between independent nodes.
### Key Benefits
- **End-to-end transaction visibility**: Track transactions from submission through consensus to ledger inclusion
- **Consensus round analysis**: Understand timing and behavior of consensus phases across validators
- **RPC performance insights**: Identify slow handlers and optimize response times
- **Network topology understanding**: Visualize message propagation patterns between peers
- **Incident debugging**: Correlate events across distributed nodes during issues
### Estimated Performance Overhead
| Metric | Overhead | Notes |
| ------------- | ---------- | ------------------------------------------------ |
| CPU | 1-3% | Span creation and attribute setting |
| Memory | <10 MB | SDK statics + batch buffer + worker thread stack |
| Network | 10-50 KB/s | Compressed OTLP export to collector |
| Latency (p99) | <2% | With proper sampling configuration |
---
## Document Structure
This implementation plan is organized into modular documents for easier navigation:
<div align="center">
```mermaid
flowchart TB
overview["📋 OpenTelemetryPlan.md<br/>(This Document)"]
subgraph fundamentals["Fundamentals"]
fund["00-tracing-fundamentals.md"]
end
subgraph analysis["Analysis & Design"]
arch["01-architecture-analysis.md"]
design["02-design-decisions.md"]
end
subgraph impl["Implementation"]
strategy["03-implementation-strategy.md"]
config["05-configuration-reference.md"]
end
subgraph deploy["Deployment & Planning"]
phases["06-implementation-phases.md"]
backends["07-observability-backends.md"]
appendix["08-appendix.md"]
end
overview --> fundamentals
overview --> analysis
overview --> impl
overview --> deploy
fund --> arch
arch --> design
design --> strategy
strategy --> config
config --> phases
phases --> backends
backends --> appendix
style overview fill:#1b5e20,stroke:#0d3d14,color:#fff,stroke-width:2px
style fundamentals fill:#00695c,stroke:#004d40,color:#fff
style fund fill:#00695c,stroke:#004d40,color:#fff
style analysis fill:#0d47a1,stroke:#082f6a,color:#fff
style impl fill:#bf360c,stroke:#8c2809,color:#fff
style deploy fill:#4a148c,stroke:#2e0d57,color:#fff
style arch fill:#0d47a1,stroke:#082f6a,color:#fff
style design fill:#0d47a1,stroke:#082f6a,color:#fff
style strategy fill:#bf360c,stroke:#8c2809,color:#fff
style config fill:#bf360c,stroke:#8c2809,color:#fff
style phases fill:#4a148c,stroke:#2e0d57,color:#fff
style backends fill:#4a148c,stroke:#2e0d57,color:#fff
style appendix fill:#4a148c,stroke:#2e0d57,color:#fff
```
</div>
---
## Table of Contents
| Section | Document | Description |
| ------- | ---------------------------------------------------------- | ---------------------------------------------------------------------- |
| **0** | [Tracing Fundamentals](./00-tracing-fundamentals.md) | Distributed tracing concepts, span relationships, context propagation |
| **1** | [Architecture Analysis](./01-architecture-analysis.md) | xrpld component analysis, trace points, instrumentation priorities |
| **2** | [Design Decisions](./02-design-decisions.md) | SDK selection, exporters, span naming, attributes, context propagation |
| **3** | [Implementation Strategy](./03-implementation-strategy.md) | Directory structure, key principles, performance optimization |
| **5** | [Configuration Reference](./05-configuration-reference.md) | xrpld config, CMake integration, Collector configurations |
| **6** | [Implementation Phases](./06-implementation-phases.md) | 5-phase timeline, tasks, risks, success metrics |
| **7** | [Observability Backends](./07-observability-backends.md) | Backend selection guide and production architecture |
| **8** | [Appendix](./08-appendix.md) | Glossary, references, version history |
---
## 0. Tracing Fundamentals
This document introduces distributed tracing concepts for readers unfamiliar with the domain. It covers what traces and spans are, how parent-child and follows-from relationships model causality, how context propagates across service boundaries, and how sampling controls data volume. It also maps these concepts to xrpld-specific scenarios like transaction relay and consensus.
➡️ **[Read Tracing Fundamentals](./00-tracing-fundamentals.md)**
---
## 1. Architecture Analysis
> **WS** = WebSocket | **TxQ** = Transaction Queue
The xrpld node consists of several key components that require instrumentation for comprehensive distributed tracing. The main areas include the RPC server (HTTP/WebSocket), Overlay P2P network, Consensus mechanism (RCLConsensus), JobQueue for async task execution, PathFinding, Transaction Queue (TxQ), fee escalation (LoadManager), ledger acquisition, validator management, and existing observability infrastructure (PerfLog, Insight/StatsD, Journal logging).
Key trace points span across transaction submission via RPC, peer-to-peer message propagation, consensus round execution, ledger building, path computation, transaction queue behavior, fee escalation, and validator health. The implementation prioritizes high-value, low-risk components first: RPC handlers provide immediate value with minimal risk, while consensus tracing requires careful implementation to avoid timing impacts.
➡️ **[Read full Architecture Analysis](./01-architecture-analysis.md)**
---
## 2. Design Decisions
> **OTLP** = OpenTelemetry Protocol | **CNCF** = Cloud Native Computing Foundation
The OpenTelemetry C++ SDK is selected for its CNCF backing, active development, and native performance characteristics. Traces are exported via OTLP/HTTP to an OpenTelemetry Collector, which provides flexible routing and sampling. OTLP/gRPC is planned future work (see design decisions §2.2.2).
Span naming follows a hierarchical `<component>.<operation>` convention (e.g., `rpc.submit`, `tx.relay`, `consensus.round`). Context propagation uses W3C Trace Context headers for HTTP and embedded Protocol Buffer fields for P2P messages. The implementation coexists with existing PerfLog and Insight observability systems through correlation IDs.
**Data Collection & Privacy**: Telemetry collects only operational metadata (timing, counts, hashes) — never sensitive content (private keys, balances, amounts, raw payloads). Account addresses are public ledger identifiers and are emitted raw; there is no redaction setting. Trace volume is reduced, where wanted, by collector-side sampling. Node operators retain full control over telemetry configuration.
➡️ **[Read full Design Decisions](./02-design-decisions.md)**
---
## 3. Implementation Strategy
The telemetry code is organized under `include/xrpl/telemetry/` for headers and `src/libxrpl/telemetry/` for implementation. Key principles include RAII-based span management via `SpanGuard` (with `discard()` for dropping unwanted spans), a `FilteringSpanProcessor` that intercepts `OnEnd()` to prevent discarded spans from entering the export pipeline, conditional compilation with `XRPL_ENABLE_TELEMETRY`, and minimal runtime overhead through batch processing and efficient sampling.
Performance optimization strategies include head sampling fixed at 100% (intentionally not configurable, so trace keep/drop decisions stay coherent across nodes), tail-based sampling at the collector for errors and slow traces to reduce volume, batch export to reduce network overhead, and conditional instrumentation that compiles to no-ops when disabled.
➡️ **[Read full Implementation Strategy](./03-implementation-strategy.md)**
---
## 5. Configuration Reference
> **OTLP** = OpenTelemetry Protocol | **APM** = Application Performance Monitoring
Configuration is handled through the `[telemetry]` section in `xrpld.cfg` with options for enabling/disabling, exporter selection, endpoint configuration, and component-level filtering. Head sampling is fixed at 1.0 (not operator-configurable); volume reduction is done by tail sampling in the collector. CMake integration includes a `XRPL_ENABLE_TELEMETRY` option for compile-time control.
OpenTelemetry Collector configurations are provided for development and production (with tail-based sampling, Tempo, and Elastic APM). Docker Compose examples enable quick local development environment setup.
➡️ **[View full Configuration Reference](./05-configuration-reference.md)**
---
## 6. Implementation Phases
The implementation spans 9 weeks across 5 phases:
| Phase | Duration | Focus | Key Deliverables |
| ----- | --------- | ------------------- | --------------------------------------------------- |
| 1 | Weeks 1-2 | Core Infrastructure | SDK integration, Telemetry interface, Configuration |
| 2 | Weeks 3-4 | RPC Tracing | HTTP context extraction, Handler instrumentation |
| 3 | Weeks 5-6 | Transaction Tracing | Protocol Buffer context, Relay propagation |
| 4 | Weeks 7-8 | Consensus Tracing | Round spans, Proposal/validation tracing |
| 5 | Week 9 | Documentation | Runbook, Dashboards, Training |
**Total Effort**: 47 person-days (2 developers working in parallel)
➡️ **[View full Implementation Phases](./06-implementation-phases.md)**
---
## 7. Observability Backends
> **APM** = Application Performance Monitoring | **GCS** = Google Cloud Storage
Grafana Tempo is recommended for all environments due to its cost-effectiveness and Grafana integration, while Elastic APM is ideal for organizations with existing Elastic infrastructure.
The recommended production architecture uses a gateway collector pattern with regional collectors performing tail-based sampling, routing traces to multiple backends (Tempo for primary storage, Elastic for log correlation, S3/GCS for long-term archive).
➡️ **[View Observability Backend Recommendations](./07-observability-backends.md)**
---
## 8. Appendix
The appendix contains a glossary of OpenTelemetry and xrpld-specific terms, references to external documentation and specifications, version history for this implementation plan, and a complete document index.
➡️ **[View Appendix](./08-appendix.md)**
---
_This document provides a comprehensive implementation plan for integrating OpenTelemetry distributed tracing into the xrpld XRP Ledger node software. For detailed information on any section, follow the links to the corresponding sub-documents._

View File

@@ -1,207 +0,0 @@
# Phase 2: RPC Tracing Completion Task List
> **Goal**: Complete RPC tracing coverage with unit tests, Grafana search filters, PathFind instrumentation, and config hardening. Build on the Phase 1c SpanGuard factory foundation to achieve production-quality RPC observability.
>
> **Scope**: Unit tests for core telemetry, Grafana Tempo search filters, PathFind RPC tracing, config validation (`std::clamp`).
>
> **Branch**: `pratik/otel-phase2-rpc-tracing` (from `pratik/otel-phase1c-rpc-integration`)
### Related Plan Documents
| Document | Relevance |
| ------------------------------------------------------------ | ------------------------------------------------------------- |
| [04-code-samples.md](./04-code-samples.md) | TraceContextPropagator (§4.4.2), RPC instrumentation (§4.5.3) |
| [02-design-decisions.md](./02-design-decisions.md) | W3C Trace Context (§2.5), span attributes (§2.4.2) |
| [06-implementation-phases.md](./06-implementation-phases.md) | Phase 2 tasks (§6.3), definition of done (§6.11.2) |
---
## Task 2.1: W3C Trace Context HTTP Header Extraction
**Status**: DEFERRED → Phase 3
**Reason**: W3C context propagation (`traceparent`/`tracestate` headers) requires a consumer — in Phase 2, RPC spans are entirely local to the node. Phase 3 introduces cross-node transaction tracing via protobuf context propagation, which is the first use case for extracted trace context. Implementing it here without a consumer would be dead code.
**Implemented in**: `pratik/otel-phase3-tx-tracing` — `TraceContextPropagator.h/.cpp`
---
## Task 2.2: Per-Category Span Creation
**Status**: COMPLETE (superseded by Phase 1c design)
**Original plan**: Add `XRPL_TRACE_PEER` and `XRPL_TRACE_LEDGER` macros.
**Actual implementation**: Phase 1c replaced all tracing macros with the `SpanGuard::span(TraceCategory, prefix, name)` factory pattern. The `TraceCategory` enum (`Rpc`, `Transactions`, `Consensus`, `Peer`, `Ledger`) serves the same conditional-creation purpose without macros. No separate task needed — the factory already supports all categories.
---
## Task 2.3: Add shouldTraceLedger() to Telemetry Interface
**Objective**: The `Setup` struct has a `traceLedger` field but there's no corresponding virtual method. Add it for interface completeness.
**What to do**:
- Edit `include/xrpl/telemetry/Telemetry.h`:
- Add `virtual bool shouldTraceLedger() const = 0;`
- Update all implementations:
- `src/libxrpl/telemetry/Telemetry.cpp` (TelemetryImpl, NullTelemetryOtel)
- `src/libxrpl/telemetry/NullTelemetry.cpp` (NullTelemetry)
**Key modified files**:
- `include/xrpl/telemetry/Telemetry.h`
- `src/libxrpl/telemetry/Telemetry.cpp`
- `src/libxrpl/telemetry/NullTelemetry.cpp`
---
## Task 2.4: Unit Tests for Core Telemetry Infrastructure
**Status**: COMPLETE
**Objective**: Add unit tests for the core telemetry abstractions to validate correctness and catch regressions.
**Implemented**:
- `src/tests/libxrpl/telemetry/TelemetryConfig.cpp`:
- Test Setup defaults (all fields have correct initial values)
- Test `makeTelemetrySetup` config parser (empty section, full section, edge cases)
- Test `samplingRatio` clamping (values outside 0.0-1.0)
- `src/tests/libxrpl/telemetry/SpanGuardFactory.cpp`:
- Test null guard methods are safe (setAttribute, setOk, setError, addEvent on null)
- Test category span returns null when telemetry disabled
- Test child/linked span null when no parent context
- Test move construction transfers ownership
- Test recordException safe on null guard
- Test discard() safe on null guard
- `src/tests/libxrpl/telemetry/main.cpp` — GTest runner
- `src/tests/libxrpl/CMakeLists.txt` — test target with optional OTel linking
---
## Task 2.5: Enhance RPC Span Attributes
**Status**: DEFERRED (low priority)
**Reason**: The high-value attributes (`command`, `version`, `role`, `status`) are already set by Phase 1c. The remaining HTTP transport-level attributes (`http.method`, `net.peer.ip`, `http.status_code`) provide limited additional insight since:
- `http.method` is always POST for JSON-RPC
- `net.peer.ip` is debug-level info available in logs
- `duration_ms` is redundant with span duration (OTel captures start/end time natively)
These can be added later if dashboard queries specifically need them. The node health attributes (Task 2.8) provide far more operational value and were prioritized instead.
---
## Task 2.6: Build Verification and Performance Baseline
**Objective**: Verify the build succeeds with and without telemetry, and establish a performance baseline.
**What to do**:
1. Build with `telemetry=ON` and verify no compilation errors
2. Build with `telemetry=OFF` and verify no regressions
3. Run existing unit tests to verify no breakage
4. Document any build issues in lessons.md
**Verification Checklist**:
- [ ] `conan install . --build=missing -o telemetry=True` succeeds
- [ ] `cmake --preset default -Dtelemetry=ON` configures correctly
- [ ] Build succeeds with telemetry ON
- [ ] Build succeeds with telemetry OFF
- [ ] Existing tests pass with telemetry ON
- [ ] Existing tests pass with telemetry OFF
---
## Task 2.8: RPC Span Attribute Enrichment — Node Health Context
**Status**: DROPPED.
Node health (`amendment_blocked`, `server_state`) is not part of the telemetry surface. Operators consume the same data via the existing `server_info` / `server_state` RPC commands, so duplicating it on traces adds storage and cardinality cost without new value. The OTel C++ SDK 1.18.0 also does not support runtime updates to the resource, ruling out resource-level emission of these dynamic-by-nature flags.
---
## Task 2.9: PathFind RPC Instrumentation
**Status**: COMPLETE
**Objective**: Trace the path_find and ripple_path_find RPC handlers to capture request latency and computation cost.
**Spans added**:
- `pathfind.request` — wraps `doPathFind()` and `doRipplePathFind()` RPC handlers
- `pathfind.compute` — wraps `PathRequest::doUpdate()` (`pathfind_fast` attr)
- `pathfind.update_all` — wraps `PathRequestManager::updateAll()` on ledger close (`pathfind_ledger_index`, `pathfind_num_requests` attrs; emitted only when active subscriptions exist)
- `pathfind.discover` — wraps the entire per-source-asset loop in `PathRequest::findPaths()` (`pathfind_search_level`, `pathfind_num_paths` attrs). One span per RPC call instead of N (one per source asset). Trade-off: per-asset breakdown is lost; storage and cardinality bounded.
**Attribute namespacing**: All pathfind attributes use the `pathfind_*` underscore form per the Phase 1c naming-spec rule 5.
**New file**: `src/xrpld/rpc/detail/PathFindSpanNames.h`
**Modified files**:
- `src/xrpld/rpc/handlers/orderbook/PathFind.cpp`
- `src/xrpld/rpc/handlers/orderbook/RipplePathFind.cpp`
- `src/xrpld/rpc/detail/PathRequest.cpp`
- `src/xrpld/rpc/detail/PathRequestManager.cpp`
- `src/xrpld/rpc/detail/Pathfinder.cpp`
---
## Task 2.10: RPC and PathFind Span Attribute Gap Fill
**Status**: COMPLETE
**Objective**: Wire up workflow-identifying attributes that enable filtering and grouping traces by request characteristics without drilling into child spans.
**Attributes added**:
| Span | Attribute | Type | Source |
| ------------------- | ---------------------------- | ------ | --------------------------------- |
| `rpc.http_request` | `request_payload_size` | int64 | `request.body().size()` |
| `rpc.process` | `is_batch` | bool | `method == "batch"` check |
| `rpc.process` | `batch_size` | int64 | `params.size()` (only when batch) |
| `rpc.ws_message` | `command` | string | `jv[command]` or `jv[method]` |
| `rpc.command.*` | `load_type` | string | `context.loadType.label()` |
| `pathfind.compute` | `pathfind_dest_currency` | string | `to_string(saDstAmount_.asset())` |
| `pathfind.discover` | `pathfind_num_source_assets` | int64 | `sourceAssets.size()` |
_Note: `pathfind_dest_amount` was removed — the destination amount is a financial value excluded by the privacy policy (design §2.4.4)._
**New attr keys**: `RpcSpanNames.h` (`isBatch`, `batchSize`, `loadType`), `PathFindSpanNames.h` (`destCurrency`, `numSourceAssets`).
**Modified files**:
- `src/xrpld/rpc/detail/RpcSpanNames.h`
- `src/xrpld/rpc/detail/PathFindSpanNames.h`
- `src/xrpld/rpc/detail/ServerHandler.cpp`
- `src/xrpld/rpc/detail/RPCHandler.cpp`
- `src/xrpld/rpc/detail/PathRequest.cpp`
---
## Summary
| Task | Description | Status | Notes |
| ---- | ------------------------------------------- | ------------------- | --------------------------------------------------------- |
| 2.1 | W3C Trace Context header extraction | Deferred → Phase 3 | No consumer in Phase 2; needs cross-node tracing |
| 2.2 | Per-category span creation | Complete (Phase 1c) | Superseded by TraceCategory enum + SpanGuard |
| 2.3 | Add shouldTraceLedger() interface method | Complete (Phase 1c) | Delivered in Phase 1c base branch |
| 2.4 | Unit tests for core telemetry | Complete | TelemetryConfig + SpanGuardFactory tests |
| 2.5 | Enhanced RPC span attributes (HTTP-level) | Deferred | Low value; span duration covers timing natively |
| 2.6 | Build verification and performance baseline | Complete | Verified in CI on Phase 1c |
| 2.7 | Grafana Tempo search filters | Complete | rpc-command, rpc-status, rpc-role filters |
| 2.8 | RPC span attribute enrichment (node health) | Dropped | Available via `server_info`/`server_state` RPC |
| 2.9 | PathFind RPC instrumentation | Complete | request, compute, update_all, discover |
| 2.10 | RPC/PathFind span attribute gap fill | Complete | Batch detection, payload size, load cost, pathfind params |
**Delivered in this branch**: Tasks 2.4, 2.7, 2.9, 2.10.
**Deferred with rationale**: Tasks 2.1 (→Phase 3), 2.5 (low priority).
**Dropped**: Task 2.8 (node health not duplicated on traces).
**Superseded**: Task 2.2 (Phase 1c SpanGuard factory covers this).

View File

@@ -1,547 +0,0 @@
# Phase 3: Transaction Tracing Task List
> **Goal**: Trace the full transaction lifecycle from RPC submission through peer relay, including cross-node context propagation via Protocol Buffer extensions. This is the WALK phase that demonstrates true distributed tracing.
>
> **Scope**: Protocol Buffer `TraceContext` message, context serialization, PeerImp transaction instrumentation, NetworkOPs processing instrumentation, HashRouter visibility, and multi-node relay context propagation.
>
> **Branch**: `pratik/otel-phase3-tx-tracing` (from `pratik/otel-phase2-rpc-tracing`)
### Related Plan Documents
| Document | Relevance |
| ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------ |
| [04-code-samples.md](./04-code-samples.md) | TraceContext protobuf (§4.4.1), PeerImp instrumentation (§4.5.1), context serialization (§4.4.2) |
| [01-architecture-analysis.md](./01-architecture-analysis.md) | Transaction flow (§1.3), key trace points (§1.6) |
| [06-implementation-phases.md](./06-implementation-phases.md) | Phase 3 tasks (§6.4), definition of done (§6.11.3) |
| [02-design-decisions.md](./02-design-decisions.md) | Context propagation design (§2.5), attribute schema (§2.4.3) |
---
## Task 3.1: Define TraceContext Protocol Buffer Message
**Objective**: Add trace context fields to the P2P protocol messages so trace IDs can propagate across nodes.
**What to do**:
- Edit `include/xrpl/proto/xrpl.proto` (or `src/xrpld/proto/ripple.proto`, wherever the proto is):
- Add `TraceContext` message definition:
```protobuf
message TraceContext {
bytes trace_id = 1; // 16-byte trace identifier
bytes span_id = 2; // 8-byte span identifier
uint32 trace_flags = 3; // bit 0 = sampled, bit 1 = random
reserved 4; // trace_state (W3C tracestate), added later
}
```
- Add `optional TraceContext trace_context = 1001;` to:
- `TMTransaction`
- `TMProposeSet` (for Phase 4 use)
- `TMValidation` (for Phase 4 use)
- Use high field numbers (1001+) to avoid conflicts with existing fields
- Regenerate protobuf C++ code
**Key modified files**:
- `include/xrpl/proto/xrpl.proto` (or equivalent)
**Reference**:
- [04-code-samples.md §4.4.1](./04-code-samples.md) — TraceContext message definition
- [02-design-decisions.md §2.5.2](./02-design-decisions.md) — Protocol buffer context propagation design
---
## Task 3.2: Implement Protobuf Context Serialization
**Objective**: Create utilities to serialize/deserialize OTel trace context to/from protobuf `TraceContext` messages.
**What to do**:
- Create `include/xrpl/telemetry/TraceContextPropagator.h` (extend from Phase 2 if exists, or add protobuf methods):
- Add protobuf-specific methods:
- `static Context extractFromProtobuf(protocol::TraceContext const& proto)` — reconstruct OTel context from protobuf fields
- `static void injectToProtobuf(Context const& ctx, protocol::TraceContext& proto)` — serialize current span context into protobuf fields
- Both methods guard behind `#ifdef XRPL_ENABLE_TELEMETRY`
- Create/extend `src/libxrpl/telemetry/TraceContextPropagator.cpp`:
- Implement extraction: read trace_id (16 bytes), span_id (8 bytes), trace_flags from protobuf, construct `SpanContext`, wrap in `Context`
- Implement injection: get current span from context, serialize its TraceId, SpanId, and TraceFlags into protobuf fields
**Key new/modified files**:
- `include/xrpl/telemetry/TraceContextPropagator.h`
- `src/libxrpl/telemetry/TraceContextPropagator.cpp`
**Reference**:
- [04-code-samples.md §4.4.2](./04-code-samples.md) — Full extract/inject implementation
---
## Task 3.3: Instrument PeerImp Transaction Handling
**Objective**: Add trace spans to the peer-level transaction receive and relay path.
**What to do**:
- Edit `src/xrpld/overlay/detail/PeerImp.cpp`:
- In `onMessage(TMTransaction)` / `handleTransaction()`:
- Extract parent trace context from incoming `TMTransaction::trace_context` field (if present)
- Create `tx.receive` span as child of extracted context (or new root if none)
- Set attributes: `tx_hash`, `peer_id`, `tx_status`
- Create the span only after `HashRouter::shouldProcess()` accepts, so a dropped duplicate produces no span
- Wrap validation call with child span `tx.validate`
- Wrap relay with `tx.relay` span
- When relaying to peers:
- Inject current trace context into outgoing `TMTransaction::trace_context`
- Set `relay_count` attribute
- Use `SpanGuard::span(TraceCategory::Transactions, "tx", "receive")` factory
(Phase 1c replaced macros with the SpanGuard factory pattern)
> **Note**: The `tx.receive` guard is `.detached()` before being moved into the
> `RcvCheckTx` job so its Scope is popped on the peer thread, not leaked to the
> worker (else later peer messages would inherit this transaction's trace).
**Key modified files**:
- `src/xrpld/overlay/detail/PeerImp.cpp`
**Reference**:
- [04-code-samples.md §4.5.1](./04-code-samples.md) — Full PeerImp instrumentation example
- [01-architecture-analysis.md §1.3](./01-architecture-analysis.md) — Transaction flow diagram
- [01-architecture-analysis.md §1.6](./01-architecture-analysis.md) — tx.receive trace point
---
## Task 3.4: Instrument NetworkOPs Transaction Processing
**Objective**: Trace the transaction processing pipeline in NetworkOPs, covering both sync and async paths.
**What to do**:
- Edit `src/xrpld/app/misc/NetworkOPs.cpp`:
- In `processTransaction()`:
- Create `tx.process` span
- Set attributes: `tx_hash`, `tx_type`, `local` (whether from RPC or peer)
- Record whether sync or async path is taken
- `.detached()` the guard before storing it in `TransactionStatus::span`,
since it is applied on a batch worker thread — this pops the Scope on the
origin thread and stops later work inheriting this transaction's trace
- In `doTransactionAsync()`:
- Capture parent context before queuing
- Create `tx.queue` span with queue depth attribute
- Add event when transaction is dequeued for processing
- In `doTransactionSync()`:
- Create `tx.process_sync` span
- Record result (applied, queued, rejected)
**Key modified files**:
- `src/xrpld/app/misc/NetworkOPs.cpp`
**Reference**:
- [01-architecture-analysis.md §1.6](./01-architecture-analysis.md) — tx.validate and tx.process trace points
- [02-design-decisions.md §2.4.3](./02-design-decisions.md) — Transaction attribute schema
---
## Task 3.5: Instrument HashRouter for Dedup Visibility
**Objective**: Make transaction deduplication visible in traces by recording HashRouter decisions as span attributes/events.
**What to do**:
- Edit `src/xrpld/overlay/detail/PeerImp.cpp` (in handleTransaction):
- After calling `HashRouter::shouldProcess()` or `addSuppressionPeer()`:
- Start the span here, not before, so only transactions this node will process are traced
- Record `tx_flags` showing current HashRouter state (SAVED, TRUSTED, etc.)
- Add `tx.first_seen` or `tx.duplicate` event
- This is NOT a modification to HashRouter itself — just recording its decisions as span attributes in the existing PeerImp instrumentation from Task 3.3.
**Key modified files**:
- `src/xrpld/overlay/detail/PeerImp.cpp` (same changes as 3.3, logically grouped)
---
## Task 3.6: Context Propagation in Transaction Relay
**Status**: COMPLETE (transaction relay). Consensus proposal/validation
propagation is deferred to Phase 4 — see "Planned (Phase 4)" below.
**Objective**: Ensure trace context flows correctly when transactions are relayed between peers, creating linked spans across nodes.
**What was done**:
- **TX send side**: `NetworkOPs::apply()` now injects the tx.process span's trace
context into the outgoing `TMTransaction` protobuf before relay, using
`telemetry::injectSpanContext()`. The receiving node's `txReceiveSpan()` (already
wired in PeerImp) extracts the parent span_id and creates the tx.receive span
as a child of the sender's tx.process span.
- **Edge cases**: Missing trace context (older peers) degrades gracefully to
standalone spans. Invalid/corrupted context is treated as absent. Trace
flags are propagated and respected.
**New infrastructure**:
- `SpanGuard::getTraceBytes()` — extracts raw trace_id/span_id/trace_flags
from a span without exposing OTel types. Safe to call from any thread.
- `PropagationHelpers.h` — `injectSpanContext(SpanGuard&, proto)` bridge
between SpanGuard and protobuf TraceContext.
- `TraceContextPropagator.h` — `injectToProtobuf(ctx, proto)` for
same-thread injection via OTel RuntimeContext.
**Key modified files**:
- `src/xrpld/app/misc/NetworkOPs.cpp` — tx relay injection
- `include/xrpl/telemetry/SpanGuard.h` — `TraceBytes` struct, `getTraceBytes()`
- `src/libxrpl/telemetry/SpanGuard.cpp` — `getTraceBytes()` implementation
- `src/xrpld/telemetry/PropagationHelpers.h` — inject helpers (new file)
**Planned (Phase 4 — not in this PR)**:
The consensus proposal/validation propagation below is Phase 4 scope and is
not implemented on this branch. It is listed here only to record the intended
design.
- **Proposal send/receive**: `RCLConsensus::Adaptor::propose()` injects the
current thread's active span context into the `TMProposeSet` protobuf via
`telemetry::injectToProtobuf()`. PeerImp creates a
`consensus.proposal.receive` span that extracts the sender's trace context
as parent (via `ConsensusReceiveTracing.h`).
- **Validation send/receive**: `RCLConsensus::Adaptor::validate()` injects
the current thread's active span context into the `TMValidation` protobuf.
PeerImp creates a `consensus.validation.receive` span that extracts the
sender's trace context as parent.
- Planned files: `src/xrpld/app/consensus/RCLConsensus.cpp` (send injection),
`src/xrpld/overlay/detail/PeerImp.cpp` (receive spans),
`src/xrpld/telemetry/ConsensusReceiveTracing.h` (receive span helpers,
new file).
**Reference**:
- [02-design-decisions.md §2.5](./02-design-decisions.md) — Context propagation design
- [04-code-samples.md §4.5.1](./04-code-samples.md) — Relay context injection pattern
---
## Task 3.7: Build Verification and Testing
**Objective**: Verify all Phase 3 changes compile and work correctly.
**What to do**:
1. Build with `telemetry=ON` — verify no compilation errors
2. Build with `telemetry=OFF` — verify no regressions
3. Run existing unit tests
4. Verify protobuf regeneration produces correct C++ code
5. Document any issues encountered
**Verification Checklist**:
- [ ] Protobuf changes generate valid C++
- [ ] Build succeeds with telemetry ON
- [ ] Build succeeds with telemetry OFF
- [ ] Existing tests pass
- [ ] No undefined symbols from new telemetry calls
---
## Task 3.8: Transaction Span Peer Version Attribute
> **Source**: [External Dashboard Parity](../docs/superpowers/specs/2026-03-30-external-dashboard-parity-design.md) — adds peer version context inspired by the community [xrpl-validator-dashboard](https://github.com/realgrapedrop/xrpl-validator-dashboard).
>
> **Upstream**: Phase 2 (RPC span infrastructure must exist).
> **Downstream**: Phase 10 (validation checks for this attribute).
**Objective**: Add the relaying peer's xrpld version to `tx.receive` spans so operators can correlate transaction issues with peer version mismatches during network upgrades.
**What to do**:
- Edit `src/xrpld/overlay/detail/PeerImp.cpp`:
- In the `tx.receive` span block (after existing `peer_id` setAttribute call):
- Add `peer_version` (string) — from `this->getVersion()`
- Only set if `getVersion()` returns a non-empty string (avoid empty-string attributes)
**New span attribute**:
| Attribute | Type | Source | Example |
| -------------- | ------ | -------------------- | --------------- |
| `peer_version` | string | `peer->getVersion()` | `"xrpld-2.4.0"` |
**Rationale**: Transaction relay is where version mismatches cause subtle serialization or validation bugs. Tracing "this tx came from a v2.3.0 peer" helps diagnose compatibility issues. The community dashboard tracks peer versions externally; this brings version awareness into the trace itself.
**Key modified files**:
- `src/xrpld/overlay/detail/PeerImp.cpp`
**Exit Criteria**:
- [ ] `tx.receive` spans carry `peer_version` attribute with a non-empty version string
- [ ] Attribute is omitted (not set to empty string) when `getVersion()` returns empty
- [ ] Attribute visible in Jaeger span detail view
---
## Task 3.9: Deterministic Transaction Trace ID
> **Upstream**: Task 3.2 (protobuf serialization), Task 3.3 (PeerImp span exists).
> **Downstream**: Phase 10 (workload validation can query by tx hash directly).
> **Pattern**: Mirrors the consensus deterministic trace ID in Phase 4a
> (`createDeterministicContext` in `RCLConsensus.cpp`), adapted for transactions.
**Objective**: Derive the trace_id for transaction spans deterministically from the
transaction hash so that all nodes handling the same transaction independently produce
spans under the same trace_id — regardless of whether protobuf context propagation
succeeds.
**Why**: The current approach creates spans with random trace_ids and relies entirely
on protobuf `TraceContext` propagation to link them. If any hop in the relay chain
drops the context (older peers, message corruption, mixed-version networks), the trace
splits and downstream spans become impossible to find. With deterministic trace_ids,
correlation is guaranteed because every node derives the same trace_id from the same
`txID`.
**Approach — deterministic trace_id + protobuf span_id propagation**:
1. Derive `trace_id = txHash[0:16]` (first 16 bytes of the 32-byte transaction hash).
2. Generate a random 8-byte `span_id` per node (each node's span is unique within
the shared trace).
3. Create the span under this deterministic context as parent.
4. **Additionally**, if protobuf `TraceContext` is present in the incoming
`TMTransaction` message, extract the sender's `span_id` and use it as the span's
parent — this preserves parent-child ordering in the trace tree.
5. If protobuf context is absent (older peer, first hop), the span still has the
correct deterministic `trace_id` — it appears as a sibling root in the same trace
rather than being lost.
This gives the best of both worlds: guaranteed cross-node correlation via deterministic
`trace_id`, plus parent-child relay ordering via protobuf `span_id` when available.
**What to do**:
- Create `createDeterministicTxContext(uint256 const& txHash)` utility function:
- Location: shared header or file-local in `PeerImp.cpp` and `NetworkOPs.cpp`
(or a shared telemetry utility if both need it).
- Pattern: identical to `createDeterministicContext(uint256 const& ledgerId)` in
`RCLConsensus.cpp` — take `txHash[0:16]` as trace_id, random span_id via
`default_prng()`, sampled flag set, `remote=false`.
- Guard behind `#ifdef XRPL_ENABLE_TELEMETRY`.
```cpp
opentelemetry::context::Context
createDeterministicTxContext(uint256 const& txHash)
{
namespace trace = opentelemetry::trace;
// First 16 bytes of the 32-byte tx hash as trace ID.
trace::TraceId traceId(
opentelemetry::nostd::span<uint8_t const, 16>(txHash.data(), 16));
// Random span_id so each node's span is unique within the trace.
uint8_t spanIdBytes[8];
auto const rval = default_prng()();
std::memcpy(spanIdBytes, &rval, sizeof(spanIdBytes));
trace::SpanId spanId(
opentelemetry::nostd::span<uint8_t const, 8>(spanIdBytes, 8));
trace::SpanContext syntheticCtx(
traceId, spanId, trace::TraceFlags(1), /* remote = */ false);
return opentelemetry::context::Context{}.SetValue(
trace::kSpanKey,
opentelemetry::nostd::shared_ptr<trace::Span>(
new trace::DefaultSpan(syntheticCtx)));
}
```
- Edit `src/xrpld/overlay/detail/PeerImp.cpp` — restructure `handleTransaction()`:
- **Move span creation after deserialization** (txID must be known first):
1. Deserialize `STTx` and get `txID` (existing code at line ~1382).
2. Create deterministic parent context: `auto detCtx = createDeterministicTxContext(txID)`.
3. If `m->has_trace_context()`: extract protobuf context via `extractFromProtobuf()`,
**combine** with deterministic trace_id — use the protobuf span_id as parent
to preserve relay ordering, but override trace_id with the deterministic one.
4. If no protobuf context: create span under `detCtx` directly.
5. Set all existing attributes (`hash`, `peerId`, `peerVersion`, etc.).
- **Combining deterministic trace_id with protobuf parent span_id**:
When both are available, construct a synthetic `SpanContext` with:
- `trace_id` = `txHash[0:16]` (deterministic)
- `span_id` = extracted from protobuf (sender's span_id → becomes parent)
- `trace_flags` = from protobuf
- `remote` = true (came from another node)
```cpp
// Pseudo-code for the combined context:
auto detTraceId = trace::TraceId(txHash.data(), 16);
auto remoteSpanId = /* from extractFromProtobuf */;
auto remoteFlags = /* from extractFromProtobuf */;
trace::SpanContext combinedCtx(
detTraceId, remoteSpanId, remoteFlags, /* remote = */ true);
// Use as parent context for the new span.
```
- Edit `src/xrpld/app/misc/NetworkOPs.cpp` — update `processTransaction()`:
- `transaction->getID()` is already available at the top of the function.
- Create deterministic parent context from `txID`.
- Create `tx.process` span under this context.
- No protobuf context to extract here (NetworkOPs is intra-node), so
deterministic context alone is sufficient.
- Add `trace_strategy` attribute to spans:
- Add `inline constexpr auto traceStrategy = "trace_strategy";`
to `TxSpanNames.h`.
- Set on each tx span: `span.setAttribute(tx_span::attr::traceStrategy, "deterministic")`.
**Key new/modified files**:
- `src/xrpld/overlay/detail/PeerImp.cpp` — restructured span creation
- `src/xrpld/app/misc/NetworkOPs.cpp` — deterministic context for tx.process
- `src/xrpld/telemetry/TxSpanNames.h` — new `traceStrategy` attribute constant
- New or shared utility for `createDeterministicTxContext()` (location TBD: could be
a shared header like `include/xrpl/telemetry/DeterministicContext.h`, or file-local
if only used in two places)
**Interaction with existing tasks**:
- **Task 3.3 (PeerImp instrumentation)**: The span creation in `handleTransaction()`
must be restructured — the span currently starts before `txID` is known. This task
moves it after deserialization.
- **Task 3.6 (Relay context propagation)**: Protobuf injection at the relay site
remains the same — `injectToProtobuf()` serializes the current span's `span_id`.
The receiver extracts it and combines with the deterministic `trace_id`.
- **Phase 4a (Consensus deterministic trace ID)**: This task follows the same pattern.
Consider extracting a shared utility (e.g., `createDeterministicContext(uint256)`)
that both consensus and transaction tracing use.
**Exit Criteria**:
- [ ] `tx.receive` and `tx.process` spans have deterministic trace_id = `txHash[0:16]`
- [ ] All nodes handling the same transaction produce spans under the same trace_id
- [x] Protobuf `span_id` propagation still works when available (parent-child ordering)
- [ ] Missing protobuf context (old peer) degrades gracefully to sibling spans, not lost traces
- [ ] `trace_strategy` attribute set to `"deterministic"` on all tx spans
- [ ] Trace queryable by tx hash (truncate hash → trace_id → direct lookup in Tempo)
**Deliverables implemented (not in original plan)**:
- **`SpanGuard::txSpan()` factory method** (`include/xrpl/telemetry/SpanGuard.h`):
Two overloads for creating transaction spans with deterministic trace IDs:
- `txSpan(category, group, name, txHash)` — standalone span (deterministic
trace_id from `txHash[0:16]`, no parent span_id).
- `txSpan(category, group, name, txHash, parentCtx)` — child span (deterministic
trace_id combined with protobuf-extracted parent span_id for relay ordering).
- **`TxTracing.h` helper functions** (`src/xrpld/telemetry/TxTracing.h`):
File-local helpers that wrap `SpanGuard::txSpan()` for the two main PeerImp call
sites:
- `txReceiveSpan(txHash, parentCtx)` — creates `tx.receive` span with
deterministic trace_id and optional protobuf parent context.
- `txProcessSpan(txHash)` — creates `tx.process` span with deterministic
trace_id only (no protobuf parent, used intra-node).
- **Note**: `TxTracing.h` includes `xrpl.pb.h` unconditionally (outside
`#ifdef XRPL_ENABLE_TELEMETRY`) because `protocol::TMTransaction` appears in
the function signatures regardless of telemetry build mode.
---
## Task 3.10: TxQ Instrumentation
**Status**: COMPLETE
**Objective**: Trace the transaction queue lifecycle — enqueue decisions, direct apply, batch clear, ledger-close accept loop, per-tx apply, and cleanup.
**Spans added**:
- `txq.enqueue` — wraps `TxQ::apply()` with tx_hash attribute
- `txq.apply_direct` — wraps `TxQ::tryDirectApply()` fast-path
- `txq.batch_clear` — wraps `TxQ::tryClearAccountQueueUpThruTx()`
- `txq.accept` — wraps `TxQ::accept()` ledger-close dequeue with queue_size attr
- `txq.accept_tx` — per-tx span inside accept loop with tx_hash, ter_code,
retries_remaining attributes
- `txq.cleanup` — wraps `TxQ::processClosedLedger()` with ledger_seq attribute
**New file**: `src/xrpld/app/misc/detail/TxQSpanNames.h`
**Modified file**: `src/xrpld/app/misc/detail/TxQ.cpp`
---
## Task 3.11: TX and TxQ Span Attribute Gap Fill
**Status**: COMPLETE
**Objective**: Add workflow-identifying attributes to transaction spans so operators can filter by transaction type and see outcomes without off-chain correlation.
**Attributes added**:
| Span | Attribute | Type | Source |
| ----------------- | -------------------- | ------ | -------------------------------------------------------------------------------------------------------------------------- |
| `tx.process` | `tx_type` | string | `TxFormats::getInstance().findByType(stx->getTxnType())->getName()` |
| `tx.process` | `fee` | int64 | `stx->getFieldAmount(sfFee).xrp().drops()` |
| `tx.process` | `sequence` | int64 | `stx->getSeqProxy().value()` |
| `tx.process` | `tx_<account field>` | string | one per top-level `STI_ACCOUNT` field, raw r-address (`tx_account`, `tx_destination`, ...); keys in `TxAccountSpanNames.h` |
| `tx.process` | `ter_result` | string | `transToken(e.result)` (set after batch application) |
| `tx.process` | `applied` | bool | `e.applied` (set after batch application) |
| `tx.receive` | `tx_type` | string | `TxFormats::getInstance().findByType(stx->getTxnType())->getName()` |
| `txq.enqueue` | `tx_type` | string | same pattern as above |
| `txq.enqueue` | `txq_status` | string | `queued` / `applied_direct` / `applied` / `failed` / `rejected` |
| `txq.enqueue` | `ter_code` | string | `transToken(directApplied->ter)` (set on the direct-apply path) |
| `txq.enqueue` | `fee_level_paid` | int64 | `getFeeLevelPaid(view, *tx).value()` |
| `txq.enqueue` | `required_fee_level` | int64 | `getRequiredFeeLevel(...).value()` |
| `txq.batch_clear` | `num_cleared` | int64 | queued txs cleared ahead of the applying tx |
| `txq.cleanup` | `expired_count` | int64 | entries dropped for passed `LastLedgerSequence` |
| `txq.accept_tx` | `txq_status` | string | `applied` / `failed` / `retried` |
| `txq.accept_tx` | `ter_code` | string | `transToken(txnResult)` (set before branching on the outcome) |
| `txq.accept` | `ledger_changed` | bool | set at end of accept loop |
**New attr keys**: `TxSpanNames.h` (`txType`, `fee`, `sequence`, `terResult`, `applied`), `TxQSpanNames.h` (`txType`).
**Modified files**:
- `src/xrpld/telemetry/TxSpanNames.h`
- `src/xrpld/app/misc/detail/TxQSpanNames.h`
- `src/xrpld/app/misc/NetworkOPs.cpp`
- `src/xrpld/overlay/detail/PeerImp.cpp`
- `src/xrpld/app/misc/detail/TxQ.cpp`
---
## Summary
| Task | Description | New Files | Modified Files | Depends On |
| ---- | ----------------------------------- | --------- | -------------- | ---------- |
| 3.1 | TraceContext protobuf message | 0 | 1 | Phase 2 |
| 3.2 | Protobuf context serialization | 1-2 | 0 | 3.1 |
| 3.3 | PeerImp transaction instrumentation | 0 | 1 | 3.2 |
| 3.4 | NetworkOPs transaction processing | 0 | 1 | Phase 2 |
| 3.5 | HashRouter dedup visibility | 0 | 1 | 3.3 |
| 3.6 | Relay context propagation | 0 | 1-2 | 3.3, 3.5 |
| 3.7 | Build verification and testing | 0 | 0 | 3.1-3.6 |
| 3.8 | TX span peer version attribute | 0 | 1 | 3.3 |
| 3.9 | Deterministic transaction trace ID | 0-1 | 3 | 3.2, 3.3 |
| 3.10 | TxQ instrumentation (6 spans) | 1 | 1 | 3.4 |
| 3.11 | TX/TxQ span attribute gap fill | 0 | 5 | 3.3, 3.10 |
**Parallel work**: Tasks 3.1 and 3.4 can start in parallel. Task 3.2 depends on 3.1. Tasks 3.3 and 3.5 depend on 3.2. Task 3.6 depends on 3.3 and 3.5. Task 3.8 depends on 3.3 (span must exist). Task 3.9 depends on 3.2 and 3.3. Task 3.10 depends on 3.4 (tx.process span must exist).
**Exit Criteria** (from [06-implementation-phases.md §6.11.3](./06-implementation-phases.md)):
- [x] Transaction traces span across nodes
- [x] Trace context in Protocol Buffer messages
- [ ] HashRouter deduplication visible in traces
- [ ] <5% overhead on transaction throughput
- [x] Deterministic trace_id: same trace_id for same tx across all nodes
- [x] Protobuf span_id propagation preserves parent-child ordering when available

View File

@@ -1,949 +0,0 @@
# Phase 4: Consensus Tracing Task List
> **Goal**: Full observability into consensus rounds — track round lifecycle, phase transitions, proposal handling, and validation. This is the RUN phase that completes the distributed tracing story.
>
> **Scope**: RCLConsensus instrumentation for round starts, phase transitions (open/establish/accept), proposal send/receive, validation handling, and correlation with transaction traces from Phase 3.
>
> **Branch**: `pratik/otel-phase4-consensus-tracing` (from `pratik/otel-phase3-tx-tracing`)
> **Note on attribute names**: the `xrpl.<domain>.<field>` keys shown below are
> written in the older dotted form for readability — it mirrors how the fully
> qualified attribute reads in a Tempo trace view. The implemented keys follow
> the convention in [CONTRIBUTING.md](../CONTRIBUTING.md#telemetry-span-attribute-naming)
> (underscore form, e.g. `consensus_round`, `consensus_mode`); the
> `*SpanNames.h` constants are the single source of truth.
### Related Plan Documents
| Document | Relevance |
| ------------------------------------------------------------ | ----------------------------------------------------------- |
| [04-code-samples.md](./04-code-samples.md) | Consensus instrumentation (§4.5.2), consensus span patterns |
| [01-architecture-analysis.md](./01-architecture-analysis.md) | Consensus round flow (§1.4), key trace points (§1.6) |
| [06-implementation-phases.md](./06-implementation-phases.md) | Phase 4 tasks (§6.5), definition of done (§6.11.4) |
| [02-design-decisions.md](./02-design-decisions.md) | Consensus attribute schema (§2.4.4) |
---
## Task 4.1: Instrument Consensus Round Start ✅
**Objective**: Create a root span for each consensus round that captures the round's key parameters.
**Status**: DONE (implemented via Task 4a.2 `startRoundTracing()` helper).
**What was done**:
- `RCLConsensus::Adaptor::startRoundTracing()` creates `consensus.round` span
via `SpanGuard::hashSpan()` (deterministic) or `SpanGuard::span()` (attribute strategy)
- Attributes set: `xrpl.consensus.ledger_id`, `xrpl.ledger.seq`,
`xrpl.consensus.mode`, `trace_strategy`, `xrpl.consensus.round_id`
- Round span stored as `roundSpan_` member in `RCLConsensus::Adaptor`
- `roundSpanContext_` snapshot captured for cross-thread span linking
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
- `src/xrpld/app/consensus/RCLConsensus.h` (span and context members)
**Reference**:
- [04-code-samples.md §4.5.2](./04-code-samples.md) — startRound instrumentation example
- [01-architecture-analysis.md §1.4](./01-architecture-analysis.md) — Consensus round flow
---
## Task 4.2: Instrument Phase Transitions ✅
**Objective**: Create child spans for each consensus phase (open, establish, accept) to show timing breakdown.
**Status**: DONE. All consensus phases are now instrumented:
- `consensus.establish` — created in `Consensus.h::startEstablishTracing()`
- `consensus.ledger_close` — created in `RCLConsensus.cpp::onClose()`
- `consensus.accept` / `consensus.accept.apply` — created in `onAccept()` / `doAccept()`
- `consensus.phase.open` — `openSpan_` member in `Consensus.h`, created in `startRoundInternal()`, ended in `closeLedger()`
**Design notes**:
- `phase` attribute — phases are distinguished by span names instead
- `phase.enter` / `phase.exit` events — not added (span start/end serves this purpose)
- `phase_duration_ms` attribute — not set (span duration captures this)
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
- `src/xrpld/consensus/Consensus.h` (template-level establish phase tracking)
**Reference**:
- [04-code-samples.md §4.5.2](./04-code-samples.md) — phaseTransition instrumentation
---
## Task 4.3: Instrument Proposal Handling ✅
**Objective**: Trace proposal send and receive to show validator coordination.
**Status**: DONE. Both send and receive paths are instrumented.
**What was done**:
- In `Adaptor::propose()`:
- Creates `consensus.proposal.send` span via `SpanGuard::span()`
- Sets `xrpl.consensus.round` attribute
- In `PeerImp::onMessage(TMProposeSet)`:
- Creates `consensus.proposal.receive` span
- Sets `trusted` attribute (bool)
**Done here** (cross-node propagation, send + receive):
- Trace context injection for `TMProposeSet::trace_context` in `propose()`
- Receive-side extraction in `PeerImp::onMessage(TMProposeSet)` via
`telemetry::proposalReceiveSpan()` (parents the receive span on the
sender's context when a valid `trace_context` is present)
**Not implemented** (deferred to Phase 4b):
- `consensus.proposal.relay` span in `share(RCLCxPeerPos)`
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
**Reference**:
- [04-code-samples.md §4.5.2](./04-code-samples.md) — peerProposal instrumentation
- [02-design-decisions.md §2.4.4](./02-design-decisions.md) — Consensus attribute schema
---
## Task 4.4: Instrument Validation Handling ✅
**Objective**: Trace validation send and receive to show ledger validation flow.
**Status**: DONE. Both send and receive paths are instrumented.
**What was done**:
- In `Adaptor::validate()` (called from `doAccept()`):
- Creates `consensus.validation.send` span via `Adaptor::createValidationSpan()`
- Uses `SpanGuard::linkedSpan()` to create a follows-from link to the round span
- Thread-safe: uses `roundSpanContext_` snapshot (captured on consensus thread,
read on jtACCEPT thread)
- Sets `xrpl.ledger.seq` and `proposing` attributes
- In `PeerImp::onMessage(TMValidation)`:
- Creates `consensus.validation.receive` span
- Sets `trusted` attribute (bool)
- Sets `xrpl.ledger.seq` attribute
**Not implemented** (deferred to Phase 4b — cross-node propagation):
- Validated ledger hash, signing time attributes on send span (see Task 4.8)
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
---
## Task 4.5: Add Consensus-Specific Attributes ✅
**Objective**: Enrich consensus spans with detailed attributes for debugging and analysis.
**Status**: DONE. All core attributes are set across various spans, including the previously missing `tx_count` and `disputes_count`.
**Implemented attributes** (across various spans):
- `xrpl.ledger.seq` — on `consensus.round`, `consensus.accept.apply`
- `xrpl.consensus.round` — on `consensus.proposal.send`
- `xrpl.consensus.mode` — on `consensus.round`, `consensus.ledger_close`
- `proposers` — on `consensus.accept`, `consensus.establish`, `consensus.update_positions`
- `converge_percent` — on `consensus.establish`, `consensus.update_positions`, `consensus.check`
- `tx_count` — on `consensus.accept.apply` span (in `doAccept()`)
- `disputes_count` — on `consensus.update_positions` span (in `updateOurPositions()`)
- `close_reason` — on `consensus.phase.open`, from `whyCloseLedger()` (`anomaly`, `others_closed`, `idle`, `normal`)
- `close_time_avalanche_state` — on `consensus.establish`, terminal regime (`init`, `mid`, `late`, `stuck`)
**Design notes**:
- `phase` — phases distinguished by span names instead
- `phase_duration_ms` — span duration captures this
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
- `src/xrpld/consensus/Consensus.h`
---
## Task 4.6: Correlate Transaction and Consensus Traces ✅
**Objective**: Link transaction traces from Phase 3 with consensus traces so you can follow a transaction from submission through consensus into the ledger.
**Status**: DONE. Transaction-consensus correlation implemented via `tx.included` events in `doAccept()`.
**What was done**:
- In `doAccept()` (RCLConsensus.cpp):
- Records `tx.included` events on the `consensus.accept.apply` span for each transaction in the accepted set
- Each event includes `xrpl.tx.id` attribute with the transaction hash
- This links consensus traces to individual transactions
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
---
## Task 4.7: Build Verification and Testing ✅
**Objective**: Verify all Phase 4 changes compile and don't affect consensus timing.
**What to do**:
1. Build with `telemetry=ON` — verify no compilation errors
2. Build with `telemetry=OFF` — verify no regressions (critical for consensus code)
3. Run existing consensus-related unit tests
4. Verify that `SpanGuard` factory methods compile to no-ops when disabled
5. Check that no consensus-critical code paths are affected by instrumentation overhead
**Verification Checklist**:
- [x] Build succeeds with telemetry ON
- [x] Build succeeds with telemetry OFF
- [x] Existing consensus tests pass
- [x] `SpanGuard` no-op implementation prevents overhead when telemetry is OFF
- [x] Phase timing instrumentation doesn't use blocking operations
---
## Task 4.8: Consensus Validation Span Enrichment — NOT DONE
> **Source**: [External Dashboard Parity](../docs/superpowers/specs/2026-03-30-external-dashboard-parity-design.md) — adds validation agreement context inspired by the community [xrpl-validator-dashboard](https://github.com/realgrapedrop/xrpl-validator-dashboard).
>
> **Upstream**: Phase 4 tasks 4.1-4.4 (span creation must exist).
> **Downstream**: Phase 7 (ValidationTracker reads these attributes), Phase 10 (validation checks).
**Objective**: Add ledger hash, validation type, and quorum data to consensus validation spans on both send and receive paths. This enables trace-level validation agreement analysis — filter by ledger hash to see which validators agreed for a given ledger.
**Status**: Not implemented. None of the enrichment attributes are set. The `consensus.validation.send` span only has `ledger.seq` and `proposing`. The `consensus.accept` span has `quorum` set to `result.proposers` (not the actual validator quorum from `app_.validators().quorum()`). No `PeerImp.cpp` changes were made.
**What to do**:
- Edit `src/xrpld/app/consensus/RCLConsensus.cpp`:
- On the `consensus.validation.send` span (in `validate()` / `doAccept()`):
- Add `xrpl.validation.ledger_hash` (string) — the ledger hash being validated
- Add `xrpl.validation.full` (bool) — whether this is a full validation (not partial)
- On the `consensus.accept` span (in `onAccept()`):
- Add `validation_quorum` (int64) — from `app_.validators().quorum()`
- Add `proposers_validated` (int64) — from `result.proposers`
- Edit `src/xrpld/overlay/detail/PeerImp.cpp`:
- On the `peer.validation.receive` span:
- Add `xrpl.peer.validation.ledger_hash` (string) — from deserialized `STValidation` object
- Add `xrpl.peer.validation.full` (bool) — from `STValidation` flags
**New span attributes**:
| Span | Attribute | Type | Source |
| --------------------------- | ---------------------------------- | ------ | --------------------------------- |
| `consensus.validation.send` | `xrpl.validation.ledger_hash` | string | Ledger hash from validate() args |
| `consensus.validation.send` | `xrpl.validation.full` | bool | Full vs partial validation |
| `peer.validation.receive` | `xrpl.peer.validation.ledger_hash` | string | From STValidation deserialization |
| `peer.validation.receive` | `xrpl.peer.validation.full` | bool | From STValidation flags |
| `consensus.accept` | `validation_quorum` | int64 | `app_.validators().quorum()` |
| `consensus.accept` | `proposers_validated` | int64 | `result.proposers` |
**Rationale**: The external dashboard's most valuable feature is validation agreement tracking. By recording the ledger hash on both outgoing and incoming validation spans, we create the raw data for agreement analysis at the trace level. Example Tempo query:
```
{name="consensus.validation.send"} | xrpl.validation.ledger_hash = "A1B2C3..."
```
Phase 7's `ValidationTracker` builds metric-level aggregation (1h/24h agreement %) on top of this data.
**Key modified files (not yet modified)**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
- `src/xrpld/overlay/detail/PeerImp.cpp`
**Exit Criteria**:
- [x] `consensus.validation.send` spans carry `ledger_hash` and `full_validation`
- [ ] `peer.validation.receive` spans carry `xrpl.peer.validation.ledger_hash` and `xrpl.peer.validation.full`
- [ ] `consensus.accept` spans carry `validation_quorum` and `proposers_validated`
- [x] Ledger hash attributes match between send and receive for the same ledger
- [ ] No impact on consensus performance
---
## Task 4.9: Consensus Span Attribute Gap Fill
**Status**: COMPLETE
**Objective**: Add workflow-critical attributes to consensus spans that enable operators to understand consensus outcomes, identify bow-out proposals, and correlate validations to specific ledgers.
**Attributes added**:
| Span | Attribute | Type | Source |
| --------------------------- | ----------------- | ------ | ------------------------------------- |
| `consensus.proposal.send` | `is_bow_out` | bool | `proposal.isBowOut()` |
| `consensus.accept` | `consensus_state` | string | `result.state` (yes/moved_on/expired) |
| `consensus.accept` | `disputes_count` | int64 | `result.disputes.size()` |
| `consensus.validation.send` | `ledger_hash` | string | `ledger.ledger->header().hash` |
**New attr keys**: `ConsensusSpanNames.h` (`isBowOut`, `ledgerHash`).
**Modified files**:
- `src/xrpld/consensus/ConsensusSpanNames.h`
- `src/xrpld/app/consensus/RCLConsensus.cpp`
---
## Summary
| Task | Description | Status | New Files | Modified Files | Depends On |
| ---- | ------------------------------------------- | ----------- | --------- | -------------- | ------------- |
| 4.1 | Consensus round start instrumentation | ✅ Done | 0 | 2 | Phase 3 |
| 4.2 | Phase transition instrumentation | ✅ Done | 0 | 1-2 | 4.1 |
| 4.3 | Proposal handling instrumentation | ✅ Done | 0 | 2 | 4.1 |
| 4.4 | Validation handling instrumentation | ✅ Done | 0 | 2 | 4.1 |
| 4.5 | Consensus-specific attributes | ✅ Done | 0 | 2 | 4.2, 4.3, 4.4 |
| 4.6 | Transaction-consensus correlation | ✅ Done | 0 | 1 | 4.2, Phase 3 |
| 4.7 | Build verification and testing | ✅ Done | 0 | 0 | 4.1-4.6 |
| 4.8 | Validation span enrichment (ext. dashboard) | ❌ Not done | 0 | 2 | 4.4 |
| 4.9 | Consensus span attribute gap fill | ✅ Done | 0 | 2 | 4.1-4.5 |
**Parallel work**: Tasks 4.2, 4.3, and 4.4 can run in parallel after 4.1 is complete. Task 4.5 depends on all three. Task 4.6 depends on 4.2 and Phase 3. Task 4.8 depends on 4.4 (validation spans must exist).
### Implemented Spans
| Span Name | Method | Key Attributes |
| --------------------------- | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `consensus.proposal.send` | `Adaptor::propose` | `xrpl.consensus.round`, `is_bow_out` |
| `consensus.ledger_close` | `Adaptor::onClose` | `xrpl.ledger.seq`, `xrpl.consensus.mode` |
| `consensus.accept` | `Adaptor::onAccept` | `proposers`, `round_time_ms`, `quorum`, `disputes_count`, `consensus_state` |
| `consensus.accept.apply` | `Adaptor::doAccept` | `close_time_ripple_epoch_s`, `close_time_correct`, `close_resolution_ms`, `consensus_state`, `proposing`, `round_time_ms`, `xrpl.ledger.seq`, `parent_close_time_ripple_epoch_s`, `close_time_self_ripple_epoch_s`, `close_time_vote_bins`, `resolution_direction` |
| `consensus.validation.send` | `Adaptor::onAccept` (via validate) | `proposing`, `ledger_hash`, `ledger_seq`, `full_validation`, `validation_sign_time` |
#### Close Time Attributes (consensus.accept.apply)
The `consensus.accept.apply` span captures ledger close time agreement details
driven by `avCT_CONSENSUS_PCT` (75% validator agreement threshold):
- **`close_time_ripple_epoch_s`** — Agreed-upon ledger close time (XRPL epoch seconds). When validators disagree (`consensusCloseTime == epoch`), this is synthetically set to `prevCloseTime + 1s`.
- **`close_time_correct`** — `true` if validators reached agreement, `false` if they "agreed to disagree" (close time forced to prev+1s).
- **`close_resolution_ms`** — Rounding granularity for close time (starts at 30s, decreases as ledger interval stabilizes).
- **`consensus_state`** — `"finished"` (normal) or `"moved_on"` (consensus failed, adopted best available).
- **`proposing`** — Whether this node was proposing.
- **`round_time_ms`** — Total consensus round duration.
- **`parent_close_time_ripple_epoch_s`** — Previous ledger's close time (XRPL epoch seconds). Enables computing close-time deltas across consecutive rounds without correlating separate spans.
- **`close_time_self_ripple_epoch_s`** — This node's own proposed close time before consensus voting.
- **`close_time_vote_bins`** — Number of distinct close-time vote bins from peer proposals. Higher values indicate less agreement among validators.
- **`resolution_direction`** — Whether close-time resolution `"increased"` (coarser), `"decreased"` (finer), or stayed `"unchanged"` relative to the previous ledger.
**Exit Criteria** (from [06-implementation-phases.md §6.11.4](./06-implementation-phases.md)):
- [x] Complete consensus round traces
- [x] Phase transitions visible (open, establish, close, accept)
- [x] Proposals and validations traced — send and receive; relay deferred to Phase 4b
- [x] Close time agreement tracked (per `avCT_CONSENSUS_PCT`)
- [x] No impact on consensus timing
- [x] Transaction-consensus correlation (Task 4.6) — `tx.included` events in doAccept
- [ ] Validation span enrichment (Task 4.8) — not implemented
---
# Phase 4a: Establish-Phase Gap Fill & Cross-Node Correlation
> **Goal**: Fill tracing gaps in the consensus establish phase (disputes, convergence,
> threshold escalation, mode changes) and establish cross-node correlation using a
> deterministic shared trace ID derived from `previousLedger.id()`.
>
> **Approach**: Direct instrumentation in `Consensus.h` and `RCLConsensus.cpp`.
> All spans use `SpanGuard` factory methods (`span()`, `hashSpan()`, `linkedSpan()`)
> with `TraceCategory::Consensus` gating. Long-lived spans (round, establish) are
> stored as `std::optional<SpanGuard>` class members. Short-lived scoped spans
> (update_positions, check) are local variables. No macros are used — all tracing
> is via direct `SpanGuard` API calls. `SpanGuard` compiles to no-ops when
> telemetry is disabled.
>
> **Branch**: `pratik/otel-phase4-consensus-tracing`
## Design: Switchable Correlation Strategy
Two strategies for cross-node trace correlation, switchable via config:
### Strategy A — Deterministic Trace ID (Default)
Derive `trace_id = SHA256(previousLedger.id())[0:16]` so all nodes in the same
consensus round share the same trace_id without P2P context propagation.
- **Pros**: All nodes appear in the same trace in Tempo/Jaeger automatically.
No collector-side post-processing needed.
- **Cons**: Overrides OTel's random trace_id generation; requires custom
`IdGenerator` or manual span context construction.
### Strategy B — Attribute-Based Correlation
Use normal random trace_id but attach `xrpl.consensus.ledger_id` as an attribute
on every consensus span. Correlation happens at query time via Tempo/Grafana
`by attribute` queries.
- **Pros**: Standard OTel trace_id semantics; no SDK customization.
- **Cons**: Cross-node correlation requires query-time joins, not automatic.
### Config
```ini
[telemetry]
# "deterministic" (default) or "attribute"
consensus_trace_strategy=deterministic
```
The C++ API to query this at runtime is `Telemetry::getConsensusTraceStrategy()`,
which returns a `std::string const&` (`"deterministic"` or `"attribute"`).
### Implementation
In `RCLConsensus::Adaptor::startRound()`:
- If `deterministic`:
1. Compute `trace_id_bytes = SHA256(prevLedgerID)[0:16]`
2. Construct `opentelemetry::trace::TraceId(trace_id_bytes)`
3. Create a synthetic `SpanContext` with this trace_id and a random span_id:
```cpp
auto traceId = opentelemetry::trace::TraceId(trace_id_bytes);
auto spanId = opentelemetry::trace::SpanId(random_8_bytes);
auto syntheticCtx = opentelemetry::trace::SpanContext(
traceId, spanId, opentelemetry::trace::TraceFlags(1), false);
```
4. Wrap in `opentelemetry::context::Context` via
`opentelemetry::trace::SetSpan(context, syntheticSpan)`
5. Call `startSpan("consensus.round", parentContext)` so the new span
inherits the deterministic trace_id.
- If `attribute`: start a normal `consensus.round` span, set
`xrpl.consensus.ledger_id = previousLedger.id()` as attribute.
Both strategies always set `xrpl.consensus.round_id` (round number) and
`xrpl.consensus.ledger_id` (previous ledger hash) as attributes.
---
## Design: Span Hierarchy
```
consensus.round (root — created in RCLConsensus::startRound, closed at accept)
│ link → previous round's SpanContext (follows-from)
│
├── consensus.establish (phaseEstablish → acceptance, in Consensus.h)
│ ├── consensus.update_positions (each updateOurPositions call)
│ │ └── consensus.dispute.resolve (per-tx dispute resolution event)
│ ├── consensus.check (each haveConsensus call)
│ └── consensus.mode_change (short-lived span in adaptor on mode transition)
│
├── consensus.accept (existing onAccept span — reparented under round)
│
└── consensus.validation.send (existing — reparented, follows-from link to round)
```
### Span Links (follows-from relationships)
| Link Source | Link Target | Rationale |
| ----------------------------------------- | -------------------------- | ------------------------------------------------------------------------------ |
| `consensus.round` (N+1) | `consensus.round` (N) | Causal chain: round N+1 exists because round N accepted |
| `consensus.validation.send` | `consensus.round` | Validation follows from the round that produced it; may outlive the round span |
| _(Phase 4b)_ Received proposal processing | Sender's `consensus.round` | Cross-node causal link via P2P context propagation |
---
## Task 4a.0: Prerequisites — Extend SpanGuard and Telemetry APIs ✅
**Objective**: Add missing API surface needed by later tasks.
**Status**: Done, but implemented differently than originally planned. The macro-based
approach (`XRPL_TRACE_CONSENSUS`, `XRPL_TRACE_ADD_EVENT`, `XRPL_TRACE_SET_ATTR`) was
**not used**. Instead, all consensus tracing uses `SpanGuard` factory methods and
direct method calls, which is cleaner and avoids macro control-flow issues.
**What was done**:
1. **`SpanGuard::addEvent()` with attributes** — implemented as planned:
```cpp
using EventAttribute = std::pair<std::string_view, std::string_view>;
void addEvent(std::string_view name,
std::initializer_list<EventAttribute> attrs);
```
Callers pass plain `string_view` pairs; the implementation converts internally.
```cpp
// Actual usage in Consensus.h::updateOurPositions():
span.addEvent(
"dispute.resolve",
{{consensus::span::attr::txId, to_string(txId)},
{consensus::span::attr::disputeOurVote, dispute.getOurVote() ? "yes" : "no"}});
```
2. **Span link support** — implemented via `SpanGuard::linkedSpan()` static factory
instead of a `Telemetry::startSpan()` overload:
```cpp
static SpanGuard linkedSpan(
std::string_view name, SpanContext const& linkTarget);
```
3. **No macros added** — `TracingInstrumentation.h` was not created. The `XRPL_TRACE_CONSENSUS`,
`XRPL_TRACE_ADD_EVENT`, and `XRPL_TRACE_SET_ATTR` macros from the original plan were
not implemented. All consensus tracing uses direct `SpanGuard` API:
- `SpanGuard::span()` — create scoped spans
- `SpanGuard::hashSpan()` — create spans with deterministic trace IDs
- `SpanGuard::linkedSpan()` — create spans with follows-from links
- `span.setAttribute()` — set attributes directly
- `span.addEvent()` — add events directly
**Key modified files**:
- `include/xrpl/telemetry/SpanGuard.h` — `addEvent()` overload, `EventAttribute` type alias
- `src/libxrpl/telemetry/SpanGuard.cpp` — `addEvent()` implementation
---
## Task 4a.1: Adaptor `getTelemetry()` Method — NOT DONE (Not Needed)
**Objective**: Give `Consensus.h` access to the telemetry subsystem without
coupling the generic template to OTel headers.
**Status**: Not implemented as specified. The `getTelemetry()` adaptor method was
not needed because `SpanGuard::span()` is a static factory method that internally
checks telemetry state via the global `Telemetry` singleton. `Consensus.h` creates
spans by calling `SpanGuard::span(TraceCategory::Consensus, ...)` directly, without
needing adaptor access. Only `RCLConsensus::Adaptor` uses `app_.getTelemetry()`
directly (for `getConsensusTraceStrategy()` in `startRoundTracing()`).
**Key insight**: The `XRPL_TRACE_*` macro approach would have required
`adaptor_.getTelemetry()`. Since macros were not used, this task became unnecessary.
---
## Task 4a.2: Switchable Round Span with Deterministic Trace ID ✅
**Objective**: Create a `consensus.round` root span in `startRound()` that uses
the switchable correlation strategy. Store span context as a member for child
spans in `Consensus.h`.
**Status**: Done. Implemented in `Adaptor::startRoundTracing()`.
**What was done**:
- `RCLConsensus::Adaptor::startRoundTracing()` helper:
- Reads `consensus_trace_strategy` via `app_.getTelemetry().getConsensusTraceStrategy()`
- **Deterministic**: uses `SpanGuard::hashSpan()` with `prevLgr.id()` data
- **Attribute**: uses `SpanGuard::span(TraceCategory::Consensus, seg::consensus, "round")`
- Sets attributes: `xrpl.consensus.ledger_id`, `xrpl.ledger.seq`, `xrpl.consensus.mode`, `trace_strategy`, `xrpl.consensus.round_id`
- Captures `roundSpanContext_` snapshot for cross-thread span linking
- Saves `prevRoundContext_` from previous round for follows-from links
- **`SpanGuard::hashSpan()` factory**: encapsulates deterministic trace ID logic:
```cpp
static SpanGuard hashSpan(
TraceCategory cat, std::string_view name,
std::uint8_t const* hashData, std::size_t hashSize);
```
Derives `trace_id = hashData[0:16]` so all nodes in the same round share
the same trace_id. Compiles to no-op when telemetry is disabled.
- `consensus_trace_strategy` config parsed in `TelemetryConfig.cpp`,
stored in `Telemetry::Setup`, accessible via `Telemetry::getConsensusTraceStrategy()`
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp` — `startRoundTracing()` implementation
- `src/xrpld/app/consensus/ConsensusSpanNames.h` — **(new)** compile-time span name and attribute key constants
- `include/xrpl/telemetry/Telemetry.h` — `consensusTraceStrategy` in Setup, `getConsensusTraceStrategy()`
- `src/libxrpl/telemetry/TelemetryConfig.cpp` — parse new config option
---
## Task 4a.3: Span Members in `Consensus.h` ✅
**Objective**: Add span storage to the `Consensus` class so that spans created
in `startRound()` (adaptor) are accessible from `phaseEstablish()`,
`updateOurPositions()`, and `haveConsensus()` (template methods).
**Status**: Done with documented plan deviation.
**What was done**:
- `establishSpan_` added to `Consensus` private members (as planned):
```cpp
std::optional<xrpl::telemetry::SpanGuard> establishSpan_;
```
- **Plan deviation**: `roundSpan_`, `prevRoundContext_`, and `roundSpanContext_`
are stored in `RCLConsensus::Adaptor` (not `Consensus.h`) because the adaptor
has access to telemetry config for the deterministic trace ID strategy.
- **No `#ifdef XRPL_ENABLE_TELEMETRY` guards**: Members use `std::optional<SpanGuard>`
and `SpanContext` which have no-op implementations when telemetry is disabled,
so `#ifdef` guards are unnecessary. The members are always present in the class
layout but incur negligible overhead.
- Includes added unconditionally to `Consensus.h`:
```cpp
#include <xrpl/telemetry/SpanGuard.h>
#include <xrpld/app/consensus/ConsensusSpanNames.h>
```
No `TracingInstrumentation.h` include (file doesn't exist; macros not used).
**Key modified files**:
- `src/xrpld/consensus/Consensus.h`
- `src/xrpld/app/consensus/RCLConsensus.h` (round span and context members)
---
## Task 4a.4: Instrument `phaseEstablish()` ✅
**Objective**: Create `consensus.establish` span wrapping the establish phase,
with attributes for convergence progress.
**Status**: Done. Implemented via three private helpers in `Consensus.h`.
**What was done**:
- `startEstablishTracing()` — creates `consensus.establish` span via
`SpanGuard::span(TraceCategory::Consensus, seg::consensus, "establish")`.
Called once at start of establish phase. No `#ifdef` guards needed —
`SpanGuard::span()` returns a no-op guard when telemetry is disabled.
- `updateEstablishTracing()` — sets attributes on each `phaseEstablish()` call:
- `converge_percent` — `convergePercent_`
- `establish_count` — `establishCounter_`
- `proposers` — `currPeerPositions_.size()`
- `endEstablishTracing()` — calls `establishSpan_.reset()` on phase exit.
**Key modified files**:
- `src/xrpld/consensus/Consensus.h` — `phaseEstablish()` method + 3 helper methods
---
## Task 4a.5: Instrument `updateOurPositions()` ✅
**Objective**: Trace each position update cycle including dispute resolution
details.
**Status**: DONE. Span, dispute events with yays/nays, and disputes_count attribute are all implemented.
**What was done**:
- Creates `consensus.update_positions` scoped span via
`SpanGuard::span(TraceCategory::Consensus, seg::consensus, "update_positions")`:
```cpp
auto span = SpanGuard::span(TraceCategory::Consensus, seg::consensus, "update_positions");
```
- Attributes set:
- `converge_percent` — current convergence
- `proposers` — `currPeerPositions_.size()`
- `have_close_time_consensus` — close time consensus state
- `close_time_threshold` — `avCT_CONSENSUS_PCT`
- `disputes_count` — number of active disputes
- Dispute events recorded via direct `span.addEvent()` call with yays/nays:
```cpp
span.addEvent(
"dispute.resolve",
{{consensus::span::attr::txId, to_string(txId)},
{consensus::span::attr::disputeOurVote, dispute.getOurVote() ? "yes" : "no"},
{consensus::span::attr::disputeYays, std::to_string(dispute.getYays())},
{consensus::span::attr::disputeNays, std::to_string(dispute.getNays())}});
```
**Not implemented**:
- `proposers_agreed` / `proposers_total` attributes — not set
**Key modified files**:
- `src/xrpld/consensus/Consensus.h` — `updateOurPositions()` method
- `src/xrpld/consensus/DisputedTx.h` — added `getYays()` / `getNays()` (currently unused)
---
## Task 4a.6: Instrument `haveConsensus()` (Threshold & Convergence) ✅
**Objective**: Trace consensus checking including threshold escalation.
**Status**: DONE. The `consensus.check` span is created with all planned attributes
including the avalanche threshold.
**What was done**:
- Creates `consensus.check` scoped span via
`SpanGuard::span(TraceCategory::Consensus, seg::consensus, "check")`:
```cpp
auto span = SpanGuard::span(TraceCategory::Consensus, seg::consensus, "check");
```
- Attributes set:
- `agree_count` — peers that agree with our position
- `disagree_count` — peers that disagree
- `converge_percent` — convergence percentage
- `have_close_time_consensus` — close time consensus state
- `threshold_percent` — set to `avCT_CONSENSUS_PCT` (75%)
- `consensus_result` — "yes", "no", or "moved_on"
- `avalanche_threshold` — the escalated weight from `getNeededWeight()` on the `consensus.update_positions` span
**Key modified files**:
- `src/xrpld/consensus/Consensus.h` — `haveConsensus()` method
---
## Task 4a.7: Instrument Mode Changes ✅
**Objective**: Trace consensus mode transitions (proposing ↔ observing,
wrongLedger, switchedLedger).
**Status**: Done.
**What was done**:
- In `RCLConsensus::Adaptor::onModeChange()`, creates a scoped span via direct
`SpanGuard::span()` call:
```cpp
auto span = telemetry::SpanGuard::span(
telemetry::TraceCategory::Consensus, telemetry::seg::consensus, "mode_change");
span.setAttribute(consensus::span::attr::modeOld, to_string(before).c_str()); // "mode_old"
span.setAttribute(consensus::span::attr::modeNew, to_string(after).c_str()); // "mode_new"
```
- `MonitoredMode::set()` in `Consensus.h` calls `adaptor_.onModeChange(before, after)`.
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp` — `onModeChange()`
---
## Task 4a.8: Reparent Existing Spans Under Round ✅
**Objective**: Make existing consensus spans (`consensus.accept`,
`consensus.accept.apply`, `consensus.validation.send`) children of the
`consensus.round` root span instead of being standalone.
**Status**: DONE. All three spans are now parented under the round span.
**What was done**:
- `consensus.validation.send` uses `SpanGuard::linkedSpan()` to create a
follows-from link to `roundSpanContext_`. This is thread-safe because
`roundSpanContext_` is a lightweight `SpanContext` snapshot captured on the
consensus thread and read on the jtACCEPT worker thread.
- `consensus.accept` and `consensus.accept.apply` now use
`SpanGuard::childSpan(name, roundSpanContext_)` instead of `SpanGuard::span()`
to explicitly parent under the round span context. This solves the cross-thread
parenting problem:
- `doAccept()` runs on the jtACCEPT worker thread (not the consensus thread)
- `childSpan()` explicitly passes the parent context, bypassing OTel's
thread-local context propagation
**Key modified files**:
- `src/xrpld/app/consensus/RCLConsensus.cpp`
---
## Task 4a.9: Build Verification and Testing ✅
**Objective**: Verify all Phase 4a changes compile cleanly with telemetry ON
and OFF, and don't affect consensus timing.
**What to do**:
1. Build with `telemetry=ON` — verify no compilation errors
2. Build with `telemetry=OFF` — verify `SpanGuard` compiles to no-ops
3. Run existing consensus unit tests
4. Verify `SpanGuard` / `SpanContext` members have negligible overhead when disabled
5. Run `pccl` pre-commit checks
**Verification Checklist**:
- [x] Build succeeds with telemetry ON
- [x] Build succeeds with telemetry OFF
- [x] Existing consensus tests pass
- [x] `SpanGuard` no-op path verified (no `#ifdef` needed — disabled at runtime)
- [x] No new virtual calls in hot consensus paths
- [x] `pccl` passes
---
## Phase 4a Summary
| Task | Description | Status | New Files | Modified Files | Depends On |
| ---- | ------------------------------------------------ | ------------------------ | --------- | -------------- | ---------- |
| 4a.0 | Prerequisites: extend SpanGuard & Telemetry APIs | ✅ Done (no macros) | 0 | 2 | Phase 4 |
| 4a.1 | Adaptor `getTelemetry()` method | ⏭️ Skipped (not needed) | 0 | 0 | Phase 4 |
| 4a.2 | Switchable round span with deterministic traceID | ✅ Done | 1 | 3 | 4a.0 |
| 4a.3 | Span members in `Consensus.h` | ✅ Done (with deviation) | 0 | 2 | — |
| 4a.4 | Instrument `phaseEstablish()` | ✅ Done | 0 | 1 | 4a.3 |
| 4a.5 | Instrument `updateOurPositions()` | ✅ Done | 0 | 2 | 4a.0, 4a.3 |
| 4a.6 | Instrument `haveConsensus()` (thresholds) | ✅ Done | 0 | 1 | 4a.3 |
| 4a.7 | Instrument mode changes | ✅ Done | 0 | 1 | — |
| 4a.8 | Reparent existing spans under round | ✅ Done | 0 | 1 | 4a.0, 4a.2 |
| 4a.9 | Build verification and testing | ✅ Done | 0 | 0 | 4a.0-4a.8 |
**Parallel work**: Tasks 4a.0 and 4a.1 can run in parallel. Tasks 4a.4, 4a.5, 4a.6, and 4a.7 can run in parallel after 4a.3 (and 4a.0 for 4a.5).
### New Spans (Phase 4a)
| Span Name | Location | Key Attributes (actually set) |
| ---------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `consensus.round` | `RCLConsensus.cpp` | `xrpl.consensus.round_id`, `xrpl.consensus.ledger_id`, `xrpl.ledger.seq`, `xrpl.consensus.mode`, `trace_strategy` |
| `consensus.phase.open` | `Consensus.h` | `start_reason`, `previous_close_agree`, `peer_positions_at_open`, `early_close_triggered` (at start); `open_duration_ms`, `peer_positions_at_close`, `tx_sets_acquired`, `close_reason`, `proposers_validated` (at close) |
| `consensus.establish` | `Consensus.h` | `converge_percent`, `establish_count`, `proposers`, `disputes_count`, `close_time_avalanche_state` |
| `consensus.update_positions` | `Consensus.h` | `converge_percent`, `proposers`, `have_close_time_consensus`, `close_time_threshold`, `disputes_count`, `avalanche_threshold` |
| `consensus.check` | `Consensus.h` | `agree_count`, `disagree_count`, `converge_percent`, `have_close_time_consensus`, `threshold_percent`, `consensus_result` |
| `consensus.mode_change` | `RCLConsensus.cpp` | `mode_old`, `mode_new` |
### New Events (Phase 4a)
| Event Name | Parent Span | Attributes (actually set) |
| ----------------- | ---------------------------- | ---------------------------------------------------------------- |
| `dispute.resolve` | `consensus.update_positions` | `xrpl.tx.id`, `dispute_our_vote`, `dispute_yays`, `dispute_nays` |
| `tx.included` | `consensus.accept.apply` | `xrpl.tx.id` |
### New Attributes (Phase 4a)
```cpp
// Round-level (on consensus.round) — ALL IMPLEMENTED
"xrpl.consensus.round_id" = int64 // Consensus round number
"xrpl.consensus.ledger_id" = string // previousLedger.id() hash
"trace_strategy" = string // "deterministic" or "attribute"
// Establish-level — IMPLEMENTED
"converge_percent" = int64 // Convergence % (0-100+)
"establish_count" = int64 // Number of establish iterations
"agree_count" = int64 // Peers that agree (haveConsensus)
"disagree_count" = int64 // Peers that disagree
"threshold_percent" = int64 // Current threshold (avCT_CONSENSUS_PCT = 75%)
"consensus_result" = string // "yes", "no", "moved_on"
"have_close_time_consensus" = bool // Close time consensus reached
"close_time_threshold" = int64 // Close time voting threshold
// Establish-level — IMPLEMENTED
"disputes_count" = int64 // Active disputes (on update_positions)
"avalanche_threshold" = int64 // Escalated weight (on update_positions)
// Establish-level — NOT IMPLEMENTED
// "proposers_agreed" = int64 // Peers agreeing with us — not set
// "proposers_total" = int64 // Total peer positions — not set (not defined)
// Mode change — ALL IMPLEMENTED
"mode_old" = string // Previous mode
"mode_new" = string // New mode
```
### Implementation Notes
- **No macros**: The planned `XRPL_TRACE_CONSENSUS`, `XRPL_TRACE_ADD_EVENT`, and
`XRPL_TRACE_SET_ATTR` macros were not implemented. All consensus tracing uses
`SpanGuard` factory methods (`span()`, `hashSpan()`, `linkedSpan()`) and direct
method calls (`setAttribute()`, `addEvent()`). This avoids macro control-flow
issues and is cleaner than the planned approach.
- **Separation of concerns**: All non-trivial telemetry code extracted to private
helpers (`startRoundTracing`, `createValidationSpan`, `startEstablishTracing`,
`updateEstablishTracing`, `endEstablishTracing`). Business logic methods contain
single-line calls to these helpers.
- **Thread safety**: `createValidationSpan()` runs on the jtACCEPT worker thread.
Instead of accessing `roundSpan_` across threads, a `roundSpanContext_` snapshot
(lightweight `SpanContext` value type) is captured on the consensus thread in
`startRoundTracing()` and read by `createValidationSpan()`. The job queue
provides the happens-before guarantee.
- **No `#ifdef` guards**: Span members use `std::optional<SpanGuard>` and `SpanContext`
which have no-op implementations when telemetry is disabled. No `#ifdef XRPL_ENABLE_TELEMETRY`
guards needed around members or includes.
- **No `getTelemetry()` adaptor method**: `SpanGuard::span()` is a static factory that
internally checks telemetry state, so `Consensus.h` doesn't need adaptor access
for span creation. Only `RCLConsensus::Adaptor` accesses `app_.getTelemetry()` directly.
- **Config validation**: `consensus_trace_strategy` is validated to be either
`"deterministic"` or `"attribute"`, falling back to `"deterministic"` for
unrecognised values.
- **Plan deviation**: `roundSpan_` is stored in `RCLConsensus::Adaptor` (not
`Consensus.h`) because the adaptor has access to telemetry config and can
implement the deterministic trace ID strategy. `establishSpan_` is correctly
in `Consensus.h` as planned.
---
# Phase 4b: Cross-Node Propagation (Future — Documentation Only)
> **Goal**: Wire `TraceContextPropagator` for P2P messages so that proposals
> and validations carry trace context between nodes. This enables true
> distributed tracing where a proposal sent by Node A creates a child span
> on Node B.
>
> **Status**: NOT IMPLEMENTED. The protobuf fields and propagator class exist
> but are not wired. This section documents the design for future work.
## Architecture
```
Node A (proposing) Node B (receiving)
───────────────── ──────────────────
consensus.round consensus.round
├── propose() ├── peerProposal()
│ └── TraceContextPropagator │ └── TraceContextPropagator
│ ::injectToProtobuf( │ ::extractFromProtobuf(
│ TMProposeSet.trace_context) │ TMProposeSet.trace_context)
│ │ └── span link → Node A's context
└── validate() └── onValidation()
└── inject into TMValidation └── extract from TMValidation
```
## Wiring Points
| Message | Inject Location | Extract Location | Protobuf Field |
| --------------- | ---------------------------------- | ----------------------------------- | -------------------------- |
| `TMProposeSet` | `Adaptor::propose()` | `PeerImp::onMessage(TMProposeSet)` | field 1001: `TraceContext` |
| `TMValidation` | `Adaptor::validate()` | `PeerImp::onMessage(TMValidation)` | field 1001: `TraceContext` |
| `TMTransaction` | `NetworkOPs::processTransaction()` | `PeerImp::onMessage(TMTransaction)` | field 1001: `TraceContext` |
## Span Link Semantics
Received messages use **span links** (follows-from), NOT parent-child:
- The receiver's processing span links to the sender's context
- This preserves each node's independent trace tree
- Cross-node correlation visible via linked traces in Tempo/Jaeger
## Interaction with Deterministic Trace ID (Strategy A)
When using deterministic trace_id (Phase 4a default), cross-node spans already
share the same trace_id. P2P propagation adds **span-level** linking:
- Without propagation: spans from different nodes appear in the same trace
(same trace_id) but without parent-child or follows-from relationships.
- With propagation: spans have explicit links showing which proposal/validation
from Node A caused processing on Node B.
## Prerequisites
- Phase 4a (this task list) — establish phase tracing must be in place
- `TraceContextPropagator` free functions (already exist in
`include/xrpl/telemetry/TraceContextPropagator.h`)
- Protobuf `TraceContext` message (already exists, field 1001)

View File

@@ -1,250 +0,0 @@
# Phase 5: Documentation & Deployment Task List
> **Goal**: Production readiness — Grafana dashboards, spanmetrics pipeline, operator runbook, alert definitions, and final integration testing. This phase ensures the telemetry system is useful and maintainable in production.
>
> **Scope**: Grafana dashboard definitions, OTel Collector spanmetrics connector, Prometheus integration, alert rules, operator documentation, and production-ready Docker Compose stack.
>
> **Branch**: `pratik/otel-phase5-docs-deployment` (from `pratik/otel-phase4-consensus-tracing`)
> **Note on attribute names**: the `xrpl.<domain>.<field>` keys shown below
> (including the collector spanmetrics dimension examples) are written in the
> older dotted form for readability — it mirrors how the fully qualified
> attribute reads in a Tempo trace view. The implemented keys follow the
> convention in [CONTRIBUTING.md](../CONTRIBUTING.md#telemetry-span-attribute-naming)
> (underscore form, e.g. `command`, `rpc_status`); the `*SpanNames.h` constants
> are the single source of truth, and the real collector dimensions must use
> those exact underscore keys (the CI naming check enforces this).
### Related Plan Documents
| Document | Relevance |
| ---------------------------------------------------------------- | -------------------------------------------------------------------------- |
| [07-observability-backends.md](./07-observability-backends.md) | Tempo setup (§7.1), Grafana dashboards (§7.6), alerts (§7.6.3) |
| [05-configuration-reference.md](./05-configuration-reference.md) | Collector config (§5.5), production config (§5.5.2), Docker Compose (§5.6) |
| [06-implementation-phases.md](./06-implementation-phases.md) | Phase 5 tasks (§6.6), definition of done (§6.11.5) |
---
## Task 5.1: Add Spanmetrics Connector to OTel Collector
**Objective**: Derive RED metrics (Rate, Errors, Duration) from trace spans automatically, enabling Grafana time-series dashboards.
**What to do**:
- Edit `docker/telemetry/otel-collector-config.yaml`:
- Add `spanmetrics` connector:
```yaml
connectors:
spanmetrics:
histogram:
explicit:
buckets: [1ms, 5ms, 10ms, 25ms, 50ms, 100ms, 250ms, 500ms, 1s, 5s]
dimensions:
- name: xrpl.rpc.command
- name: xrpl.rpc.status
- name: xrpl.consensus.phase
- name: xrpl.tx.type
```
- Add `prometheus` exporter:
```yaml
exporters:
prometheus:
endpoint: 0.0.0.0:8889
```
- Wire the pipeline:
```yaml
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch]
exporters: [debug, otlp/tempo, spanmetrics]
metrics:
receivers: [spanmetrics]
exporters: [prometheus]
```
- Edit `docker/telemetry/docker-compose.yml`:
- Expose port `8889` on the collector for Prometheus scraping
- Add Prometheus service
- Add Prometheus as Grafana datasource
**Key modified files**:
- `docker/telemetry/otel-collector-config.yaml`
- `docker/telemetry/docker-compose.yml`
**Key new files**:
- `docker/telemetry/prometheus.yml` (Prometheus scrape config)
- `docker/telemetry/grafana/provisioning/datasources/prometheus.yaml`
**Reference**:
- [POC_taskList.md §Next Steps](./POC_taskList.md) — Metrics pipeline for Grafana dashboards
---
## Task 5.2: Create Grafana Dashboards
**Objective**: Provide pre-built Grafana dashboards for RPC performance, transaction lifecycle, and consensus health.
**What to do**:
- Create `docker/telemetry/grafana/provisioning/dashboards/dashboards.yaml` (provisioning config)
- Create dashboard JSON files:
1. **RPC Performance Dashboard** (`rpc-performance.json`):
- RPC request latency (p50/p95/p99) by command — histogram panel
- RPC throughput (requests/sec) by command — time series
- RPC error rate by command — bar gauge
- Top slowest RPC commands — table
2. **Transaction Overview Dashboard** (`transaction-overview.json`):
- Transaction processing rate — time series
- Transaction latency distribution — histogram
- Suppression rate (duplicates) — stat panel
- Transaction processing path (sync vs async) — pie chart
3. **Consensus Health Dashboard** (`consensus-health.json`):
- Consensus round duration — time series
- Phase duration breakdown (open/establish/accept) — stacked bar
- Proposals sent/received per round — stat panel
- Consensus mode distribution (proposing/observing) — pie chart
- Store dashboards in `docker/telemetry/grafana/dashboards/`
**Key new files**:
- `docker/telemetry/grafana/provisioning/dashboards/dashboards.yaml`
- `docker/telemetry/grafana/dashboards/rpc-performance.json`
- `docker/telemetry/grafana/dashboards/transaction-overview.json`
- `docker/telemetry/grafana/dashboards/consensus-health.json`
**Reference**:
- [07-observability-backends.md §7.6](./07-observability-backends.md) — Grafana dashboard specifications
- [01-architecture-analysis.md §1.8.3](./01-architecture-analysis.md) — Dashboard panel examples
---
## Task 5.3: Define Alert Rules
**Objective**: Create alert definitions for key telemetry anomalies.
**What to do**:
- Create `docker/telemetry/grafana/provisioning/alerting/alerts.yaml`:
- **RPC Latency Alert**: p99 latency > 1s for any command over 5 minutes
- **RPC Error Rate Alert**: Error rate > 5% for any command over 5 minutes
- **Consensus Duration Alert**: Round duration > 10s (warn), > 30s (critical)
- **Transaction Processing Alert**: Processing rate drops below threshold
- **Telemetry Pipeline Health**: No spans received for > 2 minutes
**Key new files**:
- `docker/telemetry/grafana/provisioning/alerting/alerts.yaml`
**Reference**:
- [07-observability-backends.md §7.6.3](./07-observability-backends.md) — Alert rule definitions
---
## Task 5.4: Production Collector Configuration
**Objective**: Create a production-ready OTel Collector configuration with tail-based sampling and resource limits.
**What to do**:
- Create `docker/telemetry/otel-collector-config-production.yaml`:
- Tail-based sampling policy:
- Always sample errors and slow traces
- 10% base sampling rate for normal traces
- Always sample first trace for each unique RPC command
- Resource limits:
- Memory limiter processor (80% of available memory)
- Queued retry for export failures
- TLS configuration for production endpoints
- Health check endpoint
**Key new files**:
- `docker/telemetry/otel-collector-config-production.yaml`
**Reference**:
- [05-configuration-reference.md §5.5.2](./05-configuration-reference.md) — Production collector config
---
## Task 5.5: Operator Runbook
**Objective**: Create operator documentation for managing the telemetry system in production.
**What to do**:
- Create `docs/telemetry-runbook.md`:
- **Setup**: How to enable telemetry in xrpld
- **Configuration**: All config options with descriptions
- **Collector Deployment**: Docker Compose vs. Kubernetes vs. bare metal
- **Troubleshooting**: Common issues and resolutions
- No traces appearing
- High memory usage from telemetry
- Collector connection failures
- Sampling configuration tuning
- **Performance Tuning**: Batch size, queue size, sampling ratio guidelines
- **Upgrading**: How to upgrade OTel SDK and Collector versions
**Key new files**:
- `docs/telemetry-runbook.md`
---
## Task 5.6: Final Integration Testing
**Objective**: Validate the complete telemetry stack end-to-end.
**What to do**:
1. Start full Docker stack (Collector, Tempo, Grafana, Prometheus)
2. Build xrpld with `telemetry=ON`
3. Run in standalone mode with telemetry enabled
4. Generate RPC traffic and verify traces in Tempo
5. Verify dashboards populate in Grafana
6. Verify alerts trigger correctly
7. Test telemetry OFF path (no regressions)
8. Run full test suite
**Verification Checklist**:
- [ ] Docker stack starts without errors
- [ ] Traces appear in Tempo with correct hierarchy
- [ ] Grafana dashboards show metrics derived from spans
- [ ] Prometheus scrapes spanmetrics successfully
- [ ] Alerts can be triggered by simulated conditions
- [ ] Build succeeds with telemetry ON and OFF
- [ ] Full test suite passes
---
## Summary
| Task | Description | New Files | Modified Files | Depends On |
| ---- | ---------------------------------- | --------- | -------------- | ---------- |
| 5.1 | Spanmetrics connector + Prometheus | 2 | 2 | Phase 4 |
| 5.2 | Grafana dashboards | 4 | 0 | 5.1 |
| 5.3 | Alert definitions | 1 | 0 | 5.1 |
| 5.4 | Production collector config | 1 | 0 | Phase 4 |
| 5.5 | Operator runbook | 1 | 0 | Phase 4 |
| 5.6 | Final integration testing | 0 | 0 | 5.1-5.5 |
**Parallel work**: Tasks 5.1, 5.4, and 5.5 can run in parallel. Tasks 5.2 and 5.3 depend on 5.1. Task 5.6 depends on all others.
**Exit Criteria** (from [06-implementation-phases.md §6.11.5](./06-implementation-phases.md)):
- [ ] Dashboards deployed and showing data
- [ ] Alerts configured and tested
- [ ] Operator documentation complete
- [ ] Production collector config ready
- [ ] Full test suite passes

View File

@@ -1683,131 +1683,3 @@ validators.txt
# set to ssl_verify to 0.
[ssl_verify]
1
#-------------------------------------------------------------------------------
#
# 11. Telemetry (OpenTelemetry Tracing)
#
#-------------------------------------------------------------------------------
#
# Enables distributed tracing via OpenTelemetry. Takes effect only if tracing
# was compiled in with the Conan option `-o telemetry=True`; there is no CMake
# option for it. See docs/build/telemetry.md.
#
# [telemetry]
#
# enabled=0
#
# Enable or disable telemetry at runtime. Default: 0 (disabled).
# This key, use_tls and the trace_* keys take 0, 1, true or false, in
# any case. Any other value is a startup error that names the key.
#
# service_name=xrpld
#
# OTel resource attribute `service.name`. Default: xrpld.
# The node's network ID (from [network_id]) is automatically added
# as the `xrpl.network.id` and `xrpl.network.type` resource attributes.
#
# service_instance_id=<node_public_key>
#
# OTel resource attribute `service.instance.id`, which tells one node's
# telemetry apart from another's. Normally left unset: the node identity
# is not known when telemetry is constructed, so the server fills this in
# with its own Base58 node public key later during startup. Set it only
# to pin a stable instance name of your own choosing; doing so suppresses
# the node-public-key fallback.
# Default: the node's Base58 public key.
#
# traces_endpoint=http://localhost:4318/v1/traces
#
# The OTLP/HTTP endpoint spans are exported to. The server sends trace
# data as protobuf-encoded HTTP POST requests to this URL. The full URL
# including the signal path is used verbatim; no other endpoint is
# derived from it.
# Default: http://localhost:4318/v1/traces.
#
# --- TLS settings for the OTLP exporter connection ---
#
# use_tls=0
#
# Whether to hand tls_ca_cert to the exporter as its CA bundle. TLS is
# selected by the scheme of traces_endpoint, not by this key. So with
# enabled=1, use_tls=1 requires traces_endpoint to start with https://
# (lower case). Any other URL, including the default, makes xrpld fail
# to start with a message that names traces_endpoint.
# Default: 0 (no CA file is passed, so the exporter keeps its own
# default trust store).
#
# tls_ca_cert=
#
# Path to a PEM-encoded CA certificate bundle for verifying the
# collector's certificate. Only used when use_tls=1, and passed to the
# exporter unchanged. The path is not checked while the config is parsed,
# so a missing or unreadable file surfaces as an export failure at
# runtime rather than as a startup error.
# Default: empty (system CA store).
#
# Head sampling is intentionally fixed at 1.0 (sample everything) and is
# not configurable. A per-node sampling ratio would let nodes make
# divergent keep/drop decisions for the same distributed trace, producing
# broken/partial traces. A span with a remote parent is decided by the same
# fixed ratio, not by the peer's sampled flag. Reduce volume at the collector
# via tail sampling instead; for node-local post-hoc dropping use
# SpanGuard::discard() in code.
#
# trace_rpc=1
#
# Enable tracing for JSON-RPC and WebSocket API request handling —
# command parsing, execution, and response serialization. Default: 1.
#
# trace_transactions=1
#
# Enable tracing for the transaction lifecycle — submission, validation,
# application to ledgers, and final disposition. Default: 1.
#
# trace_consensus=1
#
# Enable tracing for the consensus round lifecycle — proposals,
# validations, mode changes, and ledger acceptance. Default: 1.
#
# trace_peer=1
#
# Enable tracing for peer-to-peer protocol messages — overlay message
# send/receive, peer handshakes, and routing. High volume; enabled
# by default. Default: 1.
#
# trace_ledger=1
#
# Enable tracing for ledger close and accept operations — ledger
# building, state hashing, and write-back to the node store. Default: 1.
#
# consensus_trace_strategy=deterministic
#
# How the consensus round span picks its trace id. Two values are
# accepted, and anything else makes xrpld fail to start.
#
# deterministic (the default, and the value to use): the trace id comes
# from the previous ledger hash, so every validator of a round reports
# into one trace and the round can be read end to end across nodes.
#
# random: experimental only, and not used. Each node invents its own
# trace id, so a single round arrives as one separate trace per node.
# Those traces can only be lined up by hand through the
# consensus_ledger_id span attribute.
#
# --- Batch processor tuning ---
#
# batch_size=512
#
# Maximum number of spans exported in a single batch. Must be at least 1
# and must not exceed max_queue_size. Default: 512.
#
# batch_delay_ms=5000
#
# Maximum delay (milliseconds) before a partial batch is flushed.
# Must be at least 1. Default: 5000 (5 seconds).
#
# max_queue_size=2048
#
# Maximum number of spans queued in memory before drops occur. Must be
# at least 1 and at least as large as batch_size. Default: 2048.
#

View File

@@ -211,41 +211,9 @@ target_link_libraries(
xrpl.libxrpl.conditions
)
# Telemetry module — OpenTelemetry distributed tracing support.
# Sources: include/xrpl/telemetry/ (headers), src/libxrpl/telemetry/ (impl).
# When telemetry=ON, links the Conan-provided umbrella target
# opentelemetry-cpp::opentelemetry-cpp (individual component targets like
# ::api, ::sdk are not available in the Conan package).
#
# Declared before its consumers (consensus, tx) because add_module isolates
# each module's headers: a module can only include xrpl/telemetry/ headers if
# it links this target, and the target must already exist at that point.
#
# Links xrpl.libxrpl.protocol PRIVATELY for sha512Half (digest.h) and the
# SField table behind TxAccountSpanNames.cpp
add_module(xrpl telemetry)
target_link_libraries(
xrpl.libxrpl.telemetry
PUBLIC xrpl.libxrpl.basics xrpl.libxrpl.beast xrpl.libxrpl.config
PRIVATE xrpl.libxrpl.protocol
)
if(telemetry)
target_link_libraries(
xrpl.libxrpl.telemetry
PUBLIC opentelemetry-cpp::opentelemetry-cpp
)
# PUBLIC, so a parent project that adds this one with add_subdirectory()
# and links this module sees the same class layouts it was built with.
# CMakeLists.txt also sets the define for every target in this project.
# Conan consumers get it from conanfile.py instead.
target_compile_definitions(
xrpl.libxrpl.telemetry
PUBLIC XRPL_ENABLE_TELEMETRY
)
endif()
add_module(xrpl tx)
target_link_libraries(xrpl.libxrpl.tx PUBLIC xrpl.libxrpl.ledger)
# Links xrpl.libxrpl.telemetry for the consensus tracing spans declared in
# include/xrpl/consensus/ConsensusSpanNames.h.
add_module(xrpl consensus)
target_link_libraries(
xrpl.libxrpl.consensus
@@ -254,13 +222,6 @@ target_link_libraries(
xrpl.libxrpl.json
xrpl.libxrpl.protocol
xrpl.libxrpl.ledger
xrpl.libxrpl.telemetry
)
add_module(xrpl tx)
target_link_libraries(
xrpl.libxrpl.tx
PUBLIC xrpl.libxrpl.ledger xrpl.libxrpl.telemetry
)
add_library(xrpl.libxrpl)
@@ -297,7 +258,6 @@ target_link_modules(
resource
server
shamap
telemetry
tx
)

View File

@@ -31,7 +31,7 @@ namespace xrpl::ledger_entries {
// builder's STObject and the wrapper's SLE.
TEST(${name}Tests, BuilderSettersRoundTrip)
{
UInt256 const index{1u};
constexpr UInt256 index{1};
% for field in fields:
auto const ${field["paramName"]}Value = ${canonical_expr(field)};
@@ -85,7 +85,7 @@ TEST(${name}Tests, BuilderSettersRoundTrip)
// from that SLE, build a new wrapper, and verify all fields (and validate()).
TEST(${name}Tests, BuilderFromSleRoundTrip)
{
UInt256 const index{2u};
constexpr UInt256 index{2};
% for field in fields:
auto const ${field["paramName"]}Value = ${canonical_expr(field)};
@@ -146,7 +146,7 @@ TEST(${name}Tests, BuilderFromSleRoundTrip)
// 3) Verify wrapper throws when constructed from wrong ledger entry type.
TEST(${name}Tests, WrapperThrowsOnWrongEntryType)
{
UInt256 const index{3u};
constexpr UInt256 index{3};
// Build a valid ledger entry of a different type
// Ticket requires: Account, OwnerNode, TicketSequence, PreviousTxnID, PreviousTxnLgrSeq
@@ -177,7 +177,7 @@ TEST(${name}Tests, WrapperThrowsOnWrongEntryType)
// 4) Verify builder throws when constructed from wrong ledger entry type.
TEST(${name}Tests, BuilderThrowsOnWrongEntryType)
{
UInt256 const index{4u};
constexpr UInt256 index{4};
// Build a valid ledger entry of a different type
% if wrong_le_include == "Ticket":
@@ -207,7 +207,7 @@ TEST(${name}Tests, BuilderThrowsOnWrongEntryType)
// 5) Build with only required fields and verify optional fields return nullopt.
TEST(${name}Tests, OptionalFieldsReturnNullopt)
{
UInt256 const index{3u};
constexpr UInt256 index{3};
% for field in required_fields:
auto const ${field["paramName"]}Value = ${canonical_expr(field)};

View File

@@ -11,14 +11,11 @@
"rocksdb/10.5.1#4a197eca381a3e5ae8adf8cffa5aacd0%1782392413.075713",
"re2/20251105#8579cfd0bda4daf0683f9e3898f964b4%1782392402.431897",
"protobuf/6.33.5#ff253ead763bd8d9904a52979cd21e81%1782392410.233933",
"opentelemetry-cpp/1.28.0#2cbf71db4e0e0535df20be305005cb2f%1785939947.292581",
"openssl/3.6.3#f806de8933e3bf6f01016c6a888cee2e%1783945160.863288",
"nudb/2.0.9#11149c73f8f2baff9a0198fe25971fc7%1782392402.297166",
"nlohmann_json/3.11.3#45828be26eb619a2e04ca517bb7b828d%1701220705.259",
"mpt-crypto/1.0.2#b313cef0c1a493eb970ad185b2e9bab7%1784285108.866483",
"lz4/1.10.0#982d9b673900f665a1da109e09c17cab%1782392402.164188",
"libiconv/1.17#9923bc6dc6f106646d6967e0039a5ada%1782392792.775744",
"libcurl/8.21.0#8c26e59c04891ba3373ea3552e18f67f%1783067699.863",
"libbacktrace/cci.20210118#a7691bfccd8caaf66309df196790a5a1%1782392402.420732",
"libarchive/3.8.7#c446109bd1f1d8ba7936c94189bc50e6%1782392403.066892",
"jemalloc/5.3.1#1fc58d55316041f10fbc1e8a2eae632a%1776700028.228",
@@ -38,15 +35,9 @@
"zlib/1.3.2#1cb806da49011867778ffb6ac7190fcb%1782392402.122708",
"strawberryperl/5.32.1.1#8d114504d172cfea8ea1662d09b6333e%1782395692.540639",
"protobuf/6.33.5#ff253ead763bd8d9904a52979cd21e81%1782392410.233933",
"pkgconf/2.5.1#93c2051284cba1279494a43a4fcfeae2%1757684701.089",
"opentelemetry-proto/1.7.0#ed6d5bd761bef0afb0ba09676420b9ea%1749461220.268",
"ninja/1.13.2#c8c5dc2a52ed6e4e42a66d75b4717ceb%1764096931.974",
"nasm/2.16.01#31e26f2ee3c4346ecd347911bd126904%1782395690.33162",
"msys2/cci.latest#d22fe7b2808f5fd34d0a7923ace9c54f%1770657326.649",
"meson/1.10.2#9d2d10681fe7fe61c788c58626c89b25%1775558003.754",
"m4/1.4.19#1727f439cf74e83826ec96d0b4904eee%1784541921.659",
"libtool/2.4.7#14e7739cc128bc1623d2ed318008e47e%1755679003.847",
"gnu-config/cci.20210814#466e9d4d7779e1c142443f7ea44b4284%1762363589.329",
"cmake/4.3.3#840cf00ea09777e05c2050a50a82c722%1782392418.696091",
"b2/5.4.2#ffd6084a119587e70f11cd45d1a386e2%1782392402.624226",
"automake/1.16.5#b91b7c384c3deaa9d535be02da14d04f%1755524470.56",
@@ -76,9 +67,6 @@
],
"lz4/[>=1.9.4 <2]": [
"lz4/1.10.0#982d9b673900f665a1da109e09c17cab"
],
"protobuf/[>=4.25.3 <7]": [
"protobuf/6.33.5#ff253ead763bd8d9904a52979cd21e81"
]
},
"config_requires": []

View File

@@ -1,5 +1,3 @@
{# Read SANITIZERS first: the next line rebinds `os` from the module to the OS name. #}
{% set sanitizers = os.getenv("SANITIZERS", "") %}
{% set os = detect_api.detect_os() %}
{% set arch = detect_api.detect_arch() %}
{% set compiler, version, compiler_exe = detect_api.detect_default_compiler() %}
@@ -51,17 +49,6 @@ tools.build:compiler_executables={'c':'{{ cc_exe }}','cpp':'{{ cxx_exe }}'}
user.package:cppstd_version=23
tools.info.package_id:confs+=["user.package:cppstd_version"]
{# opentelemetry-cpp is an external dependency: sanitizer reports from its own #}
{# code are not ours to fix, so build it without sanitizer checks. #}
{# TSan stays on, because it must see the library's atomics or it reports false races. #}
{% if sanitizers and compiler != "msvc" %}
{% set otel_sanitizer_flags = ["-fno-sanitize=all"] %}
{% if "thread" in sanitizers %}
{% set _ = otel_sanitizer_flags.append("-fsanitize=thread") %}
{% endif %}
opentelemetry-cpp/*:tools.build:cxxflags+={{ otel_sanitizer_flags }}
{% endif %}
{% if os == "Macos" %}
[buildenv]
{# os.version adds -mmacosx-version-min to compiler command lines, #}
@@ -71,15 +58,3 @@ opentelemetry-cpp/*:tools.build:cxxflags+={{ otel_sanitizer_flags }}
{# Scoped to boost/* since it is the only gap. #}
boost/*:MACOSX_DEPLOYMENT_TARGET={{ min_macos_version }}
{% endif %}
{% if os == "Windows" %}
# opentelemetry-cpp's recipe removes the `shared` option on Windows and never
# sets BUILD_SHARED_LIBS, so its upstream CMake defaults the protobuf-generated
# `opentelemetry_proto` target to a DLL (opentelemetry_proto.dll). The rest of
# the project links statically and nothing deploys that DLL next to the
# executables, so the telemetry unit test fails to start with
# STATUS_DLL_NOT_FOUND (0xC0000135). Force the dependency to build fully static
# so no runtime DLL is produced. The conf is folded into the package id so a
# fresh static binary is built instead of reusing a previously cached one.
opentelemetry-cpp/*:tools.cmake.cmaketoolchain:extra_variables={"BUILD_SHARED_LIBS": "OFF"}
opentelemetry-cpp/*:tools.info.package_id:confs+=["tools.cmake.cmaketoolchain:extra_variables"]
{% endif %}

View File

@@ -26,7 +26,6 @@ class Xrpl(ConanFile):
"rocksdb": [True, False],
"shared": [True, False],
"static": [True, False],
"telemetry": [True, False],
"tests": [True, False],
"unity": [True, False],
"xrpld": [True, False],
@@ -62,7 +61,6 @@ class Xrpl(ConanFile):
"rocksdb": True,
"shared": False,
"static": True,
"telemetry": True,
"tests": False,
"unity": False,
"xrpld": False,
@@ -148,10 +146,6 @@ class Xrpl(ConanFile):
self.requires("rocksdb/10.5.1")
self.requires("secp256k1/0.7.1", transitive_headers=True)
self.requires("sqlite3/3.53.0", force=True)
# OpenTelemetry C++ SDK for distributed tracing (optional).
# Provides OTLP/HTTP exporter, batch span processor, and trace API.
if self.options.telemetry:
self.requires("opentelemetry-cpp/1.28.0", transitive_headers=True)
self.requires("xxhash/0.8.3", transitive_headers=True)
exports_sources = (
@@ -193,7 +187,6 @@ class Xrpl(ConanFile):
tc.variables["rocksdb"] = self.options.rocksdb
tc.variables["BUILD_SHARED_LIBS"] = self.options.shared
tc.variables["static"] = self.options.static
tc.variables["telemetry"] = self.options.telemetry
tc.variables["unity"] = self.options.unity
tc.variables["xrpld"] = self.options.xrpld
tc.generate()
@@ -248,9 +241,3 @@ class Xrpl(ConanFile):
]
if self.options.rocksdb:
libxrpl.requires.append("rocksdb::librocksdb")
if self.options.telemetry:
libxrpl.requires.append("opentelemetry-cpp::opentelemetry-cpp")
# The public telemetry headers pick their class layout on this
# define, so a consumer that does not see it compiles a different
# SpanGuard than the one inside the library it links.
libxrpl.defines.append("XRPL_ENABLE_TELEMETRY")

View File

@@ -1,86 +0,0 @@
# Docker Compose stack for xrpld OpenTelemetry observability.
#
# Provides services for local development:
# - otel-collector: receives OTLP traces from xrpld, batches and
# forwards them to Tempo. Listens on ports 4317 (gRPC)
# and 4318 (HTTP).
# - tempo: Grafana Tempo tracing backend, queryable via Grafana Explore
# on port 3000. Recommended for production (S3/GCS storage, TraceQL).
# - grafana: dashboards on port 3000, pre-configured with Tempo
# datasource.
#
# Usage:
# docker compose -f docker/telemetry/docker-compose.yml up -d
#
# Configure xrpld to export traces by adding to xrpld.cfg:
# [telemetry]
# enabled=1
# traces_endpoint=http://localhost:4318/v1/traces
services:
# OpenTelemetry Collector: receives spans from xrpld via OTLP protocol,
# batches them for efficiency, and forwards to Tempo for storage.
otel-collector:
image: otel/opentelemetry-collector-contrib:0.158.0
command: ["--config=/etc/otel-collector-config.yaml"]
# Published on the host loopback only. The receivers have no auth and no
# TLS, so only processes on this host may reach them. Note this 127.0.0.1
# is the HOST interface docker listens on; the container-side bind lives in
# the collector config and is a separate choice. Upstream asks for a
# specific interface rather than 0.0.0.0 on either side (CWE-1327):
# https://opentelemetry.io/docs/security/config-best-practices/
ports:
- "127.0.0.1:4317:4317" # OTLP gRPC receiver
- "127.0.0.1:4318:4318" # OTLP HTTP receiver (xrpld sends traces here)
- "127.0.0.1:13133:13133" # Health check endpoint
volumes:
# Mount collector pipeline config (receivers → processors → exporters)
- ./otel-collector-config.yaml:/etc/otel-collector-config.yaml:ro
depends_on:
- tempo
networks:
- xrpld-telemetry
# Grafana Tempo: distributed tracing backend that stores and indexes
# spans. Queryable via TraceQL in Grafana Explore.
tempo:
image: grafana/tempo:2.9.4
command: ["-config.file=/etc/tempo.yaml"]
ports:
- "127.0.0.1:3200:3200" # Tempo HTTP API (health check, query)
volumes:
# Mount Tempo storage and ingestion config
- ./tempo.yaml:/etc/tempo.yaml:ro
# Persistent volume for trace data (WAL + blocks)
- tempo-data:/var/tempo
networks:
- xrpld-telemetry
# Grafana: visualization UI with Tempo pre-configured as a datasource.
# Anonymous admin access enabled for local development convenience.
grafana:
image: grafana/grafana:13.1.2
environment:
- GF_AUTH_ANONYMOUS_ENABLED=true # No login required for local dev
- GF_AUTH_ANONYMOUS_ORG_ROLE=Admin # Full access without auth
ports:
- "127.0.0.1:3000:3000" # Grafana web UI
volumes:
# Auto-provision Tempo datasource and search filters on startup
- ./grafana/provisioning:/etc/grafana/provisioning:ro
depends_on:
- tempo
networks:
- xrpld-telemetry
# Named volume for Tempo trace storage (WAL and compacted blocks).
# Data persists across container restarts. Remove with:
# docker compose -f docker/telemetry/docker-compose.yml down -v
volumes:
tempo-data:
# Isolated bridge network so services communicate by container name
# (e.g., the collector reaches Tempo at http://tempo:4317).
networks:
xrpld-telemetry:
driver: bridge

View File

@@ -1,216 +0,0 @@
# Grafana datasource provisioning for Grafana Tempo.
# Auto-configures Tempo as a trace data source on Grafana startup.
# Access Grafana at http://localhost:3000, then use Explore -> Tempo
# to browse xrpld traces using TraceQL.
#
# Search filters provide pre-configured dropdowns in the Explore UI.
# Each phase adds filters for the span attributes it introduces.
# Base filters — node identity, service, span name, status.
# RPC command, status, role filters.
# Path-finding request, mode and ledger filters.
# Transaction hash, local/peer origin, status.
# Consensus mode, round, ledger sequence, close time.
apiVersion: 1
datasources:
- name: Tempo
type: tempo
access: proxy
url: http://tempo:3200
uid: tempo
jsonData:
nodeGraph:
enabled: true
# Service map and traces-to-metrics require a Prometheus datasource
# (not included in this stack). These features are inactive until a
# Prometheus service is added to docker-compose.yml.
serviceMap:
datasourceUid: prometheus
tracesToMetrics:
datasourceUid: prometheus
spanStartTimeShift: "-1h"
spanEndTimeShift: "1h"
search:
filters:
# --- Node identification filters ---
# service.name: logical service name (default: "xrpld").
# Useful when running multiple service types in the same collector.
- id: service-name
tag: service.name
operator: "="
scope: resource
type: dynamic
# service.instance.id: unique node identifier — defaults to the
# node's public key (e.g., nHB1X37...). Distinguishes individual
# nodes in a multi-node cluster or network.
- id: node-id
tag: service.instance.id
operator: "="
scope: resource
type: dynamic
# service.version: xrpld build version (e.g., "2.4.0-b1").
# Filter traces from specific software releases.
- id: node-version
tag: service.version
operator: "="
scope: resource
type: dynamic
# xrpl.network.id: numeric network identifier
# (0 = mainnet, 1 = testnet, 2 = devnet, etc.).
# Derived from the [network_id] config section.
- id: network-id
tag: xrpl.network.id
operator: "="
scope: resource
type: dynamic
# xrpl.network.type: human-readable network name derived from
# network ID ("mainnet", "testnet", "devnet", "unknown").
- id: network-type
tag: xrpl.network.type
operator: "="
scope: resource
type: dynamic
# --- Span intrinsic filters ---
# name: the span operation name (e.g., "rpc.command.server_info").
# Use to find traces for a specific RPC command or subsystem.
- id: span-name
tag: name
operator: "="
scope: intrinsic
type: dynamic
# status: span completion status ("ok", "error", "unset").
# Filter for failed operations to diagnose errors.
- id: span-status
tag: status
operator: "="
scope: intrinsic
type: dynamic
# duration: span wall-clock duration. Use with ">" operator
# to find slow operations (e.g., duration > 500ms).
- id: span-duration
tag: duration
operator: ">"
scope: intrinsic
type: dynamic
# RPC tracing filters
- id: rpc-command
tag: command
operator: "="
scope: span
type: dynamic
- id: rpc-status
tag: rpc_status
operator: "="
scope: span
type: dynamic
- id: rpc-role
tag: rpc_role
operator: "="
scope: span
type: dynamic
# Path-finding filters. Only the attributes that select a request get
# a dropdown; pathfind_num_paths, pathfind_num_requests and
# pathfind_num_source_assets are measurements read off a span, not
# things an operator searches by.
- id: pathfind-source-account
tag: pathfind_source_account
operator: "="
scope: span
type: dynamic
- id: pathfind-dest-account
tag: pathfind_dest_account
operator: "="
scope: span
type: dynamic
- id: pathfind-dest-currency
tag: pathfind_dest_currency
operator: "="
scope: span
type: dynamic
- id: pathfind-fast
tag: pathfind_fast
operator: "="
scope: span
type: dynamic
# Changes every ledger, so it must be dynamic rather than static.
- id: pathfind-ledger-index
tag: pathfind_ledger_index
operator: "="
scope: span
type: dynamic
- id: pathfind-search-level
tag: pathfind_search_level
operator: ">"
scope: span
type: dynamic
# Transaction tracing filters
- id: tx-hash
tag: tx_hash
operator: "="
scope: span
type: static
# One filter per account role a transaction names. tx_account is on
# every transaction; the rest appear only on the types that carry
# that field, so they are dynamic.
- id: tx-account
tag: tx_account
operator: "="
scope: span
type: static
- id: tx-destination
tag: tx_destination
operator: "="
scope: span
type: dynamic
- id: tx-owner
tag: tx_owner
operator: "="
scope: span
type: dynamic
- id: tx-issuer
tag: tx_issuer
operator: "="
scope: span
type: dynamic
- id: tx-origin
tag: local
operator: "="
scope: span
type: dynamic
- id: tx-status
tag: tx_status
operator: "="
scope: span
type: dynamic
# Consensus tracing filters
- id: consensus-mode
tag: consensus_mode
operator: "="
scope: span
type: static
- id: consensus-round
tag: consensus_round
operator: "="
scope: span
type: dynamic
- id: consensus-ledger-seq
tag: ledger_seq
operator: "="
scope: span
type: static
- id: consensus-close-time-correct
tag: close_time_correct
operator: "="
scope: span
type: static
- id: consensus-state
tag: consensus_state
operator: "="
scope: span
type: static
- id: consensus-close-resolution
tag: close_resolution_ms
operator: "="
scope: span
type: dynamic

View File

@@ -1,74 +0,0 @@
# OpenTelemetry Collector configuration for xrpld development.
#
# Pipeline: OTLP receiver -> batch processor -> debug + Tempo.
# xrpld sends traces via OTLP/HTTP to port 4318. The collector batches
# them and forwards to Tempo via OTLP/gRPC on the Docker network. Tempo
# is queryable via Grafana Explore using TraceQL.
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 1s
send_batch_size: 100
# Deployment-tier tagging. Each collector serves ONE environment and ONE
# network, so it stamps both onto every signal it forwards. This lets a
# single Grafana stack hold data from many collectors and filter by tier.
# - deployment.environment: the collector IS the environment (local, ci,
# test, prod), so it is authoritative -> upsert (overwrite).
# - xrpl.network.type: the xrpld node knows its own chain and already
# stamps this, so the collector only fills it when absent -> insert.
# This keeps a node's real network (e.g. a local node on mainnet)
# from being overwritten by a collector's default.
# Replace the placeholder values per collector; see docker/telemetry
# tier examples.
resource/tier:
attributes:
- key: deployment.environment
value: local
action: upsert
- key: xrpl.network.type
value: mainnet
action: insert
# Strip SDK-injected resource attributes (telemetry.sdk.language/name/version).
# The OpenTelemetry SDK auto-adds these to every Resource; they carry no
# operational value and clutter the attribute set on every backend, so drop
# them here for all signals.
resource/stripsdk:
attributes:
- key: telemetry.sdk.language
action: delete
- key: telemetry.sdk.name
action: delete
- key: telemetry.sdk.version
action: delete
# No attribute hashing or redaction. Account addresses in span attributes
# (pathfind_source_account, pathfind_dest_account) are public ledger
# identifiers and are stored as emitted so they join against explorers and
# logs. Do not add an attributes/hash processor for them.
exporters:
debug:
verbosity: detailed
otlp_grpc/tempo:
endpoint: tempo:4317
tls:
insecure: true
extensions:
health_check:
endpoint: 0.0.0.0:13133
service:
extensions: [health_check]
pipelines:
traces:
receivers: [otlp]
processors: [resource/tier, resource/stripsdk, batch]
exporters: [debug, otlp_grpc/tempo]

View File

@@ -1,61 +0,0 @@
# Grafana Tempo configuration for xrpld telemetry stack.
#
# Runs in single-binary mode for local development.
# Receives traces via OTLP/gRPC from the OTel Collector and stores
# them locally. Queryable via Grafana Explore using the Tempo datasource.
#
# Search filters are configured on the Grafana datasource side
# (grafana/provisioning/datasources/tempo.yaml). Tempo auto-indexes
# all span attributes for search in single-binary mode.
#
# For production, replace local storage with S3/GCS backend and adjust
# retention via the compactor settings. See:
# https://grafana.com/docs/tempo/latest/configuration/
stream_over_http_enabled: true
server:
http_listen_port: 3200
distributor:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
ingester:
max_block_duration: 5m
compactor:
compaction:
block_retention: 1h
# Enable metrics generator for service graph and span metrics.
# Produces RED metrics (rate, errors, duration) per service/span,
# feeding Grafana's service map visualization.
metrics_generator:
registry:
external_labels:
source: tempo
storage:
path: /var/tempo/generator/wal
# Uncomment and add a Prometheus service to docker-compose.yml
# to enable remote_write for service graph metrics:
# remote_write:
# - url: http://prometheus:9090/api/v1/write
overrides:
defaults:
metrics_generator:
processors:
- service-graphs
- span-metrics
storage:
trace:
backend: local
wal:
path: /var/tempo/wal
local:
path: /var/tempo/blocks

View File

@@ -1,295 +0,0 @@
# OpenTelemetry Tracing for xrpld
This document explains how to build xrpld with OpenTelemetry distributed tracing support, configure the runtime telemetry options, and set up the observability backend to view traces.
- [OpenTelemetry Tracing for xrpld](#opentelemetry-tracing-for-xrpld)
- [Overview](#overview)
- [Building with Telemetry](#building-with-telemetry)
- [Summary](#summary)
- [Build steps](#build-steps)
- [Install dependencies](#install-dependencies)
- [Call CMake](#call-cmake)
- [Build](#build)
- [Building without telemetry](#building-without-telemetry)
- [Viewing traces locally](#viewing-traces-locally)
- [Troubleshooting](#troubleshooting)
- [Conan lockfile error](#conan-lockfile-error)
- [CMake target not found](#cmake-target-not-found)
- [No traces in Grafana](#no-traces-in-grafana)
- [Conditional compilation](#conditional-compilation)
- [Span lifetime and cross-thread handling](#span-lifetime-and-cross-thread-handling)
- [`SpanGuard` versus `ScopedSpanGuard`](#spanguard-versus-scopedspanguard)
- [Coroutine-aware context storage](#coroutine-aware-context-storage)
- [Handing a span to a job](#handing-a-span-to-a-job)
- [Why are unrelated spans in my trace?](#why-are-unrelated-spans-in-my-trace)
- [Injecting trace context into a protobuf message](#injecting-trace-context-into-a-protobuf-message)
## Overview
xrpld supports optional [OpenTelemetry](https://opentelemetry.io/) distributed tracing.
When enabled, it instruments RPC requests and the transaction lifecycle with trace spans that are exported via
OTLP/HTTP to an OpenTelemetry Collector, which forwards them to a tracing backend
such as Grafana Tempo.
Telemetry is gated twice — once at compile time and once at runtime:
- **Compile time**: The Conan option `telemetry` must be `True`. It is the only switch: there is no CMake option, because `conan install` writes the value into the generated toolchain and CMake reads it from there.
When disabled, all `SpanGuard` calls compile to inline no-ops (defined in `SpanGuard.h`)
with zero overhead — no OTel SDK dependency required.
- **Runtime**: The `[telemetry]` config section must set `enabled=1`.
When disabled at runtime, a no-op implementation is used.
## Building with Telemetry
### Summary
Follow the same instructions as mentioned in [BUILD.md](../../BUILD.md) but with the following changes:
1. Pass `-o telemetry=True` to `conan install` to pull the `opentelemetry-cpp` dependency.
2. CMake will automatically pick up `telemetry=ON` from the Conan-generated toolchain.
3. Build as usual.
---
### Build steps
```bash
cd /path/to/xrpld
rm -rf .build
mkdir .build
cd .build
```
#### Install dependencies
The `telemetry` option adds `opentelemetry-cpp/1.28.0` as a dependency.
If the Conan lockfile does not yet include this package, bypass it with `--lockfile=""`.
```bash
conan install .. \
--output-folder . \
--build missing \
--settings build_type=Debug \
-o telemetry=True \
-o tests=True \
-o xrpld=True \
--lockfile=""
```
> **Note**: The first build with telemetry may take longer as `opentelemetry-cpp`
> and its transitive dependencies are compiled from source.
#### Call CMake
The Conan-generated toolchain file sets `telemetry=ON` automatically.
No additional CMake flags are needed beyond the standard ones.
```bash
cmake .. -G Ninja \
-DCMAKE_TOOLCHAIN_FILE:FILEPATH=build/generators/conan_toolchain.cmake \
-DCMAKE_BUILD_TYPE=Debug \
-Dtests=ON -Dxrpld=ON
```
You should see in the CMake output:
```
-- OpenTelemetry tracing enabled
```
#### Build
```bash
cmake --build . --parallel $(nproc)
```
## Building without telemetry
Pass `-o telemetry=False` to `conan install`.
Omitting the option is not enough. It then resolves to whatever the recipe's current default is.
The `opentelemetry-cpp` dependency will not be downloaded,
the `XRPL_ENABLE_TELEMETRY` preprocessor define will not be set,
and all tracing macros will compile to no-ops.
The resulting binary is identical to one built before telemetry support was added.
## Viewing traces locally
[`docker/telemetry/docker-compose.yml`](../../docker/telemetry/docker-compose.yml) runs a local backend: an OpenTelemetry Collector, Grafana Tempo and Grafana. Its header comment lists each service and port.
1. Build xrpld with telemetry, as in [Building with Telemetry](#building-with-telemetry).
2. From the repository root, start the stack: `docker compose -f docker/telemetry/docker-compose.yml up -d`.
3. Set `enabled=1` in the `[telemetry]` section of your xrpld config. The default `traces_endpoint` already points at this collector, so no other key is needed. Section 11 of [`cfg/xrpld-example.cfg`](../../cfg/xrpld-example.cfg) documents every key.
4. Start xrpld.
5. Open Grafana at `http://localhost:3000` (no login), go to **Explore**, pick the **Tempo** data source, and search for `service.name` = `xrpld`.
Spans are exported in batches, so a span reaches Tempo a few seconds after the work it records.
To stop the stack, run `docker compose -f docker/telemetry/docker-compose.yml down`. Add `-v` to also delete the stored traces.
## Troubleshooting
### Conan lockfile error
If you see `ERROR: Requirement 'opentelemetry-cpp/1.28.0' not in lockfile 'requires'`,
the lockfile was generated without the telemetry dependency.
Pass `--lockfile=""` to bypass the lockfile, or regenerate it with telemetry enabled.
### CMake target not found
If CMake reports that `opentelemetry-cpp` targets are not found,
ensure you ran `conan install` with `-o telemetry=True` and that the
Conan-generated toolchain file is being used.
The Conan package provides a single umbrella target
`opentelemetry-cpp::opentelemetry-cpp` (not individual component targets).
### No traces in Grafana
Check each hop in order:
1. The CMake output shows `-- OpenTelemetry tracing enabled`. If it does not, tracing is not compiled in.
2. The xrpld log shows `Telemetry started successfully` from the `Telemetry` partition at `info` level. If it does not, `enabled=1` is not set.
3. `docker compose -f docker/telemetry/docker-compose.yml logs otel-collector` prints every span the collector receives. If spans appear there but not in Grafana, the fault is between the collector and Tempo. If they do not appear, xrpld cannot reach `traces_endpoint`.
## Conditional compilation
All OpenTelemetry SDK types are hidden behind the pimpl idiom in `SpanGuard.cpp`. When `XRPL_ENABLE_TELEMETRY` is not defined, `SpanGuard.h` provides an all-inline no-op stub class with no OTel dependencies. At runtime, if `enabled=0` is set in config (or the section is omitted), a `NullTelemetry` implementation is used that returns no-op spans.
Those two layers remove the span, but they do **not** remove the work that computes what you pass to it. The compiled-out guards are ordinary inline functions with ordinary parameters, so every argument is evaluated before the empty body is entered:
```cpp
// to_string() allocates a 64-character string even in a build with telemetry
// compiled out. The call then does nothing with it.
span.setAttribute(attr::txHash, to_string(txID).c_str());
```
Guard the work, not just the call. Testing the guard is enough: its `operator bool()` is a literal `false` when telemetry is compiled out, so the whole block is eliminated, and when telemetry is compiled in it also skips the work if tracing is switched off in config or the span's category is disabled.
```cpp
if (span)
span.setAttribute(attr::txHash, to_string(txID).c_str());
```
A span that exists but was sampled out still pays: there is no `isRecording()` to test.
The `XRPL_METRIC_*` macros are the opposite case. They expand to `do { } while (false)` and discard their arguments, so anything named only inside a macro argument list disappears on its own and needs no guard.
## Span lifetime and cross-thread handling
Telemetry exposes two RAII guards with split responsibilities, plus a
non-owning activation helper. Picking the right one is what keeps a trace's
parent/child nesting and its per-line log correlation correct.
### `SpanGuard` versus `ScopedSpanGuard`
- **`SpanGuard`** owns a span and nothing else. It is _thread-free_: it never
touches the active-context stack, so it carries no thread affinity and may be
moved to and ended on any thread. It is movable and move-assignable. Create
one with `SpanGuard::span(cat, prefix, name)`, or with
`SpanGuard::freshRoot(...)` to start a fresh trace root that ignores whatever
span is currently ambient. Reach for a plain `SpanGuard` whenever the span
must leave the context store that created it — for example when it is handed
into a job.
- **`ScopedSpanGuard`** owns a `SpanGuard` _plus_ an active OTel scope that
pushes the span onto the current context store. While it lives, the span is
the ambient parent for child spans created on that store, and log lines
emitted under it carry its `trace_id`. It is non-copyable and non-movable —
short-lived stack RAII. It offers the same factories (`freshRoot(...)`,
`childSpan(...)`). When the span must outlive the scope, convert it with
`operator SpanGuard() &&`: that pops the scope on the origin store and yields
the bare, thread-free `SpanGuard`.
Rule of thumb: use `ScopedSpanGuard` for same-thread (or same-coroutine)
nesting and log correlation; use the plain `SpanGuard` whenever the span
crosses out of the store that created it.
```mermaid
flowchart TD
SG["SpanGuard<br/>(unscoped, thread-free)<br/>owns span only; movable across threads and coroutines"]
SSG["ScopedSpanGuard<br/>(scoped, store-bound)<br/>owns a SpanGuard plus an active scope on the current store"]
SA["ScopedActivation<br/>(non-owning)<br/>activates a borrowed span; never owns or ends it"]
SSG -->|"handoff: pop scope, yield bare span"| SG
SG -->|"activate / activateIfLive"| SA
SA -.->|"borrows span, no ownership"| SG
classDef box fill:#e8f0fe,stroke:#3b5bdb,color:#111827;
class SG,SSG,SA box;
```
### Coroutine-aware context storage
The active-context stack is not a plain `thread_local`. At telemetry start
xrpld installs `CoroAwareContextStorage`, which keeps the stack in an
`xrpl::LocalValue`. Because `JobQueue::Coro::resume()` swaps the coroutine's
`LocalValue` store in and out with the coroutine, the ambient context _follows
the coroutine_ across every yield and resume — even when it resumes on a
different worker thread. A `ScopedSpanGuard` held across a coroutine yield is
therefore safe: its scope rides the coroutine and pops on the same store it was
pushed onto, so it never pops the wrong stack. Off a coroutine the `LocalValue`
transparently gives each thread its own store, so behaviour matches OTel's
default thread-local storage. This is what lets the RPC entry, process, and
command spans be scoped — for correct nesting and per-line log-trace
correlation — even though the RPC path yields.
### Handing a span to a job
The hand-off pattern is: create a thread-free `SpanGuard` at the origin (or
convert a `ScopedSpanGuard` with `operator SpanGuard() &&`), move it into the
job closure, and inside the worker body activate it non-destructively with
`telemetry::activateIfLive(handle)`. That call takes no ownership and returns a
`ScopedActivation` (a no-op if the handle is empty or the span inactive) which
makes the span the ambient context so log lines in the worker body carry its
`trace_id`. The activation neither owns nor ends the span — the owning
`SpanGuard` still controls its lifetime and ends it when the closure is
destroyed. Keep the activation confined to a synchronous, non-yielding block.
There is no detach step: a `SpanGuard` is already thread-free.
### Why are unrelated spans in my trace?
Historically a scoped guard destroyed off its origin thread popped the wrong
context stack, leaving a stale ambient span in place that later work inherited.
The current design removes that failure mode in two ways:
- **Coroutine-aware storage** makes a scope held across a coroutine yield pop on
the same store it was pushed onto, so a coroutine that resumes on another
worker never pops the wrong stack.
- **A same-store assertion** in `ScopedSpanGuard` (and `ScopedActivation`)
records the `LocalValue` store its scope was pushed onto and checks, in
debug/test/fuzzing builds, that destruction, hand-off, and discard all happen
while that same store is active — turning a genuine cross-store misuse into an
immediate assertion failure rather than a silently corrupted trace.
If a trace still shows unrelated spans nested under one operation, the usual
cause is an inbound entry point that inherited an ambient parent it should not
have. Start such an operation with `freshRoot()` so it begins a clean trace
root and never adopts whatever span happened to be active. To move a span
across a store boundary, keep it in a thread-free `SpanGuard` (or convert via
`operator SpanGuard() &&`) rather than holding a `ScopedSpanGuard` across the
boundary.
### Injecting trace context into a protobuf message
Pass the **whole message** to the injection helpers, never `*msg.mutable_trace_context()`.
On a protobuf `optional` submessage, `mutable_` allocates the submessage and sets its has-bit, and that happens at the call site before the helper runs. A caller that dereferences it therefore puts an empty `TraceContext` on the wire whenever nothing is recorded, and every receiving peer parses it only to drop it. `trace_context` is field 1001, so the wasted bytes are a 2-byte tag plus a zero length.
```cpp
// Right: the helper decides whether the submessage is created at all.
telemetry::injectSpanContext(span, msg);
// Wrong: the submessage exists before the helper can decide anything.
telemetry::injectSpanContext(span, *msg.mutable_trace_context());
```
`injectCurrentContext(msg)` does the same for whichever span is active on the calling thread, deciding via `SpanGuard::hasCurrentContext()`. Four states have to come out right:
| Build | Runtime | Result |
| ------------ | ---------------- | -------------------------------------------------------------------- |
| compiled out | n/a | no submessage; the bytes on the wire match a build without telemetry |
| compiled in | a span is active | `trace_id`, `span_id` and the trace flags are written |
| compiled in | no active span | no submessage, rather than an empty one |
| compiled in | `enabled=0` | no submessage |
`hasCurrentContext()` reads the span straight out of the runtime context. `opentelemetry::trace::GetSpan()` would be shorter, but it returns a heap-allocated `DefaultSpan` when the context holds no span — an allocation in exactly the case the predicate exists to keep free.

12
flake.lock generated
View File

@@ -2,11 +2,11 @@
"nodes": {
"nixpkgs": {
"locked": {
"lastModified": 1781173989,
"narHash": "sha256-fnzKKPvS+oieI/pTzotA5tkoM47EB1NpaBcgk4R97hE=",
"lastModified": 1791130267,
"narHash": "sha256-1sjwcQMcgAbmiip/VivBSiK8p2bTkQGOvCuO7vnwfx4=",
"owner": "NixOS",
"repo": "nixpkgs",
"rev": "8c91a71d13451abc40eb9dae8910f972f979852f",
"rev": "9013764fcc0ea99fa16cf7aa7decf8a5c3889dd2",
"type": "github"
},
"original": {
@@ -47,11 +47,11 @@
]
},
"locked": {
"lastModified": 1784611586,
"narHash": "sha256-OfqgY+0hp/zseZB7uyH0U8kIDPS4scZZCyAurEplvG0=",
"lastModified": 1791189933,
"narHash": "sha256-Y8eN7ki0Urx4Ojf4w1xSpWg8uk/3lJqK+E0Oth77rH0=",
"owner": "oxalica",
"repo": "rust-overlay",
"rev": "14f58845249f3552a89b07772626b8d3c632fa86",
"rev": "e60029353d0c48d216bc4b065168ccd4079c166f",
"type": "github"
},
"original": {

View File

@@ -1,29 +1,31 @@
#pragma once
#include <xrpl/beast/type_name.h>
#include <xrpl/beast/utility/instrumentation.h>
#include <boost/core/type_name.hpp>
#include <algorithm>
#include <atomic>
#include <cstddef>
#include <cstdint>
#include <iterator>
#include <string>
#include <type_traits>
#include <utility>
#include <vector>
namespace xrpl {
/**
* Manages all counted object types.
*
* Counters register themselves on a lock-free intrusive list maintained
* by this object when constructed. Because counters are never destroyed
* or removed, the ABA problem does not apply.
*
* The registry is iterable as a forward range.
*/
class CountedObjects
{
public:
static CountedObjects&
getInstance() noexcept;
using Entry = std::pair<std::string, int>;
using List = std::vector<Entry>;
[[nodiscard]] List
getCounts(int minimumThreshold) const;
public:
/**
* Implementation for @ref CountedObject.
@@ -33,68 +35,147 @@ public:
class Counter
{
public:
Counter(std::string name) noexcept : name_(std::move(name)), count_(0)
{
// Insert ourselves at the front of the lock-free linked list
CountedObjects& instance = CountedObjects::getInstance();
Counter* head = nullptr;
Counter(std::string name) noexcept;
do
{
head = instance.head_.load();
next_ = head;
} while (instance.head_.exchange(this) != head);
// Counters are intrusive list nodes whose addresses are published
// in the registry; they must never be copied or moved. The atomic
// members already force this, but we make it explicit.
Counter(Counter const&) = delete;
Counter&
operator=(Counter const&) = delete;
Counter(Counter&&) = delete;
Counter&
operator=(Counter&&) = delete;
++instance.count_;
}
~Counter() noexcept = default;
int
std::uint32_t
increment() noexcept
{
return ++count_;
auto const newCount = count_.fetch_add(1, std::memory_order::relaxed) + 1;
XRPL_ASSERT(newCount != 0, "xrpl::CountedObjects::Counter::increment : no overflow");
auto maxCount = maxCount_.load(std::memory_order::relaxed);
while (newCount > maxCount &&
!maxCount_.compare_exchange_weak(maxCount, newCount, std::memory_order::relaxed))
{
}
return newCount;
}
int
std::uint32_t
decrement() noexcept
{
return --count_;
auto const prev = count_.fetch_sub(1, std::memory_order::relaxed);
XRPL_ASSERT(prev != 0, "xrpl::CountedObjects::Counter::decrement : no underflow");
return prev - 1;
}
[[nodiscard]] int
getCount() const noexcept
[[nodiscard]] std::uint32_t
count() const noexcept
{
return count_.load();
return count_.load(std::memory_order::relaxed);
}
[[nodiscard]] Counter*
getNext() const noexcept
[[nodiscard]] std::uint32_t
max() const noexcept
{
return next_;
return std::max(
count_.load(std::memory_order::relaxed),
maxCount_.load(std::memory_order::relaxed));
}
[[nodiscard]] std::string const&
getName() const noexcept
name() const noexcept
{
return name_;
}
private:
friend class CountedObjects;
Counter* next_ = nullptr;
std::atomic<std::uint32_t> count_ = 0;
std::atomic<std::uint32_t> maxCount_ = 0;
std::string const name_;
std::atomic<int> count_;
Counter* next_;
};
private:
CountedObjects() noexcept;
~CountedObjects() noexcept = default;
class Iterator
{
public:
using value_type = Counter const;
using reference = value_type&;
using pointer = value_type*;
using difference_type = std::ptrdiff_t;
using iterator_category = std::forward_iterator_tag;
explicit Iterator(Counter* c = nullptr) noexcept : current_(c)
{
}
reference
operator*() const noexcept
{
return *current_;
}
pointer
operator->() const noexcept
{
return current_;
}
Iterator&
operator++() noexcept
{
current_ = current_->next_;
return *this;
}
Iterator
operator++(int) noexcept
{
auto tmp = *this;
++*this;
return tmp;
}
bool
operator==(Iterator const&) const noexcept = default;
private:
Counter* current_;
};
constexpr CountedObjects() noexcept = default;
[[nodiscard]] auto
begin() const noexcept
{
return Iterator{head_.load(std::memory_order::acquire)};
}
[[nodiscard]] auto
end() const noexcept
{
return Iterator{};
}
private:
std::atomic<int> count_;
std::atomic<Counter*> head_;
std::atomic<Counter*> head_ = nullptr;
};
/** The global counted object registry. */
inline constinit CountedObjects gCountedObjects;
inline CountedObjects::Counter::Counter(std::string name) noexcept
: next_(gCountedObjects.head_.load(std::memory_order::relaxed)), name_(std::move(name))
{
while (!gCountedObjects.head_.compare_exchange_weak(
next_, this, std::memory_order::release, std::memory_order::relaxed))
;
}
//------------------------------------------------------------------------------
/**
@@ -103,27 +184,37 @@ private:
* Derived classes have their instances counted automatically. This is used
* for reporting purposes.
*
* The constructors are private and `Object` is befriended so that the
* CRTP parameter must be the deriving class itself: a copy-paste error
* like `class B : public CountedObject<A>` fails to compile instead of
* silently polluting A's count.
*
* @note This class has no move operations by design: a derived class's
* move constructor falls back to the copy constructor for this
* base, so the newly created instance is counted. This keeps the
* invariant that count is the number of outstanding subobjects.
*
* @warning Counted objects constructed during dynamic initialization of
* other translation units may have their increments discarded when
* counter itself is dynamically initialized. Do not create counted
* objects before main() begins.
*
* @ingroup basics
*/
template <class Object>
requires std::is_class_v<Object>
class CountedObject
{
private:
static auto&
getCounter() noexcept
{
static CountedObjects::Counter kC{beast::typeName<Object>()};
return kC;
}
static inline CountedObjects::Counter counter{boost::core::type_name<Object>()};
CountedObject() noexcept
{
getCounter().increment();
counter.increment();
}
CountedObject(CountedObject const&) noexcept
{
getCounter().increment();
counter.increment();
}
CountedObject&
@@ -132,7 +223,7 @@ private:
public:
~CountedObject() noexcept
{
getCounter().decrement();
counter.decrement();
}
friend Object;

View File

@@ -3,11 +3,11 @@
#pragma once
#include <xrpl/basics/ByteUtilities.h>
#include <xrpl/beast/type_name.h>
#include <xrpl/beast/utility/instrumentation.h>
#include <boost/align.hpp>
#include <boost/container/static_vector.hpp>
#include <boost/core/type_name.hpp>
#include <boost/predef.h>
#include <algorithm>
@@ -16,6 +16,7 @@
#include <cstring>
#include <mutex>
#include <stdexcept>
#include <typeinfo>
#include <vector>
#if BOOST_OS_LINUX
@@ -335,7 +336,7 @@ public:
}) != cfg.end())
{
throw std::runtime_error(
"SlabAllocatorSet<" + beast::typeName<Type>() + ">: duplicate slab size");
"SlabAllocatorSet<" + boost::core::type_name<Type>() + ">: duplicate slab size");
}
for (auto const& c : cfg)

File diff suppressed because it is too large Load Diff

View File

@@ -1,7 +1,8 @@
#pragma once
#include <xrpl/basics/sanitizers.h>
#include <xrpl/beast/type_name.h>
#include <boost/core/type_name.hpp>
#include <exception>
#include <string>
@@ -55,7 +56,7 @@ Throw(Args&&... args)
std::is_convertible_v<E*, std::exception*>, "Exception must derive from std::exception.");
E e(std::forward<Args>(args)...);
logThrow(std::string("Throwing exception of type " + beast::typeName<E>() + ": ") + e.what());
logThrow("Throwing exception of type " + boost::core::type_name<E>() + ": " + e.what());
throw std::move(e);
}

View File

@@ -0,0 +1,200 @@
#pragma once
#include <concepts>
#include <type_traits>
#include <utility>
namespace xrpl {
/**
* Opt-in bitwise operators for scoped enumerations.
*
* Provides `&`, `|`, `^`, `~` and the corresponding compound assignment forms
* for any scoped enumeration with an unsigned underlying type that opts in by
* specializing @ref OptIn:
*
* @code
* namespace xrpl {
*
* enum class MyFlags : std::uint32_t { a = 1, b = 2, c = 4 };
*
* template <>
* struct enum_bitops::OptIn<MyFlags> : std::true_type {};
*
* } // namespace xrpl
* @endcode
*
* The operators are constexpr and noexcept, and return the enumeration type.
*
* @par Where the specialization may be declared
* In namespace xrpl, or at global scope with full qualification, and after
* the enumeration is defined but before the operators are first used. This
* cannot be done in a nested namespace or at class scope; for enumerations
* nested in a class this means after the class definition.
*
* @par Where the operators are found
* The operators are declared in namespace xrpl, so argument-dependent lookup
* will find them only for enumerations declared in xrpl or nested in a class
* declared directly in xrpl. Enumerations in a nested namespace will only be
* found from code inside xrpl, and only if no enclosing scope declares an
* operator of the same name; elsewhere they require using-declarations.
*/
namespace enum_bitops {
/**
* Types that may be opted in to the bitwise operators.
*
* Satisfied by a scoped enumeration whose underlying type is an unsigned
* integer type.
*
* @tparam T The type to test.
*/
template <typename T>
concept Eligible = std::is_scoped_enum_v<T> && std::unsigned_integral<std::underlying_type_t<T>>;
/**
* Opt-in switch for the bitwise operators.
*
* The primary template derives from std::false_type. Specialize it to derive
* from std::true_type to enable the operators for an @ref Eligible enumeration.
*
* @tparam T The scoped enumeration to opt in.
*/
template <Eligible T>
struct OptIn : std::false_type
{
};
/**
* Enumerations for which the bitwise operators are enabled.
*
* Satisfied when @p T satisfies @ref Eligible and @ref OptIn has been
* specialized for it to derive from std::true_type.
*
* @tparam T The type to test.
*/
template <typename T>
concept Candidate = Eligible<T> && OptIn<T>::value;
} // namespace enum_bitops
// These are declared in xrpl, not in enum_bitops, so that argument-dependent
// lookup finds them for enumerations whose associated namespace is xrpl.
/**
* @name Bitwise operators for opted-in scoped enumerations
*
* Each operator applies the corresponding built-in operator to the
* underlying values and converts the result back to the enumeration type.
* Both operands must have the same enumeration type; there is no implicit
* conversion to or from the underlying type.
*/
/** @{ */
/**
* Bitwise AND.
*
* @param lhs The left operand.
* @param rhs The right operand.
* @return The bits set in both @p lhs and @p rhs.
*/
template <enum_bitops::Candidate T>
constexpr T
operator&(T lhs, T rhs) noexcept
{
return static_cast<T>(std::to_underlying(lhs) & std::to_underlying(rhs));
}
/**
* Bitwise OR.
*
* @param lhs The left operand.
* @param rhs The right operand.
* @return The bits set in @p lhs, in @p rhs, or in both.
*/
template <enum_bitops::Candidate T>
constexpr T
operator|(T lhs, T rhs) noexcept
{
return static_cast<T>(std::to_underlying(lhs) | std::to_underlying(rhs));
}
/**
* Bitwise exclusive OR.
*
* @param lhs The left operand.
* @param rhs The right operand.
* @return The bits set in exactly one of @p lhs and @p rhs.
*/
template <enum_bitops::Candidate T>
constexpr T
operator^(T lhs, T rhs) noexcept
{
return static_cast<T>(std::to_underlying(lhs) ^ std::to_underlying(rhs));
}
/**
* Bitwise complement.
*
* The result has every bit of the underlying type that is clear in
* @p val, including bits that no enumerator names. It is intended for
* clearing flags, as in `flags & ~flag`.
*
* @param val The operand.
* @return The complement of @p val.
*/
template <enum_bitops::Candidate T>
constexpr T
operator~(T val) noexcept
{
return static_cast<T>(~std::to_underlying(val));
}
/**
* Bitwise AND assignment.
*
* @param lhs The value to modify.
* @param rhs The right operand.
* @return A reference to @p lhs.
*/
template <enum_bitops::Candidate T>
constexpr T&
operator&=(T& lhs, T rhs) noexcept
{
lhs = lhs & rhs;
return lhs;
}
/**
* Bitwise OR assignment.
*
* @param lhs The value to modify.
* @param rhs The right operand.
* @return A reference to @p lhs.
*/
template <enum_bitops::Candidate T>
constexpr T&
operator|=(T& lhs, T rhs) noexcept
{
lhs = lhs | rhs;
return lhs;
}
/**
* Bitwise exclusive OR assignment.
*
* @param lhs The value to modify.
* @param rhs The right operand.
* @return A reference to @p lhs.
*/
template <enum_bitops::Candidate T>
constexpr T&
operator^=(T& lhs, T rhs) noexcept
{
lhs = lhs ^ rhs;
return lhs;
}
/** @} */
} // namespace xrpl

View File

@@ -2,107 +2,212 @@
#include <xrpl/beast/utility/instrumentation.h> // IWYU pragma: keep
#include <concepts>
#include <cstdint>
#include <limits>
#include <memory>
#include <type_traits>
#include <utility>
namespace xrpl {
// safe_cast adds compile-time checks to a static_cast to ensure that
// the destination can hold all values of the source. This is particularly
// handy when the source or destination is an enumeration type.
/** Every value of @p Src can be represented by @p Dest.
template <class Src, class Dest>
concept SafeToCast = (std::is_integral_v<Src> && std::is_integral_v<Dest>) &&
(std::is_signed_v<Src> || std::is_unsigned_v<Dest>) &&
(std::is_signed_v<Src> != std::is_signed_v<Dest> ? sizeof(Dest) > sizeof(Src)
: sizeof(Dest) >= sizeof(Src));
Given two integral types, the cast is safe when the destination can
hold every possible value of the source.
template <class Dest, class Src>
Comparing the bounds requires care: we use @c std::cmp_less_equal and
@c std::cmp_greater_equal; the plain relational operators would apply
arithmetic conversions, resulting in incorrect results when comparing
across signedness.
Because @c std::cmp_* requires standard signed or unsigned integer
arguments, which excludes character types and bool, we first widen
all type bounds to the maximum-width integer type while preserving
signedness.
@note Extended integer types, like __int128 on gcc, cannot be safely
widened to a standard integer type, so the concept will reject
them.
*/
template <typename Src, typename Dest>
concept SafeToCast = std::is_integral_v<Src> && std::is_integral_v<Dest> && []() consteval {
using WideSrc = std::conditional_t<std::is_signed_v<Src>, std::intmax_t, std::uintmax_t>;
using WideDest = std::conditional_t<std::is_signed_v<Dest>, std::intmax_t, std::uintmax_t>;
// Note: this guard must be an evaluated branch, and not a
// static_assert. The lambda body is outside the immediate
// context, so a substitution-time failure here would be a
// hard error rather than leaving the concept unsatisfied.
if constexpr (sizeof(Src) > sizeof(WideSrc) || sizeof(Dest) > sizeof(WideDest))
{
return false;
}
else
{
return std::cmp_less_equal(
static_cast<WideDest>(std::numeric_limits<Dest>::min()),
static_cast<WideSrc>(std::numeric_limits<Src>::min())) &&
std::cmp_greater_equal(
static_cast<WideDest>(std::numeric_limits<Dest>::max()),
static_cast<WideSrc>(std::numeric_limits<Src>::max()));
}
}();
/** Compile-time-checked static_cast that rejects non-value preserving casts.
@note There is deliberately no enum-to-enum overload, and no overload
returning the underlying type of an enum. For the latter, use
@c std::to_underlying.
*/
/** @{ */
template <typename Dest, typename Src>
requires(std::is_integral_v<Dest> && std::is_integral_v<Src>)
constexpr Dest
safeCast(Src s) noexcept
requires(std::is_integral_v<Dest> && std::is_integral_v<Src>)
{
static_assert(
std::is_signed_v<Dest> || std::is_unsigned_v<Src>, "Cannot cast signed to unsigned");
constexpr unsigned kNotSame = std::is_signed_v<Dest> != std::is_signed_v<Src>;
static_assert(
sizeof(Dest) >= sizeof(Src) + kNotSame,
"Destination is too small to hold all values of source");
SafeToCast<Src, Dest>, "This cast is not value-preserving. Please use unsafeCast instead.");
return static_cast<Dest>(s);
}
template <class Dest, class Src>
template <typename Dest, typename Src>
requires(std::is_enum_v<Dest> && std::is_integral_v<Src>)
constexpr Dest
safeCast(Src s) noexcept
requires(std::is_enum_v<Dest> && std::is_integral_v<Src>)
{
return static_cast<Dest>(safeCast<std::underlying_type_t<Dest>>(s));
}
template <class Dest, class Src>
template <typename Dest, typename Src>
requires(std::is_integral_v<Dest> && std::is_enum_v<Src>)
constexpr Dest
safeCast(Src s) noexcept
requires(std::is_integral_v<Dest> && std::is_enum_v<Src>)
{
return safeCast<Dest>(static_cast<std::underlying_type_t<Src>>(s));
return safeCast<Dest>(std::to_underlying(s));
}
/** @} */
// unsafe_cast explicitly flags a static_cast as not necessarily able to hold
// all values of the source. It includes a compile-time check so that if
// underlying types become safe, it can be converted to a safe_cast.
/** Integral-to-integral cast that is known to be narrowing or sign-erasing.
template <class Dest, class Src>
This explicitly flags a conversion that can lose information for some
values of the source type, where the call site accepts that loss (or
truncation is the intended behavior).
The compile-time check ensures the cast remains "unsafe": if the types
involved later change such that the conversion becomes inherently
value-preserving, the static assertion fires with instructions to
migrate the call site to @ref safeCast.
If the conversion's safety depends on a runtime precondition rather
than on the types, or varies across instantiations of generic code,
use @ref checkedCast instead.
*/
/** @{ */
template <typename Dest, typename Src>
requires(std::is_integral_v<Dest> && std::is_integral_v<Src>)
constexpr Dest
unsafeCast(Src s) noexcept
requires(std::is_integral_v<Dest> && std::is_integral_v<Src>)
{
static_assert(
!SafeToCast<Src, Dest>,
"Only unsafe if casting signed to unsigned or "
"destination is too small");
!SafeToCast<Src, Dest>, "This cast is value-preserving. Please use safeCast instead.");
return static_cast<Dest>(s);
}
template <class Dest, class Src>
template <typename Dest, typename Src>
requires(std::is_enum_v<Dest> && std::is_integral_v<Src>)
constexpr Dest
unsafeCast(Src s) noexcept
requires(std::is_enum_v<Dest> && std::is_integral_v<Src>)
{
return static_cast<Dest>(unsafeCast<std::underlying_type_t<Dest>>(s));
}
template <class Dest, class Src>
template <typename Dest, typename Src>
requires(std::is_integral_v<Dest> && std::is_enum_v<Src>)
constexpr Dest
unsafeCast(Src s) noexcept
requires(std::is_integral_v<Dest> && std::is_enum_v<Src>)
{
return unsafeCast<Dest>(static_cast<std::underlying_type_t<Src>>(s));
return unsafeCast<Dest>(std::to_underlying(s));
}
/** @} */
/** Integral-to-integral cast when the caller has performed a bounds check.
This documents that a runtime precondition or external invariant, which
is not necessarily visible to the compiler, guarantees that the requested
conversion is value-preserving for the values that can actually occur.
This is primarily meant for generic code, where the same expression may
be value-preserving for one instantiation but not for another, making
both @ref safeCast and @ref unsafeCast unusable.
Unlike @ref safeCast and @ref unsafeCast, this imposes no static check
on the type relationship: the caller's claim is about runtime values,
not about types.
*/
/** @{ */
template <typename Dest, typename Src>
requires(std::is_integral_v<Dest> && std::is_integral_v<Src>)
constexpr Dest
checkedCast(Src s) noexcept
{
return static_cast<Dest>(s);
}
template <typename Dest, typename Src>
requires(std::is_enum_v<Dest> && std::is_integral_v<Src>)
constexpr Dest
checkedCast(Src s) noexcept
{
return static_cast<Dest>(checkedCast<std::underlying_type_t<Dest>>(s));
}
template <typename Dest, typename Src>
requires(std::is_integral_v<Dest> && std::is_enum_v<Src>)
constexpr Dest
checkedCast(Src s) noexcept
{
return checkedCast<Dest>(std::to_underlying(s));
}
/** @} */
/** Downcast within a class hierarchy, verified in debug builds.
Performs a static_cast down a hierarchy, but in debug builds verifies
via dynamic_cast that the object's dynamic type actually permits the
downcast. Both build modes execute the same conversion; debug builds
merely add the check.
The pointer form passes null through unchanged, as dynamic_cast does.
@note The check requires @p Src to be polymorphic; in release builds
an invalid downcast is undefined behavior on use, exactly as
with a bare static_cast.
*/
/** @{ */
template <class Dest, class Src>
requires std::is_pointer_v<Dest>
requires(
std::is_pointer_v<Dest> && std::is_polymorphic_v<Src> &&
std::derived_from<std::remove_pointer_t<Dest>, Src> &&
(std::is_const_v<std::remove_pointer_t<Dest>> || !std::is_const_v<Src>) &&
(std::is_volatile_v<std::remove_pointer_t<Dest>> || !std::is_volatile_v<Src>))
inline Dest
safeDowncast(Src* s) noexcept
{
#ifdef NDEBUG
XRPL_ASSERT(
s == nullptr || dynamic_cast<Dest>(s) != nullptr, "xrpl::safeDowncast : valid downcast");
return static_cast<Dest>(s); // NOLINT(cppcoreguidelines-pro-type-static-cast-downcast)
#else
auto* result = dynamic_cast<Dest>(s);
XRPL_ASSERT(result != nullptr, "xrpl::safeDowncast : pointer downcast is valid");
return result;
#endif
}
template <class Dest, class Src>
requires std::is_lvalue_reference_v<Dest>
requires(
std::is_lvalue_reference_v<Dest> && std::is_polymorphic_v<Src> &&
std::derived_from<std::remove_reference_t<Dest>, Src>)
inline Dest
safeDowncast(Src& s) noexcept
{
#ifndef NDEBUG
XRPL_ASSERT(
dynamic_cast<std::add_pointer_t<std::remove_reference_t<Dest>>>(&s) != nullptr,
"xrpl::safeDowncast : reference downcast is valid");
#endif
return static_cast<Dest>(s); // NOLINT(cppcoreguidelines-pro-type-static-cast-downcast)
return *safeDowncast<std::add_pointer_t<std::remove_reference_t<Dest>>>(std::addressof(s));
}
/** @} */
} // namespace xrpl

View File

@@ -4,88 +4,215 @@
#include <xrpl/beast/utility/instrumentation.h>
#include <boost/predef/architecture.h>
#include <atomic>
#include <concepts>
#include <limits>
#include <type_traits>
#ifndef __aarch64__
#if BOOST_ARCH_X86
#include <immintrin.h>
#endif
namespace xrpl {
/** An unsigned integral type suitable for use as a spinlock.
The type must be always lock-free when wrapped in std::atomic, so
that lock operations cannot themselves take a (library-level) lock.
*/
template <typename T>
concept SpinlockValueType = std::is_unsigned_v<T> && std::atomic<T>::is_always_lock_free;
/** A spinlock value type that additionally supports the atomic bitwise
operations required to pack multiple locks into a single integer.
*/
template <typename T>
concept PackedSpinlockValueType = SpinlockValueType<T> && requires(std::atomic<T>& a, T v) {
{ a.fetch_or(v) } -> std::same_as<T>;
{ a.fetch_and(v) } -> std::same_as<T>;
};
namespace detail {
/**
* Inform the processor that we are in a tight spin-wait loop.
*
* Spinlocks caught in tight loops can result in the processor's pipeline
* filling up with comparison operations, resulting in a misprediction at
* the time the lock is finally acquired, necessitating pipeline flushing
* which is ridiculously expensive and results in very high latency.
*
* This function instructs the processor to "pause" for some architecture
* specific amount of time, to prevent this.
/** Inform the processor that we are in a tight spin-wait loop.
Spinlocks caught in tight loops can result in the processor's pipeline
filling up with comparison operations, resulting in a misprediction at
the time the lock is finally acquired, necessitating pipeline flushing
which is ridiculously expensive and results in very high latency.
This function instructs the processor to "pause" for some architecture
specific amount of time, to prevent this.
*/
inline void
spinPause() noexcept
{
#ifdef __aarch64__
asm volatile("yield");
#else
#if BOOST_ARCH_X86
_mm_pause();
#elif BOOST_ARCH_ARM
asm volatile("yield" ::: "memory");
#else
#error No implementation available for spinPause to use
#endif
}
} // namespace detail
/** @{ */
/**
* Classes to handle arrays of spinlocks packed into a single atomic integer:
*
* Packed spinlocks allow for tremendously space-efficient lock-sharding
* but they come at a cost.
*
* First, the implementation is necessarily low-level and uses advanced
* features like memory ordering and highly platform-specific tricks to
* maximize performance. This imposes a significant and ongoing cost to
* developers.
*
* Second, and perhaps most important, is that the packing of multiple
* locks into a single integer which, albeit space-efficient, also has
* performance implications stemming from data dependencies, increased
* cache-coherency traffic between processors and heavier loads on the
* processor's load/store units.
*
* To be sure, these locks can have advantages but they are definitely
* not general purpose locks and should not be thought of or used that
* way. The use cases for them are likely few and far between; without
* a compelling reason to use them, backed by profiling data, it might
* be best to use one of the standard locking primitives instead. Note
* that in most common platforms, `std::mutex` is so heavily optimized
* that it can, usually, outperform spinlocks.
*
* @tparam T An unsigned integral type (e.g. std::uint16_t)
*/
//------------------------------------------------------------------------------
/**
* A class that grabs a single packed spinlock from an atomic integer.
*
* This class meets the requirements of Lockable:
* https://en.cppreference.com/w/cpp/named_req/Lockable
/** Attempt to acquire a spinlock without blocking.
@note This interface is primarily intended for one-shot attempts to
acquire the lock. Avoid calling this function directly from a
loop and use @ref spinLock instead.
@tparam T An unsigned integral type.
@param lock The atomic variable used as the lock.
@return true if the lock was acquired, false if it was already held.
*/
template <class T>
template <SpinlockValueType T>
[[nodiscard]] bool
spinTryLock(std::atomic<T>& lock) noexcept
{
// A compare-exchange is required here, not an unconditional exchange:
// a failed attempt must not modify the lock word, in case the atomic
// is shared with PackedSpinlock).
T expected = 0;
return lock.compare_exchange_strong(
expected,
std::numeric_limits<T>::max(),
std::memory_order::acquire,
std::memory_order::relaxed);
;
}
/** Acquire a spinlock, blocking until available.
Uses a TTAS (test-and-test-and-set) pattern, so waiters share the cache
line read-only, helping to avoid unnecessary coherency traffic.
@tparam T An unsigned integral type.
@param lock The atomic variable used as the lock.
*/
template <SpinlockValueType T>
void
spinLock(std::atomic<T>& lock) noexcept
{
do
{
// Relaxed ordering is sufficient for the spin: this load is only
// a filter. The acquire on the successful exchange in spinTryLock
// is what synchronizes the critical section.
while (lock.load(std::memory_order::relaxed) != 0)
detail::spinPause();
} while (!spinTryLock(lock));
}
/** Release a spinlock.
@tparam T An unsigned integral type.
@param lock The atomic variable used as the lock.
*/
template <SpinlockValueType T>
void
spinUnlock(std::atomic<T>& lock) noexcept
{
lock.store(0, std::memory_order::release);
}
//------------------------------------------------------------------------------
/** A Lockable interface to a spinlock implemented on top of an atomic.
@tparam T An unsigned integral type.
@note Using `PackedSpinlock` and `Spinlock` against the same underlying
atomic integer is possible but can result in `Spinlock` not being
able to acquire the lock during periods of high contention due to
the way the two locks operate: `Spinlock` spins and tries to grab
all the bits at once, whereas any given `PackedSpinlock` instance
only tries to grab one bit at a time. Caveat emptor.
This class meets the requirements of Lockable:
https://en.cppreference.com/w/cpp/named_req/Lockable
*/
template <SpinlockValueType T>
class Spinlock
{
std::atomic<T>& lock_;
public:
Spinlock(Spinlock const&) = delete;
Spinlock&
operator=(Spinlock const&) = delete;
/** Construct a spinlock handle.
@param lock The atomic integer to spin against.
@note For performance reasons, you should strive to have `lock` be
on a cacheline by itself.
*/
explicit Spinlock(std::atomic<T>& lock) noexcept : lock_(lock)
{
}
[[nodiscard]] bool
try_lock() noexcept // NOLINT(readability-identifier-naming)
{
return spinTryLock(lock_);
}
void
lock() noexcept
{
spinLock(lock_);
}
void
unlock() noexcept
{
spinUnlock(lock_);
}
};
//------------------------------------------------------------------------------
/** A Lockable interface to a packed spinlock implemented on top of an atomic.
Packed spinlocks offer tremendous space-efficient lock-sharding but
they come at a cost.
First, the implementation is necessarily low-level and uses advanced
features like memory ordering and highly platform-specific tricks to
maximize performance. This imposes a significant and ongoing cost to
developers.
Second, and perhaps most important, is that the packing of multiple
locks into a single integer which, albeit space-efficient, also has
performance implications stemming from data dependencies, increased
cache-coherency traffic between processors and heavier loads on the
processor's load/store units.
To be sure, these locks can have advantages but they are definitely
not general purpose locks and should not be thought of or used that
way. The use cases for them are likely few and far between; without
a compelling reason to use them, backed by profiling data, it might
be best to use one of the standard locking primitives instead. Note
that in most common platforms, `std::mutex` is so heavily optimized
that it can, usually, outperform spinlocks.
@tparam T An unsigned integral type (e.g. std::uint16_t)
This class meets the requirements of Lockable:
https://en.cppreference.com/w/cpp/named_req/Lockable
*/
template <PackedSpinlockValueType T>
class PackedSpinlock
{
// clang-format off
static_assert(std::is_unsigned_v<T>);
static_assert(std::atomic<T>::is_always_lock_free);
static_assert(
std::is_same_v<decltype(std::declval<std::atomic<T>&>().fetch_or(0)), T> &&
std::is_same_v<decltype(std::declval<std::atomic<T>&>().fetch_and(0)), T>,
"std::atomic<T>::fetch_and(T) and std::atomic<T>::fetch_and(T) are required by packed_spinlock");
// clang-format on
private:
std::atomic<T>& bits_;
T const mask_;
@@ -94,120 +221,49 @@ public:
PackedSpinlock&
operator=(PackedSpinlock const&) = delete;
/**
* A single spinlock packed inside the specified atomic
*
* @param lock The atomic integer inside which the spinlock is packed.
* @param index The index of the spinlock this object acquires.
*
* @note For performance reasons, you should strive to have `lock` be
* on a cacheline by itself.
/** Construct a packed spinlock handle for a single bit.
@param lock The atomic integer inside which the spinlock is packed.
@param index The index of the spinlock this object acquires.
@note For performance reasons, you should strive to have `lock` be
on a cacheline by itself.
*/
PackedSpinlock(std::atomic<T>& lock, int index) : bits_(lock), mask_(static_cast<T>(1) << index)
{
XRPL_ASSERT(
index >= 0 && (mask_ != 0),
"xrpl::PackedSpinlock::PackedSpinlock : valid index and mask");
}
[[nodiscard]] bool
try_lock() // NOLINT(readability-identifier-naming)
{
return (bits_.fetch_or(mask_, std::memory_order_acquire) & mask_) == 0;
}
void
lock()
{
while (!try_lock())
{
// The use of relaxed memory ordering here is intentional and
// serves to help reduce cache coherency traffic during times
// of contention by avoiding writes that would definitely not
// result in the lock being acquired.
while ((bits_.load(std::memory_order_relaxed) & mask_) != 0)
detail::spinPause();
}
}
void
unlock()
{
bits_.fetch_and(~mask_, std::memory_order_release);
}
};
/**
* A spinlock implemented on top of an atomic integer.
*
* @note Using `packed_spinlock` and `spinlock` against the same underlying
* atomic integer can result in `spinlock` not being able to actually
* acquire the lock during periods of high contention, because of how
* the two locks operate: `spinlock` will spin trying to grab all the
* bits at once, whereas any given `packed_spinlock` will only try to
* grab one bit at a time. Caveat emptor.
*
* This class meets the requirements of Lockable:
* https://en.cppreference.com/w/cpp/named_req/Lockable
*/
template <class T>
class Spinlock
{
static_assert(std::is_unsigned_v<T>);
static_assert(std::atomic<T>::is_always_lock_free);
private:
std::atomic<T>& lock_;
public:
Spinlock(Spinlock const&) = delete;
Spinlock&
operator=(Spinlock const&) = delete;
/**
* Grabs the
*
* @param lock The atomic integer to spin against.
*
* @note For performance reasons, you should strive to have `lock` be
* on a cacheline by itself.
*/
Spinlock(std::atomic<T>& lock) : lock_(lock)
PackedSpinlock(std::atomic<T>& lock, int index) noexcept
: bits_(lock), mask_([index]() {
XRPL_ASSERT(
index >= 0 && index < std::numeric_limits<T>::digits,
"xrpl::PackedSpinlock::PackedSpinlock : valid index");
return static_cast<T>(T{1} << index);
}())
{
}
[[nodiscard]] bool
try_lock() // NOLINT(readability-identifier-naming)
try_lock() noexcept // NOLINT(readability-identifier-naming)
{
T expected = 0;
return lock_.compare_exchange_weak(
expected,
std::numeric_limits<T>::max(),
std::memory_order_acquire,
std::memory_order_relaxed);
return (bits_.fetch_or(mask_, std::memory_order::acquire) & mask_) == 0;
}
void
lock()
lock() noexcept
{
while (!try_lock())
do
{
// The use of relaxed memory ordering here is intentional and
// serves to help reduce cache coherency traffic during times
// of contention by avoiding writes that would definitely not
// result in the lock being acquired.
while (lock_.load(std::memory_order_relaxed) != 0)
// of contention by avoiding writes that are unlikely to grab
// the requested lock.
while ((bits_.load(std::memory_order::relaxed) & mask_) != 0)
detail::spinPause();
}
} while (!try_lock());
}
void
unlock()
unlock() noexcept
{
lock_.store(0, std::memory_order_release);
bits_.fetch_and(~mask_, std::memory_order::release);
}
};
/** @} */
} // namespace xrpl

View File

@@ -1,47 +0,0 @@
#pragma once
#include <cstdlib>
#include <string>
#include <type_traits>
#include <typeinfo>
#ifndef _MSC_VER
#include <cxxabi.h>
#endif
namespace beast {
template <typename T>
std::string
typeName()
{
using TR = std::remove_reference_t<T>;
std::string name = typeid(TR).name();
#ifndef _MSC_VER
if (auto s = abi::__cxa_demangle(name.c_str(), nullptr, nullptr, nullptr))
{
name = s;
// NOLINTNEXTLINE(cppcoreguidelines-no-malloc)
std::free(s);
}
#endif
if (std::is_const_v<TR>)
name += " const";
if (std::is_volatile_v<TR>)
name += " volatile";
if (std::is_lvalue_reference_v<T>)
{
name += "&";
}
else if (std::is_rvalue_reference_v<T>)
{
name += "&&";
}
return name;
}
} // namespace beast

View File

@@ -2,26 +2,24 @@
#pragma once
#include <compare>
#include <concepts>
namespace beast {
/**
* Zero allows classes to offer efficient comparisons to zero.
*
* Zero is a struct to allow classes to efficiently compare with zero without
* requiring an rvalue construction.
*
* It's often the case that we have classes which combine a number and a unit.
* In such cases, comparisons like t > 0 or t != 0 make sense, but comparisons
* like t > 1 or t != 1 do not.
* like t > 1 or t != 1 do not. Comparing against kZero expresses exactly that,
* without constructing a T.
*
* The class Zero allows such comparisons to be easily made.
*
* The comparing class T either needs to have a method called signum() which
* returns a positive number, 0, or a negative; or there needs to be a signum
* function which resolves in the namespace which takes an instance of T and
* returns a positive, zero or negative number.
* A type T participates if either `t.signum()` or an unqualified `signum(t)`
* found by argument-dependent lookup returns an integer that is negative,
* zero, or positive according to the sign of t. Both `t == kZero` and
* `kZero == t` work, as do all six relational operators in either order.
*/
struct Zero
{
explicit Zero() = default;
@@ -30,115 +28,44 @@ struct Zero
inline constexpr Zero kZero{};
/**
* Default implementation of signum calls the method on the class.
* Default implementation of signum: call the member function.
*/
template <typename T>
auto
signum(T const& t)
template <class T>
requires requires(T const& t) {
{ t.signum() } -> std::integral;
}
[[nodiscard]] constexpr auto
signum(T const& t) noexcept(noexcept(t.signum()))
{
return t.signum();
}
namespace detail::zero_helper {
namespace detail {
// For argument dependent lookup to function properly, calls to signum must
// be made from a namespace that does not include overloads of the function..
/**
* A type with a usable signum: either the member-based default above, or a
* `signum(t)` overload in T's own namespace, found by ADL. A user overload
* that is a better match than the template wins, as usual.
*/
template <class T>
auto
callSignum(T const& t)
concept HasSignum = requires(T const& t) {
{ signum(t) } -> std::integral;
};
} // namespace detail
template <detail::HasSignum T>
[[nodiscard]] constexpr bool
operator==(T const& t, Zero) noexcept(noexcept(signum(t)))
{
return signum(t);
return signum(t) == 0;
}
} // namespace detail::zero_helper
// Handle operators where T is on the left side using signum.
template <typename T>
bool
operator==(T const& t, Zero)
template <detail::HasSignum T>
[[nodiscard]] constexpr std::strong_ordering
operator<=>(T const& t, Zero) noexcept(noexcept(signum(t)))
{
return detail::zero_helper::callSignum(t) == 0;
}
template <typename T>
bool
operator!=(T const& t, Zero)
{
return detail::zero_helper::callSignum(t) != 0;
}
template <typename T>
bool
operator<(T const& t, Zero)
{
return detail::zero_helper::callSignum(t) < 0;
}
template <typename T>
bool
operator>(T const& t, Zero)
{
return detail::zero_helper::callSignum(t) > 0;
}
template <typename T>
bool
operator>=(T const& t, Zero)
{
return detail::zero_helper::callSignum(t) >= 0;
}
template <typename T>
bool
operator<=(T const& t, Zero)
{
return detail::zero_helper::callSignum(t) <= 0;
}
// Handle operators where T is on the right side by
// reversing the operation, so that T is on the left side.
template <typename T>
bool
operator==(Zero, T const& t)
{
return t == kZero;
}
template <typename T>
bool
operator!=(Zero, T const& t)
{
return t != kZero;
}
template <typename T>
bool
operator<(Zero, T const& t)
{
return t > kZero;
}
template <typename T>
bool
operator>(Zero, T const& t)
{
return t < kZero;
}
template <typename T>
bool
operator>=(Zero, T const& t)
{
return t <= kZero;
}
template <typename T>
bool
operator<=(Zero, T const& t)
{
return t >= kZero;
return signum(t) <=> 0;
}
} // namespace beast

View File

@@ -63,7 +63,6 @@ struct Sections
static constexpr auto kSslVerifyDir = "ssl_verify_dir";
static constexpr auto kSslVerifyFile = "ssl_verify_file";
static constexpr auto kSweepInterval = "sweep_interval";
static constexpr auto kTelemetry = "telemetry";
static constexpr auto kTransactionQueue = "transaction_queue";
static constexpr auto kValidationSeed = "validation_seed";
static constexpr auto kValidatorKeys = "validator_keys";

View File

@@ -8,13 +8,10 @@
#include <xrpl/beast/utility/instrumentation.h>
#include <xrpl/consensus/ConsensusParms.h>
#include <xrpl/consensus/ConsensusProposal.h>
#include <xrpl/consensus/ConsensusSpanLabels.h>
#include <xrpl/consensus/ConsensusSpanNames.h>
#include <xrpl/consensus/ConsensusTypes.h>
#include <xrpl/json/json_value.h>
#include <xrpl/json/json_writer.h>
#include <xrpl/ledger/LedgerTiming.h>
#include <xrpl/telemetry/SpanGuard.h>
#include <algorithm>
#include <chrono>
@@ -27,42 +24,16 @@
#include <ranges>
#include <sstream>
#include <string>
#include <string_view>
#include <utility>
namespace xrpl {
/**
* Determines why the current ledger should close at this time.
*
* Holds the close decision. shouldCloseLedger() delegates here and adds only
* a comparison, so the logging happens once either way; call whichever suits.
* Parameters match shouldCloseLedger().
*
* @return The deciding branch, or KeepOpen if no close condition is met.
*/
LedgerCloseReason
whyCloseLedger(
bool anyTransactions,
std::size_t prevProposers,
std::size_t proposersClosed,
std::size_t proposersValidated,
std::chrono::milliseconds prevRoundTime,
std::chrono::milliseconds timeSincePrevClose,
std::chrono::milliseconds openTime,
std::chrono::milliseconds idleInterval,
ConsensusParms const& parms,
beast::Journal j,
std::unique_ptr<std::stringstream> const& clog = {});
/**
* Determines whether the current ledger should close at this time.
*
* This function should be called when a ledger is open and there is no close
* in progress, or when a transaction is received and no close is in progress.
*
* Equivalent to `whyCloseLedger(...) != LedgerCloseReason::KeepOpen`.
*
* @param anyTransactions indicates whether any transactions have been received
* @param prevProposers proposers in the last closing
* @param proposersClosed proposers who have currently closed this ledger
@@ -476,29 +447,13 @@ public:
getJson(bool full) const;
private:
/**
* Why startRoundInternal is being entered.
*
* Distinguishes the normal Initial entry (from public startRound)
* from a Recovered re-entry (from handleWrongLedger after the
* correct prior ledger was acquired mid-round). The Recovered path
* resets phase to Open within the SAME round, which is a state
* transition that would otherwise be invisible in traces — the
* recovery flag drives a `phase.recovery` event on the round span.
*/
enum class StartRoundReason : std::uint8_t {
Initial,
Recovered,
};
void
startRoundInternal(
NetClock::time_point const& now,
LedgerT::ID const& prevLedgerID,
LedgerT const& prevLedger,
ConsensusMode mode,
std::unique_ptr<std::stringstream> const& clog,
StartRoundReason reason = StartRoundReason::Initial);
std::unique_ptr<std::stringstream> const& clog);
// Change our view of the previous ledger
void
@@ -672,67 +627,6 @@ private:
// nodes that have bowed out of this consensus process
HashSet<NodeIDT> deadNodes_;
/**
* Span for the establish phase of consensus.
* Created when the ledger closes and we enter phaseEstablish;
* cleared (ended) when consensus is reached. A thread-free SpanGuard,
* emplaced and reset() on different job workers.
*/
std::optional<xrpl::telemetry::SpanGuard> establishSpan_;
/**
* Captured context of establishSpan_ (its own span context). Children
* (update_positions, check) build from this explicit context.
*/
xrpl::telemetry::SpanContext establishSpanContext_;
/**
* Span for the open phase of consensus.
* Created in startRoundInternal(); cleared (ended) in closeLedger().
* A thread-free SpanGuard, emplaced and reset() on different job workers.
*/
std::optional<xrpl::telemetry::SpanGuard> openSpan_;
/**
* Record how the open-phase span began.
*
* @param reason Which startRoundInternal() entry path created it.
* @param prevLedger The prior ledger, read before previousLedger_ is set.
*/
void
annotateOpenStart(StartRoundReason reason, LedgerT const& prevLedger);
/**
* Record what ended the open phase.
*
* @param closeReason The deciding whyCloseLedger() branch.
* @param proposersValidated Trusted peers already past the prior ledger.
*/
void
annotateOpenClose(LedgerCloseReason closeReason, std::size_t proposersValidated);
/**
* Create the establish-phase span if not yet active.
* Called on each phaseEstablish() invocation; no-op while span is live.
*/
void
startEstablishTracing();
/**
* Overwrite convergence metrics on the establish span each iteration.
* Final span attributes always reflect the last state before consensus.
*/
void
updateEstablishTracing();
/**
* End the establish span, recording its terminal regime.
* Also called from startRoundInternal() on a wrongLedger recovery, so a
* round that never reaches Accepted still reports the regime it reached.
*/
void
endEstablishTracing();
// Journal for debugging
beast::Journal const j_;
};
@@ -796,52 +690,13 @@ Consensus<Adaptor>::startRoundInternal(
LedgerT::ID const& prevLedgerID,
LedgerT const& prevLedger,
ConsensusMode mode,
std::unique_ptr<std::stringstream> const& clog,
StartRoundReason const reason)
std::unique_ptr<std::stringstream> const& clog)
{
// Recovery path: handleWrongLedger acquired the correct prior ledger
// and re-entered startRoundInternal mid-round. The roundSpan_ owned by
// the adaptor is still the SAME span as before — startRoundTracing is
// not called on the recovery path, so we record the recovery as an
// event on the surviving round span. Pass empty phaseLabel to leave
// consensus_phase unchanged (the actual phase reset to open is marked
// separately by the new openSpan_ being emplaced below).
if (reason == StartRoundReason::Recovered)
{
adaptor_.onPhaseEvent(telemetry::consensus::span::event::phaseRecovery, "");
}
phase_ = ConsensusPhase::Open;
JLOG(j_.debug()) << "transitioned to ConsensusPhase::Open ";
CLOG(clog) << "startRoundInternal transitioned to ConsensusPhase::Open, "
"previous ledgerID: "
<< prevLedgerID << ", seq: " << prevLedger.seq() << ". ";
// End establishSpan_ so a wrongLedger recovery mid-establish doesn't leak
// the prior round's span into the new one (startEstablishTracing
// early-returns when establishSpan_ is populated). Via
// endEstablishTracing() so the recovered round still records its terminal
// regime; closeTimeAvalancheState_ is not reset until further down.
endEstablishTracing();
// Child of the round span via its captured context: parent phase.open
// explicitly under roundSpanContext_. An invalid round context (round span
// not yet created) yields a null guard. openSpan_ is a thread-free
// SpanGuard emplaced here on one job worker and reset() on another; no
// scope to strip.
openSpan_.emplace(
telemetry::SpanGuard::childSpan(
telemetry::consensus::span::phaseOpen, adaptor_.roundSpanContext()));
annotateOpenStart(reason, prevLedger);
// On the Recovered path, fire phase.open here because startRoundTracing
// (which fires it for the Initial path) is not called on re-entry. On
// the Initial path this is a no-op because the round span hasn't been
// created yet — the phase.open event is fired later by startRoundTracing
// after the new round span is in place.
if (reason == StartRoundReason::Recovered)
{
adaptor_.onPhaseEvent(
telemetry::consensus::span::event::phaseOpen,
telemetry::consensus::span::val::phaseOpen);
}
mode_.set(mode, adaptor_);
now_ = now;
prevLedgerID_ = prevLedgerID;
@@ -865,21 +720,10 @@ Consensus<Adaptor>::startRoundInternal(
playbackProposals();
CLOG(clog) << "number of peer proposals,previous proposers: " << currPeerPositions_.size()
<< ',' << prevProposers_ << ". ";
// We may be falling behind, don't wait for the timer
// consider closing the ledger immediately
bool const closeImmediately = currPeerPositions_.size() > (prevProposers_ / 2);
// Annotate before the timerEntry() below, which can end this span.
if (openSpan_ && *openSpan_)
{
namespace cs = telemetry::consensus::span;
// Head start after playbackProposals() replayed the buffered
// positions. Pairs with peer_positions_at_close.
openSpan_->setAttribute(
cs::attr::peerPositionsAtOpen, static_cast<int64_t>(currPeerPositions_.size()));
openSpan_->setAttribute(cs::attr::earlyCloseTriggered, closeImmediately);
}
if (closeImmediately)
if (currPeerPositions_.size() > (prevProposers_ / 2))
{
// We may be falling behind, don't wait for the timer
// consider closing the ledger immediately
CLOG(clog) << "consider closing the ledger immediately. ";
timerEntry(now_, clog);
}
@@ -1103,9 +947,6 @@ Consensus<Adaptor>::simulate(
result_->proposers = prevProposers_ = currPeerPositions_.size();
prevRoundTime_ = result_->roundTime.read();
phase_ = ConsensusPhase::Accepted;
adaptor_.onPhaseEvent(
telemetry::consensus::span::event::phaseAccepted,
telemetry::consensus::span::val::phaseAccepted);
adaptor_.onForceAccept(
*result_, previousLedger_, closeResolution_, rawCloseTimes_, mode_.get(), getJson(true));
// NOLINTEND(bugprone-unchecked-optional-access)
@@ -1254,13 +1095,7 @@ Consensus<Adaptor>::handleWrongLedger(
{
JLOG(j_.info()) << "Have the consensus ledger " << prevLedgerID_;
CLOG(clog) << "Have the consensus ledger " << prevLedgerID_ << ". ";
startRoundInternal(
now_,
lgrId,
*newLedger,
ConsensusMode::SwitchedLedger,
clog,
StartRoundReason::Recovered);
startRoundInternal(now_, lgrId, *newLedger, ConsensusMode::SwitchedLedger, clog);
}
else
{
@@ -1368,23 +1203,20 @@ Consensus<Adaptor>::phaseOpen(std::unique_ptr<std::stringstream> const& clog)
<< ", previous ledger close time resolution: "
<< previousLedger_.closeTimeResolution().count() << "ms. ";
// Decide if we should close the ledger. whyCloseLedger() so the deciding
// branch can be recorded; it logs the same, so only one is called.
LedgerCloseReason const closeReason = whyCloseLedger(
anyTransactions,
prevProposers_,
proposersClosed,
proposersValidated,
prevRoundTime_,
sinceClose,
openTime_.read(),
idleInterval,
adaptor_.parms(),
j_,
clog);
if (closeReason != LedgerCloseReason::KeepOpen)
// Decide if we should close the ledger
if (shouldCloseLedger(
anyTransactions,
prevProposers_,
proposersClosed,
proposersValidated,
prevRoundTime_,
sinceClose,
openTime_.read(),
idleInterval,
adaptor_.parms(),
j_,
clog))
{
annotateOpenClose(closeReason, proposersValidated);
CLOG(clog) << "closing ledger. ";
closeLedger(clog);
}
@@ -1524,8 +1356,6 @@ Consensus<Adaptor>::phaseEstablish(std::unique_ptr<std::stringstream> const& clo
XRPL_ASSERT(result_, "xrpl::Consensus::phaseEstablish : result is set");
// NOLINTBEGIN(bugprone-unchecked-optional-access) assert above
startEstablishTracing();
++peerUnchangedCounter_;
++establishCounter_;
@@ -1553,8 +1383,6 @@ Consensus<Adaptor>::phaseEstablish(std::unique_ptr<std::stringstream> const& clo
updateOurPositions(clog);
updateEstablishTracing();
// Nothing to do if too many laggards or we don't have consensus.
if (shouldPause(clog) || !haveConsensus(clog))
return;
@@ -1572,26 +1400,7 @@ Consensus<Adaptor>::phaseEstablish(std::unique_ptr<std::stringstream> const& clo
adaptor_.updateOperatingMode(currPeerPositions_.size());
prevProposers_ = currPeerPositions_.size();
prevRoundTime_ = result_->roundTime.read();
endEstablishTracing();
{
namespace cs = telemetry::consensus::span;
if (result_->state == ConsensusState::Yes)
{
adaptor_.onOutcomeEvent(cs::event::outcomeYes);
}
else if (result_->state == ConsensusState::MovedOn)
{
adaptor_.onOutcomeEvent(cs::event::outcomeMovedOn);
}
else if (result_->state == ConsensusState::Expired)
{
adaptor_.onOutcomeEvent(cs::event::outcomeExpired);
}
}
phase_ = ConsensusPhase::Accepted;
adaptor_.onPhaseEvent(
telemetry::consensus::span::event::phaseAccepted,
telemetry::consensus::span::val::phaseAccepted);
JLOG(j_.debug()) << "transitioned to ConsensusPhase::Accepted";
adaptor_.onAccept(
*result_,
@@ -1611,26 +1420,7 @@ Consensus<Adaptor>::closeLedger(std::unique_ptr<std::stringstream> const& clog)
// We should not be closing if we already have a position
XRPL_ASSERT(!result_, "xrpl::Consensus::closeLedger : result is not set");
// Annotate the open-phase span with end-of-phase metadata before
// ending it, so a Tempo/Jaeger query for consensus.phase.open shows
// how long the phase ran and how many peer positions arrived during it.
openTime_.tick(clock_.now());
if (openSpan_ && *openSpan_)
{
namespace cs = telemetry::consensus::span;
openSpan_->setAttribute(
cs::attr::openDurationMs, static_cast<int64_t>(openTime_.read().count()));
openSpan_->setAttribute(
cs::attr::peerPositionsAtClose, static_cast<int64_t>(currPeerPositions_.size()));
// Read before our own position is added below, so this counts only
// what peers shared.
openSpan_->setAttribute(cs::attr::txSetsAcquired, static_cast<int64_t>(acquired_.size()));
}
openSpan_.reset();
phase_ = ConsensusPhase::Establish;
adaptor_.onPhaseEvent(
telemetry::consensus::span::event::phaseEstablish,
telemetry::consensus::span::val::phaseEstablish);
JLOG(j_.debug()) << "transitioned to ConsensusPhase::Establish";
rawCloseTimes_.self = now_;
peerUnchangedCounter_ = 0;
@@ -1687,18 +1477,6 @@ Consensus<Adaptor>::updateOurPositions(std::unique_ptr<std::stringstream> const&
// We must have a position if we are updating it
XRPL_ASSERT(result_, "xrpl::Consensus::updateOurPositions : result is set");
// NOLINTBEGIN(bugprone-unchecked-optional-access) assert above
using namespace telemetry;
// Child of the establish span via its captured context (establishSpan_ is
// a thread-free SpanGuard, so parent explicitly via its context). A null
// context — the establish phase has not started — yields a null guard, so
// the setAttribute calls below are no-ops.
auto span = SpanGuard::childSpan(consensus::span::updatePositions, establishSpanContext_);
span.setAttribute(
consensus::span::attr::convergePercent, static_cast<int64_t>(convergePercent_));
span.setAttribute(
consensus::span::attr::proposers, static_cast<int64_t>(currPeerPositions_.size()));
span.setAttribute(
consensus::span::attr::disputesCount, static_cast<int64_t>(result_->disputes.size()));
ConsensusParms const& parms = adaptor_.parms();
// Compute a cutoff time
@@ -1758,24 +1536,6 @@ Consensus<Adaptor>::updateOurPositions(std::unique_ptr<std::stringstream> const&
// now a no
mutableSet->erase(txId);
}
// The event exists only for the span, so it is guarded on the
// span being active. Unguarded, every dispute that flips
// position builds a 64-character tx hash plus two number
// strings, on every establish tick.
if (span)
{
auto const yaysStr = std::to_string(dispute.getYays());
auto const naysStr = std::to_string(dispute.getNays());
span.addEvent(
consensus::span::event::disputeResolve,
{{consensus::span::attr::txId, to_string(txId)},
{consensus::span::attr::disputeOurVote,
dispute.getOurVote() ? std::string_view{consensus::span::val::yes}
: std::string_view{consensus::span::val::no}},
{consensus::span::attr::disputeYays, yaysStr},
{consensus::span::attr::disputeNays, naysStr}});
}
}
}
@@ -1800,8 +1560,6 @@ Consensus<Adaptor>::updateOurPositions(std::unique_ptr<std::stringstream> const&
if (newState)
closeTimeAvalancheState_ = *newState;
CLOG(clog) << "neededWeight " << neededWeight << ". ";
span.setAttribute(
consensus::span::attr::avalancheThreshold, static_cast<int64_t>(neededWeight));
int participants = currPeerPositions_.size();
if (mode_.get() == ConsensusMode::Proposing)
@@ -1856,10 +1614,6 @@ Consensus<Adaptor>::updateOurPositions(std::unique_ptr<std::stringstream> const&
}
}
span.setAttribute(consensus::span::attr::haveCloseTimeConsensus, haveCloseTimeConsensus_);
span.setAttribute(
consensus::span::attr::closeTimeThreshold, static_cast<int64_t>(parms.avCtConsensusPct));
if (!ourNewSet &&
((consensusCloseTime != asCloseTime(result_->position.closeTime())) ||
result_->position.isStale(ourCutoff)))
@@ -1911,10 +1665,6 @@ Consensus<Adaptor>::haveConsensus(std::unique_ptr<std::stringstream> const& clog
// Must have a stance if we are checking for consensus
XRPL_ASSERT(result_, "xrpl::Consensus::haveConsensus : has result");
// NOLINTBEGIN(bugprone-unchecked-optional-access) assert above
using namespace telemetry;
// Child of the establish span via its captured context (establishSpan_ is
// a thread-free SpanGuard, so parent explicitly via its context).
auto span = SpanGuard::childSpan(consensus::span::check, establishSpanContext_);
// CHECKME: should possibly count unacquired TX sets as disagreeing
int agree = 0, disagree = 0;
@@ -1973,38 +1723,6 @@ Consensus<Adaptor>::haveConsensus(std::unique_ptr<std::stringstream> const& clog
j_,
clog);
// Set span attributes before the early-return branches below so the
// consensus.check span carries diagnostic data even when consensus is
// not reached (the No / Expired paths return early).
span.setAttribute(consensus::span::attr::agreeCount, static_cast<int64_t>(agree));
span.setAttribute(consensus::span::attr::disagreeCount, static_cast<int64_t>(disagree));
span.setAttribute(
consensus::span::attr::convergePercent, static_cast<int64_t>(convergePercent_));
span.setAttribute(consensus::span::attr::haveCloseTimeConsensus, haveCloseTimeConsensus_);
span.setAttribute(
consensus::span::attr::thresholdPercent,
static_cast<int64_t>(adaptor_.parms().avCtConsensusPct));
span.setAttribute(
consensus::span::attr::proposersFinished, static_cast<int64_t>(currentFinished));
span.setAttribute(consensus::span::attr::consensusStalled, stalled);
span.setAttribute(
consensus::span::attr::establishCount, static_cast<int64_t>(establishCounter_));
std::string_view stateStr = consensus::span::val::no;
if (result_->state == ConsensusState::Yes)
{
stateStr = consensus::span::val::yes;
}
else if (result_->state == ConsensusState::MovedOn)
{
stateStr = consensus::span::val::movedOn;
}
else if (result_->state == ConsensusState::Expired)
{
stateStr = consensus::span::val::expired;
}
span.setAttribute(consensus::span::attr::consensusResult, stateStr);
if (result_->state == ConsensusState::No)
{
CLOG(clog) << "No consensus. ";
@@ -2167,89 +1885,4 @@ Consensus<Adaptor>::asCloseTime(NetClock::time_point raw) const
return roundCloseTime(raw, closeResolution_);
}
template <class Adaptor>
void
Consensus<Adaptor>::annotateOpenStart(StartRoundReason const reason, LedgerT const& prevLedger)
{
if (!openSpan_ || !*openSpan_)
return;
namespace cs = telemetry::consensus::span;
openSpan_->setAttribute(
cs::attr::startReason,
reason == StartRoundReason::Recovered ? std::string_view{cs::val::startRecovered}
: std::string_view{cs::val::startInitial});
// From the parameter: previousLedger_ is not assigned until later.
openSpan_->setAttribute(cs::attr::previousCloseAgree, prevLedger.closeAgree());
}
template <class Adaptor>
void
Consensus<Adaptor>::annotateOpenClose(
LedgerCloseReason const closeReason,
std::size_t const proposersValidated)
{
// Called before closeLedger() ends the span, and only on the closing tick,
// so each attribute is written once per round.
if (!openSpan_ || !*openSpan_)
return;
namespace cs = telemetry::consensus::span;
openSpan_->setAttribute(cs::attr::closeReason, cs::closeReasonLabel(closeReason));
openSpan_->setAttribute(cs::attr::proposersValidated, static_cast<int64_t>(proposersValidated));
}
template <class Adaptor>
void
Consensus<Adaptor>::startEstablishTracing()
{
if (establishSpan_)
return;
// Child of the round span via its captured context: parent establish
// explicitly under roundSpanContext_. An invalid round context (round span
// not yet created) yields a null guard.
establishSpan_.emplace(
telemetry::SpanGuard::childSpan(
telemetry::consensus::span::establish, adaptor_.roundSpanContext()));
// Capture the establish context; children (update_positions, check) parent
// to it explicitly via establishSpanContext_. establishSpan_ is a
// thread-free SpanGuard reset() on a different worker than it is emplaced
// on -- no scope work.
if (*establishSpan_)
{
establishSpanContext_ = establishSpan_->spanContext();
}
}
template <class Adaptor>
void
Consensus<Adaptor>::updateEstablishTracing()
{
if (!establishSpan_)
return;
namespace cs = telemetry::consensus::span;
establishSpan_->setAttribute(cs::attr::convergePercent, static_cast<int64_t>(convergePercent_));
establishSpan_->setAttribute(cs::attr::establishCount, static_cast<int64_t>(establishCounter_));
establishSpan_->setAttribute(
cs::attr::proposers, static_cast<int64_t>(currPeerPositions_.size()));
if (result_)
{
establishSpan_->setAttribute(
cs::attr::disputesCount, static_cast<int64_t>(result_->disputes.size()));
}
}
template <class Adaptor>
void
Consensus<Adaptor>::endEstablishTracing()
{
// Terminal convergence regime, recorded once before the span ends.
if (establishSpan_ && *establishSpan_)
{
namespace cs = telemetry::consensus::span;
establishSpan_->setAttribute(
cs::attr::closeTimeAvalancheState, cs::avalancheStateLabel(closeTimeAvalancheState_));
}
establishSpan_.reset();
establishSpanContext_ = telemetry::SpanContext{};
}
} // namespace xrpl

View File

@@ -1,86 +0,0 @@
#pragma once
/**
* Enum-to-label mappings for consensus span attribute values.
*
* These mappings live in their own header so ConsensusSpanNames.h stays
* dependency-free like its siblings: the span-name and attribute-key
* constants are included by overlay and app translation units that have no
* use for the consensus enums, while these mappings are needed only by
* Consensus.h.
*
* ConsensusSpanNames.h (constants only, no domain deps)
* ^
* | includes
* ConsensusSpanLabels.h --includes--> ConsensusParms.h, ConsensusTypes.h
* ^
* | includes
* Consensus.h
*/
#include <xrpl/consensus/ConsensusParms.h>
#include <xrpl/consensus/ConsensusSpanNames.h>
#include <xrpl/consensus/ConsensusTypes.h>
#include <string_view>
namespace xrpl::telemetry::consensus::span {
/**
* Map a close-time avalanche state to its `avalanche_state` label.
*
* The regime escalates Init -> Mid -> Late -> Stuck, raising the close-time
* agreement threshold at each step.
*
* @param state The state held by Consensus::closeTimeAvalancheState_.
* @return The wire label; one of val::avalanche*.
*
* @note No default arm, so a new enumerator is a -Wswitch warning; the
* fall-through returns "unknown" rather than a plausible-looking regime.
*/
[[nodiscard]] constexpr std::string_view
avalancheStateLabel(ConsensusParms::AvalancheState const state)
{
switch (state)
{
case ConsensusParms::AvalancheState::Init:
return val::avalancheInit;
case ConsensusParms::AvalancheState::Mid:
return val::avalancheMid;
case ConsensusParms::AvalancheState::Late:
return val::avalancheLate;
case ConsensusParms::AvalancheState::Stuck:
return val::avalancheStuck;
}
return val::unknown;
}
/**
* Map a ledger-close decision to its `close_reason` label.
*
* @param reason The value returned by whyCloseLedger().
* @return The wire label; one of val::close*.
*
* @note No default arm, so a new enumerator is a -Wswitch warning; the
* fall-through returns "unknown". `keep_open` is mapped but never emitted.
*/
[[nodiscard]] constexpr std::string_view
closeReasonLabel(LedgerCloseReason const reason)
{
switch (reason)
{
case LedgerCloseReason::KeepOpen:
return val::closeKeepOpen;
case LedgerCloseReason::Anomaly:
return val::closeAnomaly;
case LedgerCloseReason::OthersClosed:
return val::closeOthersClosed;
case LedgerCloseReason::Idle:
return val::closeIdle;
case LedgerCloseReason::Normal:
return val::closeNormal;
}
return val::unknown;
}
} // namespace xrpl::telemetry::consensus::span

View File

@@ -1,391 +0,0 @@
#pragma once
/**
* Compile-time span name constants for consensus tracing.
*
* Used by RCLConsensus (app), Consensus.h (template), and PeerImp
* (overlay) for consensus lifecycle spans.
* Built on StaticStr/join() from SpanNames.h.
*
* ## Span Hierarchy
*
* Root span created in Adaptor::startRoundTracing(). In "deterministic"
* strategy the trace-id is derived from the previous ledger hash so all
* nodes tracing the same round share a trace.
*
* consensus.round [main thread, root]
* | Created: Adaptor::startRoundTracing()
* | Attrs: consensus_ledger_id, ledger_seq, trace_strategy,
* | consensus_round_id; consensus_mode from
* | Adaptor::onModeChange()
* |
* +-- consensus.phase.open [main thread, child]
* | Created: Consensus::startRoundInternal()
* | Ended: Consensus::closeLedger()
* | Attrs: start_reason, previous_close_agree, peer_positions_at_open,
* | early_close_triggered (at start); open_duration_ms,
* | peer_positions_at_close, tx_sets_acquired, close_reason,
* | proposers_validated (at close; absent if the round is
* | recovered or simulated, neither of which reaches
* | closeLedger())
* |
* +-- consensus.proposal.send [main thread]
* | Created: Adaptor::propose()
* | Attrs: consensus_round (proposeSeq)
* |
* +-- consensus.ledger_close [main thread]
* | Created: Adaptor::onClose()
* | Attrs: ledger_seq, consensus_mode
* |
* +-- consensus.establish [main thread, child]
* | Created: Consensus::startEstablishTracing()
* | Ended: Consensus::phaseEstablish() on accept
* | Attrs: converge_percent, establish_count, proposers,
* | disputes_count (overwritten each iteration);
* | close_time_avalanche_state (terminal, at end)
* |
* +-- consensus.update_positions [main thread]
* | Created: Consensus::updateOurPositions()
* | Attrs: converge_percent, proposers, disputes_count
* | Events: per-dispute vote details (tx_id, our_vote, yays, nays)
* |
* +-- consensus.check [main thread]
* | Created: Consensus::haveConsensus()
* | Attrs: agree/disagree counts, threshold_percent, result
* |
* +-- consensus.accept [main thread, child of round]
* | Created: Adaptor::makeAcceptSpan(), shared_ptr kept alive
* | until doAccept() completes on jtACCEPT thread
* | Attrs: proposers, round_time_ms, quorum
* | |
* | +-- consensus.accept.apply [jtACCEPT thread, child of accept]
* | Created: Adaptor::doAccept(), scoped: the txq spans doAccept
* | goes on to create nest under it; the tx apply-stage
* | spans are hash-derived roots and do not
* | Attrs: ledger_seq, close_time_ripple_epoch_s, close_time_correct,
* | close_resolution_ms, consensus_state, proposing, round_time_ms,
* | parent_close_time_ripple_epoch_s, close_time_self_ripple_epoch_s,
* | close_time_vote_bins, resolution_direction, tx_count
* | Events: tx.included (per tx, attrs: tx_id)
* |
* +~~~ consensus.validation.send [jtACCEPT thread, linked]
* | Created: Adaptor::createValidationSpan() (follows-from link)
* | Attrs: ledger_seq, proposing
* |
* +-- consensus.mode_change [main thread]
* Created: Adaptor::onModeChange(), only when the mode moves
* Attrs: mode_old, mode_new
*
* Standalone spans (no parent, created per-message in overlay):
*
* consensus.proposal.receive [PeerImp I/O thread]
* Created: PeerImp::onMessage(TMProposeSet)
*
* consensus.validation.receive [PeerImp I/O thread]
* Created: PeerImp::onMessage(TMValidation)
*
* Legend:
* +-- child-of relationship (same trace)
* +~~~ follows-from link (separate sub-tree, causal link)
*/
#include <xrpl/telemetry/SpanNames.h>
namespace xrpl::telemetry::consensus::span {
// ===== Span name segments ====================================================
namespace part {
inline constexpr auto proposal = makeStr("proposal");
inline constexpr auto validation = makeStr("validation");
inline constexpr auto accept = makeStr("accept");
inline constexpr auto phase = makeStr("phase");
} // namespace part
namespace op {
inline constexpr auto round = makeStr("round");
inline constexpr auto proposalSend = join(part::proposal, makeStr("send"));
inline constexpr auto ledgerClose = makeStr("ledger_close");
inline constexpr auto establish = makeStr("establish");
inline constexpr auto updatePositions = makeStr("update_positions");
inline constexpr auto check = makeStr("check");
inline constexpr auto accept = makeStr("accept");
inline constexpr auto acceptApply = join(part::accept, makeStr("apply"));
inline constexpr auto validationSend = join(part::validation, makeStr("send"));
inline constexpr auto modeChange = makeStr("mode_change");
inline constexpr auto proposalReceive = join(part::proposal, makeStr("receive"));
inline constexpr auto validationReceive = join(part::validation, makeStr("receive"));
inline constexpr auto phaseOpen = join(part::phase, makeStr("open"));
} // namespace op
// ===== Full span names (prefix.op) ===========================================
inline constexpr auto round = join(seg::consensus, op::round);
inline constexpr auto proposalSend = join(seg::consensus, op::proposalSend);
inline constexpr auto ledgerClose = join(seg::consensus, op::ledgerClose);
inline constexpr auto establish = join(seg::consensus, op::establish);
inline constexpr auto updatePositions = join(seg::consensus, op::updatePositions);
inline constexpr auto check = join(seg::consensus, op::check);
inline constexpr auto accept = join(seg::consensus, op::accept);
inline constexpr auto acceptApply = join(seg::consensus, op::acceptApply);
inline constexpr auto validationSend = join(seg::consensus, op::validationSend);
inline constexpr auto modeChange = join(seg::consensus, op::modeChange);
inline constexpr auto proposalReceive = join(seg::consensus, op::proposalReceive);
inline constexpr auto validationReceive = join(seg::consensus, op::validationReceive);
inline constexpr auto phaseOpen = join(seg::consensus, op::phaseOpen);
// ===== Attribute keys ========================================================
namespace attr {
/**
* Canonical shared constants (defined in SpanNames.h). `ledgerHash` and
* `fullValidation` are shared with the peer.validation.receive span — same
* concept, same key, distinguished by span name (not an emitter prefix).
*/
using ::xrpl::telemetry::attr::closeResolutionMs;
using ::xrpl::telemetry::attr::closeTimeCorrect;
using ::xrpl::telemetry::attr::closeTimeRippleEpochS;
using ::xrpl::telemetry::attr::fullValidation;
using ::xrpl::telemetry::attr::ledgerHash;
using ::xrpl::telemetry::attr::ledgerSeq;
/**
* Domain-qualified attrs (rule 5 — bare name ambiguous across domains).
* Use `<domain>_<field>` underscore form for TraceQL ergonomics.
*/
inline constexpr auto ledgerId = makeStr("consensus_ledger_id");
/**
* Consensus mode. On consensus.round it is written by onModeChange, the point
* at which the engine applies the mode; on consensus.ledger_close the engine
* passes the mode in.
*/
inline constexpr auto mode = makeStr("consensus_mode");
inline constexpr auto round = makeStr("consensus_round");
inline constexpr auto roundId = makeStr("consensus_round_id");
/**
* Current phase name attached to consensus.round; updated on each
* phase transition event (open/establish/accepted).
*/
inline constexpr auto consensusPhase = makeStr("consensus_phase");
/**
* Boolean flag set on consensus.check when checkConsensus reports stalled.
*/
inline constexpr auto consensusStalled = makeStr("consensus_stalled");
/**
* Domain-owned bare attrs.
*/
inline constexpr auto proposers = makeStr("proposers");
inline constexpr auto roundTimeMs = makeStr("round_time_ms");
inline constexpr auto proposing = makeStr("proposing");
/**
* Round continuity / context attrs (set on consensus.round at round start).
*/
inline constexpr auto previousProposers = makeStr("previous_proposers");
inline constexpr auto previousRoundTimeMs = makeStr("previous_round_time_ms");
inline constexpr auto previousLedgerSeq = makeStr("previous_ledger_seq");
inline constexpr auto closeTimeResolutionMs = makeStr("close_time_resolution_ms");
/**
* Open-phase start metadata (set on consensus.phase.open at creation).
*
* A handleWrongLedger recovery emits a SECOND phase.open span under the same
* round, so `start_reason` is what tells the two apart.
*/
inline constexpr auto startReason = makeStr("start_reason");
inline constexpr auto previousCloseAgree = makeStr("previous_close_agree");
inline constexpr auto peerPositionsAtOpen = makeStr("peer_positions_at_open");
inline constexpr auto earlyCloseTriggered = makeStr("early_close_triggered");
/**
* Open-phase end metadata (set on consensus.phase.open before reset).
*
* `proposers_validated` counts validators of the previous ledger, unlike
* `proposers_finished` below, which counts those already past it.
*/
inline constexpr auto openDurationMs = makeStr("open_duration_ms");
inline constexpr auto peerPositionsAtClose = makeStr("peer_positions_at_close");
inline constexpr auto txSetsAcquired = makeStr("tx_sets_acquired");
inline constexpr auto closeReason = makeStr("close_reason");
inline constexpr auto proposersValidated = makeStr("proposers_validated");
/**
* Ledger-close inputs.
*/
inline constexpr auto txCountOpen = makeStr("tx_count_open");
/**
* Establish/check additional state.
*/
inline constexpr auto proposersFinished = makeStr("proposers_finished");
/**
* Establish-phase end metadata.
*
* The terminal close-time regime, set once. Qualified because DisputedTx
* tracks a SECOND, per-transaction avalanche; this is not that one. The
* derived `avalanche_threshold` cannot be inverted back to it.
*/
inline constexpr auto closeTimeAvalancheState = makeStr("close_time_avalanche_state");
/**
* Accept/apply enrichment.
*/
inline constexpr auto disputesResolvedCount = makeStr("disputes_resolved_count");
/**
* Validation send/receive enrichment. (`full_validation` is shared — see the
* `using` re-export above.)
*/
inline constexpr auto validationSignTime = makeStr("validation_sign_time");
/**
* Receive-side hash prefixes for cross-peer correlation.
*/
inline constexpr auto prevLedgerPrefix = makeStr("prev_ledger_prefix");
inline constexpr auto positionHashPrefix = makeStr("position_hash_prefix");
/**
* "consensus_state" — domain-qualified (collides with other domains' state).
*/
inline constexpr auto consensusState = makeStr("consensus_state");
/**
* Close-time instants, both NetClock readings in whole seconds since the XRP
* Ledger epoch (2000-01-01T00:00:00Z) — see `closeTimeRippleEpochS` in
* SpanNames.h for why the epoch is spelled into the key.
*
* `parentCloseTimeRippleEpochS` is the previous ledger's close time;
* `closeTimeSelfRippleEpochS` is this node's own close-time vote for the round,
* so the pair shows how far the node's position sat from the ledger it built on.
*
* `closeTimeVoteBins` is not a time: it holds the number of distinct close-time
* positions seen from peers this round.
*/
inline constexpr auto parentCloseTimeRippleEpochS = makeStr("parent_close_time_ripple_epoch_s");
inline constexpr auto closeTimeSelfRippleEpochS = makeStr("close_time_self_ripple_epoch_s");
inline constexpr auto closeTimeVoteBins = makeStr("close_time_vote_bins");
inline constexpr auto resolutionDirection = makeStr("resolution_direction");
inline constexpr auto convergePercent = makeStr("converge_percent");
inline constexpr auto establishCount = makeStr("establish_count");
inline constexpr auto avalancheThreshold = makeStr("avalanche_threshold");
inline constexpr auto closeTimeThreshold = makeStr("close_time_threshold");
inline constexpr auto haveCloseTimeConsensus = makeStr("have_close_time_consensus");
inline constexpr auto agreeCount = makeStr("agree_count");
inline constexpr auto disagreeCount = makeStr("disagree_count");
inline constexpr auto thresholdPercent = makeStr("threshold_percent");
/**
* "consensus_result" — domain-qualified (collides with generic result).
*/
inline constexpr auto consensusResult = makeStr("consensus_result");
inline constexpr auto quorum = makeStr("quorum");
inline constexpr auto traceStrategy = makeStr("trace_strategy");
inline constexpr auto modeOld = makeStr("mode_old");
inline constexpr auto modeNew = makeStr("mode_new");
/**
* "is_bow_out" — whether this proposal is a bow-out (resigning from round).
*/
inline constexpr auto isBowOut = makeStr("is_bow_out");
/**
* Transaction/dispute attrs used in consensus accept spans.
*/
inline constexpr auto txId = makeStr("tx_id");
inline constexpr auto disputeOurVote = makeStr("dispute_our_vote");
inline constexpr auto disputeYays = makeStr("dispute_yays");
inline constexpr auto disputeNays = makeStr("dispute_nays");
inline constexpr auto txCount = makeStr("tx_count");
inline constexpr auto disputesCount = makeStr("disputes_count");
/**
* Trust flag (is the message origin a trusted UNL validator). Qualified by
* message type, shared with the peer.{proposal,validation}.receive spans:
* consensus.proposal.receive uses `proposal_trusted`, consensus.validation.
* receive uses `validation_trusted`. Same concept on both emitters → same key.
*/
inline constexpr auto proposalTrusted = makeStr("proposal_trusted");
inline constexpr auto validationTrusted = makeStr("validation_trusted");
/**
* "validation_receive_status" — which exit the inbound validation took on
* consensus.validation.receive. Set once per exit, so a dropped validation
* (microseconds) is separable from a queued one (job wait plus
* checkValidation); without it the span reports two unrelated latency
* distributions and every quantile over it is meaningless.
*
* Deliberately NOT `validation_status`: that key belongs to
* consensus.validation.accept and carries what the validation store did
* (`ValStatus`). One key with two value domains would make any aggregation
* that does not also filter on span name meaningless.
*/
inline constexpr auto validationReceiveStatus = makeStr("validation_receive_status");
} // namespace attr
// ===== Event names ===========================================================
namespace event {
/**
* "dispute.resolve"
*/
inline constexpr auto disputeResolve = join(makeStr("dispute"), makeStr("resolve"));
/**
* "tx.included" — one per transaction of the agreed consensus set, recorded
* before the ledger is built. A transaction that then fails to apply still
* has an event, so this is a superset of the accepted ledger's contents.
*/
inline constexpr auto txIncluded = join(makeStr("tx"), makeStr("included"));
/**
* Phase transition events — fired on consensus.round at each transition
* so the round-level span carries a complete timeline of phase changes,
* including the handleWrongLedger recovery edge that re-enters Open.
*/
inline constexpr auto phaseOpen = join(makeStr("phase"), makeStr("open"));
inline constexpr auto phaseEstablish = join(makeStr("phase"), makeStr("establish"));
inline constexpr auto phaseAccepted = join(makeStr("phase"), makeStr("accepted"));
inline constexpr auto phaseRecovery = join(makeStr("phase"), makeStr("recovery"));
/**
* Outcome events — fired on consensus.round at the establish→accepted
* transition so the path that drove acceptance is queryable.
*/
inline constexpr auto outcomeYes = join(makeStr("outcome"), makeStr("yes"));
inline constexpr auto outcomeMovedOn = join(makeStr("outcome"), makeStr("moved_on"));
inline constexpr auto outcomeExpired = join(makeStr("outcome"), makeStr("expired"));
} // namespace event
// ===== Attribute values ======================================================
namespace val {
inline constexpr auto finished = makeStr("finished");
inline constexpr auto movedOn = makeStr("moved_on");
inline constexpr auto yes = makeStr("yes");
inline constexpr auto no = makeStr("no");
inline constexpr auto expired = makeStr("expired");
inline constexpr auto increased = makeStr("increased");
inline constexpr auto decreased = makeStr("decreased");
inline constexpr auto unchanged = makeStr("unchanged");
// consensus_phase attribute values (the phase the round is entering).
inline constexpr auto phaseOpen = makeStr("open");
inline constexpr auto phaseEstablish = makeStr("establish");
inline constexpr auto phaseAccepted = makeStr("accepted");
// start_reason values (how startRoundInternal was entered).
inline constexpr auto startInitial = makeStr("initial");
inline constexpr auto startRecovered = makeStr("recovered");
// close_time_avalanche_state values, one per AvalancheState enumerator.
inline constexpr auto avalancheInit = makeStr("init");
inline constexpr auto avalancheMid = makeStr("mid");
inline constexpr auto avalancheLate = makeStr("late");
inline constexpr auto avalancheStuck = makeStr("stuck");
// Sentinel for an unmapped enumerator, matching to_string(ConsensusPhase).
inline constexpr auto unknown = makeStr("unknown");
// close_reason values, one per LedgerCloseReason enumerator. keep_open is
// never emitted: the attribute is only set on the path that closes.
inline constexpr auto closeKeepOpen = makeStr("keep_open");
inline constexpr auto closeAnomaly = makeStr("anomaly");
inline constexpr auto closeOthersClosed = makeStr("others_closed");
inline constexpr auto closeIdle = makeStr("idle");
inline constexpr auto closeNormal = makeStr("normal");
// validation_receive_status values, one per exit of the receive path.
inline constexpr auto validationQueued = makeStr("queued");
inline constexpr auto validationDroppedDiverged = makeStr("dropped_diverged");
inline constexpr auto validationDroppedLoad = makeStr("dropped_load");
/**
* The validation was dropped because the job queue is stopping. The server
* stops its job queue while it shuts down. From then on JobQueue::addJob()
* declines new jobs. The check job never runs.
*/
inline constexpr auto validationDroppedQueueStopping = makeStr("dropped_queue_stopping");
} // namespace val
} // namespace xrpl::telemetry::consensus::span

View File

@@ -80,32 +80,10 @@ to_string(ConsensusMode m)
}
}
/**
* Title Case display name for telemetry attributes and dashboards.
* Separate from to_string() which is used in logs and must remain stable.
*/
inline std::string
toDisplayString(ConsensusMode m)
{
switch (m)
{
case ConsensusMode::Proposing:
return "Proposing";
case ConsensusMode::Observing:
return "Observing";
case ConsensusMode::WrongLedger:
return "Wrong Ledger";
case ConsensusMode::SwitchedLedger:
return "Switched Ledger";
default:
return "Unknown";
}
}
/**
* Phases of consensus for a single ledger round.
*
* @code
* @code
* "close" "accept"
* open ------- > establish ---------> accepted
* ^ | |
@@ -154,41 +132,6 @@ to_string(ConsensusPhase p)
}
}
/**
* Why the open ledger should, or should not, close right now.
*
* Returned by whyCloseLedger(); shouldCloseLedger() reduces it to a bool.
*
* @note KeepOpen is not a close reason. Compare against it rather than
* treating the enum as a flag.
*/
enum class LedgerCloseReason : std::uint8_t {
/**
* No close condition is met yet.
*/
KeepOpen,
/**
* Timings out of range; close defensively.
*/
Anomaly,
/**
* More than half the network has closed or validated.
*/
OthersClosed,
/**
* Nothing waiting and the idle interval elapsed.
*/
Idle,
/**
* Transactions waiting and both minimum-open floors met.
*/
Normal,
};
/**
* Measures the duration of phases of consensus
*/

View File

@@ -197,24 +197,6 @@ public:
[[nodiscard]] json::Value
getJson() const;
/**
* Number of peers voting yes.
*/
[[nodiscard]] int
getYays() const
{
return yays_;
}
/**
* Number of peers voting no.
*/
[[nodiscard]] int
getNays() const
{
return nays_;
}
private:
int yays_{0}; //< Number of yes votes
int nays_{0}; //< Number of no votes

View File

@@ -1,7 +1,7 @@
#pragma once
#include <xrpl/basics/base_uint.h>
#include <xrpl/basics/safe_cast.h>
#include <xrpl/basics/enum_bitops.h>
#include <xrpl/beast/utility/instrumentation.h>
#include <xrpl/ledger/OwnerCounts.h>
#include <xrpl/ledger/ReadView.h>
@@ -22,80 +22,39 @@
namespace xrpl {
// Bitwise flag enum with existing operator overloads
// NOLINTNEXTLINE(cppcoreguidelines-use-enum-class)
enum ApplyFlags : std::uint32_t {
TapNone = 0x00,
enum class ApplyFlags : std::uint32_t {
None = 0x00,
// This is a local transaction with the
// fail_hard flag set.
TapFailHard = 0x10,
FailHard = 0x10,
// This is not the transaction's last pass
// Transaction can be retried, soft failures allowed
TapRetry = 0x20,
Retry = 0x20,
// Transaction came from a privileged source
TapUnlimited = 0x400,
Unlimited = 0x400,
// Transaction is executing as part of a batch
TapBatch = 0x800,
Batch = 0x800,
// Transaction shouldn't be applied
// Signatures shouldn't be checked
TapDryRun = 0x1000,
DryRun = 0x1000,
// Transaction is being preflighted as the payload of a
// TransactionProposalCreate. Its signatures are collected on-ledger
// afterward, so signature-presence checks (e.g. Batch signer matching)
// are skipped at proposal-creation time (On-Chain Cosigner spec
// §5.3.1.2).
TapProposal = 0x2000
Proposal = 0x2000
};
constexpr ApplyFlags
operator|(ApplyFlags const& lhs, ApplyFlags const& rhs)
template <>
struct enum_bitops::OptIn<ApplyFlags> : std::true_type
{
return safeCast<ApplyFlags>(
safeCast<std::underlying_type_t<ApplyFlags>>(lhs) |
safeCast<std::underlying_type_t<ApplyFlags>>(rhs));
}
static_assert((TapFailHard | TapRetry) == safeCast<ApplyFlags>(0x30u), "ApplyFlags operator |");
static_assert((TapRetry | TapFailHard) == safeCast<ApplyFlags>(0x30u), "ApplyFlags operator |");
constexpr ApplyFlags
operator&(ApplyFlags const& lhs, ApplyFlags const& rhs)
{
return safeCast<ApplyFlags>(
safeCast<std::underlying_type_t<ApplyFlags>>(lhs) &
safeCast<std::underlying_type_t<ApplyFlags>>(rhs));
}
static_assert((TapFailHard & TapRetry) == TapNone, "ApplyFlags operator &");
static_assert((TapRetry & TapFailHard) == TapNone, "ApplyFlags operator &");
constexpr ApplyFlags
operator~(ApplyFlags const& flags)
{
return safeCast<ApplyFlags>(~safeCast<std::underlying_type_t<ApplyFlags>>(flags));
}
static_assert(~TapRetry == safeCast<ApplyFlags>(0xFFFFFFDFu), "ApplyFlags operator ~");
inline ApplyFlags
operator|=(ApplyFlags& lhs, ApplyFlags const& rhs)
{
lhs = lhs | rhs;
return lhs;
}
inline ApplyFlags
operator&=(ApplyFlags& lhs, ApplyFlags const& rhs)
{
lhs = lhs & rhs;
return lhs;
}
};
//------------------------------------------------------------------------------

View File

@@ -1,11 +1,17 @@
#pragma once
#include <xrpl/basics/base_uint.h>
#include <xrpl/beast/utility/Journal.h>
#include <xrpl/ledger/ApplyView.h>
#include <xrpl/ledger/ReadView.h>
#include <xrpl/ledger/entries/SLEBase.h>
#include <xrpl/protocol/Indexes.h>
#include <xrpl/protocol/LedgerFormats.h>
#include <xrpl/protocol/SField.h>
#include <xrpl/protocol/STVector256.h>
#include <cstddef>
#include <optional>
namespace xrpl {
@@ -25,6 +31,31 @@ public:
: Base(keylet::skip(), view, j)
{
}
/**
* Looks up a hash `diff` slots back from the most recent entry in the
* sfHashes vector (diff == 0 is the most recent entry).
*
* Callers decide which skip list entry to read and how to translate a
* target ledger sequence into `diff`; this only does the bounds check
* and vector indexing shared by both the recent (stride 1) and distant
* (stride 256) skip lists.
*
* @param diff how many slots back from the most recent hash to look up;
* 0 is the most recent hash.
* @return the hash at that slot, or std::nullopt if the entry does not
* exist or `diff` is out of range.
*/
[[nodiscard]] std::optional<UInt256>
hashAt(std::size_t diff) const
{
if (!this->exists())
return std::nullopt;
STVector256 const vec = (*this)->getFieldV256(sfHashes);
if (vec.size() > diff)
return vec[vec.size() - diff - 1];
return std::nullopt;
}
};
using LedgerHashesEntryR = LedgerHashesEntry<ReadView>;

View File

@@ -1,8 +1,10 @@
#pragma once
#include <xrpl/basics/base_uint.h>
#include <xrpl/nodestore/NodeObject.h>
#include <memory>
#include <optional>
namespace xrpl::node_store {
@@ -23,7 +25,7 @@ public:
/**
* Construct the decoded blob from raw data.
*/
DecodedBlob(void const* key, void const* value, int valueBytes);
DecodedBlob(std::optional<uint256> key, void const* value, int valueBytes);
/**
* Determine if the decoding was successful.
@@ -43,10 +45,10 @@ public:
private:
bool success_{false};
void const* key_;
uint256 key_;
NodeObjectType objectType_{NodeObjectType::Unknown};
unsigned char const* objectData_{nullptr};
int dataBytes_;
int dataBytes_ = 0;
};
} // namespace xrpl::node_store

View File

@@ -1,5 +1,6 @@
#pragma once
#include <xrpl/basics/base_uint.h>
#include <xrpl/beast/utility/instrumentation.h>
#include <xrpl/nodestore/NodeObject.h>
@@ -37,11 +38,6 @@ namespace xrpl::node_store {
class EncodedBlob
{
/**
* The 32-byte key of the serialized object.
*/
std::array<std::uint8_t, 32> key_{};
/**
* A pre-allocated buffer for the serialized object.
*
@@ -65,6 +61,11 @@ class EncodedBlob
*/
std::uint8_t* const ptr_;
/**
* The 32-byte key of the serialized object.
*/
uint256 key_;
public:
explicit EncodedBlob(std::shared_ptr<NodeObject> const& obj)
: size_([&obj]() {
@@ -76,11 +77,12 @@ public:
return obj->getData().size() + 9;
}())
, ptr_((size_ <= payload_.size()) ? payload_.data() : new std::uint8_t[size_])
, key_(obj->getHash())
{
std::fill_n(ptr_, 8, std::uint8_t{0});
ptr_[8] = static_cast<std::uint8_t>(obj->getType());
std::copy_n(obj->getData().data(), obj->getData().size(), ptr_ + 9);
std::copy_n(obj->getHash().data(), obj->getHash().size(), key_.data());
}
~EncodedBlob()
@@ -95,10 +97,10 @@ public:
delete[] ptr_;
}
[[nodiscard]] void const*
[[nodiscard]] uint256 const&
getKey() const noexcept
{
return static_cast<void const*>(key_.data());
return key_;
}
[[nodiscard]] std::size_t

View File

@@ -84,26 +84,6 @@ message TMPublicKey {
// If you want to send an amount that is greater than any single address of yours
// you must first combine coins from one address to another.
// Trace context for OpenTelemetry distributed tracing across nodes.
// Uses W3C Trace Context format internally.
//
// Field numbering note: this message is embedded as field 1001 on
// TMTransaction, TMProposeSet, and TMValidation. Field numbers >= 1000
// are reserved for optional, observability-only additions that must not
// collide with protocol-semantic fields (which historically use 1-99).
// Older peers that do not understand field 1001 will simply ignore it
// per protobuf wire-format rules, preserving backwards compatibility.
message TraceContext {
optional bytes trace_id = 1; // 16-byte trace identifier
optional bytes span_id = 2; // 8-byte parent span identifier
optional uint32 trace_flags = 3; // bit 0 = sampled, bit 1 = random
// Field 4 is held for trace_state (W3C tracestate). It will be added
// here later, with a size limit.
reserved 4;
reserved "trace_state";
}
enum TransactionStatus {
tsNEW = 1; // origin node did/could not validate
tsCURRENT = 2; // scheduled to go in this ledger
@@ -120,9 +100,6 @@ message TMTransaction {
required TransactionStatus status = 2;
optional uint64 receiveTimestamp = 3;
optional bool deferred = 4; // not applied to open ledger
// Optional trace context for OpenTelemetry distributed tracing
optional TraceContext trace_context = 1001;
}
message TMTransactions {
@@ -171,9 +148,6 @@ message TMProposeSet {
// Number of hops traveled
optional uint32 hops = 12 [deprecated = true];
// Optional trace context for OpenTelemetry distributed tracing
optional TraceContext trace_context = 1001;
}
enum TxSetStatus {
@@ -211,9 +185,6 @@ message TMValidation {
// Number of hops traveled
optional uint32 hops = 3 [deprecated = true];
// Optional trace context for OpenTelemetry distributed tracing
optional TraceContext trace_context = 1001;
}
// An array of Endpoint messages

View File

@@ -61,14 +61,22 @@ parseBase58(std::string const& s);
/**
* A special account that's used as the "issuer" for XRP.
*/
AccountID const&
xrpAccount();
constexpr inline AccountID const&
xrpAccount() noexcept
{
static constexpr AccountID kAccount(beast::kZero);
return kAccount;
}
/**
* A placeholder for empty accounts.
*/
AccountID const&
noAccount();
constexpr inline AccountID const&
noAccount() noexcept
{
static constexpr AccountID kAccount = xrpAccount().next();
return kAccount;
}
/**
* Convert hex or base58 string to AccountID.
@@ -80,10 +88,10 @@ bool
toIssuer(AccountID&, std::string const&);
// DEPRECATED Should be checking the currency or native flag
inline bool
isXRP(AccountID const& c)
constexpr inline bool
isXRP(AccountID const& c) noexcept
{
return c == beast::kZero;
return c == xrpAccount();
}
// DEPRECATED
@@ -101,22 +109,6 @@ operator<<(std::ostream& os, AccountID const& x)
return os;
}
/**
* Initialize the global cache used to map AccountID to base58 conversions.
*
* The cache is optional and need not be initialized. But because conversion
* is expensive (it requires a SHA-256 operation) in most cases the overhead
* of the cache is worth the benefit.
*
* @param count The number of entries the cache should accommodate. Zero will
* disable the cache, releasing any memory associated with it.
*
* @note The function will only initialize the cache the first time it is
* invoked. Subsequent invocations do nothing.
*/
void
initAccountIdCache(std::size_t count);
} // namespace xrpl
//------------------------------------------------------------------------------

View File

@@ -1,10 +1,10 @@
#pragma once
#include <xrpl/basics/contract.h>
#include <xrpl/beast/type_name.h>
#include <xrpl/protocol/SOTemplate.h>
#include <boost/container/flat_map.hpp>
#include <boost/core/type_name.hpp>
#include <algorithm>
#include <cstddef>
@@ -84,7 +84,7 @@ public:
* Derived classes will load the object with all the known formats.
*/
private:
KnownFormats() : name_(beast::typeName<Derived>())
KnownFormats() : name_(boost::core::type_name<Derived>())
{
}

View File

@@ -33,7 +33,20 @@ public:
STBitString(SField const& n);
STBitString(value_type const& v);
template <typename Tag>
requires(!std::is_void_v<Tag>)
STBitString(BaseUInt<Bits, Tag> const& v) : value_(v)
{
}
STBitString(SField const& n, value_type const& v);
template <typename Tag>
requires(!std::is_void_v<Tag>)
STBitString(SField const& n, BaseUInt<Bits, Tag> const& v) : STBase(n), value_(v)
{
}
STBitString(SerialIter& sit, SField const& name);
[[nodiscard]] SerializedTypeID
@@ -166,7 +179,7 @@ template <typename Tag>
void
STBitString<Bits>::setValue(BaseUInt<Bits, Tag> const& v)
{
value_ = v;
value_ = value_type{v};
}
template <int Bits>

View File

@@ -2,6 +2,7 @@
#include <xrpl/basics/CountedObject.h>
#include <xrpl/basics/UnorderedContainers.h>
#include <xrpl/basics/enum_bitops.h>
#include <xrpl/beast/utility/instrumentation.h>
#include <xrpl/json/json_value.h>
#include <xrpl/protocol/AccountID.h>
@@ -15,6 +16,7 @@
#include <cstdint>
#include <memory>
#include <optional>
#include <type_traits>
#include <utility>
#include <vector>
@@ -22,18 +24,8 @@ namespace xrpl {
class STPathElement final : public CountedObject<STPathElement>
{
unsigned int type_;
AccountID accountID_;
PathAsset assetID_;
AccountID issuerID_;
bool isOffer_;
std::size_t hashValue_;
public:
// Bitwise values (typeCurrency | typeMPT)
// NOLINTNEXTLINE(cppcoreguidelines-use-enum-class)
enum Type {
enum class Type : std::uint8_t {
TypeNone = 0x00,
TypeAccount = 0x01, // Rippling through an account (vs taking an offer).
TypeCurrency = 0x10, // Currency follows.
@@ -45,6 +37,8 @@ public:
// Combination of all types.
};
using enum Type;
STPathElement();
STPathElement(STPathElement const&) = default;
STPathElement&
@@ -62,42 +56,42 @@ public:
bool forceAsset = false);
STPathElement(
unsigned int uType,
Type type,
AccountID const& account,
PathAsset const& asset,
AccountID const& issuer);
[[nodiscard]] std::uint32_t
getNodeType() const;
[[nodiscard]] Type
getNodeType() const noexcept;
[[nodiscard]] bool
isOffer() const;
isOffer() const noexcept;
[[nodiscard]] bool
isAccount() const;
isAccount() const noexcept;
[[nodiscard]] bool
hasIssuer() const;
hasIssuer() const noexcept;
[[nodiscard]] bool
hasCurrency() const;
hasCurrency() const noexcept;
[[nodiscard]] bool
hasMPT() const;
hasMPT() const noexcept;
[[nodiscard]] bool
hasAsset() const;
hasAsset() const noexcept;
[[nodiscard]] bool
isNone() const;
isNone() const noexcept;
// Nodes are either an account ID or a offer prefix. Offer prefixs denote a
// class of offers.
[[nodiscard]] AccountID const&
getAccountID() const;
getAccountID() const noexcept;
[[nodiscard]] PathAsset const&
getPathAsset() const;
getPathAsset() const noexcept;
[[nodiscard]] Currency const&
getCurrency() const;
@@ -106,17 +100,35 @@ public:
getMPTID() const;
[[nodiscard]] AccountID const&
getIssuerID() const;
getIssuerID() const noexcept;
[[nodiscard]] bool
isType(Type const& pe) const;
isType(Type pe) const noexcept;
bool
operator==(STPathElement const& t) const;
friend bool
operator==(STPathElement const& lhs, STPathElement const& rhs) noexcept
{
return lhs.isType(TypeAccount) == rhs.isType(TypeAccount) &&
lhs.hashValue_ == rhs.hashValue_ && lhs.accountID_ == rhs.accountID_ &&
lhs.assetID_ == rhs.assetID_ && lhs.issuerID_ == rhs.issuerID_;
}
private:
static std::size_t
getHash(STPathElement const& element);
Type type_;
AccountID accountID_;
PathAsset assetID_;
AccountID issuerID_;
bool isOffer_;
std::size_t hashValue_;
};
template <>
struct enum_bitops::OptIn<STPathElement::Type> : std::true_type
{
};
template <class Hasher>
@@ -124,7 +136,7 @@ void
hash_append(Hasher& h, STPathElement const& e) noexcept
{
using beast::hash_append;
hash_append(h, (e.getNodeType() & STPathElement::TypeAccount) != 0u);
hash_append(h, e.isType(STPathElement::TypeAccount));
hash_append(h, e.getAccountID());
hash_append(h, e.getPathAsset());
hash_append(h, e.getIssuerID());
@@ -404,11 +416,11 @@ inline STPathElement::STPathElement(
}
inline STPathElement::STPathElement(
unsigned int uType,
Type type,
AccountID const& account,
PathAsset const& asset,
AccountID const& issuer)
: type_(uType)
: type_(type)
, accountID_(account)
, assetID_(asset)
, issuerID_(issuer)
@@ -423,56 +435,56 @@ inline STPathElement::STPathElement(
hashValue_ = getHash(*this);
}
inline std::uint32_t
STPathElement::getNodeType() const
inline STPathElement::Type
STPathElement::getNodeType() const noexcept
{
return type_;
}
inline bool
STPathElement::isOffer() const
STPathElement::isOffer() const noexcept
{
return isOffer_;
}
inline bool
STPathElement::isAccount() const
STPathElement::isAccount() const noexcept
{
return !isOffer();
}
inline bool
STPathElement::isType(Type const& pe) const
STPathElement::isType(Type pe) const noexcept
{
return (type_ & pe) != 0u;
return (type_ & pe) != STPathElement::TypeNone;
}
inline bool
STPathElement::hasIssuer() const
STPathElement::hasIssuer() const noexcept
{
return isType(STPathElement::TypeIssuer);
}
inline bool
STPathElement::hasCurrency() const
STPathElement::hasCurrency() const noexcept
{
return isType(STPathElement::TypeCurrency);
}
inline bool
STPathElement::hasMPT() const
STPathElement::hasMPT() const noexcept
{
return isType(STPathElement::TypeMpt);
}
inline bool
STPathElement::hasAsset() const
STPathElement::hasAsset() const noexcept
{
return isType(STPathElement::TypeAsset);
}
inline bool
STPathElement::isNone() const
STPathElement::isNone() const noexcept
{
return getNodeType() == STPathElement::TypeNone;
}
@@ -480,13 +492,13 @@ STPathElement::isNone() const
// Nodes are either an account ID or a offer prefix. Offer prefixs denote a
// class of offers.
inline AccountID const&
STPathElement::getAccountID() const
STPathElement::getAccountID() const noexcept
{
return accountID_;
}
inline PathAsset const&
STPathElement::getPathAsset() const
STPathElement::getPathAsset() const noexcept
{
return assetID_;
}
@@ -504,18 +516,11 @@ STPathElement::getMPTID() const
}
inline AccountID const&
STPathElement::getIssuerID() const
STPathElement::getIssuerID() const noexcept
{
return issuerID_;
}
inline bool
STPathElement::operator==(STPathElement const& t) const
{
return (type_ & TypeAccount) == (t.type_ & TypeAccount) && hashValue_ == t.hashValue_ &&
accountID_ == t.accountID_ && assetID_ == t.assetID_ && issuerID_ == t.issuerID_;
}
// ------------ STPath ------------
inline STPath::STPath(std::vector<STPathElement> p) : path_(std::move(p))

View File

@@ -596,7 +596,7 @@ template <std::size_t Bits, class Tag>
BaseUInt<Bits, Tag>
SerialIter::getBitString()
{
auto const n = Bits / 8;
constexpr auto n = BaseUInt<Bits, Tag>::size();
if (remain_ < n)
Throw<std::runtime_error>("invalid SerialIter getBitString");
@@ -607,7 +607,7 @@ SerialIter::getBitString()
used_ += n;
remain_ -= n;
return BaseUInt<Bits, Tag>::fromVoid(x);
return BaseUInt<Bits, Tag>{std::span<std::uint8_t const, n>{x, n}};
}
} // namespace xrpl

View File

@@ -14,7 +14,7 @@ namespace xrpl {
// Various protocol and system specific constant globals.
/* The name of the system. */
static inline std::string const&
inline std::string const&
systemName()
{
static std::string const kName = "xrpld";
@@ -52,11 +52,10 @@ isLegalAmountSigned(XRPAmount const& amount)
}
/* The currency code for the native currency. */
static inline std::string const&
inline std::string
systemCurrencyCode()
{
static std::string const kCode = "XRP";
return kCode;
return "XRP";
}
/**

View File

@@ -61,26 +61,40 @@ using Domain = BaseUInt<256>;
/**
* XRP currency.
*/
Currency const&
xrpCurrency();
constexpr inline Currency const&
xrpCurrency() noexcept
{
static constexpr Currency const kCurrency(beast::kZero);
return kCurrency;
}
/**
* A placeholder for empty currencies.
*/
Currency const&
noCurrency();
constexpr inline Currency const&
noCurrency() noexcept
{
static constexpr Currency const kCurrency(xrpCurrency().next());
return kCurrency;
}
/**
* We deliberately disallow the currency that looks like "XRP" because too
* many people were using it instead of the correct XRP currency.
*
* Note that this doesn't catch "xRP" or "xrp" or other case variations.
*/
Currency const&
badCurrency();
inline bool
isXRP(Currency const& c)
constexpr inline Currency const&
badCurrency() noexcept
{
return c == beast::kZero;
static constexpr Currency kCurrency{"0000000000000000000000005852500000000000"};
return kCurrency;
}
constexpr inline bool
isXRP(Currency const& c) noexcept
{
return c == xrpCurrency();
}
/**
@@ -98,7 +112,7 @@ to_string(Currency const& c);
* to rewrite some unit test code.
*/
bool
toCurrency(Currency&, std::string const&);
toCurrency(Currency&, std::string_view);
/**
* Tries to convert a string to a Currency, returns noCurrency() on failure.
@@ -107,8 +121,7 @@ toCurrency(Currency&, std::string const&);
* unfortunate; changing this will require very careful checking
* everywhere and may mean having to rewrite some unit test code.
*/
Currency
toCurrency(std::string const&);
Currency toCurrency(std::string_view);
inline std::ostream&
operator<<(std::ostream& os, Currency const& x)

View File

@@ -8,6 +8,7 @@
#include <array>
#include <cstddef>
#include <cstdint>
#include <span>
#include <type_traits>
namespace xrpl {
@@ -182,7 +183,7 @@ public:
operator result_type() noexcept
{
auto const digest = Sha512Hasher::result_type(h_);
return result_type::fromVoid(digest.data());
return result_type{std::span{digest}.first<result_type::size()>()};
}
private:

View File

@@ -392,6 +392,7 @@ JSS(max_ledger); // in/out: LedgerCleaner
JSS(max_queue_size); // out: TxQ
JSS(max_spend_drops); // out: AccountInfo
JSS(max_spend_drops_total); // out: AccountInfo
JSS(maximum);
JSS(mean); // out: get_aggregate_price
JSS(median); // out: get_aggregate_price
JSS(median_fee); // out: TxQ

View File

@@ -8,6 +8,7 @@
#include <cstdint>
#include <cstring>
#include <span>
namespace xrpl::nft {
@@ -98,7 +99,8 @@ getTaxon(UInt256 const& id)
inline AccountID
getIssuer(UInt256 const& id)
{
return AccountID::fromVoid(id.data() + 4);
return AccountID{
std::span<unsigned char const, uint256::size()>{id}.subspan<4, AccountID::size()>()};
}
} // namespace xrpl::nft

View File

@@ -2,13 +2,10 @@
#include <xrpl/basics/base_uint.h>
#include <string_view>
namespace xrpl::nft {
// NFT directory pages order their contents based only on the low 96 bits of
// the NFToken value. This mask provides easy access to the necessary mask.
constexpr UInt256 kPageMask(
std::string_view("0000000000000000000000000000000000000000ffffffffffffffffffffffff"));
constexpr UInt256 kPageMask{"0000000000000000000000000000000000000000ffffffffffffffffffffffff"};
} // namespace xrpl::nft

View File

@@ -1,132 +0,0 @@
#pragma once
/**
* A coroutine-aware OpenTelemetry runtime-context storage.
*
* OpenTelemetry's default ThreadLocalContextStorage keeps the active-context
* stack in a plain static thread_local, so the ambient span does NOT follow a
* JobQueue::Coro across yield/resume — a scope pushed on one worker is stranded
* when the coroutine resumes on another. This storage instead keeps the stack
* in an xrpl::LocalValue, which JobQueue::Coro::resume() swaps in and out with
* the coroutine (see Coro.ipp). The active context therefore rides the
* coroutine: scopes held across yield are safe, and per-line log-trace
* correlation (Log.cpp reads RuntimeContext::GetCurrent()) is retained.
*
* Off a coroutine, LocalValue transparently provides a per-thread store, so
* behaviour is identical to the default thread-local storage.
*
* +-------------------------------------------------+
* | CoroAwareContextStorage |
* | (opentelemetry RuntimeContextStorage) |
* +-------------------------------------------------+
* | - stack_ : LocalValue<vector<Context>> |
* +-------------------------------------------------+
* | + GetCurrent() : Context |
* | + Attach(ctx) : unique_ptr<Token> |
* | + Detach(token): bool |
* +-------------------------------------------------+
* | backed by (coro-aware)
* +------------------------+
* | xrpl::LocalValue store |
* | (swapped by Coro) |
* +------------------------+
*
* Install once at telemetry start via
* opentelemetry::context::RuntimeContext::SetRuntimeContextStorage(), BEFORE
* any span is created (SDK requirement).
*
* @note Thread-safety: each thread/coroutine sees its own LocalValue store, so
* the stack is never shared across threads — no locking needed. The storage
* object itself is stateless apart from the LocalValue handle.
* @note Known limitation: install once, before the first span; resetting the
* storage while spans exist is undefined behaviour (SDK). A span created on a
* raw worker thread and later moved onto a coroutine does not retro-attach to
* the coroutine's store — not a pattern here (spans are created inside their
* own coro/job body).
*
* Example 1 — install at telemetry start (primary use):
* @code
* using opentelemetry::context::RuntimeContext;
* RuntimeContext::SetRuntimeContextStorage(
* opentelemetry::nostd::shared_ptr<
* opentelemetry::context::RuntimeContextStorage>(
* new xrpl::telemetry::CoroAwareContextStorage()));
* @endcode
*
* Example 2 — a scope held across a coroutine yield stays correct:
* @code
* // Inside a JobQueue::Coro body:
* auto span = ScopedSpanGuard(TraceCategory::Rpc, "rpc", "process");
* context.coro->yield(); // may resume on another worker
* // span is still the ambient context here — the storage rode the coro.
* @endcode
*
* Example 3 — edge case: off a coroutine, behaves like thread-local storage:
* @code
* // On a plain worker thread (no coro): each thread has its own stack,
* // identical to opentelemetry's default ThreadLocalContextStorage.
* auto span = ScopedSpanGuard(TraceCategory::Ledger, "ledger", "build");
* @endcode
*/
#ifdef XRPL_ENABLE_TELEMETRY
#include <xrpl/basics/LocalValue.h>
#include <opentelemetry/context/context.h>
#include <opentelemetry/context/runtime_context.h>
#include <opentelemetry/nostd/unique_ptr.h>
#include <vector>
namespace xrpl::telemetry {
class CoroAwareContextStorage : public opentelemetry::context::RuntimeContextStorage
{
public:
CoroAwareContextStorage() = default;
// Non-copyable and non-movable: the LocalValue store is keyed by the
// address of stack_, so copying or moving would mis-key the store and
// strand its entries. Only ever new-ed once and installed on the SDK, so
// these are for intent/safety rather than a live bug.
CoroAwareContextStorage(CoroAwareContextStorage const&) = delete;
CoroAwareContextStorage&
operator=(CoroAwareContextStorage const&) = delete;
CoroAwareContextStorage(CoroAwareContextStorage&&) = delete;
CoroAwareContextStorage&
operator=(CoroAwareContextStorage&&) = delete;
/**
* @return the current (top-of-stack) context for this coro/thread.
*/
opentelemetry::context::Context
GetCurrent() noexcept override;
/**
* Push a context frame onto this coro/thread's stack.
* @param context the context to make current.
* @return a token that Detach() uses to pop back to the prior frame.
*/
opentelemetry::nostd::unique_ptr<opentelemetry::context::Token>
Attach(opentelemetry::context::Context const& context) noexcept override;
/**
* Pop the stack back through the frame the token refers to.
* @param token a token returned by Attach().
* @return true if the frame was found and detached.
*/
bool
Detach(opentelemetry::context::Token& token) noexcept override;
private:
/**
* The active-context stack, stored coro-locally. LocalValue hands back the
* stack for whichever store (coro or thread) is currently installed.
*/
LocalValue<std::vector<opentelemetry::context::Context>> stack_;
};
} // namespace xrpl::telemetry
#endif // XRPL_ENABLE_TELEMETRY

View File

@@ -1,188 +0,0 @@
#pragma once
#ifdef XRPL_ENABLE_TELEMETRY
#include <opentelemetry/sdk/trace/id_generator.h>
#include <opentelemetry/trace/span_id.h>
#include <opentelemetry/trace/trace_id.h>
#include <array>
#include <cstdint>
#include <memory>
namespace xrpl::telemetry {
/**
* OTel IdGenerator that can mint a deterministic (hash-derived) trace_id.
*
* By default the OTel SDK generates random trace_ids, so a span derived from
* a stable hash (e.g. a transaction id) cannot become a real trace root: the
* SDK either inherits the active span as parent or invents a random root id.
* This generator lets a caller pin the trace_id of the next forced-root span
* to a chosen 16-byte value, so hash-derived spans line up across nodes into
* one trace. When no value is pinned it behaves exactly like the default
* random generator.
*
* The pinned value lives in a thread-local slot set by PendingTraceId (below).
* Only GenerateTraceId() consults that slot, and only on the SDK's no-parent
* (root) branch; span_ids are always random. is_random_ is false so the W3C
* random-trace-id flag is not set on these deterministic ids.
*
* Dependency / data-flow diagram:
*
* +-----------------------------------------------------------+
* | DeterministicIdGenerator |
* | (IdGenerator) |
* +-----------------------------------------------------------+
* | - random_ : unique_ptr<IdGenerator> (random delegate) |
* +-----------------------------------------------------------+
* | GenerateTraceId(): |
* | thread-local pending id set? --yes--> return that id |
* | --no --> random_ ... |
* | GenerateSpanId(): always ---------------> random_ ... |
* +-----------------------------------------------------------+
* ^ |
* sets/clears| delegates|
* | v
* +----------------+ +--------------------+
* | PendingTraceId | | RandomIdGenerator |
* | (RAII guard) | | (delegate) |
* +----------------+ +--------------------+
*
* @note Thread safety: the pending trace_id is a file-local thread_local, so
* each thread sees only its own pinned value and no synchronization is needed.
* The pending id is consumed ONLY by GenerateTraceId(), which the SDK calls
* solely when a span has no valid parent; GenerateSpanId() never reads it and
* is always random.
* @note Limitation: minting a deterministic root only works when the caller
* forces the SDK's root branch (e.g. by starting the span with a
* Context{kIsRootSpanKey, true}); otherwise the SDK reuses the parent's
* trace_id and never calls GenerateTraceId(). The primary such caller is
* SpanGuard::hashSpan(), which forces the root branch to mint deterministic
* per-object trace roots.
*
* Example 1 - primary use, inside a forced-root hash span (shown as
* pseudocode):
* @code
* // std::array<std::uint8_t, 16> id = deriveTraceIdFromHash(txHash);
* // PendingTraceId const pending{id}; // pin id for this thread
* // auto root = Context{kIsRootSpanKey, true}; // force the no-parent branch
* // auto guard = telemetry.startSpan("tx.process", root);
* // // GenerateTraceId() returns `id`; the guard's trace_id == id.
* @endcode
*
* Example 2 - edge case, a normal child span never consults the pending id:
* @code
* // With an active parent span, startSpan() inherits the parent's trace_id
* // and the SDK does NOT call GenerateTraceId(), so no PendingTraceId is used.
* // auto child = parentGuard.childSpan(rpc_span::prefix::command); // random/parent trace_id
* @endcode
*/
class DeterministicIdGenerator final : public opentelemetry::sdk::trace::IdGenerator
{
/**
* Random generator the deterministic path falls back to. Used for every
* span_id and for any trace_id when no PendingTraceId is active.
*/
std::unique_ptr<opentelemetry::sdk::trace::IdGenerator> random_;
public:
/**
* Build a generator with is_random_ = false and a random delegate.
*/
DeterministicIdGenerator();
/**
* @return the thread's pending trace_id if a PendingTraceId is active,
* otherwise a fresh random trace_id from the delegate. Consuming the
* pending id clears it so it applies to exactly one root span.
*/
opentelemetry::trace::TraceId
GenerateTraceId() noexcept override;
/**
* @return a fresh random span_id. Never uses the pending trace_id.
*/
opentelemetry::trace::SpanId
GenerateSpanId() noexcept override;
};
/**
* RAII guard that pins a deterministic trace_id for the next forced-root span
* started on this thread.
*
* Mirrors DiscardScope: it sets a thread-local pending trace_id on
* construction and clears it on destruction, so the pinned id stays confined
* to the guard's scope and cannot leak onto a later span. On destruction it
* asserts that the id was actually consumed by GenerateTraceId() — if it was
* not, the SDK took a branch other than the root branch the caller intended,
* which is a bug worth catching in debug/test builds.
*
* Wrap ONLY the single forced-root startSpan() call that must receive the
* deterministic id. Non-copyable and non-movable: its sole purpose is the
* scoped lifetime of the pending id.
*
* @note Thread safety: the pending id is thread-local, so a guard on one
* thread never affects another. Construct and destroy the guard on the same
* thread as the startSpan() call it wraps.
* @note Limitation: the consumed-assert only holds when the wrapped span is
* started on the SDK root branch (Context{kIsRootSpanKey, true}); wrapping a
* child span would leave the id unconsumed and trip the assert. Nesting two
* guards on one thread is unsupported: the inner guard consumes/clears first,
* so the outer id is silently dropped and its destructor trips the assert.
*
* Example 1 - primary use, wrapping a forced-root span start (pseudocode):
* @code
* // {
* // PendingTraceId const pending{id}; // pin for this scope
* // auto root = Context{kIsRootSpanKey, true};
* // auto guard = telemetry.startSpan(name, root);// consumes the pinned id
* // } // ~PendingTraceId asserts the id was consumed, then clears it
* @endcode
*
* Example 2 - edge case, misuse detection: wrapping a non-root span leaves the
* id unconsumed, so ~PendingTraceId trips XRPL_ASSERT in debug/test builds.
*/
class PendingTraceId
{
public:
/**
* Pin @p id as the pending trace_id for this thread.
* @param id The 16-byte trace_id the next forced-root span should adopt.
*/
explicit PendingTraceId(std::array<std::uint8_t, 16> const& id) noexcept;
/**
* Assert the pinned id was consumed by a root span, then clear it so it
* never leaks onto the next span (even in release, where the assert is a
* no-op).
*/
~PendingTraceId() noexcept;
PendingTraceId(PendingTraceId const&) = delete;
PendingTraceId&
operator=(PendingTraceId const&) = delete;
PendingTraceId(PendingTraceId&&) = delete;
PendingTraceId&
operator=(PendingTraceId&&) = delete;
};
/**
* Running count of deterministic trace_ids that were pinned by a
* PendingTraceId but never consumed by a forced-root span before the guard was
* destroyed. Bumped once per silent drop.
*
* ~PendingTraceId also asserts on this condition, but the assert is a no-op in
* release builds, so this monotonic counter is the only signal that a
* deterministic trace root was lost there. It is process-wide (summed across
* all threads) and intended for a future diagnostic metric or debugger
* inspection; reading it has no side effects.
*
* @return the total number of unconsumed deterministic trace_id drops so far.
*/
[[nodiscard]] std::uint64_t
unconsumedDeterministicIdDrops() noexcept;
} // namespace xrpl::telemetry
#endif // XRPL_ENABLE_TELEMETRY

View File

@@ -1,95 +0,0 @@
#pragma once
/**
* Thread-local discard signaling between SpanGuard and the span processor.
*
* SpanGuard::discard() wants to drop a span without sending it to the
* exporter. The OTel SDK calls SpanProcessor::OnEnd() synchronously on the
* same thread that calls Span::End(), so a thread-local flag set just before
* End() and read inside OnEnd() lets FilteringSpanProcessor drop the span
* before it enters the batch export queue.
*
* This side-channel avoids inspecting the Recordable's internals (which vary
* by exporter type — SpanData vs OtlpRecordable).
*
* The flag is a *private* thread-local member of DiscardScope, mutated only by
* its constructor and destructor. This gives real access control rather than a
* naming convention: no code that includes this header can flip the flag
* directly — it can only enter a DiscardScope (which sets and clears the flag
* over its own lifetime) and observe the state via DiscardScope::isActive().
* Binding set/clear to a scope also means the flag cannot leak onto the next
* span even if End() were to throw.
*
* Kept in a separate header to avoid transitive include bloat: SpanGuard.h
* only needs this signaling, not the full Telemetry.h with BasicConfig/Journal.
*
* Usage:
* @code
* // In SpanGuard::discard():
* {
* DiscardScope discardScope; // flag set for this scope only
* span->End(); // OnEnd() runs synchronously, sees flag
* } // flag cleared here, unconditionally
*
* // In FilteringSpanProcessor::OnEnd():
* if (DiscardScope::isActive())
* return; // drop the span
* @endcode
*
* @note Thread safety: the flag is thread-local, so each thread observes only
* its own discard signal — no synchronization is required.
*
* @see SpanGuard::discard(), FilteringSpanProcessor (FilteringSpanProcessor.h)
*/
namespace xrpl::telemetry {
/**
* RAII guard that marks the current thread's span for discard.
*
* Sets the thread-local discard flag on construction and clears it on
* destruction, so a span ended within the guard's scope is dropped by
* FilteringSpanProcessor::OnEnd() while the flag stays confined to that scope.
* The flag is private and mutated only here, so no other code can set it.
* Non-copyable and non-movable — its sole purpose is the scoped flag lifetime.
*/
class DiscardScope
{
public:
DiscardScope() noexcept
{
discarding = true;
}
~DiscardScope()
{
discarding = false;
}
DiscardScope(DiscardScope const&) = delete;
DiscardScope&
operator=(DiscardScope const&) = delete;
DiscardScope(DiscardScope&&) = delete;
DiscardScope&
operator=(DiscardScope&&) = delete;
/**
* @return true if the current thread is inside a DiscardScope, i.e. the
* span ending now should be dropped rather than exported. Read by
* FilteringSpanProcessor::OnEnd().
*/
[[nodiscard]] static bool
isActive() noexcept
{
return discarding;
}
private:
/**
* Thread-local discard flag. Private, so only this class's ctor/dtor can
* mutate it; observers use isActive(). One instance per thread.
*/
inline static thread_local bool discarding = false;
};
} // namespace xrpl::telemetry

View File

@@ -1,138 +0,0 @@
#pragma once
#ifdef XRPL_ENABLE_TELEMETRY
#include <opentelemetry/sdk/trace/processor.h>
#include <opentelemetry/sdk/trace/recordable.h>
#include <opentelemetry/trace/span_context.h>
#include <chrono>
#include <memory>
namespace xrpl::telemetry {
/**
* Span processor that sits in front of the batch processor. It drops
* discarded spans, and it makes every span export each attribute key once.
*
* The OTLP recordable appends every SetAttribute() call without looking up
* the key. Without this processor a key set twice is exported twice, and a
* reader that takes the first value misses the update. So the processor
* wraps each recordable it hands to the SDK. The wrapper keeps the last value
* for each key and writes each key once, in first-set order, when the span
* ends. Every other call goes straight to the delegate's recordable.
*
* Dependency diagram:
*
* +------------------------+ delegate_ +--------------------+
* | FilteringSpanProcessor |------------->| BatchSpanProcessor |
* +------------------------+ +--------------------+
* | reads | makes
* v v
* DiscardScope (DiscardFlag.h) OtlpRecordable
*
* Data flow for one span:
*
* MakeRecordable() --> wrapper around the delegate's recordable
* SetAttribute() --> wrapper keeps an owned copy, last value wins
* other calls --> forwarded to the delegate's recordable at once
* OnEnd() --> discarded? drop the span
* --> else write each kept key once, then pass the
* delegate's recordable to the delegate
*
* The discard check reads a thread-local flag, not the span's attributes,
* because the recordable type depends on the exporter and has no common
* getter. The flag works because Span::End() calls OnEnd() on its own thread.
*
* @code
* // Production: wrap the batch processor before building the provider.
* auto processor = std::make_unique<FilteringSpanProcessor>(std::move(batchProcessor));
*
* // A key set twice is exported once, as "accepted".
* auto span = processor->MakeRecordable();
* span->SetAttribute("consensus_phase", "open");
* span->SetAttribute("consensus_phase", "accepted");
* processor->OnEnd(std::move(span));
*
* // Edge case: a span ended inside a DiscardScope never reaches the delegate.
* {
* DiscardScope discardScope;
* processor->OnEnd(processor->MakeRecordable());
* }
* @endcode
*
* @note Thread safety: OnStart() and OnEnd() may run on many threads at once.
* Each wrapper belongs to one span, and the SDK span serializes calls to it.
* The discard flag is thread-local.
* @note Limitations: every distinct key is kept until the span ends, so memory
* grows with the number of distinct keys, and the delegate's attribute count
* limit applies when the keys are written. Attributes inside one event or link
* are forwarded as given. An attribute a delegate sets on the recordable in its
* own OnStart() skips the merge.
*/
class FilteringSpanProcessor final : public opentelemetry::sdk::trace::SpanProcessor
{
/**
* Receives every span that is not discarded. The batch processor in
* production.
*/
std::unique_ptr<opentelemetry::sdk::trace::SpanProcessor> delegate_;
public:
/**
* @param delegate Processor that receives every span that is not discarded.
*/
explicit FilteringSpanProcessor(
std::unique_ptr<opentelemetry::sdk::trace::SpanProcessor> delegate);
/**
* Asks the delegate for a recordable and wraps it.
*
* @return the wrapper, or nullptr when the delegate made none. The SDK then
* records nothing for the span.
*/
std::unique_ptr<opentelemetry::sdk::trace::Recordable>
MakeRecordable() noexcept override;
/**
* Tells the delegate a span started, passing the delegate's own recordable.
*
* @param span Recordable from MakeRecordable(). Any other recordable is
* passed on as it is.
* @param parentContext Context of the span's parent.
*/
void
OnStart(
opentelemetry::sdk::trace::Recordable& span,
opentelemetry::trace::SpanContext const& parentContext) noexcept override;
/**
* Drops the span if it was discarded. Otherwise writes each kept attribute
* once and passes the delegate's own recordable to the delegate.
*
* @param span Recordable from MakeRecordable(). Any other recordable is
* passed on as it is.
*/
void
OnEnd(std::unique_ptr<opentelemetry::sdk::trace::Recordable>&& span) noexcept override;
/**
* @param timeout Longest time to wait for the delegate.
* @return the delegate's result.
*/
bool
ForceFlush(
std::chrono::microseconds timeout = std::chrono::microseconds::max()) noexcept override;
/**
* @param timeout Longest time to wait for the delegate.
* @return the delegate's result.
*/
bool
Shutdown(
std::chrono::microseconds timeout = std::chrono::microseconds::max()) noexcept override;
};
} // namespace xrpl::telemetry
#endif // XRPL_ENABLE_TELEMETRY

View File

@@ -1,40 +0,0 @@
#pragma once
#ifdef XRPL_ENABLE_TELEMETRY
#include <opentelemetry/sdk/trace/sampler.h>
#include <memory>
namespace xrpl::telemetry {
/**
* Build the head sampler that TelemetryImpl installs on its TracerProvider.
*
* The result is a ParentBased sampler. A span with no parent or a remote
* parent is decided by one TraceIdRatioBasedSampler, so a peer's sampled flag
* cannot turn our spans off or on. A span with a local parent follows it.
*
* parent of the new span delegate
* ---------------------- ------------------------
* none (root span) --> TraceIdRatioBased(ratio)
* remote, sampled --> TraceIdRatioBased(ratio)
* remote, not sampled --> TraceIdRatioBased(ratio)
* local, sampled --> AlwaysOn
* local, not sampled --> AlwaysOff
*
* The ratio sampler decides from the trace id alone, so nodes that use the
* same ratio make the same choice for every trace.
*
* @note The sampler has no mutable state, so the SDK may call ShouldSample()
* on many threads at once.
*
* @param ratio Fraction of trace ids to sample, from 0.0 (none) to 1.0 (all).
* @return A new sampler that owns its delegates.
*/
[[nodiscard]] std::unique_ptr<opentelemetry::sdk::trace::Sampler>
makeHeadSampler(double ratio);
} // namespace xrpl::telemetry
#endif // XRPL_ENABLE_TELEMETRY

View File

@@ -1,62 +0,0 @@
#pragma once
/**
* Account-address redaction helper for telemetry span attributes.
*
* A single pure helper that turns a string into a short, stable,
* obfuscated token, for a span attribute whose value should not be stored
* in the clear and is hard to guess.
*
* Not applied to any span today. Account addresses are public ledger
* identifiers, so the path-finding spans emit them raw (see
* PathFindSpanNames.h) and no collector processor hashes them. Use this
* helper only for a value that is genuinely private, and document the
* reason at the attribute constant.
*
* The returned token is the first 16 hex characters (lowercase) of the
* SHA-512Half digest of the input. It is deterministic (same input
* always maps to the same token) so spans for one value still correlate
* across nodes and restarts.
*
* The hash is unsalted, so it is obfuscation, not a secrecy guarantee.
* It hides a value only when that value is hard to guess: for an input
* drawn from a small or enumerable set, such as an account address, an
* observer can rebuild the value->token mapping by lookup, which is why
* account addresses are emitted raw instead. A salt is intentionally
* omitted because it would break cross-node/restart correlation, which is
* the reason for hashing rather than dropping.
*
* @note This function is pure and reentrant: it holds no global state,
* performs no I/O, and is safe to call concurrently from any thread.
*
* Usage example:
* @code
* #include <xrpl/telemetry/Redaction.h>
* using namespace xrpl::telemetry;
*
* auto const token = redactAccount(value); // 16 lowercase hex chars
* @endcode
*
* Edge case (empty input yields empty output):
* @code
* assert(redactAccount("") == "");
* @endcode
*/
#include <string>
#include <string_view>
namespace xrpl::telemetry {
/**
* Hash a value into a short, stable, obfuscated token.
*
* @param addr The value to redact. Named for its original use on account
* addresses; any string can be passed.
* @return The first 16 lowercase hex characters of sha512Half(addr),
* or an empty string when @p addr is empty.
*/
[[nodiscard]] std::string
redactAccount(std::string_view addr);
} // namespace xrpl::telemetry

File diff suppressed because it is too large Load Diff

View File

@@ -1,182 +0,0 @@
#pragma once
/**
* Compile-time string concatenation utility and shared telemetry constants.
*
* Provides StaticStr<N> — a compile-time string buffer that implicitly
* converts to std::string_view — and join() for dot-separated concatenation.
* Module-specific span names (e.g. RPC, consensus) live in their respective
* modules and build upon these shared primitives.
*
* @note These constants are NOT guarded by XRPL_ENABLE_TELEMETRY because
* call sites reference them even when SpanGuard methods are no-ops
* (the no-op stubs still accept string_view parameters). The compiler
* elides all inline constexpr values whose only uses are in dead code.
*
* @note Json::StaticString (jss.h) is a pointer wrapper without
* concatenation support. boost::static_string is not constexpr.
* StaticStr<N> exists specifically for compile-time dot-join composition.
*
* Naming conventions (see spec 2026-05-13-span-attr-naming-design):
* - Per-span attribute keys: bare field name (span name carries the domain).
* - Collision qualifier: <domain>_<field> when bare name collides across
* domains or with OTel reserved `status` (e.g. rpc_status, grpc_status).
* - Shared cross-span attributes: <domain>_<field> (underscore) form
* (e.g. tx_hash, peer_id, ledger_seq, consensus_round).
* - Resource attribute keys: xrpl.<subsystem>.<field> (dotted) form is
* RESERVED for process-identity attributes set once at startup on the
* OTel resource (e.g. xrpl.network.id, xrpl.network.type). Do not use
* this form for span attributes — it parses awkwardly in TraceQL and
* blurs the resource/span scope distinction.
* - Span prefixes: <subsystem>[.<component>].
*/
#include <cstddef>
#include <string_view>
namespace xrpl::telemetry {
// ===== Compile-time string utility =========================================
/**
* Fixed-size character buffer for compile-time string operations.
* Implicitly converts to std::string_view at zero cost.
*/
template <std::size_t N>
struct StaticStr
{
char data[N + 1]{};
static constexpr std::size_t size = N;
constexpr StaticStr() = default;
constexpr explicit StaticStr(char const (&str)[N + 1])
{
for (std::size_t i = 0; i <= N; ++i)
data[i] = str[i];
}
constexpr
operator std::string_view() const noexcept
{
return {data, N};
}
};
/**
* Deduction guide: StaticStr from string literal.
*/
template <std::size_t N>
StaticStr(char const (&)[N]) -> StaticStr<N - 1>;
/**
* Create a StaticStr from a string literal.
*/
template <std::size_t N>
constexpr auto
makeStr(char const (&str)[N])
{
return StaticStr<N - 1>(str);
}
/**
* Concatenate two StaticStr values with a dot separator.
*/
template <std::size_t A, std::size_t B>
constexpr auto
join(StaticStr<A> const& lhs, StaticStr<B> const& rhs)
{
constexpr std::size_t len = A + 1 + B; // lhs + '.' + rhs
StaticStr<len> result;
std::size_t pos = 0;
for (std::size_t i = 0; i < A; ++i)
result.data[pos++] = lhs.data[i];
result.data[pos++] = '.';
for (std::size_t i = 0; i < B; ++i)
result.data[pos++] = rhs.data[i];
result.data[pos] = '\0';
return result;
}
// ===== Shared root segments ================================================
namespace seg {
inline constexpr auto xrpl = makeStr("xrpl");
inline constexpr auto rpc = makeStr("rpc");
inline constexpr auto tx = makeStr("tx");
inline constexpr auto consensus = makeStr("consensus");
inline constexpr auto peer = makeStr("peer");
inline constexpr auto ledger = makeStr("ledger");
inline constexpr auto network = makeStr("network");
inline constexpr auto link = makeStr("link");
} // namespace seg
// ===== Shared attribute keys (used across modules) =========================
namespace attr {
inline constexpr auto networkId = join(join(seg::xrpl, seg::network), makeStr("id"));
inline constexpr auto networkType = join(join(seg::xrpl, seg::network), makeStr("type"));
/**
* Canonical shared attrs (rule 5 — <domain>_<field> underscore form).
*
* Per the naming convention header note: shared cross-span attribute
* keys use the underscore form, reserving the dotted xrpl.<domain>.<field>
* form for resource attributes set on the OTel resource at startup.
* Defined once here, aliased by domain-specific headers. These are
* literal underscore-joined names, not dot-joined via `join()`, since
* `join()` always inserts `.` between its arguments.
*/
inline constexpr auto txHash = makeStr("tx_hash");
inline constexpr auto peerId = makeStr("peer_id");
inline constexpr auto ledgerSeq = makeStr("ledger_seq");
/**
* Shared close-time attrs — bare names, reused by consensus and ledger.
*
* `closeTimeRippleEpochS` carries a NetClock reading: whole seconds since the
* XRP Ledger epoch (2000-01-01T00:00:00Z), never the Unix epoch. The key names
* both the unit and the epoch because neither is recoverable from the value.
* A consumer rendering it as wall-clock time must first add kEpochOffset
* (946684800 seconds, see basics/chrono.h); read as a Unix timestamp instead,
* it lands roughly 30 years early.
*
* `closeResolutionMs` is a duration, not an instant — the granularity the
* close time is rounded to. NetClock resolution is whole seconds, so this
* value is always a multiple of 1000.
*/
inline constexpr auto closeTimeRippleEpochS = makeStr("close_time_ripple_epoch_s");
inline constexpr auto closeTimeCorrect = makeStr("close_time_correct");
inline constexpr auto closeResolutionMs = makeStr("close_resolution_ms");
/**
* Shared validation attrs — reused by the consensus and peer validation
* spans. Same concept, same key on every span; the span name tells them
* apart, so neither is emitter-prefixed. `ledgerHash` is a ledger-object
* property (bare, like ledgerSeq); `fullValidation` is the is-full-validation
* flag. Never the dotted xrpl. form (reserved for resource attrs).
*/
inline constexpr auto ledgerHash = makeStr("ledger_hash");
inline constexpr auto fullValidation = makeStr("full_validation");
/**
* Shared "ledger being worked on" attrs — the open/tentative or in-flight
* consensus-build ledger a transaction is applied into, NOT an established or
* validated ledger (that is `ledgerSeq`, set on ledger.build / consensus.round).
* Named after the RPC field `ledger_current_index` and the `currentLedgerSeq`
* log usage. Reused by the tx lifecycle, apply-pipeline, and TxQ spans so a
* transaction's work can be correlated to the ledger it targeted.
* `currentLedgerHash` is the current view's parent-ledger hash, which equals the
* consensus.round deterministic trace-id seed on the consensus-build path.
*/
inline constexpr auto currentLedgerSeq = makeStr("current_ledger_seq");
inline constexpr auto currentLedgerHash = makeStr("current_ledger_hash");
} // namespace attr
// ===== Shared attribute values =============================================
namespace attr_val {
inline constexpr auto success = makeStr("success");
inline constexpr auto error = makeStr("error");
} // namespace attr_val
} // namespace xrpl::telemetry

View File

@@ -1,495 +0,0 @@
#pragma once
/**
* Abstract interface for OpenTelemetry distributed tracing.
*
* Provides the Telemetry base class that all components use to create trace
* spans. Three concrete implementations exist, selected at construction time
* by makeTelemetry():
*
* - TelemetryImpl (Telemetry.cpp): real OTel SDK integration, compiled
* only when XRPL_ENABLE_TELEMETRY is defined and enabled at runtime.
* - NullTelemetry (NullTelemetry.cpp): no-op stub used when telemetry is
* disabled at compile time or runtime.
* - NullTelemetryOtel (Telemetry.cpp): no-op stub that still depends on
* the OTel API (used during transition or for testing).
*
* Inheritance / dependency diagram:
*
* +--------------------+
* | Telemetry | (abstract, this file)
* | <<interface>> |
* +---------+----------+
* |
* +---------+-----------+-------------------+
* | | |
* +---+------------+ +-----+---------+ +------+----------+
* | TelemetryImpl | | NullTelemetry | | NullTelemetryOtel|
* | (Telemetry.cpp)| |(NullTelemetry | | (Telemetry.cpp) |
* | OTel SDK | | .cpp) | | noop w/ OTel API |
* +----------------+ +---------------+ +------------------+
*
* The Setup struct holds all configuration parsed from the [telemetry]
* section of xrpld.cfg. See TelemetryConfig.cpp for the parser and
* cfg/xrpld-example.cfg for the available options.
*
* OTel SDK headers are conditionally included behind XRPL_ENABLE_TELEMETRY
* so that builds without telemetry have zero dependency on opentelemetry-cpp.
*
* Usage examples:
*
* 1. Root span at a subsystem entry point (typical usage):
* @code
* #include <xrpld/rpc/detail/RpcSpanNames.h>
* using namespace xrpl::telemetry;
*
* // In an RPC handler dispatch:
* auto guard = SpanGuard::span(
* TraceCategory::Rpc, rpc_span::prefix::command, commandName);
* guard.setAttribute(rpc_span::attr::command, commandName);
* // ... process request
* // guard destructor automatically ends the span on scope exit
* @endcode
*
* 2. Child span for a sub-operation (scoped child):
* @code
* auto parent = ScopedSpanGuard(
* TraceCategory::Rpc, rpc_span::prefix::rpc, rpc_span::op::process);
* {
* auto child = parent.childSpan(rpc_span::prefix::command);
* child.setAttribute(rpc_span::attr::version, static_cast<int64_t>(apiVersion));
* // child ends here
* }
* @endcode
* childSpan() parents to the ambient scope, so the parent must be a
* ScopedSpanGuard. A plain SpanGuard is not ambient: pass its spanContext()
* to childSpan(name, ctx) instead. childSpan() takes the name verbatim, so
* pass a full dotted constant, never a bare op:: suffix.
*
* 3. Unrelated span (cross-scope, same thread):
* @code
* // gRPC and RPC handlers can be active simultaneously
* auto grpcSpan = SpanGuard::span(
* TraceCategory::Rpc, grpc_span::prefix::grpc, grpc_span::attr::method);
* auto rpcSpan = SpanGuard::span(
* TraceCategory::Rpc, rpc_span::prefix::command, commandName);
* // both spans end on scope exit
* @endcode
*
* 4. Cross-thread context propagation:
* @code
* // Thread A: take a handle to the parent span's own context
* auto ctx = parentGuard.spanContext();
*
* // Thread B: create child span with explicit parent
* auto child = SpanGuard::childSpan(rpc_span::prefix::command, ctx);
* @endcode
*
* @note Thread safety: The Telemetry interface is safe for concurrent reads
* (isEnabled, shouldTrace*, getTracer, startSpan) after start() completes.
* setServiceInstanceId() must be called before start() and is not thread-safe.
* The OTel SDK's TracerProvider and Tracer are internally thread-safe.
*/
#include <xrpl/beast/utility/Journal.h>
#include <xrpl/config/BasicConfig.h>
#include <atomic>
#include <chrono>
#include <cstdint>
#include <memory>
#include <string>
#ifdef XRPL_ENABLE_TELEMETRY
#include <opentelemetry/context/context.h>
#include <opentelemetry/nostd/shared_ptr.h>
#include <opentelemetry/trace/span.h>
#include <opentelemetry/trace/span_metadata.h>
#include <opentelemetry/trace/tracer.h>
// std::string_view appears only in the telemetry-enabled declarations below.
#include <string_view>
#endif
namespace xrpl::telemetry {
#ifdef XRPL_ENABLE_TELEMETRY
/**
* OTel instrumentation scope (tracer) name. Identifies this library as the
* source of spans; distinct from the `service.name` resource attribute
* (Setup::serviceName), which is config-overridable.
*/
inline constexpr std::string_view kTracerName{"xrpld"};
#endif
/**
* How a consensus round span picks its trace id.
*
* consensus_trace_strategy (xrpld.cfg)
* |
* v
* makeTelemetrySetup() ──> Setup::consensusTraceStrategy
* |
* v
* RCLConsensus::Adaptor::startRoundTracing()
* |
* +-- Deterministic ──> SpanGuard::hashSpan(prev ledger hash)
* +-- Random ──> SpanGuard::span() / linkedSpan()
*
* Deterministic is the strategy in use. Every validator of a round hashes the
* same previous ledger id, so all of them land in one trace.
*
* Random is experimental and not used. Each node would invent its own trace
* id, so one round would arrive as one trace per node, joinable only by the
* `consensus_ledger_id` attribute.
*
* @code
* // Branch on the strategy rather than on a string.
* if (telemetry.getConsensusTraceStrategy() == ConsensusTraceStrategy::Random)
* span = SpanGuard::span(TraceCategory::Consensus, seg::consensus, op::round);
* else
* span = SpanGuard::hashSpan(TraceCategory::Consensus, name, id.data(), id.kBytes);
*
* // Edge case: the value also goes on a span attribute, so it needs its
* // config spelling back.
* span.setAttribute(attr::traceStrategy, strategyName(ConsensusTraceStrategy::Random));
* @endcode
*
* @note Adding an enumerator means adding a spelling to strategyName() below
* and to the parser in TelemetryConfig.cpp. Both switch without a default, so
* the compiler catches a missed one.
*/
enum class ConsensusTraceStrategy : std::uint8_t { Deterministic, Random };
/**
* Config spelling of a consensus trace strategy.
*
* This is the same text `consensus_trace_strategy` accepts, and it is what
* goes on the `trace_strategy` span attribute, so the two cannot drift.
*
* @param strategy Strategy to name.
* @return "deterministic" or "random", pointing at a string literal.
*/
[[nodiscard]] constexpr char const*
strategyName(ConsensusTraceStrategy strategy)
{
switch (strategy)
{
case ConsensusTraceStrategy::Deterministic:
return "deterministic";
case ConsensusTraceStrategy::Random:
return "random";
}
// The switch covers every enumerator. This return only satisfies the
// compiler, which cannot rule out a value outside the enumeration.
return "deterministic";
}
class Telemetry
{
/**
* Global singleton pointer, set by start()/stop() in the active
* implementation. Allows SpanGuard factory methods to access the
* Telemetry instance without callers passing it explicitly.
*
* Atomic with acquire/release ordering: start()/stop() store on
* the initialization thread, factory methods load on worker threads.
* @see setInstance(), getInstance()
*/
inline static std::atomic<Telemetry*> instance{nullptr};
public:
/**
* Get the global Telemetry instance.
* @return Pointer to the active instance, or nullptr if not started.
*/
[[nodiscard]] static Telemetry*
getInstance()
{
return instance.load(std::memory_order_acquire);
}
/**
* Set the global Telemetry instance.
* Called by start()/stop() in concrete implementations.
* Tests can call this with a mock to override the global instance.
* @param t Pointer to the Telemetry instance, or nullptr to clear.
*/
static void
setInstance(Telemetry* t)
{
instance.store(t, std::memory_order_release);
}
/**
* Configuration parsed from the [telemetry] section of xrpld.cfg.
*
* All fields have sensible defaults so the section can be minimal
* or omitted entirely. See TelemetryConfig.cpp for the parser.
*/
struct Setup
{
/**
* Master switch: true to enable tracing at runtime.
*/
bool enabled = false;
/**
* OTel resource attribute `service.name`.
*/
std::string serviceName = "xrpld";
/**
* OTel resource attribute `service.version` (set from BuildInfo).
*/
std::string serviceVersion;
/**
* OTel resource attribute `service.instance.id` (defaults to node
* public key).
*/
std::string serviceInstanceId;
/**
* Full OTLP/HTTP URL where spans are sent, including the signal path.
* Used verbatim: no other endpoint is derived from it.
*/
std::string tracesEndpoint = "http://localhost:4318/v1/traces";
/**
* Whether to use TLS for the exporter connection.
*/
bool useTls = false;
/**
* Path to a CA certificate bundle for TLS verification.
*/
std::string tlsCertPath;
/**
* Head-based sampling ratio. Intentionally fixed at 1.0 (sample
* everything) and NOT read from config. A per-node ratio would let
* nodes make divergent keep/drop decisions for the same distributed
* trace, producing broken/partial traces. makeHeadSampler() applies
* this ratio to spans with no parent or a remote parent, so a peer's
* sampled flag cannot turn our spans off or on; spans with a local
* parent follow it. Volume reduction is delegated to the collector's
* tail sampling; for node-local post-hoc dropping see
* SpanGuard::discard().
*/
static constexpr double samplingRatio = 1.0;
/**
* Maximum number of spans per batch export.
*/
std::uint32_t batchSize = 512;
/**
* Delay between batch exports.
*/
std::chrono::milliseconds batchDelay = std::chrono::milliseconds{5000};
/**
* Maximum number of spans queued before dropping.
*/
std::uint32_t maxQueueSize = 2048;
/**
* Network identifier, added as an OTel resource attribute.
*/
std::uint32_t networkId = 0;
/**
* Network type label (e.g. "mainnet", "testnet", "devnet").
*/
std::string networkType = "mainnet";
/**
* Enable tracing for transaction processing.
*/
bool traceTransactions = true;
/**
* Enable tracing for consensus rounds.
*/
bool traceConsensus = true;
/**
* Enable tracing for RPC request handling.
*/
bool traceRpc = true;
/**
* Enable tracing for peer-to-peer messages (enabled by default;
* high volume).
*/
bool tracePeer = true;
/**
* Enable tracing for ledger close/accept.
*/
bool traceLedger = true;
/**
* How a consensus round span picks its trace id.
*
* Read from `consensus_trace_strategy`. Deterministic is the strategy
* in use; Random is experimental. See ConsensusTraceStrategy.
*/
ConsensusTraceStrategy consensusTraceStrategy = ConsensusTraceStrategy::Deterministic;
};
virtual ~Telemetry() = default;
/**
* Update the service instance ID (OTel resource attribute
* `service.instance.id`).
*
* Must be called before start(). The node public key is not available
* when Telemetry is constructed (during the ApplicationImp member
* initializer list), so this setter allows Application::setup() to
* inject the identity once nodeIdentity_ is known.
*
* @param id The node's base58-encoded public key or custom identifier.
*/
virtual void
setServiceInstanceId([[maybe_unused]] std::string const& id)
{
// Default no-op for NullTelemetry implementations.
}
/**
* Initialize the tracing pipeline (exporter, processor, provider).
* Call after construction.
*/
virtual void
start() = 0;
/**
* Flush pending spans and shut down the tracing pipeline.
* Call before destruction.
*/
virtual void
stop() = 0;
/**
* @return true if this instance is actively exporting spans.
*/
[[nodiscard]] virtual bool
isEnabled() const = 0;
/**
* @return true if transaction processing should be traced.
*/
[[nodiscard]] virtual bool
shouldTraceTransactions() const = 0;
/**
* @return true if consensus rounds should be traced.
*/
[[nodiscard]] virtual bool
shouldTraceConsensus() const = 0;
/**
* @return true if RPC request handling should be traced.
*/
[[nodiscard]] virtual bool
shouldTraceRpc() const = 0;
/**
* @return true if peer-to-peer messages should be traced.
*/
[[nodiscard]] virtual bool
shouldTracePeer() const = 0;
/**
* @return true if ledger close/accept should be traced.
*/
[[nodiscard]] virtual bool
shouldTraceLedger() const = 0;
/**
* @return How a consensus round span picks its trace id.
*/
[[nodiscard]] virtual ConsensusTraceStrategy
getConsensusTraceStrategy() const = 0;
#ifdef XRPL_ENABLE_TELEMETRY
/**
* Get or create a named tracer instance.
*
* @param name Tracer name used to identify the instrumentation library.
* @return A shared pointer to the Tracer.
*/
[[nodiscard]] virtual opentelemetry::nostd::shared_ptr<opentelemetry::trace::Tracer>
getTracer(std::string_view name = kTracerName) = 0;
/**
* Start a new span on the current thread's context.
*
* The span becomes a child of the current active span (if any) via
* OpenTelemetry's context propagation.
*
* @param name Span name (typically "rpc.command.<cmd>").
* @param kind The span kind (defaults to kInternal). Possible values:
* - kInternal: default, in-process operation
* - kServer: incoming synchronous request (e.g. RPC)
* - kClient: outgoing synchronous request
* - kProducer: async message send (e.g. peer broadcast)
* - kConsumer: async message receive
* @return A shared pointer to the new Span.
*/
[[nodiscard]] virtual opentelemetry::nostd::shared_ptr<opentelemetry::trace::Span>
startSpan(
std::string_view name,
opentelemetry::trace::SpanKind kind = opentelemetry::trace::SpanKind::kInternal) = 0;
/**
* Start a new span with an explicit parent context.
*
* Use this overload when the parent span is not on the current
* thread's context stack (e.g. cross-thread trace propagation).
*
* @param name Span name.
* @param parentContext The parent span's context.
* @param kind The span kind (defaults to kInternal).
* @return A shared pointer to the new Span.
*/
[[nodiscard]] virtual opentelemetry::nostd::shared_ptr<opentelemetry::trace::Span>
startSpan(
std::string_view name,
opentelemetry::context::Context const& parentContext,
opentelemetry::trace::SpanKind kind = opentelemetry::trace::SpanKind::kInternal) = 0;
#endif
};
/**
* Create a Telemetry instance.
*
* Returns a TelemetryImpl when setup.enabled is true, or a
* NullTelemetry no-op stub otherwise.
*
* @param setup Configuration from the [telemetry] config section.
* @param journal Journal for log output during initialization.
*/
std::unique_ptr<Telemetry>
makeTelemetry(Telemetry::Setup const& setup, beast::Journal journal);
/**
* Parse the [telemetry] config section into a Setup struct.
*
* @param section The [telemetry] config section.
* @param nodePublicKey Node public key, used as default instance ID.
* @param version Build version string.
* @param networkId Network identifier from [network_id] config
* (0 = mainnet, 1 = testnet, 2 = devnet).
* @return A populated Setup struct with defaults for missing values.
*/
Telemetry::Setup
makeTelemetrySetup(
Section const& section,
std::string const& nodePublicKey,
std::string const& version,
std::uint32_t networkId);
} // namespace xrpl::telemetry

View File

@@ -1,116 +0,0 @@
#pragma once
/**
* Utilities for trace context propagation across nodes.
*
* Provides serialization/deserialization of OTel trace context to/from
* Protocol Buffer TraceContext messages (P2P cross-node propagation).
* Wired into the P2P message flow via PropagationHelpers.h for
* TMTransaction, TMProposeSet, and TMValidation messages.
*
* Only compiled when XRPL_ENABLE_TELEMETRY is defined.
*
* @see PropagationHelpers.h (high-level inject helpers),
* TxTracing.h (transaction receive-side extraction),
* ConsensusReceiveTracing.h (proposal/validation receive-side).
*/
#ifdef XRPL_ENABLE_TELEMETRY
#include <xrpl/proto/xrpl.pb.h>
#include <xrpl/telemetry/TraceContextValidation.h>
#include <opentelemetry/context/context.h>
#include <opentelemetry/nostd/shared_ptr.h>
#include <opentelemetry/nostd/span.h>
#include <opentelemetry/trace/context.h>
#include <opentelemetry/trace/default_span.h>
#include <opentelemetry/trace/span.h>
#include <opentelemetry/trace/span_context.h>
#include <opentelemetry/trace/span_id.h>
#include <opentelemetry/trace/span_metadata.h>
#include <opentelemetry/trace/trace_flags.h>
#include <opentelemetry/trace/trace_id.h>
#include <cstdint>
namespace xrpl::telemetry {
// The wire limits in TraceContextValidation.h must match the OTel types the
// peer's bytes are copied into here.
static_assert(kTraceIdSize == opentelemetry::trace::TraceId::kSize);
static_assert(kSpanIdSize == opentelemetry::trace::SpanId::kSize);
static_assert(kKnownTraceFlags == opentelemetry::trace::TraceFlags::kAllW3CTraceContext2Flags);
/**
* Extract OTel context from a protobuf TraceContext message.
*
* @param proto The protobuf TraceContext received from a peer.
* @return An OTel Context with the extracted parent span, or an empty
* context if the protobuf fields are missing or invalid.
*/
[[nodiscard]] inline opentelemetry::context::Context
extractFromProtobuf(protocol::TraceContext const& proto)
{
namespace trace = opentelemetry::trace;
// Reject a malformed context (bad ids or flags) from the peer before
// trusting it as a parent. See TraceContextValidation.h.
if (!isValidTraceContext(proto))
{
return opentelemetry::context::Context{};
}
auto const* rawTraceId = reinterpret_cast<std::uint8_t const*>(proto.trace_id().data());
auto const* rawSpanId = reinterpret_cast<std::uint8_t const*>(proto.span_id().data());
trace::TraceId const traceId(
opentelemetry::nostd::span<std::uint8_t const, kTraceIdSize>(rawTraceId, kTraceIdSize));
trace::SpanId const spanId(
opentelemetry::nostd::span<std::uint8_t const, kSpanIdSize>(rawSpanId, kSpanIdSize));
trace::TraceFlags const flags(traceFlagsByte(proto));
trace::SpanContext const spanCtx(traceId, spanId, flags, /* remote = */ true);
return opentelemetry::context::Context{}.SetValue(
trace::kSpanKey,
opentelemetry::nostd::shared_ptr<trace::Span>(new trace::DefaultSpan(spanCtx)));
}
/**
* Inject the current span's trace context into a protobuf TraceContext.
*
* @param ctx The OTel context containing the span to propagate.
* @param proto The protobuf TraceContext to populate.
*/
inline void
injectToProtobuf(opentelemetry::context::Context const& ctx, protocol::TraceContext& proto)
{
namespace trace = opentelemetry::trace;
auto const span = trace::GetSpan(ctx);
if (!span)
return;
auto const& spanCtx = span->GetContext();
if (!spanCtx.IsValid())
return;
// Serialize trace_id (16 bytes)
auto const& traceId = spanCtx.trace_id();
proto.set_trace_id(traceId.Id().data(), trace::TraceId::kSize);
// Serialize span_id (8 bytes)
auto const& spanId = spanCtx.span_id();
proto.set_span_id(spanId.Id().data(), trace::SpanId::kSize);
// Serialize flags
proto.set_trace_flags(spanCtx.trace_flags().flags());
// TODO: add a trace_state field to the protobuf TraceContext (field 4
// is reserved for it in xrpl.proto) with a size limit, then write it
// here and read it in extractFromProtobuf above.
}
} // namespace xrpl::telemetry
#endif // XRPL_ENABLE_TELEMETRY

View File

@@ -1,235 +0,0 @@
#pragma once
/**
* Validation and clean-up of peer-supplied trace context.
*
* A protobuf TraceContext arrives inside untrusted peer messages
* (TMTransaction, TMProposeSet, TMValidation). A peer can send a malformed
* or all-zero trace_id/span_id, or trace_flags that do not fit 8 bits,
* either by accident or to pollute traces. This header is the single place
* those checks live, so every receive site agrees on what "valid" means.
*
* Validity follows the OpenTelemetry spec and W3C Trace Context: a trace_id
* is 16 bytes, a span_id is 8 bytes, an all-zero id is invalid, and
* trace_flags is an 8-bit field in which only the sampled and random bits
* are defined.
*
* Dependency diagram:
*
* P2P message parsed (ProtocolMessage.h)
* |
* v
* sanitizeTraceContext() --- invalid ---> trace_context cleared
* |
* valid (undefined flag bits zeroed)
* |
* v
* receive sites: isValidTraceContext(), or isValidSpanId()
* and isSameTraceId() where the trace_id is derived locally
* |
* +--- valid ----> build child span from peer context
* |
* +--- invalid --> ignore peer context, start fresh span
*
* This header depends only on the generated protobuf type. It pulls in no
* OpenTelemetry headers, so the parse-time check also runs in builds
* without telemetry, and receive sites that use SpanGuard keep SpanGuard's
* encapsulation of OTel types.
*
* @note Thread-safe: the predicates are pure, and sanitizeTraceContext()
* changes only the message it is given.
*
* Usage:
* @code
* // Once per parsed message, before any handler sees it:
* sanitizeTraceContext(*msg);
*
* // Full parent context (both ids come from the peer):
* if (isValidTraceContext(msg->trace_context()))
* buildChildSpan(...);
*
* // Local trace_id (derived from a hash, not taken from the peer).
* // The peer's span is the parent only inside that same trace:
* if (tc.has_span_id() && isValidSpanId(tc.span_id()) &&
* isSameTraceId(tc.trace_id(), localTraceId))
* buildChildSpan(..., traceFlagsByte(tc));
* @endcode
*/
#include <xrpl/proto/xrpl.pb.h>
#include <algorithm>
#include <cstddef>
#include <cstdint>
#include <limits>
#include <span>
#include <string_view>
namespace xrpl::telemetry {
/**
* Size in bytes of a trace_id.
*/
inline constexpr std::size_t kTraceIdSize = 16;
/**
* Size in bytes of a span_id.
*/
inline constexpr std::size_t kSpanIdSize = 8;
/**
* Largest valid trace_flags value. W3C trace-flags is an 8-bit field.
*/
inline constexpr std::uint32_t kMaxTraceFlags = std::numeric_limits<std::uint8_t>::max();
/**
* The trace_flags bits kept from a peer: sampled (bit 0) and random
* (bit 1). W3C Trace Context says the other bits must be zero.
*/
inline constexpr std::uint32_t kKnownTraceFlags = 0x03;
/**
* True if the bytes are a valid OTel trace_id: kTraceIdSize bytes, not all
* zero.
*
* @param traceId The raw trace_id bytes from a protobuf TraceContext.
* @return true if usable as a trace identifier, false otherwise.
*/
[[nodiscard]] inline bool
isValidTraceId(std::string_view traceId)
{
return traceId.size() == kTraceIdSize &&
std::ranges::any_of(traceId, [](char c) { return c != 0; });
}
/**
* True if the bytes are a valid OTel span_id: kSpanIdSize bytes, not all
* zero.
*
* @param spanId The raw span_id bytes from a protobuf TraceContext.
* @return true if usable as a span identifier, false otherwise.
*/
[[nodiscard]] inline bool
isValidSpanId(std::string_view spanId)
{
return spanId.size() == kSpanIdSize &&
std::ranges::any_of(spanId, [](char c) { return c != 0; });
}
/**
* True if trace_flags is absent or fits the 8-bit W3C field.
*
* @param tc The protobuf TraceContext received from a peer.
* @return false only when trace_flags is set above kMaxTraceFlags.
*/
[[nodiscard]] inline bool
isValidTraceFlags(protocol::TraceContext const& tc)
{
return !tc.has_trace_flags() || tc.trace_flags() <= kMaxTraceFlags;
}
/**
* True if the context carries a usable parent: a valid trace_id, a valid
* span_id, and trace_flags that fit 8 bits.
*
* Use this where both ids are taken from the peer (consensus receive,
* generic extraction). The transaction path derives its trace_id locally
* from the txID, so it checks isValidSpanId() and isSameTraceId() instead.
*
* @param tc The protobuf TraceContext received from a peer.
* @return true if both ids are present and valid and the flags fit, false
* otherwise.
*/
[[nodiscard]] inline bool
isValidTraceContext(protocol::TraceContext const& tc)
{
return tc.has_trace_id() && isValidTraceId(tc.trace_id()) && tc.has_span_id() &&
isValidSpanId(tc.span_id()) && isValidTraceFlags(tc);
}
/**
* True if a peer's trace_id equals the trace_id of the span being started.
*
* A parent and its child share one trace. A receive site that derives its
* trace_id locally takes the peer's span as the parent only when this holds.
*
* @param traceId The raw trace_id bytes from a protobuf TraceContext. An
* absent field reads as empty.
* @param localTraceId The trace_id of the span being started.
* @return true if traceId is exactly the kTraceIdSize bytes of
* localTraceId, false otherwise.
*/
[[nodiscard]] inline bool
isSameTraceId(std::string_view traceId, std::span<std::byte const, kTraceIdSize> localTraceId)
{
return std::ranges::equal(std::as_bytes(std::span(traceId)), localTraceId);
}
/**
* The trace_flags byte to pass to OpenTelemetry, with undefined bits
* cleared.
*
* @param tc The protobuf TraceContext received from a peer.
* @return trace_flags masked with kKnownTraceFlags, or 0 when absent.
*/
[[nodiscard]] inline std::uint8_t
traceFlagsByte(protocol::TraceContext const& tc)
{
if (!tc.has_trace_flags())
return 0;
return static_cast<std::uint8_t>(tc.trace_flags() & kKnownTraceFlags);
}
/**
* A protobuf message with an optional trace_context field.
*/
template <class Message>
concept HasTraceContext = requires(Message& msg) {
msg.has_trace_context();
msg.trace_context();
msg.clear_trace_context();
msg.mutable_trace_context();
};
/**
* Clean the trace context of a message received from a peer.
*
* Drops the whole context when isValidTraceContext() rejects it, and
* zeroes the undefined trace_flags bits otherwise. Call it once, when the
* message is parsed, so every handler and every relay of the message sees
* the cleaned context.
*
* @param msg A message just parsed from a peer. A message without a
* trace_context is left without one.
*/
template <HasTraceContext Message>
void
sanitizeTraceContext(Message& msg)
{
if (!msg.has_trace_context())
return;
auto const& tc = msg.trace_context();
if (!isValidTraceContext(tc))
{
msg.clear_trace_context();
return;
}
if (tc.has_trace_flags())
msg.mutable_trace_context()->set_trace_flags(tc.trace_flags() & kKnownTraceFlags);
}
/**
* Clean the trace context of every transaction in a batch.
*
* @param batch A TMTransactions message just parsed from a peer.
*/
inline void
sanitizeTraceContext(protocol::TMTransactions& batch)
{
for (auto& tx : *batch.mutable_transactions())
sanitizeTraceContext(tx);
}
} // namespace xrpl::telemetry

View File

@@ -1,172 +0,0 @@
#pragma once
/**
* Span attribute keys for the account-typed fields of a transaction.
*
* A transaction names one or more accounts: the sender in `Account`, and
* depending on the type a `Destination`, `Owner`, `Issuer`, `Holder` and so
* on. The tx.process span emits every one it finds as its own attribute, so
* an account can be searched for in traces whatever role it played. An
* account address is a public ledger identifier, so each is emitted as the
* raw r-address and never hashed.
*
* One key per protocol field: `tx_` followed by the field's JSON name in
* lower snake case. The full table is the initializer in
* src/libxrpl/telemetry/TxAccountSpanNames.cpp; the common ones are
*
* STTx field span attribute key
* ----------------- ------------------
* Account tx_account
* Destination tx_destination
* Owner tx_owner
* Issuer tx_issuer
* RegularKey tx_regular_key
* NFTokenMinter tx_nftoken_minter
*
* Only fields that some transaction format carries at top level have a key.
* Account-typed fields that appear only in ledger entries or inner objects
* (LowSponsor, LockingChainDoor, ...) map to nullopt. A library test walks
* TxFormats and fails when a format gains an account field with no key.
*
* Why this header lives in libxrpl rather than beside TxSpanNames.h: the
* mapping is keyed by protocol fields and its completeness is checked from
* TxFormats, which a library test can reach and a daemon header cannot.
*
* Data flow:
*
* NetworkOPs::processTransaction (src/xrpld)
* │ for each top-level field with getSType() == STI_ACCOUNT
* ▼
* accountFieldAttributeKey(field.getFName()) (this header)
* │ the key, or nullopt for a field with no key
* ▼
* span->setAttribute(key, field.getText())
*
* @code
* // Primary use: emit every account the transaction names. An empty
* // account field is skipped so it is not rendered as the zero address.
* for (auto const& field : stx)
* {
* if (field.getSType() != STI_ACCOUNT || field.isDefault())
* continue;
* if (auto const key = telemetry::accountFieldAttributeKey(field.getFName()))
* span.setAttribute(*key, toBase58(stx.getAccountID(field.getFName())));
* }
* @endcode
*
* @code
* // Edge case: a field that is not a top-level transaction account has
* // no key, so a caller must test the optional before using it.
* accountFieldAttributeKey(sfFee); // == std::nullopt
* accountFieldAttributeKey(sfLowSponsor); // == std::nullopt
* @endcode
*
* @note Only top-level fields are covered. Accounts nested in Signers, in a
* Batch's inner transactions, or as the issuer inside an Amount are not
* emitted.
* @note accountFieldAttributeKey() is thread-safe. Its table is built once
* on first use and is read-only afterwards.
*/
#include <xrpl/telemetry/SpanNames.h>
#include <optional>
#include <string_view>
namespace xrpl {
class SField;
} // namespace xrpl
namespace xrpl::telemetry {
namespace tx_account_span::attr {
/**
* "tx_account" — the sending account (`Account`). Every transaction has one.
*/
inline constexpr auto account = makeStr("tx_account");
/**
* "tx_destination" — the receiving account (`Destination`).
*/
inline constexpr auto destination = makeStr("tx_destination");
/**
* "tx_owner" — the owner of the object acted on (`Owner`).
*/
inline constexpr auto owner = makeStr("tx_owner");
/**
* "tx_issuer" — the issuer named by the transaction (`Issuer`).
*/
inline constexpr auto issuer = makeStr("tx_issuer");
/**
* "tx_authorize" — the account being authorised (`Authorize`).
*/
inline constexpr auto authorize = makeStr("tx_authorize");
/**
* "tx_unauthorize" — the account whose authorisation is removed (`Unauthorize`).
*/
inline constexpr auto unauthorize = makeStr("tx_unauthorize");
/**
* "tx_regular_key" — the regular key being set (`RegularKey`).
*/
inline constexpr auto regularKey = makeStr("tx_regular_key");
/**
* "tx_nftoken_minter" — the authorised NFToken minter (`NFTokenMinter`).
*/
inline constexpr auto nftokenMinter = makeStr("tx_nftoken_minter");
/**
* "tx_holder" — the token holder acted on (`Holder`).
*/
inline constexpr auto holder = makeStr("tx_holder");
/**
* "tx_delegate" — the delegate signing on the sender's behalf (`Delegate`).
*/
inline constexpr auto delegate = makeStr("tx_delegate");
/**
* "tx_sponsor" — the account paying the fee or reserve (`Sponsor`).
*/
inline constexpr auto sponsor = makeStr("tx_sponsor");
/**
* "tx_sponsee" — the account being sponsored (`Sponsee`).
*/
inline constexpr auto sponsee = makeStr("tx_sponsee");
/**
* "tx_counterparty" — the other party to a loan (`Counterparty`).
*/
inline constexpr auto counterparty = makeStr("tx_counterparty");
/**
* "tx_counterparty_sponsor" — the counterparty's sponsor (`CounterpartySponsor`).
*/
inline constexpr auto counterpartySponsor = makeStr("tx_counterparty_sponsor");
/**
* "tx_subject" — the subject of a credential (`Subject`).
*/
inline constexpr auto subject = makeStr("tx_subject");
/**
* "tx_other_chain_source" — the source account on the other chain (`OtherChainSource`).
*/
inline constexpr auto otherChainSource = makeStr("tx_other_chain_source");
/**
* "tx_other_chain_destination" — destination on the other chain (`OtherChainDestination`).
*/
inline constexpr auto otherChainDestination = makeStr("tx_other_chain_destination");
/**
* "tx_attestation_signer_account" — the attestation signer (`AttestationSignerAccount`).
*/
inline constexpr auto attestationSignerAccount = makeStr("tx_attestation_signer_account");
/**
* "tx_attestation_reward_account" — attestation reward account (`AttestationRewardAccount`).
*/
inline constexpr auto attestationRewardAccount = makeStr("tx_attestation_reward_account");
} // namespace tx_account_span::attr
/**
* Look up the span attribute key for an account-typed transaction field.
*
* @param field The protocol field, as returned by STBase::getFName().
* @return The `tx_*` key for a top-level transaction account field, or
* nullopt when the field is not account-typed or is carried only by ledger
* entries and inner objects.
*/
[[nodiscard]] std::optional<std::string_view>
accountFieldAttributeKey(SField const& field);
} // namespace xrpl::telemetry

View File

@@ -51,7 +51,8 @@ public:
beast::Journal journal = beast::Journal{beast::Journal::getNullSink()})
: ApplyContext(registry, base, std::nullopt, tx, preclaimResult, baseFee, flags, journal)
{
XRPL_ASSERT((flags & TapBatch) == 0, "Batch apply flag should not be set");
XRPL_ASSERT(
(flags & ApplyFlags::Batch) == ApplyFlags::None, "Batch apply flag should not be set");
}
std::reference_wrapper<ServiceRegistry> registry;

View File

@@ -59,10 +59,11 @@ public:
, parentBatchId(parentBatchId)
, j(j)
{
XRPL_ASSERT((flags & TapBatch) == TapBatch, "Batch apply flag should be set");
XRPL_ASSERT(
(flags & ApplyFlags::Batch) == ApplyFlags::Batch, "Batch apply flag should be set");
XRPL_ASSERT_IF(
(flags & TapProposal) != TapNone,
(flags & TapDryRun) != TapNone,
(flags & ApplyFlags::Proposal) != ApplyFlags::None,
(flags & ApplyFlags::DryRun) != ApplyFlags::None,
"xrpl::PreflightContext : proposal preflight implies dry run");
}
@@ -74,10 +75,11 @@ public:
beast::Journal j = beast::Journal{beast::Journal::getNullSink()})
: registry(registry), tx(tx), rules(std::move(rules)), flags(flags), j(j)
{
XRPL_ASSERT((flags & TapBatch) == 0, "Batch apply flag should not be set");
XRPL_ASSERT(
(flags & ApplyFlags::Batch) == ApplyFlags::None, "Batch apply flag should not be set");
XRPL_ASSERT_IF(
(flags & TapProposal) != TapNone,
(flags & TapDryRun) != TapNone,
(flags & ApplyFlags::Proposal) != ApplyFlags::None,
(flags & ApplyFlags::DryRun) != ApplyFlags::None,
"xrpl::PreflightContext : proposal preflight implies dry run");
}
@@ -116,7 +118,7 @@ public:
, j(j)
{
XRPL_ASSERT(
parentBatchId.has_value() == ((flags & TapBatch) == TapBatch),
parentBatchId.has_value() == ((flags & ApplyFlags::Batch) == ApplyFlags::Batch),
"Parent Batch ID should be set if batch apply flag is set");
}
@@ -129,7 +131,8 @@ public:
beast::Journal j = beast::Journal{beast::Journal::getNullSink()})
: PreclaimContext(registry, view, preflightResult, tx, flags, std::nullopt, j)
{
XRPL_ASSERT((flags & TapBatch) == 0, "Batch apply flag should not be set");
XRPL_ASSERT(
(flags & ApplyFlags::Batch) == ApplyFlags::None, "Batch apply flag should not be set");
}
PreclaimContext&

View File

@@ -41,7 +41,7 @@ struct ApplyResult
inline bool
isTecClaimHardFail(TER ter, ApplyFlags flags)
{
return isTecClaim(ter) && ((flags & TapRetry) == 0u);
return isTecClaim(ter) && ((flags & ApplyFlags::Retry) == ApplyFlags::None);
}
/**

View File

@@ -1,150 +0,0 @@
#pragma once
/**
* Compile-time span name constants for the transaction apply pipeline.
*
* Defines the span names and attribute keys used by the three apply-pipeline
* stages — preflight, preclaim, and transactor (apply) — that run inside the
* library (`src/libxrpl/tx/`). Built on the StaticStr/join() primitives from
* <xrpl/telemetry/SpanNames.h>.
*
* Why a separate header from TxSpanNames.h:
* TxSpanNames.h lives under src/xrpld/ (daemon) and serves the overlay/app
* lifecycle spans (tx.receive, tx.process). Library code (applySteps.cpp,
* Transactor.cpp) must not depend on daemon headers, so the apply-pipeline
* constants live here instead. The attribute strings ("tx_type",
* "ter_result", "applied") intentionally match TxSpanNames.h so the collector
* spanmetrics connector aggregates both sets under the same dimensions.
*
* Span hierarchy (deterministic trace_id derived from txID[0:16]):
*
* The three stages run sequentially and often on different threads, so they
* do not auto-parent. Each uses a hash-derived trace_id keyed on the same
* transaction id, placing all three under one trace without context
* propagation. A transaction that hard-fails preflight or preclaim never
* reaches the transactor span — the stage attribute identifies where it
* stopped.
*
* +-----------------------------------------------------------+
* | trace_id = txID[0:16] |
* | |
* | +-------------------+ +------------------+ +-------+ |
* | | tx.preflight | | tx.preclaim | | tx. | |
* | | stage=preflight |-->| stage=preclaim |-->| trans | |
* | | tx_type | | tx_type | | actor | |
* | | ter_result | | ter_result | | stage=| |
* | +-------------------+ +------------------+ | apply | |
* | stateless checks ledger-aware checks +-------+ |
* | (signature, fields) (sequence, fee) applies |
* +-----------------------------------------------------------+
*
* Usage:
* @code
* #include <xrpl/tx/detail/TxApplySpanNames.h>
* using namespace telemetry;
*
* // preflight() / preclaim() use hashSpan with a full span name:
* auto span = SpanGuard::hashSpan(
* TraceCategory::Transactions, tx_apply_span::preflight,
* txID.data(), txID.kBytes);
* span.setAttribute(tx_apply_span::attr::stage, tx_apply_span::val::preflight);
* span.setAttribute(tx_apply_span::attr::terResult, transToken(ter).c_str());
* @endcode
*
* @code
* // Transactor::operator() also uses hashSpan on the same txID so it
* // co-traces with preflight and preclaim under one trace_id:
* auto span = SpanGuard::hashSpan(
* TraceCategory::Transactions, tx_apply_span::transactor,
* txID.data(), txID.kBytes);
* span.setAttribute(tx_apply_span::attr::stage, tx_apply_span::val::apply);
* @endcode
*/
#include <xrpl/telemetry/SpanNames.h>
namespace xrpl::telemetry::tx_apply_span {
// ===== Span operation suffixes =============================================
namespace op {
/**
* "preflight" — stateless transaction checks (suffix form).
*/
inline constexpr auto preflight = makeStr("preflight");
/**
* "preclaim" — ledger-aware checks before fee claim (suffix form).
*/
inline constexpr auto preclaim = makeStr("preclaim");
/**
* "transactor" — the apply stage (suffix form, used with span()).
*/
inline constexpr auto transactor = makeStr("transactor");
} // namespace op
// ===== Full span names (tx.<op>) ===========================================
/**
* "tx.preflight" — full name for hashSpan() at the preflight stage.
*/
inline constexpr auto preflight = join(seg::tx, op::preflight);
/**
* "tx.preclaim" — full name for hashSpan() at the preclaim stage.
*/
inline constexpr auto preclaim = join(seg::tx, op::preclaim);
/**
* "tx.transactor" — full name for hashSpan() at the apply stage. Shares the
* txID-derived trace_id so it co-traces with tx.preflight and tx.preclaim.
*/
inline constexpr auto transactor = join(seg::tx, op::transactor);
// ===== Attribute keys ======================================================
namespace attr {
/**
* Shared "ledger being worked on" attrs (defined in SpanNames.h). Set on
* tx.preclaim and tx.transactor (both run against a view whose seq() is the
* ledger being applied into). tx.preflight is stateless (no view) and is the
* documented exception — it carries neither.
*/
using ::xrpl::telemetry::attr::currentLedgerHash;
using ::xrpl::telemetry::attr::currentLedgerSeq;
/**
* "stage" — which apply-pipeline stage this span represents. Drives the
* collector spanmetrics `stage` dimension for per-stage RED metrics.
*/
inline constexpr auto stage = makeStr("stage");
/**
* "tx_type" — transaction type name (e.g., "Payment", "OfferCreate").
* Matches tx_span::attr::txType so both share the spanmetrics dimension.
*/
inline constexpr auto txType = makeStr("tx_type");
/**
* "ter_result" — engine result code after the stage (e.g., "tesSUCCESS").
*/
inline constexpr auto terResult = makeStr("ter_result");
/**
* "applied" — whether the transaction was applied to the ledger (apply only).
*/
inline constexpr auto applied = makeStr("applied");
} // namespace attr
// ===== Attribute values (stage names) ======================================
namespace val {
/**
* "preflight" — value of the stage attribute on tx.preflight.
*/
inline constexpr auto preflight = makeStr("preflight");
/**
* "preclaim" — value of the stage attribute on tx.preclaim.
*/
inline constexpr auto preclaim = makeStr("preclaim");
/**
* "apply" — value of the stage attribute on tx.transactor.
*/
inline constexpr auto apply = makeStr("apply");
} // namespace val
} // namespace xrpl::telemetry::tx_apply_span

View File

@@ -2,144 +2,144 @@ Detected OS: macos (Darwin arm64)
Core build tools:
✅ cmake
cmake version 4.1.2
/nix/store/gvabsb4yqb5xsqzqph54rijnn4zpihnp-cmake-4.1.2/bin/cmake
cmake version 4.4.3
/nix/store/q85csxf4s4shx89zif097h1ql4ax1rrz-cmake-4.4.3/bin/cmake
✅ conan
Conan version 2.28.1
/nix/store/9jiyxmkpwmn6dcqs0765s83riw3l5ail-conan-2.28.1/bin/conan
Conan version 2.32.0
/nix/store/921jqsgbilixmr3xchj95si0jl8pi13z-conan-2.32.0/bin/conan
✅ git
git version 2.54.0
/nix/store/a14yxcqvv9x2l9mllgpirzhvz93pgprg-git-2.54.0/bin/git
git version 2.55.0
/nix/store/gw7c7m0dwca5lg4152x9lpmxpj171bzw-git-2.55.0/bin/git
✅ python3
Python 3.13.13
/nix/store/ygxqin6ydzjfawywqpp5pal8wv6sf5bh-python3-3.13.13/bin/python3.13
Python 3.14.7
/nix/store/0kcws5blwbdyxqxjd1w8hjipz61xny3k-python3-3.14.7/bin/python3.14
Development tooling:
✅ ccache
ccache version 4.13.6
/nix/store/57davyvs6p6dkrl3svzwg1ph18wsy4cz-ccache-4.13.6/bin/ccache
/nix/store/ydr5nlzxp1djb5256y0zzqa87vz55gsk-ccache-4.13.6/bin/ccache
✅ clang
clang version 22.1.7
/nix/store/192glrb2cldvziyf3378mzjqbzx3ih4g-clang-wrapper-22.1.7/bin/clang
clang version 22.1.8
/nix/store/dakcxgwz4ixkk6ps50b7vn3dvblyzsnf-clang-wrapper-22.1.8/bin/clang
✅ clang-22
clang version 22.1.7
/nix/store/rbap7zqq7mw00fyqa02p5rj7gqjp4w5i-clang-22/bin/clang-22
clang version 22.1.8
/nix/store/w826sqb98nnym6kzackg829664iqdlxj-clang-22/bin/clang-22
✅ clang++
clang version 22.1.7
/nix/store/192glrb2cldvziyf3378mzjqbzx3ih4g-clang-wrapper-22.1.7/bin/clang++
clang version 22.1.8
/nix/store/dakcxgwz4ixkk6ps50b7vn3dvblyzsnf-clang-wrapper-22.1.8/bin/clang++
✅ clang++-22
clang version 22.1.7
/nix/store/v9haf787f7bcz0mq1sad4bpyx21pj6li-clang++-22/bin/clang++-22
clang version 22.1.8
/nix/store/wnwyw3yfqrhpxij65bl3pa1x65q2zg3m-clang++-22/bin/clang++-22
✅ ClangBuildAnalyzer
ClangBuildAnalyzer 1.6.0
/nix/store/4l50ds9fa2mkvh7wg8qzrlbmjs12sb8l-clangbuildanalyzer-1.6.0/bin/ClangBuildAnalyzer
/nix/store/c7jjnw29ra78jdlzjww2m374izhy05n8-clangbuildanalyzer-1.6.0/bin/ClangBuildAnalyzer
✅ curl
curl 8.20.0 (aarch64-apple-darwin25.3.0) libcurl/8.20.0 OpenSSL/3.6.2 zlib/1.3.2 libssh2/1.11.1 nghttp2/1.69.0 mit-krb5/1.22.1
/nix/store/kclq0czaxvsgh4ym9ld7b6iwy50l1snk-curl-8.20.0-bin/bin/curl
curl 8.22.0 (aarch64-apple-darwin25.6.0) libcurl/8.22.0 OpenSSL/3.5.8 zlib/1.3.2 libssh2/1.11.1 nghttp2/1.70.0 mit-krb5/1.22.2
/nix/store/pq1h17abqy6q5y6k5dn9qj0fwb89flba-curl-8.22.0-bin/bin/curl
✅ file
file-5.47
/nix/store/dax63li7wwcbqxxkkgzc4g2rx7d4w86x-file-5.47/bin/file
file-5.48
/nix/store/y44vpziwbzfm8020dzlqqbybyyy7lb6g-file-5.48/bin/file
✅ less
less 692 (PCRE2 regular expressions)
/nix/store/lvr16y75r1pdxpdv0aph5ak2yd0hkvqm-less-692/bin/less
less 710 (PCRE2 regular expressions)
/nix/store/3yn3dx5fwg5i0inqqzwnxq7f6lwhpjqa-less-710/bin/less
✅ make
GNU Make 4.4.1
/nix/store/8wwiw8pwyhrkzyq28hqzxfl4z84lks81-gnumake-4.4.1/bin/make
/nix/store/8l4wgxmhdxfvc9g18y76bnxl1ghgcp7n-gnumake-4.4.1/bin/make
✅ netstat
present
/nix/store/qsd1kzqb0ahrk433vmyl245gp623j19s-network_cmds-730.80.3/bin/netstat
/nix/store/57fqvgdav8lrmzndkwm7qrgksl01yj23-network_cmds-741.100.2/bin/netstat
✅ ninja
1.13.2
/nix/store/bqykhrblarkj4fl0hz2mf8ngwfv6x6bz-ninja-1.13.2/bin/ninja
/nix/store/infm4vlx71q83bqz0k559w1gn4b0rq4i-ninja-1.13.2/bin/ninja
✅ perl
v5.42.0
/nix/store/js13ri9fvm0ajk1fpd3acigys2a9whdv-perl-5.42.0/bin/perl
v5.42.3
/nix/store/7rvr6qjfzk2ybzn1imjw6c6qq1n1a9cr-perl-5.42.3/bin/perl
✅ pkg-config
0.29.2
/nix/store/lzrwr375jqhhbca116kja96xf1md83l8-pkg-config-wrapper-0.29.2/bin/pkg-config
/nix/store/nhgylvj98klqggy54h7gbc3b2ia4a00y-pkg-config-wrapper-0.29.2/bin/pkg-config
✅ vim
VIM - Vi IMproved 9.2 (2026 Feb 14, compiled Jan 01 1980 00:00:00)
/nix/store/6vbkykg92w603c0sw3mkk7p7mfaawbns-vim-9.2.0389/bin/vim
/nix/store/akaa9kqm3b8rlkv61kvsqr2l2j70rvxb-vim-9.2.1001/bin/vim
✅ zip
Zip 3.0
/nix/store/z6ph729vcakbvz3wh8ln1wk6mi06w487-zip-3.0/bin/zip
/nix/store/dvawr36npnad4xa6dxi2ss0fw44aa82r-zip-3.0/bin/zip
✅ clang-apply-replacements
clang-apply-replacements version 22.1.7
/nix/store/vzyyjf3cm1hbj9wcr2qcb66x6j98zpy7-clang-tools-22.1.7/bin/clang-apply-replacements
clang-apply-replacements version 22.1.8
/nix/store/5bij4zn161lagwyndmrr21vmxjg054nv-clang-tools-22.1.8/bin/clang-apply-replacements
✅ clang-apply-replacements-22
clang-apply-replacements version 22.1.7
/nix/store/m3ii69rca4077lf4wlk7m3jcag1fs577-clang-apply-replacements-22/bin/clang-apply-replacements-22
clang-apply-replacements version 22.1.8
/nix/store/w4bz034k4l0w719rrmbwnqcgyjbswy81-clang-apply-replacements-22/bin/clang-apply-replacements-22
✅ clang-format
clang-format version 22.1.7
/nix/store/vzyyjf3cm1hbj9wcr2qcb66x6j98zpy7-clang-tools-22.1.7/bin/clang-format
clang-format version 22.1.8
/nix/store/5bij4zn161lagwyndmrr21vmxjg054nv-clang-tools-22.1.8/bin/clang-format
✅ clang-format-22
clang-format version 22.1.7
/nix/store/4fawqy6ngqcsqd2ygyyzm93q0xy3f5gs-clang-format-22/bin/clang-format-22
clang-format version 22.1.8
/nix/store/w7z346l0r8y36b6lc5i7jzj3w0581l2y-clang-format-22/bin/clang-format-22
✅ clang-tidy
LLVM version 22.1.7
/nix/store/vzyyjf3cm1hbj9wcr2qcb66x6j98zpy7-clang-tools-22.1.7/bin/clang-tidy
LLVM version 22.1.8
/nix/store/5bij4zn161lagwyndmrr21vmxjg054nv-clang-tools-22.1.8/bin/clang-tidy
✅ clang-tidy-22
LLVM version 22.1.7
/nix/store/jqw4280saixaxxihwdba9ldm2fsm6dr3-clang-tidy-22/bin/clang-tidy-22
LLVM version 22.1.8
/nix/store/0lv25qd3ddgjm3ndn70q61jfr6krlkl7-clang-tidy-22/bin/clang-tidy-22
✅ dot
dot - graphviz version 12.2.1 (0)
/nix/store/ijb4fbnqa6wzlpqnhb6q9knqpf7qqn5z-graphviz-12.2.1/bin/dot
dot - graphviz version 15.1.1 (0)
/nix/store/mc99a6bpbk2waym2cnw6ifmxli4ndb4x-graphviz-15.1.1/bin/dot
✅ doxygen
1.16.1
/nix/store/kbryjdpq9jizjb0ws0nzbf2h2ymbdiwm-doxygen-1.16.1/bin/doxygen
1.17.0
/nix/store/kz43yaj0wg6fx6cbrcmi2442z1ramy0c-doxygen-1.17.0/bin/doxygen
✅ gcovr
gcovr 8.4
/nix/store/wn8jiyh9p0bybs96s4163qp3k8vfmczx-python3.13-gcovr-8.4/bin/gcovr
/nix/store/czj1vxsizrbba280gg22cys40qxkqjgr-python3.14-gcovr-8.4/bin/gcovr
✅ gh
gh version 2.94.0 (nixpkgs)
/nix/store/fhnpw0hs0gjms1ha6ap02jq7rx13gkbp-gh-2.94.0/bin/gh
gh version 2.102.0 (2026-09-30)
/nix/store/s8bahl6zssvry93ihivfmk1l6c58byrr-gh-2.102.0/bin/gh
✅ git-cliff
git-cliff 2.13.1
/nix/store/cy0wwhgxa7yvrz97zydbq6sqmixc90fq-git-cliff-2.13.1/bin/git-cliff
git-cliff 2.14.2
/nix/store/k4m5yk0qb7hfslwsq1zjajz1d94j12bj-git-cliff-2.14.2/bin/git-cliff
✅ git-lfs
git-lfs/3.7.1 (3.7.1; darwin arm64; go 1.26.3)
/nix/store/k9r7zjfjplqa4d5s71cqvf2iv73jd9mc-git-lfs-3.7.1/bin/git-lfs
git-lfs/3.8.0 (3.8.0; darwin arm64; go 1.26.8)
/nix/store/wzixjr3yb40a13m2q654hm3bdjn6sal3-git-lfs-3.8.0/bin/git-lfs
✅ gpg
gpg (GnuPG) 2.4.9
/nix/store/cgh6iwzz5jgx9z5whka4vgj210i6npc6-gnupg-2.4.9/bin/gpg
/nix/store/0ldfwnwx385i5kh7cyvxx4mkvgjhs5rh-gnupg-2.4.9/bin/gpg
✅ pre-commit
pre-commit 4.5.1
/nix/store/z3cca68620w0w10f090szgzdnmh1waf2-pre-commit-4.5.1/bin/pre-commit
pre-commit 4.6.2
/nix/store/ivh54xypg9qvd3ha9lis0i8p4nkf8n7a-pre-commit-4.6.2/bin/pre-commit
✅ run-clang-tidy
usage: run-clang-tidy [-h] [-allow-enabling-alpha-checkers]
/nix/store/4x28x911z2f9y7adqlh3qspp4a16dig7-run-clang-tidy/bin/run-clang-tidy
/nix/store/3xss1mm3mh1hzcmg182na6ki1p5rnxf1-run-clang-tidy/bin/run-clang-tidy
✅ run-clang-tidy-22
usage: run-clang-tidy [-h] [-allow-enabling-alpha-checkers]
/nix/store/x8iymrh76sk5q91ryg5pa7i32s6gfh34-run-clang-tidy-22/bin/run-clang-tidy-22
/nix/store/w68r9z07hcq4fwbfyq3yvc7qxm8aqfbl-run-clang-tidy-22/bin/run-clang-tidy-22
Rust toolchain:
✅ cargo
cargo 1.97.1 (c980f4866 2026-06-30)
/nix/store/bnfk1sl4s9angb0vj1cj9a5y5zvqinwy-rust-minimal-1.97.1/bin/cargo
/nix/store/gklji8yq8yrx1k039c10kwz60vqf2dxj-rust-minimal-1.97.1/bin/cargo
✅ cargo-audit
cargo-audit-audit 0.22.1
/nix/store/snwkga2f5gyf404h7mmp9wriwxb8v65f-cargo-audit-0.22.1/bin/cargo-audit
cargo-audit-audit 0.22.2
/nix/store/7jgrwj557l91940aj0ag4i924ima1cqn-cargo-audit-0.22.2/bin/cargo-audit
✅ cargo-llvm-cov
cargo-llvm-cov 0.8.5
/nix/store/fpiqdh91gwyxalqp409ynm0s0g086w7w-cargo-llvm-cov-0.8.5/bin/cargo-llvm-cov
cargo-llvm-cov 0.9.0
/nix/store/5q1h285z8zzkw2qxig9a0b1avcsk1cf0-cargo-llvm-cov-0.9.0/bin/cargo-llvm-cov
✅ cargo-nextest
cargo-nextest 0.9.137
/nix/store/ylz7m947mhkgsp6i7611id3s3gcd58nq-cargo-nextest-0.9.137/bin/cargo-nextest
cargo-nextest 0.9.146
/nix/store/69v3hx3k597hcxd2zy7yad7fg0aw5yxd-cargo-nextest-0.9.146/bin/cargo-nextest
✅ clippy-driver
clippy 0.1.97 (8bab26f4f6 2026-07-14)
/nix/store/bnfk1sl4s9angb0vj1cj9a5y5zvqinwy-rust-minimal-1.97.1/bin/clippy-driver
/nix/store/gklji8yq8yrx1k039c10kwz60vqf2dxj-rust-minimal-1.97.1/bin/clippy-driver
✅ rust-analyzer
rust-analyzer 1.97.1 (8bab26f4 2026-07-14)
/nix/store/j6apc5pmd0giy15da9p650r8zklslmvi-rust-analyzer-preview-1.97.1-aarch64-apple-darwin/bin/rust-analyzer
/nix/store/30rbyk8gx7v53fxqhs4pp0fvxfjycspr-rust-analyzer-preview-1.97.1-aarch64-apple-darwin/bin/rust-analyzer
✅ rust-nightly
rustc 1.99.0-nightly (87e5904f5 2026-07-20)
/nix/store/fqpjz4l0nsnji8b2pz57mnj0akbp6hcl-rust-nightly/bin/rust-nightly
rustc 1.101.0-nightly (282215592 2026-10-04)
/nix/store/a9552rf3pylcnkijc1z0hf1lcnv50djq-rust-nightly/bin/rust-nightly
✅ rustc
rustc 1.97.1 (8bab26f4f 2026-07-14)
/nix/store/bnfk1sl4s9angb0vj1cj9a5y5zvqinwy-rust-minimal-1.97.1/bin/rustc
/nix/store/gklji8yq8yrx1k039c10kwz60vqf2dxj-rust-minimal-1.97.1/bin/rustc
✅ rustfmt
rustfmt 1.9.0-stable (8bab26f4f6 2026-07-14)
/nix/store/5ymwgr9jqjz7zzbmj0j5vqbwcd3kp0vm-rustfmt-preview-1.97.1-aarch64-apple-darwin/bin/rustfmt
/nix/store/0jmdipyn8266j9kqx2xxm2qkxlrlrjfc-rustfmt-preview-1.97.1-aarch64-apple-darwin/bin/rustfmt
Skipping git-over-HTTPS check (CHECK_TOOLS_SKIP_CLONE is set).

View File

@@ -21,7 +21,8 @@ let
)
);
toolchain = if pkgs.stdenv.isLinux then linux.toolchain else (darwin.toolchain ++ [ darwinEnv ]);
toolchain =
if pkgs.stdenv.hostPlatform.isLinux then linux.toolchain else (darwin.toolchain ++ [ darwinEnv ]);
in
{
default = pkgs.buildEnv {

View File

@@ -20,15 +20,16 @@ let
# Custom-glibc stdenvs, matching the CI environment. darwin has no custom
# glibc, so there they fall back to the plain nixpkgs stdenvs.
customGccStdenv = if pkgs.stdenv.isLinux then linux.gccStdenv else plainGccStdenv;
customClangStdenv = if pkgs.stdenv.isLinux then linux.clangStdenv else plainClangStdenv;
customGccStdenv = if pkgs.stdenv.hostPlatform.isLinux then linux.gccStdenv else plainGccStdenv;
customClangStdenv =
if pkgs.stdenv.hostPlatform.isLinux then linux.clangStdenv else plainClangStdenv;
# gcov matching each gcc shell, so `-Dcoverage=ON` builds work in the shell.
plainGcov = mkGcov {
name = "plain";
cc = gccPackage.cc;
};
customGccGcov = if pkgs.stdenv.isLinux then linux.gcov else plainGcov;
customGccGcov = if pkgs.stdenv.hostPlatform.isLinux then linux.gcov else plainGcov;
# Whole directory: init.sh locates the profiles relative to itself.
conanDir = ../conan;
@@ -51,7 +52,7 @@ let
# Not sdkEnv: a shell's stdenv already sets that up. Prepended so the stub
# beats the nixpkgs libresolv this shell's tooling drags in.
darwinLibresolvHook = pkgs.lib.optionalString pkgs.stdenv.isDarwin (
darwinLibresolvHook = pkgs.lib.optionalString pkgs.stdenv.hostPlatform.isDarwin (
pkgs.lib.concatLines (
pkgs.lib.mapAttrsToList (
name: value: ''export ${name}="${value} ''${${name}:-}"''
@@ -126,7 +127,7 @@ let
in
rec {
# macOS: Nix Clang. Linux: Nix GCC.
default = if pkgs.stdenv.isDarwin then clang else gcc;
default = if pkgs.stdenv.hostPlatform.isDarwin then clang else gcc;
# gcc/clang use the custom-glibc toolchain, matching CI. On darwin there is no
# custom glibc, so they fall back to the plain nixpkgs toolchain.
@@ -171,7 +172,7 @@ rec {
# The *-plain shells (stock nixpkgs toolchain) exist only on Linux: on darwin
# gcc/clang are already plain, so these would be redundant and are omitted, which
# makes `nix develop .#gcc-plain` fail there rather than silently aliasing gcc.
// pkgs.lib.optionalAttrs pkgs.stdenv.isLinux {
// pkgs.lib.optionalAttrs pkgs.stdenv.hostPlatform.isLinux {
gcc-plain = makeShell {
shellName = "gcc-plain";
stdenv = plainGccStdenv;

View File

@@ -1,39 +0,0 @@
#include <xrpl/basics/CountedObject.h>
#include <algorithm>
namespace xrpl {
CountedObjects&
CountedObjects::getInstance() noexcept
{
static CountedObjects kInstance;
return kInstance;
}
CountedObjects::CountedObjects() noexcept : count_(0), head_(nullptr)
{
}
CountedObjects::List
CountedObjects::getCounts(int minimumThreshold) const
{
List counts;
// When other operations are concurrent, the count
// might be temporarily less than the actual count.
counts.reserve(count_.load());
for (auto* ctr = head_.load(); ctr != nullptr; ctr = ctr->getNext())
{
if (ctr->getCount() >= minimumThreshold)
counts.emplace_back(ctr->getName(), ctr->getCount());
}
std::ranges::sort(counts);
return counts;
}
} // namespace xrpl

View File

@@ -13,8 +13,8 @@
namespace xrpl {
LedgerCloseReason
whyCloseLedger(
bool
shouldCloseLedger(
bool anyTransactions,
std::size_t prevProposers,
std::size_t proposersClosed,
@@ -27,7 +27,6 @@ whyCloseLedger(
beast::Journal j,
std::unique_ptr<std::stringstream> const& clog)
{
// Log text keeps the shouldCloseLedger name: consumers match on it.
CLOG(clog) << "shouldCloseLedger params anyTransactions: " << anyTransactions
<< ", prevProposers: " << prevProposers << ", proposersClosed: " << proposersClosed
<< ", proposersValidated: " << proposersValidated
@@ -48,7 +47,7 @@ whyCloseLedger(
JLOG(j.warn()) << ss.str();
CLOG(clog) << "closing ledger: " << ss.str() << ". ";
return LedgerCloseReason::Anomaly;
return true;
}
if ((proposersClosed + proposersValidated) > (prevProposers / 2))
@@ -56,16 +55,14 @@ whyCloseLedger(
// If more than half of the network has closed, we close
JLOG(j.trace()) << "Others have closed";
CLOG(clog) << "closing ledger because enough others have already. ";
return LedgerCloseReason::OthersClosed;
return true;
}
if (!anyTransactions)
{
// Only close at the end of the idle interval
CLOG(clog) << "no transactions, returning. ";
return timeSincePrevClose >= idleInterval // normal idle
? LedgerCloseReason::Idle
: LedgerCloseReason::KeepOpen;
return timeSincePrevClose >= idleInterval; // normal idle
}
// Preserve minimum ledger open time
@@ -73,7 +70,7 @@ whyCloseLedger(
{
JLOG(j.debug()) << "Must wait minimum time before closing";
CLOG(clog) << "not closing because under ledgerMIN_CLOSE. ";
return LedgerCloseReason::KeepOpen;
return false;
}
// Don't let this ledger close more than twice as fast as the previous
@@ -83,40 +80,12 @@ whyCloseLedger(
{
JLOG(j.debug()) << "Ledger has not been open long enough";
CLOG(clog) << "not closing because not open long enough. ";
return LedgerCloseReason::KeepOpen;
return false;
}
// Close the ledger
CLOG(clog) << "no reason to not close. ";
return LedgerCloseReason::Normal;
}
bool
shouldCloseLedger(
bool anyTransactions,
std::size_t prevProposers,
std::size_t proposersClosed,
std::size_t proposersValidated,
std::chrono::milliseconds prevRoundTime,
std::chrono::milliseconds timeSincePrevClose,
std::chrono::milliseconds openTime,
std::chrono::milliseconds idleInterval,
ConsensusParms const& parms,
beast::Journal j,
std::unique_ptr<std::stringstream> const& clog)
{
return whyCloseLedger(
anyTransactions,
prevProposers,
proposersClosed,
proposersValidated,
prevRoundTime,
timeSincePrevClose,
openTime,
idleInterval,
parms,
j,
clog) != LedgerCloseReason::KeepOpen;
return true;
}
bool

View File

@@ -80,10 +80,10 @@ BookDirs::const_iterator::operator++()
XRPL_ASSERT(index_ != kZero, "xrpl::BookDirs::const_iterator::operator++ : nonzero index");
if (!cdirNext(*view_, curKey_, sle_, entry_, index_))
{
if (index_ == 0)
if (index_ == kZero)
curKey_ = view_->succ(++curKey_, nextQuality_).value_or(kZero);
if (index_ != 0 || curKey_ == kZero)
if (index_ != kZero || curKey_ == kZero)
{
curKey_ = key_;
entry_ = 0;

View File

@@ -11,6 +11,7 @@
#include <xrpl/beast/utility/instrumentation.h>
#include <xrpl/ledger/LedgerTiming.h>
#include <xrpl/ledger/ReadView.h>
#include <xrpl/ledger/entries/FeeSettingsEntry.h>
#include <xrpl/nodestore/NodeObject.h>
#include <xrpl/protocol/Feature.h>
#include <xrpl/protocol/Fees.h>
@@ -261,7 +262,7 @@ Ledger::Ledger(Ledger const& prevLedger, NetClock::time_point closeTime)
{
header_.seq = prevLedger.header_.seq + 1;
header_.parentCloseTime = prevLedger.header_.closeTime;
header_.hash = prevLedger.header().hash + UInt256(1);
header_.hash = prevLedger.header().hash.next();
header_.drops = prevLedger.header().drops;
header_.closeTimeResolution = prevLedger.header_.closeTimeResolution;
header_.parentHash = prevLedger.header().hash;
@@ -560,7 +561,7 @@ Ledger::setup()
try
{
if (auto const sle = read(keylet::feeSettings()))
if (auto const sle = FeeSettingsEntryR(*this))
{
bool oldFees = false;
bool newFees = false;

View File

@@ -9,6 +9,7 @@
#include <xrpl/beast/utility/instrumentation.h>
#include <xrpl/ledger/ApplyView.h>
#include <xrpl/ledger/ReadView.h>
#include <xrpl/ledger/entries/LedgerHashesEntry.h>
#include <xrpl/ledger/helpers/AccountRootHelpers.h>
#include <xrpl/ledger/helpers/CredentialHelpers.h>
#include <xrpl/ledger/helpers/DirectoryHelpers.h>
@@ -361,17 +362,16 @@ hashOfSeq(ReadView const& ledger, LedgerIndex seq, beast::Journal journal)
if (int const diff = ledger.seq() - seq; diff <= 256)
{
// Within 256...
auto const hashIndex = ledger.read(keylet::skip());
LedgerHashesEntryR const hashIndex(ledger, journal);
if (hashIndex)
{
XRPL_ASSERT(
hashIndex->getFieldU32(sfLastLedgerSequence) == (ledger.seq() - 1),
"xrpl::hashOfSeq : matching ledger sequence");
STVector256 vec = hashIndex->getFieldV256(sfHashes);
if (vec.size() >= diff)
return vec[vec.size() - diff];
if (auto const hash = hashIndex.hashAt(diff - 1))
return hash;
JLOG(journal.warn()) << "Ledger " << ledger.seq() << " missing hash for " << seq << " ("
<< vec.size() << "," << diff << ")";
<< hashIndex->getFieldV256(sfHashes).size() << "," << diff << ")";
}
else
{
@@ -387,16 +387,14 @@ hashOfSeq(ReadView const& ledger, LedgerIndex seq, beast::Journal journal)
}
// in skiplist
auto const hashIndex = ledger.read(keylet::skip(seq));
LedgerHashesEntryR const hashIndex(keylet::skip(seq), ledger, journal);
if (hashIndex)
{
auto const lastSeq = hashIndex->getFieldU32(sfLastLedgerSequence);
XRPL_ASSERT(lastSeq >= seq, "xrpl::hashOfSeq : minimum last ledger");
XRPL_ASSERT((lastSeq & 0xff) == 0, "xrpl::hashOfSeq : valid last ledger");
auto const diff = (lastSeq - seq) >> 8;
STVector256 vec = hashIndex->getFieldV256(sfHashes);
if (vec.size() > diff)
return vec[vec.size() - diff - 1];
if (auto const hash = hashIndex.hashAt((lastSeq - seq) >> 8))
return hash;
}
JLOG(journal.warn()) << "Can't get seq " << seq << " from " << ledger.seq() << " error";
return std::nullopt;

View File

@@ -506,7 +506,7 @@ pseudoAccountAddress(ReadView const& view, UInt256 const& pseudoOwnerKey)
RipeshaHasher rsh;
auto const hash = sha512Half(i, view.header().parentHash, pseudoOwnerKey);
rsh(hash.data(), hash.size());
AccountID const ret = AccountID::fromRaw(static_cast<RipeshaHasher::result_type>(rsh));
AccountID const ret{static_cast<RipeshaHasher::result_type>(rsh)};
if (!view.read(keylet::account(ret)))
return ret;
}

View File

@@ -1,7 +1,6 @@
#include <xrpl/nodestore/detail/DecodedBlob.h>
#include <xrpl/basics/Blob.h>
#include <xrpl/basics/base_uint.h>
#include <xrpl/basics/safe_cast.h>
#include <xrpl/beast/utility/instrumentation.h>
#include <xrpl/nodestore/NodeObject.h>
@@ -12,8 +11,13 @@
namespace xrpl::node_store {
DecodedBlob::DecodedBlob(void const* key, void const* value, int valueBytes) : key_(key)
DecodedBlob::DecodedBlob(std::optional<UInt256> key, void const* value, int valueBytes)
{
if (!key)
return;
key_ = key.value();
/* Data format:
Bytes
@@ -63,7 +67,7 @@ DecodedBlob::createObject()
{
Blob data(objectData_, objectData_ + dataBytes_);
object = NodeObject::createObject(objectType_, std::move(data), UInt256::fromVoid(key_));
object = NodeObject::createObject(objectType_, std::move(data), key_);
}
return object;

View File

@@ -38,6 +38,7 @@
#include <functional>
#include <memory>
#include <optional>
#include <span>
#include <sstream>
#include <stdexcept>
#include <string>
@@ -216,7 +217,7 @@ public:
[&hash, pno, &status](void const* data, std::size_t size) {
nudb::detail::buffer bf;
auto const result = nodeobjectDecompress(data, size, bf);
DecodedBlob decoded(hash.data(), result.first, result.second);
DecodedBlob decoded(hash, result.first, result.second);
if (!decoded.wasOk())
{
status = Status::DataCorrupt;
@@ -240,7 +241,7 @@ public:
nudb::error_code ec;
nudb::detail::buffer bf;
auto const result = nodeobjectCompress(e.getData(), e.getSize(), bf);
db.insert(e.getKey(), result.first, result.second, ec);
db.insert(e.getKey().data(), result.first, result.second, ec);
if (ec && ec != nudb::error::key_exists)
Throw<nudb::system_error>(ec);
}
@@ -288,14 +289,25 @@ public:
Throw<nudb::system_error>(ec);
nudb::visit(
dp,
[&](void const* key,
[&](void const* keyData,
std::size_t keyBytes,
void const* data,
std::size_t size,
nudb::error_code&) {
auto key = uint256::fromRaw(
std::span{static_cast<unsigned char const*>(keyData), keyBytes});
if (!key) [[unlikely]]
{
ec = make_error_code(nudb::error::invalid_key_size);
return;
}
nudb::detail::buffer bf;
auto const result = nodeobjectDecompress(data, size, bf);
DecodedBlob decoded(key, result.first, result.second);
if (!decoded.wasOk())
{
ec = make_error_code(nudb::error::missing_value);

View File

@@ -38,6 +38,7 @@
#include <format>
#include <functional>
#include <memory>
#include <span>
#include <stdexcept>
#include <string>
@@ -111,6 +112,9 @@ public:
if (!getIfExists(keyValues, Keys::kPath, name))
Throw<std::runtime_error>("Missing path in RocksDBFactory backend");
if (keyBytes != uint256::size())
Throw<std::runtime_error>("Incorrect key size: expected 32 bytes");
rocksdb::BlockBasedTableOptions tableOptions;
options.env = env;
@@ -292,7 +296,7 @@ public:
if (getStatus.ok())
{
DecodedBlob decoded(hash.data(), string.data(), string.size());
DecodedBlob decoded(hash, string.data(), string.size());
if (decoded.wasOk())
{
@@ -317,8 +321,8 @@ public:
}
else
{
status = static_cast<Status>(
static_cast<int>(Status::CustomCode) + unsafeCast<int>(getStatus.code()));
status = checkedCast<Status>(
safeCast<int>(Status::CustomCode) + safeCast<int>(getStatus.code()));
JLOG(journal.error()) << getStatus.ToString();
}
@@ -347,7 +351,7 @@ public:
EncodedBlob const encoded(e);
wb.Put(
rocksdb::Slice(reinterpret_cast<char const*>(encoded.getKey()), keyBytes),
rocksdb::Slice(reinterpret_cast<char const*>(encoded.getKey().data()), keyBytes),
rocksdb::Slice(
reinterpret_cast<char const*>(encoded.getData()), encoded.getSize()));
}
@@ -375,9 +379,12 @@ public:
for (it->SeekToFirst(); it->Valid(); it->Next())
{
if (it->key().size() == keyBytes)
if (it->key().size() == keyBytes && it->key().size() == uint256::size())
{
DecodedBlob decoded(it->key().data(), it->value().data(), it->value().size());
DecodedBlob decoded(
uint256::fromRaw(std::span{it->key().data(), it->key().size()}),
it->value().data(),
it->value().size());
if (decoded.wasOk())
{

View File

@@ -8,6 +8,8 @@
#include <xrpl/protocol/digest.h>
#include <xrpl/protocol/tokens.h>
#include <algorithm>
#include <array>
#include <atomic>
#include <cstdint>
#include <cstring>
@@ -19,95 +21,143 @@
namespace xrpl {
namespace detail {
namespace {
/**
* Caches the base58 representations of AccountIDs
/** If true, we implement a cache for account ID lookups.
This is generally a good idea and can improve performance, especially
when serving client requests.
The cache may generate warnings under some sanitizers; if you want to
disable caching, set this variable to false.
*/
class AccountIdCache
constexpr bool cacheAccountIDs = true;
/** Represents a cached entry for the given AccountID.
We keep the size and alignment of the structure at 64 bytes to
make sure it fits within a cache line and avoid false sharing.
*/
struct alignas(64) CachedAccountID
{
private:
struct CachedAccountID
{
AccountID id;
char encoding[40] = {0};
};
/** The sequence lock for the entry.
// The actual cache
std::vector<CachedAccountID> cache_;
An even value means that the data is stable, whereas an odd
value indicates the data is being actively modified.
// We use a hash function designed to resist algorithmic complexity attacks
HardenedHash<> hasher_;
We use a 32-bit value to ensure that the counter will never
return to an observed value during a read window.
*/
std::atomic<std::uint32_t> seq{0};
// 64 spinlocks, packed into a single 64-bit value
std::atomic<std::uint64_t> locks_ = 0;
/** The account id. */
AccountID id{};
public:
AccountIdCache(std::size_t count) : cache_(count)
{
// This is non-binding, but we try to avoid wasting memory that
// is caused by overallocation.
cache_.shrink_to_fit();
}
/** The length of the cached encoding.
std::string
toBase58(AccountID const& id)
{
auto const index = hasher_(id) % cache_.size();
This field disambiguates collisions between ACCOUNT_ZERO and
never-written-to slots, since they would otherwise share the
same encoded bits.
*/
std::uint8_t len = 0;
PackedSpinlock sl(locks_, index % 64);
{
std::scoped_lock const lock(sl);
// The check against the first character of the encoding ensures
// that we don't mishandle the case of the all-zero account:
if (cache_[index].encoding[0] != 0 && cache_[index].id == id)
return cache_[index].encoding;
}
auto ret = encodeBase58Token(TokenType::AccountID, id.data(), id.size());
XRPL_ASSERT(ret.size() <= 38, "xrpl::detail::AccountIdCache : maximum result size");
{
std::scoped_lock const lock(sl);
cache_[index].id = id;
std::strcpy(cache_[index].encoding, ret.c_str());
}
return ret;
}
/** The actual string encoding. */
char encoding[39] = {0};
};
} // namespace detail
static_assert(sizeof(CachedAccountID) == 64 && alignof(CachedAccountID) == 64);
static std::unique_ptr<detail::AccountIdCache> gAccountIdCache;
// The actual cache to use: a simple, direct-mapped, best-effort construct.
//
// We do not use a hardened hash for this because an AccountID is already
// just cryptographically generated (i.e. uniformly distributed) bits and
// there is no amplification, since collisions do not carry an additional
// cost.
//
// The size is a "nice" power of 2 because it makes the calculation of an
// index easier.
constinit std::array<CachedAccountID, 65536> gAccountIdCache{};
void
initAccountIdCache(std::size_t count)
{
if (!gAccountIdCache && count != 0)
gAccountIdCache = std::make_unique<detail::AccountIdCache>(count);
}
} // namespace
std::string
toBase58(AccountID const& v)
{
if (gAccountIdCache)
return gAccountIdCache->toBase58(v);
// Since the input bits are generated by a cryptographically secure
// hash construct, we can just pick any subset of them, no muss, no
// fuss.
auto const index = [&v]() -> std::uint16_t {
auto const data = v.data();
return encodeBase58Token(TokenType::AccountID, v.data(), v.size());
return static_cast<std::uint16_t>(data[0]) | (static_cast<std::uint16_t>(data[1]) << 8);
}();
if constexpr (cacheAccountIDs)
{
auto& slot = gAccountIdCache[index];
// This is the optimistic, store-free read path: a collision with a
// concurrent writer will be detected because the seq recheck fails
// and we fall through to the slow path.
//
// These unsynchronized reads of the slot fields are, formally, data
// races and, technically speaking, result in undefined behavior and
// a tool like TSAN might flag them. These should not have an impact
// on the correctness of the returned data: the sequence lock checks
// ensure we discard any data that was in flux.
if (auto const s0 = slot.seq.load(std::memory_order::acquire);
(s0 & 1) == 0 && slot.len != 0 && slot.id == v)
{
char buf[38];
auto const len = std::min<std::size_t>(slot.len, sizeof(buf));
std::memcpy(buf, slot.encoding, len);
// Prevents the loads above from sinking below the recheck; the
// acquire on s0 constrains later operations, not earlier ones.
std::atomic_thread_fence(std::memory_order::acquire);
// Everything read from the slot — the gates above and the bytes in
// buf — precedes this recheck and is validated by it. No slot access
// may follow it.
if (slot.seq.load(std::memory_order::relaxed) == s0)
return std::string(buf, len);
}
}
auto ret = encodeBase58Token(TokenType::AccountID, v.data(), v.size());
XRPL_ASSERT(!ret.empty() && ret.size() <= 38, "xrpl::toBase58 : maximum result size");
if constexpr (cacheAccountIDs)
{
auto& slot = gAccountIdCache[index];
// The CAS to an odd value excludes other writers; if it fails, another
// thread wrote or is writing into this slot. That is not a problem: we
// just don't store our result into the cache this time.
if (auto s0 = slot.seq.load(std::memory_order::relaxed); (s0 & 1) == 0 &&
slot.seq.compare_exchange_strong(s0, s0 + 1, std::memory_order::acquire))
{
slot.id = v;
slot.len = static_cast<std::uint8_t>(ret.size());
std::copy_n(ret.data(), ret.size(), slot.encoding);
// The acquire on the CAS above keeps the writes from becoming visible
// before the seqlock transitions to an odd value, which means readers
// can detect an in-progress write. This release publishes them before
// the seqlock transitions to an even value.
slot.seq.store(s0 + 2, std::memory_order::release);
}
}
return ret;
}
template <>
std::optional<AccountID>
parseBase58(std::string const& s)
{
auto const result = decodeBase58Token(s, TokenType::AccountID);
if (result.size() != AccountID::kBytes)
return std::nullopt;
return AccountID::fromRaw(result);
return AccountID::fromRaw(decodeBase58Token(s, TokenType::AccountID));
}
//------------------------------------------------------------------------------
@@ -148,25 +198,9 @@ parseBase58(std::string const& s)
AccountID
calcAccountID(PublicKey const& pk)
{
static_assert(AccountID::kBytes == sizeof(RipeshaHasher::result_type));
RipeshaHasher rsh;
rsh(pk.data(), pk.size());
return AccountID::fromRaw(static_cast<RipeshaHasher::result_type>(rsh));
}
AccountID const&
xrpAccount()
{
static AccountID const kAccount(beast::kZero);
return kAccount;
}
AccountID const&
noAccount()
{
static AccountID const kAccount(1);
return kAccount;
return AccountID{static_cast<RipeshaHasher::result_type>(rsh)};
}
bool

View File

@@ -185,8 +185,9 @@ getBookBase(Book const& book)
UInt256
getQualityNext(UInt256 const& uBase)
{
static constexpr UInt256 kNextQuality(
"0000000000000000000000000000000000000000000000010000000000000000");
static constexpr UInt256 kNextQuality{
"0000000000000000000000000000000000000000000000010000000000000000"};
return uBase + kNextQuality;
}
@@ -429,9 +430,9 @@ payChannel(AccountID const& src, AccountID const& dst, SeqProxy const& seq) noex
Keylet
nftokenPageMin(AccountID const& owner)
{
std::array<std::uint8_t, 32> buf{};
std::memcpy(buf.data(), owner.data(), owner.size());
return {ltNFTOKEN_PAGE, UInt256::fromRaw(buf)};
std::array<std::uint8_t, UInt256::size()> buf{};
std::ranges::copy(owner, buf.begin());
return {ltNFTOKEN_PAGE, UInt256{buf}};
}
Keylet

View File

@@ -295,11 +295,9 @@ verify(PublicKey const& publicKey, Slice const& m, Slice const& sig) noexcept
NodeID
calcNodeID(PublicKey const& pk)
{
static_assert(NodeID::kBytes == sizeof(RipeshaHasher::result_type));
RipeshaHasher h;
h(pk.data(), pk.size());
return NodeID::fromRaw(static_cast<RipeshaHasher::result_type>(h));
return NodeID{static_cast<RipeshaHasher::result_type>(h)};
}
} // namespace xrpl

View File

@@ -148,12 +148,12 @@ STAmount::STAmount(SerialIter& sit, SField const& name) : STBase(name)
}
Issue issue;
issue.currency = sit.get160();
issue.currency = Currency{sit.get160()};
if (isXRP(issue.currency))
Throw<std::runtime_error>("invalid native currency");
issue.account = sit.get160();
issue.account = AccountID{sit.get160()};
if (isXRP(issue.account))
Throw<std::runtime_error>("invalid native account");

Some files were not shown because too many files have changed in this diff Show More