mirror of
https://github.com/XRPLF/rippled.git
synced 2026-09-29 16:28:08 +00:00
b3eba92e2748eecdfd147bc2b91c6d0b6091384f
49 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
185547efd2 | Merge branch 'pratik/otel-sync-diagnostics' into pratik/otel-sync-diagnostics-freshen-fix | ||
|
|
39c14e71f7 | Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics | ||
|
|
d5c2964fea |
test(telemetry): drop the suppressed requirement from tx.receive
tx.receive no longer carries a suppressed attribute: the span is created only once the node has decided to process the transaction, so there is no dropped copy for the attribute to describe. The validator fails a span that is missing a required attribute, so leaving it listed turns the telemetry-validation leg red. |
||
|
|
44d1994d6a |
refactor(nodestore): scrub site details, own the phase label strings, fix the test overload
Review follow-up on the freshen lock-hold fix: - Drop host names, dates and one-site figures from the new comments, harness notes and docs; explain the mechanism in general terms. - RotationPhase stores its stage and cache labels as owned std::string, not std::string_view: the ctor still takes views so the label constants pass without a copy, but a member view would dangle if a caller ever passed a temporary. freshenCache/recordFreshen take the cache name by std::string_view (read-only, call-scoped). - The new DatabaseRotating test called fetchNodeObject through the derived type, whose private override hides the public base method; call it through Database& instead. This was the dev-box build break. - freshenCache reports the exact fetched count when a health abort cuts it short, and stops labelling the per-partition hold 'getKeys'. - Remove a [[maybe_unused]] that silenced no warning (the build sets -Wno-unused-parameter and disables misc-unused-parameters). |
||
|
|
b1345fff8d |
docs(telemetry): describe the rotation stall without internal host names
The reference doc, span-harness notes and histogram-bucket comments named the internal AWS dev box and dates while explaining why the rotation phases are timed. Reword to the general mechanism (a multi-second freeze at the copy-walk to freshen boundary on a populated node); the specific hosts, dates and trace ids stay in the task notes. |
||
|
|
614c1a39ad |
fix(nodestore): bound the rotation freshen's cache lock hold and measure its yield
The online-delete rotation's cache freshen called TaggedCache::getKeys(),
which held the cache mutex while copying every key. On the dev box's 26
million entry tree-node cache that hold lasted 5-6 s, froze every job
that touches the cache, and dropped the RocksDB node out of sync once per
rotation: each "getKeys held the lock" warning was followed within 1-5 s
by "View of consensus changed" (5 of 5 rotations on 2026-09-15).
Copy the keys one map partition at a time instead. TaggedCache gains
forEachKeyPartition(), which holds the mutex only while one partition's
keys are copied and runs the callback with the mutex released, so the
longest hold shrinks by the partition count (8 on the dev box). The
freshen.keys rotation phase no longer exists as one step, so its span,
stage value, harness entries and docs are removed; the per-partition hold
still shows on the cache lock-hold peak gauge.
Measure what the freshen achieves, which no existing signal did.
DatabaseRotating gains duplicateCopyForwardTotal(), counting archive
copies made on duplicate fetches (the rotation's own copy walk and
freshen); copyForwardTotal() deliberately excludes those. The freshen
phase records rotation_freshen_keys_total{cache,outcome} and stamps
key_count, cache and keys_copied on its span; the copy phase stamps
nodes_copied. A warn log line per freshen reports the same numbers, and
the ledger-sync-health dashboard gets a Rotation Freshen Yield panel.
Log the "STATE->" operating-mode change at warn instead of info. It is
the only record of a mode change with an exact timestamp; the
state_changes_total counter is scraped once a minute and cannot order a
flap against a multi-second event.
Tests: five GTests for forEachKeyPartition (every key once, empty cache,
mutex free during the callback, concurrent insert, lock-hold peak), three
for duplicateCopyForwardTotal over two memory backends, one for the new
counter's series, and the new name literals.
|
||
|
|
b0cea67aed |
fix(telemetry): address final-review + CI clang-tidy findings
CI's clang-tidy leg flagged eight include-cleaner errors and three misc-const-correctness / readability-convert-member-functions-to-static / modernize-use-designated-initializers issues, all inside WP-B6's own code. Fixed as follows: - `MetricsRegistry.h`: `#include <opentelemetry/metrics/observer_result.h>` for ObserverResult; `observeCacheLockHoldPeaks` is now `static` because it touches neither instance state nor telemetry members. - `SHAMapStoreImp.h`: adds direct includes for `<cstddef>`, `<string_view>` and `<xrpl/telemetry/SpanNames.h>` (the StaticStr provider). `seconds` in `RotationPhase::~RotationPhase` is `[[maybe_unused]]` so a `-DXRPL_ENABLE_TELEMETRY=0` build under `-Werror` keeps compiling. - `SHAMapStoreImp.cpp`: direct includes for `SHAMapStoreSpanNames.h`, `SpanGuard.h`, `SpanNames.h`; `RotationPhase` locals that never call `setAttribute` are declared `const`; `RotationOutcome` uses designated initialisers. Final-review findings (WP-B6-rotation-stall-tracing.md, "What to check when reviewing"): - Panels 74 and 75 on `ledger-sync-health.json` still carried panel 41's description, axisLabel, Source and Keywords copy; rewritten to describe rotation phase duration and cache lock hold respectively. - `consensus_view_change_total` and the `view.change` round-span event were emitted but not registered with the harness. Added the counter to `not_asserted.metrics_excluded` (workload-gated) and annotated the `consensus.round` span note with the event and its two attribute keys. Not fixed (parked, see progress ledger): - The reviewer's second Important finding — a plan/code contradiction on the consensus counter — was based on a misread of the plan; the plan's "Rejected alternatives" table lists a new `TraceCategory::Nodestore` and the getKeys() fix, not the consensus counter. No action. - The Minor note about `sweep()`'s peak including lock-acquire time and `getKeys()`'s not: `sweep()` acquires and releases the lock via a `scoped_lock`, so `noteLockHold` still runs after the release and the numbers are comparable. No action. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
9adb6a255d |
test(telemetry): register the rotation spans and stall metrics with the harness
Adds `nodestore.rotate` and its eight phase children to expected_spans.json,
all `optional: true` because the 5-node localhost harness cluster never reaches
`online_delete`. Their parent-child relationships are asserted but skip-marked
so a run without a rotation stays green.
Adds `cache_metrics{metric="treenode_lock_hold_peak_us"|"fullbelow_lock_hold_peak_us"}`
to the asserted sync_diagnostics group -- both are observable and always emit,
even at zero. Puts `rotation_phase_duration_seconds` and `jobq_stall_total` in
`not_asserted.metrics_excluded`; both are workload-gated.
On the Cloud collector, adds an `ottl_condition` policy that keeps any trace
carrying a span whose name matches `^nodestore\.rotate`, so the 0.5% probabilistic
tail sampler cannot drop a rotation trace. Sampler is OR'd across policies.
|
||
|
|
521f00a484 |
merge: bring the lock-free ValidationTracker forward from phase10-workload-validation
Only the workload README conflicted; the tracker, config and test files merged clean, which closes the chain from phase-7. |
||
|
|
478b3e4b07 |
docs(telemetry): correct stale claims and citations in the harness docs
The workload README contradicted itself on --skip-loki: one bullet said CI always passes it and so the two log-correlation checks are never exercised, another said the workflow no longer passes it. The workflow mentions the flag nowhere, so the first was the stale half. Other claims checked against the tree and corrected: - both the README and the plan doc described the push trigger as filtered on branch names. The workflow has no branches filter, deliberately, because GitHub ANDs branches with paths - the plan doc printed 6 of the workflow's 12 paths globs, and claimed the workflow was 367 lines against an actual 451. The glob block is now generated from the workflow, and the line count dropped rather than restated - rpcNOT_SUPPORTED does not exist anywhere in the tree. The symbol is RpcNotSupported, and the refusal sites are RipplePathFind.cpp:59-60 and PathFind.cpp:50-51, not :48-49 and :39 - RCLConsensus.cpp:666 and :663 are not log or event lines; the tx.included event is at :720 and the per-transaction debug log at :715 - LedgerMaster.cpp:463 is fixIndex, not the ledger.store span, which is at :470 - ServerHandler.cpp:705 is inside makeJsonError; processRequest is at :718 - file counts: docker/telemetry/workload/ is 25 files, include/xrpl/telemetry/ 13 - the optional-span bullet named five causes covering 10 of 16 entries, omitting the txq.* family and the WebSocket handshake - the /api/v1/series choice was attributed to stale StatsD gauges; this harness runs no StatsD A line number in run-full-validation.sh was cited in five places and drifts on every edit to that file, so those now name the file only. The keygen helper's header records what production does instead -- validator-keys-tool create_keys then create_token, keeping the master key off the node -- and why a disposable cluster does not. |
||
|
|
198207eee4 | merge: bring the close-time attr harness fix forward from phase10-workload-validation | ||
|
|
f80faea85d |
fix(telemetry): match the renamed close-time attrs in expected_spans.json
The close-time span attributes name their unit and epoch: close_time_ripple_epoch_s, parent_close_time_ripple_epoch_s and close_time_self_ripple_epoch_s. The harness inventory still required the unsuffixed keys, so the attribute checks for consensus.accept.apply and ledger.build failed on every validation run while the spans themselves were correct. Rename the four required_attributes entries to the keys the code emits. |
||
|
|
3e4b5c71ff |
merge: bring the traces_endpoint rename forward from phase-10
Three conflicts, all between this branch's own sync-diagnostics work and phase-10's older versions. Resolved to this branch in each case, since it owns the newer content: - InboundLedger.h keeps the missing-node and receive-depth gauges and the fuller acquire-span contract. - MetricsRegistry.cpp keeps the namespaced label:: constants. - LedgerMaster.cpp keeps makeLedgerTraceSpan(), which joins the store and validate spans into one per-ledger trace by hash. LedgerMaster.cpp needed a second pass. The automatic merge had kept both sides outside the conflict markers, nesting phase-10's older promotion block inside this branch's `if (!pubLedger_)` — so setValidated, setFull and setValidLedger would each have run twice. Taking this branch's file wholesale removes the duplicate; brace balance and a single "Advancing accepted ledger" confirm it. That resolution drops two things phase-10 was carrying into this file: the storeSpan/validateSpan guard names, and the explicit scope that keeps the one-in-256 flag-ledger check outside the ledger.validate measurement. Both are re-applied on this branch in the next commit; the scope needs a variable-lifetime check that does not belong in a merge. |
||
|
|
3a63a17548 |
docs(telemetry): stop the harness contract narrating its own revisions
Notes across the workload contract described earlier versions of themselves, or cited commits that only exist inside this chain. A squash merge publishes none of it, so each reference resolves nowhere. Notes that described their own earlier text: - expected_spans.json: 'this note previously concluded', 'this note previously said', 'Un-skipped 2026-08-26', 'the reason had simply gone stale for two weeks' and 'the claim this entry carried' are replaced by the standing reason each entry holds. The wildcard pairs now say the validator globs the child via _span_name_matches(), and state the literal-collapse failure as what a different validator WOULD do rather than as history. - regression-thresholds.json: 'an earlier version of this note wrongly claimed', 'the earlier version oversold it' and 'an earlier note called that' become the cautions themselves -- do not reason from 'every ladder step is at least 2x', do not oversell the backstop, do not read a false fire as a missing override. - test_check_regression_bounds.py: the docstring gives the reason a literal is wrong here, not the story of two tests that once hard-coded one. Baseline-refresh history rewritten as measurement: - README.md, baselines/README.md, telemetry-runbook.md and regression-metrics.json no longer attribute threshold moves to 'the 2026-08-26 refresh'. The evidence is kept as measurement -- span.tx.apply.p50 has read 0.7917 ms and 0.00597 ms on the same workload, 132x apart; job.acceptLedger.running.p95 has measured a 5.74x floor on one baseline and 16.28x on another -- which is what supports the claim that a single-run baseline cannot bound these keys. Two chain-only commit ids removed, |
||
|
|
fa9f75d4e7 |
docs(telemetry): explain the absent path-finding load without the change story
The harness contract described how path-finding load came to be absent rather than why it is absent. The load exists on no branch before this one, so a squash merge publishes no revision that ever issued it: the 2026-08-25 date resolves nowhere, and 'removing it', 'used to satisfy' and 'has now cleared' compare against a state a reader cannot reach. - README: state that DEFAULT_WEIGHTS carries no ripple_path_find entry, and give the error floor as what WOULD happen if it did, rather than what removing it fixed. 'Putting it back' becomes 'Enabling it'. - expected_spans.json: the pathfind.request and pathfind.compute notes, and both hierarchy skip_reasons, now put the span's presence in the conditional -- the parent would appear if the RPC were issued, because the ScopedSpanGuard at RipplePathFind.cpp:35 sits above the rpcNOT_SUPPORTED guard at :48-49. - expected_metrics.json: the rpc_method_errored_total, pathfind_fast and pathfind_full notes drop the date and keep both independent reasons the metrics stay absent. - regression-thresholds.json: span.ledger.store 'is excluded from' the gated surface rather than 'was removed from' it. The reasoning is unchanged: pathfinding is off because Config.cpp:725-726 zeroes pathSearchMax when [validation_seed] is present, a refused call still exports an error span, and at a 3% weight that is a ~3% STATUS_CODE_ERROR floor. Documentation and JSON note strings only, no behaviour change. |
||
|
|
a214db3a90 |
docs(telemetry): remove pre-squash and plan-internal references from sync diagnostics
Comments across the sync-diagnostic work described earlier revisions of the same change, or cited identifiers a reader of the merged tree cannot resolve. Prior-state comparisons rewritten in the present tense: - MetricsRegistry.cpp carried two adjacent paragraphs prescribing opposite behaviour for a disabled quorum, one publishing int64 max and one omitting the series. The code omits it; the superseded paragraph is gone and the surviving reason SIZE_MAX must not be cast is kept. - MallocTrim, LedgerMaster, LedgerReplayTask, TransactionAcquire, Application: say what the signal is the only record of, rather than what was 'previously trace-only', 'not logged at all here' or 'used to sit inside if (debug())'. - LedgerMaster.h and SpanGuardScope: without an explicit join each ledger's spans WOULD be separate traces -- not that they were 'before this'. - Handshake: the message is forwarded byte for byte, not 'byte-identical to the previous behaviour', and the helper throws rather than 'throws as before'. - MetricNames: quorum_disabled is a separate boolean rather than a sentinel, stated without what the state 'used to be encoded by'. - LedgerMaster.cpp no longer claims to mirror the unl_quorum gauge; it does not. That gauge omits the series while this stores int64 max. - 'Split out of' / 'Split from' become 'Kept separate from' in five places. Plan-internal identifiers removed: - All 24 WP-Ax / WP-Bx work-package labels across the telemetry tests, the collector configs, tempo.yaml and the expected_* inventories. They are defined in no file in the repo, so they resolve nowhere once merged. - The two references to OpenTelemetryPlan/, which does not reach develop, now point at docs/telemetry-glossary.md 'Fresh-node sync diagnostics'. Comments and JSON note strings only, no behaviour change. |
||
|
|
e9856897ec | Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics | ||
|
|
53cc08aa52 |
fix(telemetry): assert real ancestry in the span hierarchy check
The check reported span.hierarchy.<parent>-><child> and a message reading "Found <child> as child of <parent>" on the strength of both names appearing somewhere in the same trace. A span parented by something unrelated passed, so the one property the check exists to prove was never tested. It now walks the child's parentSpanId chain looking for a span matching the parent name. Ancestry rather than a direct edge, because all 21 declared relationships are worded as the parent containing the child, so a scope appearing in between is a refactor and not a broken relationship. Span ids are compared as opaque strings: both fields come from the same Tempo response and share its encoding, so nothing here depends on whether that is hex or base64. Co-occurrence is still the search filter, which is what lets a conditional child be found in an older trace instead of only the newest ones. Verdicts are separated because they send the reader to different places: a child that is present but not under the parent is a hierarchy bug, a chain running into a span the trace lacks is one that never reached Tempo, and an unusable parent span is neither. A definite negative outranks an indefinite one, and one trace proving ancestry settles the relationship. Tests cover each verdict plus the cross-trace and cyclic-chain cases, and each one was checked against the specific defect it names. The runner now fails when it collects no tests and reports SystemExit, both of which otherwise produce a silent pass. The pathfind.request skip_reason said only the child side handles globs. Both sides do now; the blocker is the literal parent name in the Tempo query, so the skip itself stands. |
||
|
|
39fa18e898 |
test(telemetry): assert the tx-tree acquire hierarchy again
The sampling fix that just merged forward removes the only reason this was skipped. The check no longer inspects the three newest parent traces; it asks Tempo for traces containing both parent and child. Worth recording why this phase was the one that failed while its two siblings passed, because the original assertion treated all three as equivalent and they are not. InboundLedger.cpp opens each phase only when that piece is still needed: header on !haveHeader_ (:672), astree in the else of haveState_ (:689), txtree in the else of haveTransactions_ (:698). A node acquiring a ledger here almost always lacks the account-state tree, so astree opens on essentially every acquire. But it usually already holds the transaction set -- every node sees the same relayed transactions and builds the same set -- so txtree opens on a minority of acquires. The child was always emitting, 5 traces of its own on the run that failed; it just was not in the three most recent acquires. That is now all three sampling-caused skips retired: txq.accept -> txq.accept_tx and this one asserted, and txq.enqueue -> txq.batch_clear narrowed to its real remaining cause, a child that never fires under this workload at all. Contract on this branch: 24 relationships, 19 asserted, 5 skipped, and zero spans declaring a parent without an entry. The five are the two pathfind pairs and the pathfind.request parent (pathfinding disabled and no path-finding RPC issued), rpc.ws_message -> rpc.process (not a code relationship -- rpc.process is a child of rpc.http_request), and txq.batch_clear. None is a sampling artifact. Verification: JSON parses; the four validator tests pass after the merge; 0 unaccounted parentings; counters still 48 span types and 74 unique attributes; otel-naming exits 0. Whether this holds against a live Tempo is what the run this push triggers decides -- the stub proves the query shape, not the corpus. |
||
|
|
ce18bb3317 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings in the hierarchy-check sampling fix: the check now asks Tempo for traces containing both parent and child rather than inspecting the three newest parent traces, so a child conditional on a state the workload rarely reaches is found wherever it occurred. Merged clean, no conflicts, no resolution decisions. This unblocks ledger.acquire -> ledger.acquire.txtree on this branch, which was skipped for exactly that sampling problem and is un-skipped in the next commit. |
||
|
|
a87d772f40 |
fix(telemetry): find a span hierarchy where it happened, not only where it is newest
The hierarchy check searched the parent span and inspected the three newest traces it returned. That is wrong whenever the child is conditional on a state the workload only sometimes reaches: the parent fires constantly, so its newest traces are the ones LEAST likely to carry a rare child. Three relationships had been skipped as unassertable for exactly this, and in none of them was the child missing -- each emitted traces of its own and simply was not in the three most recent parent traces. The check now issues a second query, a TraceQL trace-level conjunction of the parent and child name predicates, and inspects those traces. Tempo searches its whole retention for co-occurrence instead of leaving the answer to which traces happen to be newest. The parent-only query is kept and still runs first, so "the parent stopped being emitted" stays a distinct failure from "the parent is there but the child never co-occurs" -- they mean different things to whoever reads the report, and collapsing them would lose that. The returned traces are still verified with _span_name_matches rather than the query result being trusted on its own. Tempo has already guaranteed co-occurrence, so this is redundant on the happy path; it is kept because it keeps the glob semantics in one place and means a wrongly built query cannot silently pass. _traceql_name_predicate handles the wildcard contracts. TraceQL has no glob operator, so `rpc.command.*` is sent as name=~"rpc\.command\..*" with the dots escaped -- unescaped they would match any character in those positions, which is the looseness _span_name_matches exists to avoid. Two entries follow from the fix. txq.accept -> txq.accept_tx is asserted again: its child is created inside the queued-transaction loop behind `if (feeLevelPaid >= requiredFeeLevel)` (TxQ.cpp:1530) while the parent fires on every close (:1499), which was the whole reason it failed. txq.enqueue -> txq.batch_clear stays skipped but for ONE reason now instead of two -- its child never fires at all under this workload, needing an account with a supersedable batch, so it is purely a workload gap and needs nothing further from the validator. The third, ledger.acquire -> ledger.acquire.txtree, lives on the sync-diagnostics branch and is un-skipped there once this merges forward. Written test-first, and the first test this module has had. The failing test reproduces the exact CI message, "txq.accept_tx not found in txq.accept traces", against a stubbed Tempo whose corpus holds the child only in a trace outside the newest three. Three sibling tests guard the ways this could be "fixed" wrongly: an absent child must still fail, a missing parent must still name the parent rather than the child, and a wildcard child must be satisfied by any family member. The stub records the queries issued, so the conjunction is asserted rather than assumed. A stub rather than a live Tempo because the behaviour under test is which traces the check ASKS FOR -- a passing query against real data proves the data co-operated, not that the query was right. The first run of those tests failed for the wrong reason: my stub's name-predicate regex also matched the resource.service.name="xrpld" term every query carries and so demanded a span literally named "xrpld". Fixed in the stub, with the lookbehind commented as load-bearing, before touching production code. Verification: 4/4 tests pass, and the failing one was watched failing first with the production message; the issued queries were printed and confirmed to contain the conjunction; validate_telemetry.py compiles; expected_spans.json parses; 21 relationships, 16 asserted and 5 skipped; counters still 41 span types; otel-naming exits 0. Three unrelated files in this worktree are another party's live work and were deliberately left unstaged. |
||
|
|
1f8b69a3f0 |
fix(telemetry): skip the tx-tree acquire hierarchy, which the tx set being present hides
Asserting all three ledger.acquire phase parentings treated them as equally
conditional. They are not. Run 33002568549 failed on
ledger.acquire -> ledger.acquire.txtree -- "ledger.acquire.txtree not found in
ledger.acquire traces", the single failure in 279 checks -- while header and
astree passed. Skipped rather than left red.
Not a missing span: txtree reports 5 traces of its own on that same run, one per
node. InboundLedger.cpp opens each phase only when that piece is still needed:
header on !haveHeader_ (:672), astree in the else of haveState_ (:689), txtree in
the else of haveTransactions_ (:698). A node acquiring a ledger here almost always
lacks the account-state tree, so astree opens on essentially every acquire and its
assertion holds. But it usually already HOLDS the transaction set, because every
node sees the same relayed transactions and builds the same set, so
haveTransactions_ is true and no txtree phase opens at all. It fires only on the
minority of acquires where the set was genuinely missing, and with
_validate_parent_child sampling the 3 newest parent traces
(validate_telemetry.py:803) those are not the ones sampled.
This is the third entry skipped for one underlying cause, after
txq.accept -> txq.accept_tx and txq.enqueue -> txq.batch_clear: a child that is
conditional on a state the harness rarely reaches, met by newest-N sampling of the
parent. Preferring parent traces that CONTAIN the child would retire all three at
once, and that is now the highest-value change left in this harness -- recorded in
each of the three reasons so whoever picks it up finds the whole set.
The rest of the run supports the other changes. Both rpc.command.* hierarchies
un-skipped in
|
||
|
|
f2cc740c42 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
6d17df083f |
test(telemetry): account for every declared span parenting
Each span entry documents its parent, and a separate list holds the pairs the validator actually checks. Three parentings were declared on the span entries and absent from that list entirely, so they were neither asserted nor recorded as unassertable -- silently missing rather than deliberately skipped. All three are now listed, skipped, each with the reason that actually applies. Every span declaring a parent now has an entry: the count went from 3 unaccounted to 0. txq.enqueue -> txq.batch_clear is conditional and narrowly so. The child is created in TxQ::tryClearAccountQueueUpThruTx (TxQ.cpp:550), which needs one account holding several queued transactions AND an arriving transaction that supersedes the batch. txq-burst produces queueing but arranges no such shape, and it has never been observed on a run. It would also meet the sampling limit that forced the txq.accept_tx skip, so fixing the sampling addresses both at once. rpc.command.* -> pathfind.request is the one skip caused by a wildcard PARENT rather than by a missing span, and the asymmetry is worth recording: _validate_parent_child inserts the parent name literally into its Tempo query (:801), so a wildcard parent matches nothing, while the CHILD side globs through _span_name_matches (:826-828). That is exactly why rpc.ws_message -> rpc.command.* can be asserted and this cannot. pathfind.compute -> pathfind.discover has both ends absent, for the reason the pathfind.compute entry already sets out at length: pathfinding is disabled on every harness node because Config.cpp:725-726 zeroes pathSearchMax when a [validation_seed] is present, and since 2026-08-25 no path-finding RPC is issued either. Listed so the family is fully accounted for rather than partly silent. No assertion is added or removed here -- this is accounting. The plan task that prompted it also assumed the pathfind.compute skip reason was stale and needed correcting; it is not, it already names both blockers and corrects an older liquidity-based reason, so that half of the task was a defect in my plan rather than in the file. Verification: JSON parses; 21 relationships, 15 asserted and 6 skipped; no duplicates; 0 spans declaring a parent without an entry, down from 3; counters still 41 span types; otel-naming exits 0; pre-commit clean. |
||
|
|
a215ab7bb1 |
fix(telemetry): assert the rpc.command hierarchies, whose skips described dead code
Both rpc.command.* relationships were skipped on the claim that
_validate_parent_child collapses a wildcard child to one literal name via
child_name.replace("*", "server_info"). That code does not exist.
|
||
|
|
143abfd8f6 |
fix(telemetry): stop requiring what the harness cannot guarantee
Two entries in the span contract could fail on healthy behaviour. `ledger.serve` is emitted when this node answers another node's request for ledger data, and was mandatory. A node only answers if a peer asks, and a `TMGetLedger` request is constructed in exactly two places -- src/xrpld/app/ledger/detail/InboundLedger.cpp and src/xrpld/app/ledger/detail/TransactionAcquire.cpp -- whose `ledger.acquire` and `txset.acquire` spans are both already marked optional. A mandatory check therefore rested on an optional cause. Marked optional, with the dependency named in the note so it is promoted together with them rather than alone. `peer.dial` covers one outbound connect attempt and required `outcome` and `duration_ms`. Both are set only in `reportOutcome()` (ConnectAttempt.cpp:158-199). The teardown path sets neither on purpose: an attempt destroyed during overlay shutdown, or one whose connect was aborted, ends its span in `~ConnectAttempt` (ConnectAttempt.cpp:89-102), whose own comment states that a span ending with no `outcome` is the honest record of a dial that never concluded. Since `peer.dial` is a freshRoot, each dial is its own trace with exactly one instance of the span, and the validator inspects only the most recent trace -- so one newly-aborted dial fails CI while the node is behaving correctly. `remote_endpoint` stays required; it is set at construction on every path. Verified: both call sites read in current code; `TMGetLedger` construction confined to those two files; the counters recomputed -- 48 span types (matches len(spans)) and 74 unique attributes, unchanged, because `outcome` and `duration_ms` are still required by `txset.acquire` and the `ledger.acquire` family. Not verified: that a run with these entries relaxed still exercises both spans, which only a CI dispatch can show. Neither entry can now fail on healthy behaviour, so a green run proves less than before by design. |
||
|
|
e5d7b2a4b0 |
fix(telemetry): skip the txq accept-pass hierarchy, which sampling cannot assert
|
||
|
|
92e988c8b8 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
504138dd9c |
test(telemetry): assert all three ledger-acquire phase hierarchies
The acquire span opens three phase children -- header, account-state tree and transaction tree -- and the contract documented all three parentings while checking none of them. astree had an entry marked skipped; header and txtree had no entry at all. The skip reason was that the parent is optional, because a healthy 5-node cluster agreeing from genesis rarely back-fills history, so the hierarchy check would fail against a parent with no traces. Run 32969481032 refutes the premise: ledger.acquire reported 5 traces and so did each of the three children, one per node. The parenting was never the uncertain part -- beginPhaseSpan() parents through the acquire span's own captured SpanContext rather than the ambient thread context, so it holds whichever worker opens a phase. Asserting these matters because of what the phases are for. A fresh sync is dominated by the account-state tree, and the flat parent span cannot separate that from the much smaller transaction tree or from the header wait that gates both. If a phase stops nesting under the acquire it still emits, still carries its missing-node count and its timeout flag, and nothing else in this harness notices -- but the trace stops answering which phase the sync is stuck in, which is the whole reason these spans exist. Routed here rather than to phase-10 because phase-10 has no ledger.acquire.header or ledger.acquire.txtree span at all; the phase children were introduced on this branch. Verification: JSON parses; 10 relationships, 6 asserted and 4 skipped, no duplicates, every non-wildcard child resolves to a declared span entry; counters unchanged at 48 span types and 74 unique attributes; otel-naming exits 0; pre-commit clean. Edited by surgical text replacement -- a first attempt used a json.dumps round-trip and reflowed the whole file, 148 insertions against 47 deletions with unrelated compact arrays expanded and unrelated notes rewritten; that was reverted and redone, and the churn is now 11 against 3. Whether a child is findable INSIDE the parent's fetched trace is what CI will decide. |
||
|
|
c531ac569b |
fix(telemetry): trigger the workload by what changed, and assert the span tree
Two problems, both about coverage this workflow claims to have and does not. The push trigger gated on branch NAME as well as path, and GitHub ANDs the two. Branch names are not something this repository controls, so a push to any branch outside "pratik/otel-phase*", "feature/otel-*" or "feature/telemetry-*" was never dispatched -- not queued, not skipped, no run to look at. That is not a theoretical gap: two rounds of harness fixes on pratik/otel-sync-diagnostics produced no signal at all before anyone noticed the workflow had never started. The branches filter is removed; the paths already express the real question. The path list was also incomplete in a way that matters more than it looks. The span-name and metric-name headers are the wire contract this harness asserts against by literal string, and the convention colocates each one with the class it serves -- so eight of the ten *SpanNames.h headers live under consensus/, overlay/, app/ledger/, app/main/, app/misc/, rpc/ and tx/, none of which was matched. Renaming a span constant therefore compiled clean, emptied the assertions and triggered nothing. Matched now by filename, "**/*SpanNames.h" and "**/*MetricNames.h", so future headers are covered wherever they land. Added for the same reason: include/xrpl/beast/insight (the interface headers decide what the collector can publish, so they move the metric surface as surely as the implementation), src/tests/libxrpl/telemetry (the GTests pinning those constants), and the two checker directories that gate this surface in CI. Second, the span hierarchy. Each span entry documents its parent, and a separate list holds the pairs the validator actually checks in Tempo. Those had drifted apart: 18 parentings were documented, 7 were checked. A span that stops nesting under its parent -- which is what a detached guard does -- leaves every span and every attribute intact, so no other check in this harness notices; the trace simply stops being readable as one operation. Eleven pairs are added, each one where both ends emitted on a real run: rpc.http_request -> rpc.process, the three txq parentings, and seven consensus ones under consensus.round and consensus.establish. Fourteen of eighteen are now asserted; the four still skipped are the wildcard rpc.command.* families and pathfind.compute. Three notes were also factually wrong, all repeating one mistake. They said rpc.process and rpc.http_request cannot appear because that path is HTTP-only while the load generator is WebSocket-only. The premise is right, the conclusion is not: both appear on every run, five traces each, because run-full-validation.sh polls each node's HTTP port with curl for readiness and validated-ledger progress (:449, :502). Those polls take the HTTP path. A reader acting on the old text would have gone looking for a way to make the harness speak HTTP that it already speaks. The rpc.process -> rpc.command.* skip reason inherited the same error and additionally claimed the WebSocket equivalent is "asserted above instead", which it is not -- that one is skipped for the same wildcard limitation. All three now state the real blocker, which is that _validate_parent_child resolves a wildcard child to a single literal probe. Both HTTP spans stay optional rather than being promoted: the curl polls are harness scaffolding, not workload, and a future change to how the script waits for a node could legitimately remove them. Verification: JSON parses; 18 relationships, no duplicates, every non-wildcard endpoint resolves to a declared span entry; counters still 41 span types and 62 unique attributes; workflow YAML parses, has no branches key, keeps workflow_dispatch, and every new glob was checked against the tracked file list with a matcher that reproduces GitHub's ** semantics; otel-naming exits 0; pre-commit clean on both files. The eleven new assertions are proven only to the extent that both ends emitted on run 32969481032 -- that a child is findable INSIDE the parent's fetched trace is what CI will now decide. |
||
|
|
493475a9d4 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
42a72863bb |
fix(telemetry): make the sync-diagnostics metric gate actually assert
The sync_diagnostics group asserted nothing. assert_sync_diagnostics_metrics called _check_prometheus_metric with five positional arguments against a six-parameter signature: `report` landed in `deadline` and `sem` was omitted entirely, so the call raised TypeError before a single metric was queried. Neither run_validation nor main catches anything, so the traceback propagated, run-full-validation.sh recorded the non-zero exit as a validation failure, and the four phases ordered after it -- dashboards, both parity checks and log-trace correlation -- never ran at all. Reproduced directly: TypeError, zero checks recorded. Even with the arity corrected the group would still have passed silently, because _check_prometheus_metric RETURNS its CheckResult rather than recording it and the value was discarded. Both halves are fixed by adopting the fan-out validate_metrics already uses: one shared deadline, a concurrency semaphore, gather, then report.add per result. The same call now records 55 checks where it previously recorded none. With the gate live, the inventory it guards had to be made honest. Two metrics could never have passed it. unl_fetch_total is emitted only from ValidatorSite::reportFetchOutcome, which indexes sites_[siteIdx]; sites_ comes from [validator_list_sites], and the harness writes a static [validators] file with no list site anywhere, so no fetch outcome is ever reported. handshake_negotiation_fail_total needs a rejected handshake, and no reject path was found to be reachable between identical localhost nodes. Both move to not_asserted.metrics_excluded, which is where the file's own description says workload-gated names belong. Eleven further conditional metrics -- the acquire, replay, disconnect, serve, jump and sweep counters -- were documented only inside free-text notes; they move to the same map. That matters beyond tidiness: _accounted_metric_names harvests metrics_excluded keys, so a name recorded only in prose is reported as unaccounted, and a prose note cannot be linted at all. Two metrics were wrongly excluded. rotation_state's callback gates only on dynamic_cast<DatabaseRotating*>, and online_delete=256 is set by both the cfg template and run-full-validation.sh, so SHAMapStoreImp builds a DatabaseRotatingImp, the cast succeeds, and both sub-series are observed on every collection tick. The note claiming the harness could not produce them conflated "no rotation runs" with "no series published"; the first is true and bounds the values, the second is false. Both are now asserted at value 0, where absence rather than the zero is the regression, and the note is corrected. The four new histograms listed only _bucket, or _bucket and _count. Each now lists _sum as well, matching the rpc_method_us and job_queued_us convention, so an exporter regression that drops one series cannot pass. On the span side, ledger.validate and ledger.store are the two ends of the per_ledger trace-join group, and the join is computed by hashing ledger_hash -- yet neither required it. Both spans take it unconditionally from makeLedgerTraceSpan, so requiring it is free, and without it a lost join key surfaces only as "spans landed in separate traces", naming the consequence instead of the cause. Deliberately unchanged: ledger.serve stays required and peer.dial keeps its current required attributes, though both look unsafe -- ledger.serve can only fire if an optional span fires first, and peer.dial's destructor exit sets neither outcome nor duration_ms. Those weaken assertions rather than add coverage, so they are reported rather than changed here. Verification: TypeError reproduced before the fix and absent after, with 55 checks recorded; both JSON files parse; no name is both asserted and excluded and none is duplicated; the declared span counters remain consistent at 48 and 74, proven by injecting an extra attribute and watching the check fail; check_otel_naming.py exits 0, and Rule K was proven to read these entries by injecting a bogus name in an owned family and observing exit 1; pre-commit passes on all three files; the levelization baseline is unchanged. NOT compiled -- no C++ changed. |
||
|
|
59a0595a6e |
fix(telemetry): stop the workload harness issuing refused path-finding RPC
Every node the harness starts is a validator, and validators disable pathfinding: Config.cpp:725-726 zeroes pathSearchMax whenever a [validation_seed] or [validator_token] section is present, and run-full-validation.sh writes [validation_seed] into every generated node cfg (:308) with no [path_search] section to put the default back. So doRipplePathFind refused every call at RipplePathFind.cpp:48-49 and the 3% ripple_path_find weight bought no coverage at all. It was not free either. The pathfind.request guard is constructed at RipplePathFind.cpp:35, above that refusal, so each refused call still exported a span, and the enclosing rpc.command.ripple_path_find span carried rpc_status=error. That put a steady 3% error floor into span_calls_total for STATUS_CODE_ERROR: any error-rate threshold derived from harness data before this change was measuring the harness rather than xrpld, and needs re-deriving. Removing the load makes pathfind.request unreachable, so it moves from required to optional in expected_spans.json; without that the span check would fail on every run. Three notes in that file and three in expected_metrics.json made claims that are now false, two of them citing line numbers this commit deletes; all six are corrected. The runbook required/optional count moves 26/15 to 25/16. Two facts a future reader needs. First, the weights previously summed to 103, not 100, so every percentage the docstring stated was wrong: health checks were really 38.8%, not 40%. Dropping the 3 makes the sum exactly 100 and every stated percentage correct for the first time. expected_spans.json also carried live arithmetic off the old total, "25/103 ... roughly 43%", now 25/100 and 42%. Second, baselines/baseline-timings.json was captured WITH this load. Only span.rpc.ws_message p50/p95/p99 of the 25 gated keys sees the RPC mix, and their trip points sit 3.1x to 5.9x above baseline, so the gate will not fire. But a timing baseline is workload-specific and its profile field still reads full-validation, so nothing will flag the drift: refresh it from the next CI run's timings artifact. Pathfinding now has no coverage in this harness at all. The workload README section "Pathfinding is not exercised" records that cost, the manual verification route, and a four-step restore recipe in which steps 1 and 2 alone only reinstate the error floor. |
||
|
|
7d35eb872a |
docs(telemetry): correct the harness contract's pathfinding and gauge reasons
The assert / do-not-assert decisions in expected_metrics.json and expected_spans.json were all correct, but several recorded reasons were not. Pathfinding is disabled outright on every harness node: Config.cpp:725-726 zeroes pathSearchMax whenever a [validation_seed] or [validator_token] section is present, run-full-validation.sh writes [validation_seed] for every node and has no [path_search] override, and both handlers return rpcNOT_SUPPORTED before constructing a PathRequest. - pathfind_full_milliseconds no longer claims a probabilistic path, nor prescribes an explicit ledger index, which cannot help: the config gate fires before the ledger parameter is read. - pathfind_fast_milliseconds keeps its hasCompletion argument but now leads with the config gate, which is the operative blocker. - The pathfind.compute and pathfind.discover notes and the pathfind.request to pathfind.compute skip reason no longer blame missing liquidity. pathfind.update_all now records why its request list stays empty. - statsd_gauges states the arming precondition: a beast gauge is only as safe as an observable gauge when its object exists before Application.cpp:1570, where onCollectionReady arms the registered gauges exactly once. - Alert wiring claims softened: every rule in rules.yaml is paused. - The per job type gauge group loses its bogus poll bandwidth reason, and its regex claim is corrected: there is no running state regex, so 30 of those gauges have no consumer at all. - overlay_peer_disconnects has one query consumer, not two. - Cloud dashboard copies dropped from consumer counts: that tree is ignored by git and has no tracked files. - rpc_method_errored_total explains that a refused RPC is a normal return, not a throw, so the refusals above do not make it fire. - 09-data-collection-reference.md no longer claims a Prometheus name query in expected_metrics.json. No behaviour change: the flattened check name list is byte identical. |
||
|
|
c5829df68f |
feat(telemetry): record which consensus rounds requested a tx-set fetch
A tx-set fetch carried no key tying it to the consensus round that needed the set, so attributing a stalled fetch to a round meant guessing from timestamps. One fetch is wanted by many rounds -- it is keyed by set hash, survives the round sweep, and the round never blocks on it -- so a single parent, link or attribute cannot describe the relationship. Instead the fetch span records one timestamped event per requesting round, carrying the round's parent-ledger hash and the ledger it is building. Both attribute keys already existed in the shared telemetry namespace with exactly this meaning, and the existing addEvent API is used as-is, so no new telemetry surface is added and the whole feature compiles out with telemetry disabled. The event fires once per round rather than once per peer proposal, keyed on the round's parent-ledger hash: that distinguishes rounds started on different forks at the same height, which a ledger-height compare cannot. A mid-round wrong-ledger recovery re-enters consensus without re-caching the round identity, so a fetch begun after that switch is attributed to the pre-switch round; the limitation is documented where the values are cached. Also fixes the fetch span's end time, which depended on when the C++ object was destroyed. Three of the four exits that stop pursuing a fetch -- the set arriving from elsewhere, the round sweep, and shutdown -- ended the span only via the destructor, so the recorded duration included however long any reference happened to be held. Each exit now ends the span itself, plus cancel() and container teardown, and the destructor asserts the span is already closed rather than closing it: a fallback that can never legitimately fire should fail loudly instead of hiding a missed exit. An abandoned fetch also no longer asks peers for a set nobody wants, which previously led to charging those peers for answering our own request. Verification: pre-commit and TIDY=1 clang-tidy pass; levelization is unchanged. NOT compiled -- the branch is blocked by a gcc-15 internal compiler error in the unrelated xrpl.libxrpl.rdb unity translation unit. Runtime behaviour is unasserted: xrpl_tests links only xrpl.libxrpl, so TransactionAcquire is unreachable from GTest; the added tests cover the new span-name and attribute constants only. |
||
|
|
cb88a12883 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Nine conflicts, resolved as follows. src/xrpld/app/ledger/detail/InboundLedger.cpp -- kept this branch's version. phase10 sets the span's outcome/timeouts/peer_count attributes inline at each exit; this branch replaced that with the idempotent finalizeAcquireSpan(), called on all four exits (init, done, give-up, destructor). Taking phase10's blocks would have set the outcome twice against a helper documented as not overwriting what the real exit recorded. phase10's comment explains why peer_count must not be read in a destructor; the helper solves that structurally by taking std::optional<std::size_t> and being passed std::nullopt from there. src/xrpld/telemetry/MetricsRegistry.cpp -- kept metric::ledgerEconomy over phase10's "ledger_economy" literal. This branch added the naming check that requires constants for converted families, so the literal would regress it. Took phase10's comment cleanup. src/xrpld/telemetry/MetricsRegistry.h -- kept registerRotationStateGauge(), which only exists here, and took phase10's removal of the stale task-number comment. validate_telemetry.py -- combined both. phase10 replaced serial metric polling with a concurrent fan-out on one shared deadline, because 58 metrics x 45 s of additive timeout overran the CI budget; that is kept. Its target list filters on SKIPPED_METRIC_GROUPS rather than the two literals it hardcoded, so the sync_diagnostics group stays owned by assert_sync_diagnostics_metrics() instead of being polled and reported twice. Both SYNC_DIAGNOSTICS_GROUP and METRIC_POLL_CONCURRENCY are needed and both are kept. check_otel_naming.py -- both sides extend the rule docstring. Took phase10's fuller Rule E text (doc discovery, allow-dotted markers) and re-appended rules I/J/K/L, which exist only here. expected_metrics.json -- the two sides add disjoint sibling groups, so both are kept: sync_diagnostics alongside node_health_gauges, overlay_reduce_relay, overlay_overflow, validation_lifetime_counters and not_asserted. Both dashboard uids are kept, giving 16 asserted uids against 16 dashboards on disk. expected_spans.json -- kept this branch's span set, a superset that adds the acquire phase spans, ledger.serve, txset.acquire and peer.dial, and expands ledger.acquire's required attributes. Took phase10's description, which documents what the totals mean, and its note on how the RPC wildcard span is created. total_span_types and total_unique_attributes are recomputed for the union: 48 and 74, since each side's figure counted only its own spans. Docs: took phase10's more accurate wording on what the dashboard check actually covers, and corrected the dashboard count from 15 to 16 where the merge made it stale. Verified: no conflict markers remain, both JSON contracts parse, both Python files compile, asserted dashboard uids match the dashboards on disk exactly, and the OTel naming check reports all layers consistent. |
||
|
|
22e440aee1 |
fix(telemetry): correct the phase-10 validation harness against the code
The harness manifests asserted things the code cannot produce and missed most of what it does. Two assertions were failing every run, and the metric set covered 16 of the ~41 emitted names. expected_spans.json: rpc.process was required with rpc.ws_message as its parent, but it is created only in ServerHandler::processRequest() on the HTTP path, so a WebSocket-only workload never produces it -- it is now optional and parented to rpc.http_request, and the rpc.process -> rpc.command.* edge is skipped with the real reason instead of a coroutine-context-loss diagnosis that was never the cause. Adds the missing rpc.ws_upgrade span, corrects four parents (consensus.mode_change, pathfind.request, and update_positions/check, which are children of consensus.establish rather than consensus.round), and demotes conditionally-set attributes out of required_attributes so a healthy run stops failing. Counts recomputed from the file: 41 span types, 62 unique required attributes. expected_metrics.json: 16 -> 52 asserted entries across the job-queue, RPC method, reduce-relay, overflow and validation families, plus the fifteenth dashboard uid. Metrics the harness workload cannot exercise -- erroring RPC, ledger-mismatch, TxQ overflow, and the lazily-created getobject_* instruments -- are listed in a not_asserted group the validator skips, rather than as assertions that would fail on a healthy node. The workflow's push trigger listed two globs matching nothing (include/xrpl/basics/Telemetry*.h, src/xrpld/app/misc/Telemetry*), so no C++ telemetry change ever triggered validation. Replaced with the paths the code actually lives in, including src/libxrpl/beast/insight/** for the insight export path the harness depends on. The four inert workflow_dispatch inputs are now labelled UNUSED rather than looking like working knobs. Docs: the workload README described a StatsD dirty-flag mechanism under a member name that does not exist, on a code path the harness never uses -- it sets [insight] server=otel, so gauges export through an observable-gauge callback every cycle. Adds the missing txq-burst phase, reconciles three different dashboard counts, and drops "posts summary to PR", which the workflow has no permission to do. The runbook's phase-10 section loses the last sampling_ratio reference (not a config key), gains a Regression Gate and CI subsection covering the gate that can fail CI, and its compose-logs command now names the workload compose file. cmake --preset default is left for a separate change: no CMakePresets.json is tracked, so it is wrong everywhere it appears. Also drops the dead exporter=otlp_http key the harness wrote into every node config, and stops capture_timings.py defaulting --profile to a profile that does not exist. |
||
|
|
b9497d05da |
fix(telemetry): correct the dial-outcome diagnosis and harden the site label
Adversarial validation of the previous commit found one of its two code fixes
was diagnosed wrongly and the other incomplete. Both are corrected here, along
with the layers the first pass missed.
1. The new dial outcome was named for the wrong condition. It was added as
`duplicate` on the belief that PeerFinder had already granted a slot for the
address. It has not: `Logic::onConnected` contains exactly ONE false-returning
path and it is the self-connect check, which logs "Logic dropping as self
connect" (include/xrpl/peerfinder/detail/Logic.h). The duplicate check lives
in `newOutboundSlot`, evaluated before a ConnectAttempt exists, so a real
duplicate can never reach this branch.
That mattered beyond the name: the previous commit told operators the outcome
was benign churn to ignore, when it actually reports a local misconfiguration
-- this node has its own address in [ips_fixed] or behind its advertised
endpoint, and every dial to it is wasted. Renamed to `self_connection`,
reusing the slug `handshake_negotiation_fail_total` already publishes for the
same fault so it reads identically on both signals, and every description
corrected to say so. The fail() string now reads "Self connection" too.
The first pass also missed three enforcement and contract sites: the
ConnectAttempt.h Doxygen state machine (which still mapped the slot branch
onto tls_fail), the LedgerSpanNames unit test (which pinned exactly five
values over a std::array<..., 5> and so left the new member untested), and the
span-derived twin panel plus two reference docs that still published the old
five-value domain.
2. The credential-free site label was incomplete twice over.
- It appended the port, and `Resource::Resource` DEFAULTS that to 443/https
and 80/http when the config omits one. The label would have become
`https://vl.ripple.com:443/` where Grafana Cloud currently holds
`https://vl.ripple.com`, silently renaming the series for every deployment
already scraping this metric. Verified against live label values before and
after; the port is now omitted.
- parseUrl's path group is `(/.*)?`, greedy to end of string, so a query or
fragment lands inside `path`. A list URL authenticated by `?token=...` would
have leaked exactly as userinfo did. The path is now truncated at the first
'?' or '#'.
Also updated the MetricNames.h usage example, which still taught the raw-URI
pattern to the next author, and the 09-doc row that described the label as the
configured URI.
3. Rule J hardening from the same review: `classify_instrument_kind` returns an
`other` sentinel for a non-factory macro, and storing it in the kind set could
render a future conflict as "created as counter and other". The sentinel is
now skipped, keeping it doing what it already did -- matching no shape rule.
Added a second regression test whose input the pre-fix code reported as CLEAN
(gauge-then-histogram on a `_us` name), so the guard is proven by a 0-vs-1
difference and not only by a changed message. Both new tests were run against
a reconstructed last-wins implementation and both fail against it.
Documented the conflict class in the Rule J rows of the checker README and
CONTRIBUTING, which previously described only the suffix conventions.
Verified: naming checker exits 0 with Rule J passing all 40 real names; 140
checker tests pass; 15 dashboards validate; both workload JSON files parse;
clang-tidy over the full compile database reports no finding on any changed line
of ConnectAttempt.cpp or ValidatorSite.cpp; pre-commit passes.
Not verified: not compiled. The label change adds string truncation and the
outcome rename touches a constexpr used across three translation units, so CI's
build remains the first real check on both.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f649670ef7 |
fix(telemetry): address the PR review findings
Four defects from the automated review on PR #7875, each verified against the current tree before fixing (one further comment, the row-63 dashboard overlap, was already fixed by an earlier commit and needed nothing). 1. Rule J could not detect an instrument-kind mismatch. instrument_kinds() wrote `kinds[wire] = ...`, so a wire name created through two different factories kept only the kind visited last and whichever emit site the file walk reached last silently decided the verdict. It now collects a set per name and reports the conflict itself -- one name exporting two instruments is the defect, and no suffix can be correct for both. Added a regression test that builds a name as both a counter and an observable gauge and asserts the message names both. 2. A duplicate connection was reported as `tls_fail`. The TLS handshake had in fact succeeded; PeerFinder simply already held a slot for that address, which is ordinary churn on a healthy node. Conflating the two made a rising `tls_fail` unreadable -- it could mean unreachable peers or merely a busy PeerFinder, and those need opposite responses. Added a distinct `duplicate` outcome and carried the widened vocabulary through every place that enumerates it: the panel description, both filter descriptions, the runbook branch table, the runbook outcome list and the expected_spans note. The `dial_outcome` template variable is a label_values() query, so it picks the new value up on its own. 3. ConnectAttempt::onShutdown had no `operation_aborted` guard, unlike the five other handlers in the same file. A clean teardown was therefore counted as `upgrade_fail`, inflating that outcome on any node shutting down with dials in flight. 4. ValidatorSite used the raw configured URI as a Prometheus label. [validator_list_sites] accepts credentials in the URI and ParsedUrl keeps them in username/password, so a configured `https://user:pass@host` would have copied the secret into a metric label and on into the collector, Prometheus and every dashboard. The label is now rebuilt from scheme, host, port and path -- everything needed to tell one site apart, and nothing more. Verified: naming checker exits 0 with Rule J still passing all 40 real instrument names; its unit tests now number 139 and all pass; 15 dashboards validate; both workload JSON files parse; clang-tidy over the full compile database reports no finding on either changed .cpp; pre-commit passes. Not verified: not compiled. Item 4 introduces string concatenation and item 2 a new constexpr, so CI's build is the first real check on both. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
22fd5e8601 |
feat(telemetry): make the sync board readable, and time real node writes (WP-B4)
The board and runbook had grown by append across eight work packages, so they read in the order the work was done rather than the order a node progresses. This is the coherence pass; it adds no new instrumentation. - Dashboard: 52 panels regrouped from two rows into nine that follow the fresh-start sequence — bootstrap, peer supply, sync state, acquire and SHAMap fetch, job queue, quorum and publish, terminal blockers, then back-fill and spans collapsed since they answer conditional questions. Layout only: no title, query or description changed. - Runbook: the flat step list becomes a decision tree branching on the observed symptom, with the amendment-block check first because it is terminal. Each branch names the panels, what healthy and unhealthy look like, and what to conclude. The existing steps are kept as the detail bodies. - Reference table: every signal name re-checked against the code and every named panel against the board; four stale panel references fixed. - Validation: every signal is now either asserted or covered by a note explaining why a five-node local cluster cannot produce it. Also fixes the write-latency signal, which was inert on a real node: the store duration was only recorded on the database-import path, while the two production store implementations did not time themselves, so an ordinary node reported a write count with no latency. Both now time the backend write, which is the disk work this signal exists to expose. Without it the "existing database syncs slower than a fresh one" diagnosis had no primary signal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
41b818b55b |
feat(telemetry): join a ledger's spans into one trace, add round histogram (WP-B3)
A slow fresh-sync ledger produced spans scattered across threads with no way to relate them. They now share a trace id derived from the ledger's own hash, the one value every participating site already holds, so nothing new is plumbed across threads. This is the pattern the transaction pipeline already uses for its tx id. Joined: ledger.validate, ledger.store, and a new consensus.validation.accept recorded when a trusted validation arrives. In Tempo, searching one ledger hash returns them together, so an operator can tell whether the ledger was slow to arrive, slow to be accepted, or slow to be stored. They are siblings rather than a chain because the accept gate is entered from three different threads, so no fixed parent order exists. consensus.validation.accept also records why an arriving validation did or did not advance the gate, which makes "validations arrive but are all rejected" visible for the first time. consensus_round_duration_ms turns the existing round-time span attribute into a histogram, so a fleet trend needs a metric query rather than raw trace inspection. An explicit bucket view is required, not optional: the SDK default tops out at ten seconds while consensus abandons a round at two minutes, so slow rounds would all fall in one bucket and every quantile would read exactly ten seconds. Cost is one record per round. Record layer: the histogram is native and needs no collector change. The two new bounded attributes are added as span-metric dimensions to both collector configs. The ledger hash stays out of them, since a per-ledger dimension mints a series per ledger; it is indexed in Tempo as the join key. The ledger.acquire span is not joined yet, because that file was being changed concurrently. It is registered as an optional member of the join group so nothing fails, and switching it is a one-line follow-up. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0f223ee11 |
fix(telemetry): always finalize the ledger.acquire span (WP-B1)
The span only recorded an outcome on the normal completion path. An acquire that stalled and was later swept ended with no outcome at all, and its duration stretched to the sweep interval rather than the real fetch time. So the one case these signals exist to catch, a fetch that never finishes, was the one case that could not be traced, and aggregate outcome and timeout rates read low exactly when nodes are stuck. - Adds an abandoned outcome value for the swept-while-fetching case. - Routes every exit through one idempotent finalizer, so a span is finalized exactly once whether it completes, fails, short-circuits on local data, or is destroyed mid-fetch. The destructor path cannot throw. - Adds the ledger hash to the span and backfills the sequence once known, since by-hash acquires start without one and could not otherwise be tied to a specific ledger. - Record layer: outcome stays a span-metrics dimension in both collector configs, which drift apart if only one is edited. The ledger hash is indexed in Tempo for trace search instead, because a per-ledger value as a metric dimension would mint a new series every ledger. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
01dae36fb1 |
test(telemetry): align workload harness with shared attr names
- peer.validation.receive now asserts the shared bare ledger_hash / full_validation keys (was the dotted xrpl.ledger.hash and validation_full); PARITY_SPAN_ATTRS checks both on the peer span too. - Fix a span-name drift: the per-transaction accept span is txq.accept_tx (op::acceptTx = "accept_tx"), not txq.accept.tx — the old assertion never matched and was silently skipped as optional. - Drop the "intentionally dotted" notes; there is no dotted span attribute. |
||
|
|
899fc3c912 |
fix(telemetry): drop conditional fee attrs from txq.enqueue required set
The txq.enqueue span sets tx_hash, tx_type, and txq_status on every code path (TxQ.cpp:746/748/751), but fee_level_paid and required_fee_level are set only on the fee-evaluated path (TxQ.cpp:895-898), which is reached after the rejected and applied_direct early exits. They are therefore not guaranteed on every txq.enqueue span, so requiring them caused the validation harness to fail whenever a txq.enqueue span took an early-exit path. Remove the two conditional attributes from required_attributes and document why in the span note. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fd1c8c6060 |
fix(telemetry): resolve Phase 10 validation failures surfaced by first full CI run
The print-env CI fix let the Telemetry Stack Validation job build and run the workload harness end-to-end for the first time. It reported 129/136 checks passing; this commit fixes the 7 real failures plus a latent regression-gate bug. Validation-suite fixes (verified against the CI run's actual emission + live node): - expected_metrics.json: the beast::insight job-depth gauge is `xrpld_jobq_job_count`, not `xrpld_job_count` (the latter is a Phase 9 OTel counter). Reverted the prior rename. Removed the statsd_histograms block (`xrpld_rpc_time`/`xrpld_rpc_size`): these RPC timers do not emit under the WS workload (0 series in CI). - expected_spans.json: `tx_status` is only set on suppressed/known-bad receives, so it is no longer a required attribute of every `tx.receive`. Marked `pathfind.compute` and `pathfind.discover` optional and the `pathfind.request -> pathfind.compute` hierarchy as skip — the self-to-self XRP probe returns before computing paths in a fresh cluster with no liquidity, so only `pathfind.request` fires. Regression-gate bug (telemetry-validation.yml "Print regression summary"): - `jq -e` exits non-zero when its filter result is boolean false — the normal case for a populated (non-placeholder) baseline — which was misreported as "Failed to parse baseline JSON" and failed the job. Dropped `-e` (kept `-r`) so a non-zero exit genuinely means malformed JSON. The optional-span handling and regression comparison both worked correctly in the CI run (txq.* / pathfind.update_all skipped-when-absent, 0 regressions detected). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cb9fce6890 |
fix(telemetry): align Phase 10 workload harness with current OTel recording surface + fix CI
The Phase 10 validation harness had drifted from the code's recording surface and the telemetry-validation CI job was failing before it could build. CI fix (telemetry-validation.yml): - Replace nonexistent local action ./.github/actions/print-env with the remote XRPLF/actions/print-build-env (the build-xrpld job failed in 56s on this). - Sync prepare-runner and upload-artifact action SHAs to the canonical workflow. Recording-surface reconciliation (docker/telemetry/workload/): - Migrate span attributes from dotted xrpl.<domain>.<field> to the bare/underscore form introduced by the 2026-05-13 span-attr naming redesign (tx_hash, peer_id, ledger_seq, consensus_mode, consensus_round, full_validation, quorum, ...). Dotted xrpl.ledger.hash is retained only on peer.validation.receive (shared constant), while consensus.validation.send uses bare ledger_hash. - Fix attribute placement: tx.apply carries tx_count/tx_failed (not ledger_seq); ledger.build carries ledger_seq/close_* (not tx_count/tx_failed). - Replace the phantom rpc.request span with the real WS root rpc.ws_message; drop the never-emitted duration_ms; rebuild the parent-child map accordingly. - Add the new spans the code emits: apply-pipeline stage spans (tx.preflight/preclaim/transactor with stage/tx_type/ter_result), txq.*, consensus sub-spans (round/establish/update_positions/check/phase.open), ledger.acquire, grpc.*, pathfind.*. Conditional spans are marked optional so they are skipped (not failed) when the workload does not exercise them. - validate_telemetry.py: service.name and Loki job label rippled -> xrpld; fix PARITY_SPAN_ATTRS (rename the 4 real attrs, drop the 3 that are metrics not span attrs); add optional-span handling that skips missing optional spans while still validating attributes when present. - expected_metrics.json: rippled_ -> xrpld_ on all beast::insight/overlay metrics, xrpld_job_count, the 15 on-disk xrpld-* dashboard UIDs, and the real bare spanmetrics dimension labels. - regression-metrics.json + baseline-timings.json: rpc.request -> rpc.ws_message. Metrics pipeline fix: - Switch node [insight] config from server=statsd/prefix=rippled to server=otel + /v1/metrics endpoint + prefix=xrpld across run-full-validation.sh, xrpld-validator.cfg.template, benchmark.sh and the workload compose. The collector has no StatsD receiver, so system metrics only reach Prometheus over OTLP. Synthetic load for new spans: - Add ripple_path_find to the RPC load generator (drives pathfind.* spans). - Add a high-TPS txq-burst workload phase to force fee escalation (drives txq.*). All facts verified against the *SpanNames.h headers and a live xrpld node + collector (Tempo service.name=xrpld, tx.preflight attrs [stage,ter_result,tx_type], 279 xrpld_ Prometheus metrics and zero rippled_). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
815e2b1f5d |
refactor(telemetry): fix remaining old attr refs in tests, docs, workload
- Update Telemetry.h doc example: xrpl.rpc.command -> command. - Update SpanGuardFactory.cpp test: use new bare attr names. - Update TESTING.md: rename attr refs in span table + PromQL example. - Update expected_spans.json: all attrs match simplified naming. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
e63ca4c495 |
fix(telemetry): fix dashboard UID and add parity attributes to expected_spans
- Remove duplicate 'system-node-health' UID from expected_metrics.json (already covered by 'rippled-system-node-health') - Add parity span attributes to expected_spans.json: node health on rpc.command.*, validation hash/full on consensus.validation.send, quorum/proposers on consensus.accept, validation hash/full on peer.validation.receive Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
5de8c520d1 |
Phase 10: Workload validation - synthetic load generation and telemetry checks
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |