mirror of
https://github.com/XRPLF/rippled.git
synced 2026-09-28 07:48:01 +00:00
ce18bb331769a779c296c5c8c9c5ecb7e6eb2ef8
16772 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ce18bb3317 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings in the hierarchy-check sampling fix: the check now asks Tempo for traces containing both parent and child rather than inspecting the three newest parent traces, so a child conditional on a state the workload rarely reaches is found wherever it occurred. Merged clean, no conflicts, no resolution decisions. This unblocks ledger.acquire -> ledger.acquire.txtree on this branch, which was skipped for exactly that sampling problem and is un-skipped in the next commit. |
||
|
|
a87d772f40 |
fix(telemetry): find a span hierarchy where it happened, not only where it is newest
The hierarchy check searched the parent span and inspected the three newest traces it returned. That is wrong whenever the child is conditional on a state the workload only sometimes reaches: the parent fires constantly, so its newest traces are the ones LEAST likely to carry a rare child. Three relationships had been skipped as unassertable for exactly this, and in none of them was the child missing -- each emitted traces of its own and simply was not in the three most recent parent traces. The check now issues a second query, a TraceQL trace-level conjunction of the parent and child name predicates, and inspects those traces. Tempo searches its whole retention for co-occurrence instead of leaving the answer to which traces happen to be newest. The parent-only query is kept and still runs first, so "the parent stopped being emitted" stays a distinct failure from "the parent is there but the child never co-occurs" -- they mean different things to whoever reads the report, and collapsing them would lose that. The returned traces are still verified with _span_name_matches rather than the query result being trusted on its own. Tempo has already guaranteed co-occurrence, so this is redundant on the happy path; it is kept because it keeps the glob semantics in one place and means a wrongly built query cannot silently pass. _traceql_name_predicate handles the wildcard contracts. TraceQL has no glob operator, so `rpc.command.*` is sent as name=~"rpc\.command\..*" with the dots escaped -- unescaped they would match any character in those positions, which is the looseness _span_name_matches exists to avoid. Two entries follow from the fix. txq.accept -> txq.accept_tx is asserted again: its child is created inside the queued-transaction loop behind `if (feeLevelPaid >= requiredFeeLevel)` (TxQ.cpp:1530) while the parent fires on every close (:1499), which was the whole reason it failed. txq.enqueue -> txq.batch_clear stays skipped but for ONE reason now instead of two -- its child never fires at all under this workload, needing an account with a supersedable batch, so it is purely a workload gap and needs nothing further from the validator. The third, ledger.acquire -> ledger.acquire.txtree, lives on the sync-diagnostics branch and is un-skipped there once this merges forward. Written test-first, and the first test this module has had. The failing test reproduces the exact CI message, "txq.accept_tx not found in txq.accept traces", against a stubbed Tempo whose corpus holds the child only in a trace outside the newest three. Three sibling tests guard the ways this could be "fixed" wrongly: an absent child must still fail, a missing parent must still name the parent rather than the child, and a wildcard child must be satisfied by any family member. The stub records the queries issued, so the conjunction is asserted rather than assumed. A stub rather than a live Tempo because the behaviour under test is which traces the check ASKS FOR -- a passing query against real data proves the data co-operated, not that the query was right. The first run of those tests failed for the wrong reason: my stub's name-predicate regex also matched the resource.service.name="xrpld" term every query carries and so demanded a span literally named "xrpld". Fixed in the stub, with the lookbehind commented as load-bearing, before touching production code. Verification: 4/4 tests pass, and the failing one was watched failing first with the production message; the issued queries were printed and confirmed to contain the conjunction; validate_telemetry.py compiles; expected_spans.json parses; 21 relationships, 16 asserted and 5 skipped; counters still 41 span types; otel-naming exits 0. Three unrelated files in this worktree are another party's live work and were deliberately left unstaged. |
||
|
|
da35290f27 |
test(telemetry): compile the span-name and stall-rule tests in every build
Both files gated their whole contents on XRPL_ENABLE_TELEMETRY, so 54 tests were skipped whenever telemetry was compiled out. The stated reason was that only a telemetry build puts `src/` on this target's include path, but that include path is unconditional, so the tests were reachable all along. Nothing in either file needs the OpenTelemetry SDK. The span-name and outcome headers hold constants and constexpr functions with no telemetry guards, LoadManager::evaluateStall is a static constexpr member, and the handful of guard assertions construct a default SpanGuard, which is inactive in either configuration. SpanNames.h documents this contract for its own constants. |
||
|
|
fa2a09c758 |
perf(telemetry): increment the copy-forward total only when it is read
copyForwardTotal_ is a second atomic increment beside copyForwardCount_ on the same event, kept only so a metric never goes backwards: rotate() zeroes the per-rotation tally for its log line, which leaves that counter unusable as a rate. Its only reader is MetricsRegistry.cpp:1153, through copyForwardTotal(). Guard the increment. During a rotation window every archive-served non-duplicate read pays for it, and with telemetry compiled out there is nothing to read it back. copyForwardCount_ is untouched: rotate() exchanges it for the "copied forward N archive-served reads" warning, which is real logging, not instrumentation. The virtual and the member stay declared unconditionally, so the nodestore interface has the same shape in every configuration. |
||
|
|
0ad3587462 |
perf(telemetry): skip the dial clock and span when nothing records them
dialStart_, outcomeReported_ and dialSpan_ are telemetry-only. dialStart_ is read only by the two elapsed-time computations in reportOutcome(); outcomeReported_ is written and read only there; dialSpan_ is opened in run(), ended in reportOutcome() and reset in the destructor, and read nowhere else. outcomeReported_ is not load-bearing for anything but telemetry. It is a first-call-wins latch over the histogram, the counter and the span attributes. Every terminal path calls close() or fail() itself, beside its reportOutcome() call rather than inside it, so suppressing a second report cannot suppress any teardown. Guard the three sites: the clock and span setup in run(), the whole body of reportOutcome(), and the destructor's reset(). Per outbound dial that removes a steady_clock reading, an optional emplace and reset of a span handle, and the latch write. The members stay declared in every configuration so the class has one shape; only the writes are compiled out. SpanGuard.h, SpanNames.h, MetricMacros.h and the cstdint header move behind the guard with the code that names them. ConnectAttempt.h still includes SpanGuard.h for the member. |
||
|
|
d3c1fc67ce |
perf(telemetry): skip the per-second stall bookkeeping nobody reads
updateStallState() is wholly telemetry: it applies evaluateStall(), stores the result in currentStallSeconds_ and bumps stallEventCount_. Those two members have exactly one reader each, MetricsRegistry.cpp:1986 and :2023, both inside the registry's own XRPL_ENABLE_TELEMETRY region, reached through getCurrentStallSeconds() and getStallEventCount(), which nothing else calls. The monitor thread ran it once per second for the life of the process. Guard the body, not the members or the accessors: a member set that differs between build configurations is the hazard that once made a test mock abstract. evaluateStall() stays where it is, being a public constexpr rule with its own GTest coverage in SyncStateSignals.cpp. The atomic header moves behind the same guard, as the relaxed memory orders are named only in the guarded body; LoadManager.h includes it for the members. |
||
|
|
27dc3236d1 |
perf(telemetry): build the UNL fetch site label only when recorded
reportFetchOutcome() exists only to label unl_fetch_total. It reads the parsed URI parts, copies the domain, erases any userinfo with an rfind, joins scheme, host and optional port, then appends a substr of the path -- two std::string allocations and several copies -- and that label has no other reader. It ran on every validator-list fetch, so about once per site every five minutes, whether or not anything could record the counter. Guard the whole body with XRPL_ENABLE_TELEMETRY rather than change the signature: the two failure call sites pass a compile-time constant, so an empty body is all they need. The success call site is guarded too, because its to_string(bestDisposition()) builds a std::string that only the label consumes. bestDisposition() itself keeps running, since lastRefreshStatus stores it. MetricMacros.h moves behind the same guard, as the macro is now named only inside the guarded body. MetricNames.h stays unconditional, because the fetch handlers name the outcome constants either way. |
||
|
|
8521b96d85 |
fix(telemetry): stop an incomplete capture becoming the committed baseline
The regression baseline is bootstrapped by copying a CI artifact. The workflow
tested only that timings.json existed, then printed it verbatim under a heading
inviting the reader to paste it in as the new baseline.
capture_timings.py writes that file and only then enforces --min-capture-ratio,
so an incomplete capture leaves a file that exists but covers fewer keys than
the contract declares. The verdict lived in CAPTURE_EXIT, a shell variable local
to run-full-validation.sh that no other program could read. So on a placeholder
baseline plus a thin capture, CI offered an incomplete artifact as the next
baseline, and pasting it narrowed the gate with nothing reporting that it had.
That is the failure shape this harness keeps producing: a degraded result that
looks exactly like a good one.
The artifact now carries its own completeness, next to metrics:
"capture": { "declared": 20, "captured": 20, "min_ratio": 0.5, "complete": true }
complete is the same condition the producer exits 0 on, computed once with the
exit code read off it, so the flag and the status cannot drift apart. Any
consumer can now tell a complete capture from a thin one, not just CI.
Both paste-me paths refuse rather than warn: the workflow prints the counts and
an error annotation with no JSON, and the comparator explains on stderr while
leaving stdout empty, so a redirect cannot produce a plausible-looking file. A
warning above a copyable block is still a copyable block, and a reader who has
just hit a red gate is already predisposed to re-baseline. A missing capture
block fails closed.
Refusal is scoped to bootstrapping a baseline, not to comparing against one, so
artifacts captured before this change still replay: verified against the run the
current baseline came from, which carries no capture block and still reports 0
regressions. An injected regression is still caught, and the gated surface is
unchanged at 20 keys with 5 excluded.
|
||
|
|
1f8b69a3f0 |
fix(telemetry): skip the tx-tree acquire hierarchy, which the tx set being present hides
Asserting all three ledger.acquire phase parentings treated them as equally
conditional. They are not. Run 33002568549 failed on
ledger.acquire -> ledger.acquire.txtree -- "ledger.acquire.txtree not found in
ledger.acquire traces", the single failure in 279 checks -- while header and
astree passed. Skipped rather than left red.
Not a missing span: txtree reports 5 traces of its own on that same run, one per
node. InboundLedger.cpp opens each phase only when that piece is still needed:
header on !haveHeader_ (:672), astree in the else of haveState_ (:689), txtree in
the else of haveTransactions_ (:698). A node acquiring a ledger here almost always
lacks the account-state tree, so astree opens on essentially every acquire and its
assertion holds. But it usually already HOLDS the transaction set, because every
node sees the same relayed transactions and builds the same set, so
haveTransactions_ is true and no txtree phase opens at all. It fires only on the
minority of acquires where the set was genuinely missing, and with
_validate_parent_child sampling the 3 newest parent traces
(validate_telemetry.py:803) those are not the ones sampled.
This is the third entry skipped for one underlying cause, after
txq.accept -> txq.accept_tx and txq.enqueue -> txq.batch_clear: a child that is
conditional on a state the harness rarely reaches, met by newest-N sampling of the
parent. Preferring parent traces that CONTAIN the child would retire all three at
once, and that is now the highest-value change left in this harness -- recorded in
each of the three reasons so whoever picks it up finds the whole set.
The rest of the run supports the other changes. Both rpc.command.* hierarchies
un-skipped in
|
||
|
|
fef1443a65 |
docs(telemetry): correct what the span-liveness guard actually skips
The comment claimed the guard skips work for a span that is "not being recorded", which reads as sampling awareness. It has none: operator bool() is impl_ != nullptr, and the span factories return an empty guard only when telemetry is absent, disabled at runtime, or the trace category is off. A span that exists but was sampled out still pays. There is no isRecording() in the telemetry API, so the guard is still the strongest available; only the justification was overstated. |
||
|
|
f2cc740c42 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
6d17df083f |
test(telemetry): account for every declared span parenting
Each span entry documents its parent, and a separate list holds the pairs the validator actually checks. Three parentings were declared on the span entries and absent from that list entirely, so they were neither asserted nor recorded as unassertable -- silently missing rather than deliberately skipped. All three are now listed, skipped, each with the reason that actually applies. Every span declaring a parent now has an entry: the count went from 3 unaccounted to 0. txq.enqueue -> txq.batch_clear is conditional and narrowly so. The child is created in TxQ::tryClearAccountQueueUpThruTx (TxQ.cpp:550), which needs one account holding several queued transactions AND an arriving transaction that supersedes the batch. txq-burst produces queueing but arranges no such shape, and it has never been observed on a run. It would also meet the sampling limit that forced the txq.accept_tx skip, so fixing the sampling addresses both at once. rpc.command.* -> pathfind.request is the one skip caused by a wildcard PARENT rather than by a missing span, and the asymmetry is worth recording: _validate_parent_child inserts the parent name literally into its Tempo query (:801), so a wildcard parent matches nothing, while the CHILD side globs through _span_name_matches (:826-828). That is exactly why rpc.ws_message -> rpc.command.* can be asserted and this cannot. pathfind.compute -> pathfind.discover has both ends absent, for the reason the pathfind.compute entry already sets out at length: pathfinding is disabled on every harness node because Config.cpp:725-726 zeroes pathSearchMax when a [validation_seed] is present, and since 2026-08-25 no path-finding RPC is issued either. Listed so the family is fully accounted for rather than partly silent. No assertion is added or removed here -- this is accounting. The plan task that prompted it also assumed the pathfind.compute skip reason was stale and needed correcting; it is not, it already names both blockers and corrects an older liquidity-based reason, so that half of the task was a defect in my plan rather than in the file. Verification: JSON parses; 21 relationships, 15 asserted and 6 skipped; no duplicates; 0 spans declaring a parent without an entry, down from 3; counters still 41 span types; otel-naming exits 0; pre-commit clean. |
||
|
|
c3e4c4244a |
docs(telemetry): explain why the overlay dial metrics report one fewer series
overlay_connect_total and the three overlay_dial_latency_ms series come back with 4 series per run while sibling families such as dns_resolve_* come back with 5, one per node. That was unexplained, so anyone reading the group had to choose between suspecting the exporter and re-deriving the cause. It is a topology artefact and nothing is wrong. OverlayImpl::connect asks peerFinder().newOutboundSlot for a slot and returns early when it gets a null one (OverlayImpl.cpp:464-470), before it constructs the ConnectAttempt that emits both signals (:472). Every node is seeded to dial every other node -- run-full-validation.sh:324-331 builds IPS_FIXED from all NUM_NODES-1 peers and :373-374 writes it into [ips] -- so all five nodes do try. In a full mesh each pair is dialled from both ends, and the node whose peer got there first is refused an outbound slot for an address it already holds inbound: no ConnectAttempt, so neither the counter nor the histogram. dns_resolve_* reports 5 because reportDnsResolve fires inside the resolver handler (OverlayImpl.cpp:603), which runs before any slot allocation. The note also records why the four entries assert series presence rather than a count: hard-coding 4 would bake today's mesh into the contract and break on any cluster-size change, while gaining nothing -- and it says that fewer than 4 would be worth investigating, since that means a node did not dial at all. Verified in code: the early return and its position relative to the ConnectAttempt, the [ips] construction, and the resolver call site. Not verified against a run: which node is missing on any given run, because dial ordering is not controlled and the identity is not expected to be stable. No assertion changed -- this commit adds documentation only. |
||
|
|
a215ab7bb1 |
fix(telemetry): assert the rpc.command hierarchies, whose skips described dead code
Both rpc.command.* relationships were skipped on the claim that
_validate_parent_child collapses a wildcard child to one literal name via
child_name.replace("*", "server_info"). That code does not exist.
|
||
|
|
143abfd8f6 |
fix(telemetry): stop requiring what the harness cannot guarantee
Two entries in the span contract could fail on healthy behaviour. `ledger.serve` is emitted when this node answers another node's request for ledger data, and was mandatory. A node only answers if a peer asks, and a `TMGetLedger` request is constructed in exactly two places -- src/xrpld/app/ledger/detail/InboundLedger.cpp and src/xrpld/app/ledger/detail/TransactionAcquire.cpp -- whose `ledger.acquire` and `txset.acquire` spans are both already marked optional. A mandatory check therefore rested on an optional cause. Marked optional, with the dependency named in the note so it is promoted together with them rather than alone. `peer.dial` covers one outbound connect attempt and required `outcome` and `duration_ms`. Both are set only in `reportOutcome()` (ConnectAttempt.cpp:158-199). The teardown path sets neither on purpose: an attempt destroyed during overlay shutdown, or one whose connect was aborted, ends its span in `~ConnectAttempt` (ConnectAttempt.cpp:89-102), whose own comment states that a span ending with no `outcome` is the honest record of a dial that never concluded. Since `peer.dial` is a freshRoot, each dial is its own trace with exactly one instance of the span, and the validator inspects only the most recent trace -- so one newly-aborted dial fails CI while the node is behaving correctly. `remote_endpoint` stays required; it is set at construction on every path. Verified: both call sites read in current code; `TMGetLedger` construction confined to those two files; the counters recomputed -- 48 span types (matches len(spans)) and 74 unique attributes, unchanged, because `outcome` and `duration_ms` are still required by `txset.acquire` and the `ledger.acquire` family. Not verified: that a run with these entries relaxed still exercises both spans, which only a CI dispatch can show. Neither entry can now fail on healthy behaviour, so a green run proves less than before by design. |
||
|
|
06792a508d |
perf(telemetry): read the acquire peer count only when the span records it
finalizeAcquireSpan() took the peer count as an argument, so all three real exit paths called getPeerCount() before entering it. That walks the acquire's peer set calling findPeerByShortID for each one, taking the Overlay lock every time, and the value is used only to set one span attribute -- so an acquire whose span was never recorded paid for the whole walk. Pass whether the lookup is safe instead of its result, and make the call at its point of use, inside the span-active branch. The destructor keeps passing false for the reason it always had: it can run under the InboundLedgers collection lock, where taking the Overlay lock underneath would be unsafe. The two getPeerCount() calls that drive peer recruitment are untouched; they are real logic, not instrumentation. |
||
|
|
e5d7b2a4b0 |
fix(telemetry): skip the txq accept-pass hierarchy, which sampling cannot assert
|
||
|
|
08026b46b9 |
fix(telemetry): make this branch's files compile clean with telemetry off
The metric macros discard their arguments when telemetry is compiled out, so anything named only as a macro argument disappears in that build. That produced fifteen errors across these files. - guard MetricNames.h in the nine files whose only uses of it are macro arguments; the files that pass those constants as ordinary function arguments still need it unconditionally - drop the prevMode local in setMode, reading the mode being left inline in the macro argument so nothing is computed when telemetry is off - compile out the emit helper in recordBatchOutcome and its three calls, which exist only to report per-outcome counters - drop two includes the telemetry-off test block never used - suppress the static and const suggestions on four methods whose bodies only record metrics; each reads the app_ member when telemetry is enabled |
||
|
|
d5f910cfca |
docs(telemetry): the tx-set fetch now has a span, so stop saying it has none
The consensus round-flow diagram marked `acquireTxSet -> gotTxSet` as having no span, and styled it as a step with no instrumentation. That was true when the diagram was written and stopped being true on this branch, which added the span. The fetch is now `txset.acquire`, carrying one `round.request` event per round that asked for the set. The `gotTxSet` delivery still has no span of its own, so a set that arrives too late for the round to use it leaves no trace -- worth stating, because that is the case an operator goes looking for. Corrected on this branch rather than upstream on purpose: phase-9 and phase-10 carry the same diagram but not the span, so "no span" is accurate there. Moving the fix upstream would describe instrumentation those branches do not have. |
||
|
|
92e988c8b8 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
504138dd9c |
test(telemetry): assert all three ledger-acquire phase hierarchies
The acquire span opens three phase children -- header, account-state tree and transaction tree -- and the contract documented all three parentings while checking none of them. astree had an entry marked skipped; header and txtree had no entry at all. The skip reason was that the parent is optional, because a healthy 5-node cluster agreeing from genesis rarely back-fills history, so the hierarchy check would fail against a parent with no traces. Run 32969481032 refutes the premise: ledger.acquire reported 5 traces and so did each of the three children, one per node. The parenting was never the uncertain part -- beginPhaseSpan() parents through the acquire span's own captured SpanContext rather than the ambient thread context, so it holds whichever worker opens a phase. Asserting these matters because of what the phases are for. A fresh sync is dominated by the account-state tree, and the flat parent span cannot separate that from the much smaller transaction tree or from the header wait that gates both. If a phase stops nesting under the acquire it still emits, still carries its missing-node count and its timeout flag, and nothing else in this harness notices -- but the trace stops answering which phase the sync is stuck in, which is the whole reason these spans exist. Routed here rather than to phase-10 because phase-10 has no ledger.acquire.header or ledger.acquire.txtree span at all; the phase children were introduced on this branch. Verification: JSON parses; 10 relationships, 6 asserted and 4 skipped, no duplicates, every non-wildcard child resolves to a declared span entry; counters unchanged at 48 span types and 74 unique attributes; otel-naming exits 0; pre-commit clean. Edited by surgical text replacement -- a first attempt used a json.dumps round-trip and reflowed the whole file, 148 insertions against 47 deletions with unrelated compact arrays expanded and unrelated notes rewritten; that was reverted and redone, and the churn is now 11 against 3. Whether a child is findable INSIDE the parent's fetched trace is what CI will decide. |
||
|
|
c531ac569b |
fix(telemetry): trigger the workload by what changed, and assert the span tree
Two problems, both about coverage this workflow claims to have and does not. The push trigger gated on branch NAME as well as path, and GitHub ANDs the two. Branch names are not something this repository controls, so a push to any branch outside "pratik/otel-phase*", "feature/otel-*" or "feature/telemetry-*" was never dispatched -- not queued, not skipped, no run to look at. That is not a theoretical gap: two rounds of harness fixes on pratik/otel-sync-diagnostics produced no signal at all before anyone noticed the workflow had never started. The branches filter is removed; the paths already express the real question. The path list was also incomplete in a way that matters more than it looks. The span-name and metric-name headers are the wire contract this harness asserts against by literal string, and the convention colocates each one with the class it serves -- so eight of the ten *SpanNames.h headers live under consensus/, overlay/, app/ledger/, app/main/, app/misc/, rpc/ and tx/, none of which was matched. Renaming a span constant therefore compiled clean, emptied the assertions and triggered nothing. Matched now by filename, "**/*SpanNames.h" and "**/*MetricNames.h", so future headers are covered wherever they land. Added for the same reason: include/xrpl/beast/insight (the interface headers decide what the collector can publish, so they move the metric surface as surely as the implementation), src/tests/libxrpl/telemetry (the GTests pinning those constants), and the two checker directories that gate this surface in CI. Second, the span hierarchy. Each span entry documents its parent, and a separate list holds the pairs the validator actually checks in Tempo. Those had drifted apart: 18 parentings were documented, 7 were checked. A span that stops nesting under its parent -- which is what a detached guard does -- leaves every span and every attribute intact, so no other check in this harness notices; the trace simply stops being readable as one operation. Eleven pairs are added, each one where both ends emitted on a real run: rpc.http_request -> rpc.process, the three txq parentings, and seven consensus ones under consensus.round and consensus.establish. Fourteen of eighteen are now asserted; the four still skipped are the wildcard rpc.command.* families and pathfind.compute. Three notes were also factually wrong, all repeating one mistake. They said rpc.process and rpc.http_request cannot appear because that path is HTTP-only while the load generator is WebSocket-only. The premise is right, the conclusion is not: both appear on every run, five traces each, because run-full-validation.sh polls each node's HTTP port with curl for readiness and validated-ledger progress (:449, :502). Those polls take the HTTP path. A reader acting on the old text would have gone looking for a way to make the harness speak HTTP that it already speaks. The rpc.process -> rpc.command.* skip reason inherited the same error and additionally claimed the WebSocket equivalent is "asserted above instead", which it is not -- that one is skipped for the same wildcard limitation. All three now state the real blocker, which is that _validate_parent_child resolves a wildcard child to a single literal probe. Both HTTP spans stay optional rather than being promoted: the curl polls are harness scaffolding, not workload, and a future change to how the script waits for a node could legitimately remove them. Verification: JSON parses; 18 relationships, no duplicates, every non-wildcard endpoint resolves to a declared span entry; counters still 41 span types and 62 unique attributes; workflow YAML parses, has no branches key, keeps workflow_dispatch, and every new glob was checked against the tracked file list with a matcher that reproduces GitHub's ** semantics; otel-naming exits 0; pre-commit clean on both files. The eleven new assertions are proven only to the extent that both ends emitted on run 32969481032 -- that a child is findable INSIDE the parent's fetched trace is what CI will now decide. |
||
|
|
44fd31f7cd |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
f13524c93c |
fix(telemetry): map every harness failure to exit 2, and always capture timings
Two defects raised in review of PR 6519, both about the harness misreporting its own state. The script documents exit 2 for an infrastructure failure and routes that through die(), but eleven commands were unguarded, so under set -euo pipefail a failure aborted with the tool's own status instead. Measured before the fix: docker compose exited 125, the key generator 7, a jq read 5, and several others 1 -- which the table defines as "checks failed", so an infrastructure problem was reported as a validation result. Two of the eleven are worth naming. A trailing option with no value (--nodes at the end of the command line) exited 1 because set -u aborted on the unset positional, now unified through one require_value helper. And report_stopped_nodes, which runs immediately before a die, contained an unguarded pipeline that tripped errexit, so the die never ran and a crashed cluster reported 1 -- the script failed to report the exact condition the contract exists for. Commands whose failure is genuinely tolerated were left alone. The seed read also gained a value check, because jq prints the string "null" and exits 0 for a missing key, so testing only the exit status cannot see it. Step 6 said it "ALWAYS captures timings (so CI always has an artifact from which to bootstrap/refresh the committed baseline)" while the capture sat inside the --skip-regression guard. The comment stated the intent and the code was the bug: that artifact is the only route to a refreshed baseline, and the workflow reads it unconditionally to print the paste-me block. Capture now always runs and only the comparison is gated. A capture failure still surfaces, folding into the exit code only when the gate is active, so --skip-regression cannot start failing runs that previously passed. Note a non-zero capture status does not mean the file is absent: capture_timings writes it and then fails the minimum-ratio check, so the artifact exists but is incomplete. The messages say incomplete rather than missing, so nobody goes looking for a file that is already there. The runbook's matching claims are corrected in the same commit: it said --skip-regression skips the capture, and its exit-code summary predated the uniform mapping. |
||
|
|
a4fedeceed |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
29673de531 |
fix(telemetry): correlate WebSocket replies, and stop silent placeholder data
Eight defects found in review of PR 6519. Twenty-four review threads reported them; ten were duplicates of one another and four were wrong about the code. The two that corrupt data. Both WebSocket clients reuse the socket after a recv timeout, and the library queues the late reply, so the NEXT request reads the previous response. In rpc_load_generator that misattributes latency, and the skew is permanent rather than one-off. In tx_submitter it is worse: a submit that reads an account_info reply freezes that account's sequence number and every later transaction for it fails. Both now correlate replies by request id under a single overall deadline, with a counter rather than a wall clock, since time.time() is not monotonic and collides within a tick. workload_orchestrator never cleared its fixed report paths, so a run that produced no report silently adopted the previous run's totals -- breaking the invariant evaluate_exit_gate documents. Reproduced by planting a stale total and watching it appear in a later summary. collect_system_metrics reported placeholders as if they were measurements. The consensus mean used bc with a || echo 0 fallback that neither warned nor cleared METRICS_COMPLETE, unlike every sibling path; it now uses awk, already a hard dependency here, which removes the failure mode instead of reporting it. Note this moves the mean from truncation to rounding, at most 1 ms on a value of about 45 s. Unmeasurable TPS now warns and clears the flag too. All four curl probes gained a timeout, not just the one the review named -- an unresponsive node could hang any of them. The orchestrator's help text claimed 18-dashboard coverage; there are 15 on disk, 15 uids in the contract, and the profile already said 15. The tx_submitter docstring listed twelve transaction types where ten exist, and claimed issued-currency payments that build_payment never sends. Two suggested patches were deliberately not taken. A recursive delete of the report parent sits in a per-task function and would delete earlier phases' reports mid-run, and recursively remove a caller-supplied --report-dir. The stale microsecond axis label on ledger-data-sync is real but belongs to phase 9, which carries a byte-identical copy of that dashboard, so fixing it here would leave that PR wrong and guarantee a conflict. |
||
|
|
a734da8b33 |
test(telemetry): recapture the baseline and stop gating what variance dominates
Refreshes baselines/baseline-timings.json from run 32964262700 at |
||
|
|
6cb02a1b40 |
fix(telemetry): make the Loki diagnostic count real, and Tempo errors visible
Three defects in the harness's own instrumentation, all of the same shape: a failure that reads as an absence. The Loki diagnostic reported "unavailable entries" rather than a count. It issued an unaggregated count_over_time, and because the filelog regex_parser leaves message and timestamp as log-record attributes, Loki's OTLP path turns those into structured metadata, which joins a metric query's label set. The query therefore produced one series per log line and Loki answered HTTP 400, maximum number of series reached. A second bug hid the first: the JSON helper never checked resp.status, so Loki's own explanation arrived as a mimetype complaint instead. Both fixed, in the Python and the shell twin, and verified against a real loki 3.7.6 including a genuine-zero control so that zero stays distinguishable from unavailable. _tempo_search and _tempo_get_trace called resp.json() with no status check, so any non-2xx became "0 traces" or "0 spans" -- the same class of bug as the span.name tag returning 200 with an empty list. A 404 on /api/traces/<id> legitimately means "not indexed yet", so that stays an absence and every other non-200 now raises. log.trace_id_cross_reference queried Tempo once, with no retry, while the metric checks share a poll deadline for exactly this race. It now polls on the existing METRIC_POLL_TIMEOUT_SEC/INTERVAL, so a trace that has not yet been indexed is retried rather than reported missing. The window stays at 4 hours and the assertion is unchanged. |
||
|
|
3d61ceae6e |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
33956ec240 |
fix(test): alias the second namespace the merged test file needs
|
||
|
|
8418d474a7 |
fix(telemetry): read span names from the Tempo name intrinsic
The span reverse-coverage check has never evaluated. It reported "no span
names were reported (backend unreachable or empty)" on a run where Tempo
demonstrably held data -- the same run resolved a logged trace id to 32
spans.
Root cause: the tag-values query asked for `span.name`. A span's name is a
TraceQL intrinsic, not a span-scoped attribute, so `span.name` resolves to
an attribute nothing sets. Tempo answers 200 with an empty tagValues list,
which is indistinguishable from an empty backend and never raises, so the
surrounding try/except stayed silent.
Verified against tempo 2.9.4 holding exactly one span named
probe.reverse.coverage, with the collector in front of it:
/api/v2/search/tag/span.name/values -> {"tagValues":[]}
/api/v2/search/tag/name/values -> that span's name
/api/v2/search/tag/resource.service.name/values -> xrpld
The third line is the control: the span was in Tempo, so the first line's
emptiness was the wrong tag rather than no data. Cross-checked against a
populated Tempo elsewhere, whose span scope lists real attributes
(command, ledger_seq, tx_hash) and no name tag at all, while the bare
intrinsic returns the whole span inventory.
This is pre-existing, not a regression in the reverse check: the same URL
fed the operations diagnostic before that check existed, and the last
green run before it also logged "Tempo operations (0 total)". The check
faithfully reported an empty input; the input was broken.
The neighbouring resource.service.name query is correctly scoped and is
left alone.
|
||
|
|
1e341d5413 |
fix(test): repair the merged ConsensusSpanNames test file
The add/add resolution in
|
||
|
|
493475a9d4 |
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Brings phase-10 up to
|
||
|
|
394ed2cbc0 |
fix(telemetry): correct the sync-diagnostics harness rationales and gaps
Review of |
||
|
|
42a72863bb |
fix(telemetry): make the sync-diagnostics metric gate actually assert
The sync_diagnostics group asserted nothing. assert_sync_diagnostics_metrics called _check_prometheus_metric with five positional arguments against a six-parameter signature: `report` landed in `deadline` and `sem` was omitted entirely, so the call raised TypeError before a single metric was queried. Neither run_validation nor main catches anything, so the traceback propagated, run-full-validation.sh recorded the non-zero exit as a validation failure, and the four phases ordered after it -- dashboards, both parity checks and log-trace correlation -- never ran at all. Reproduced directly: TypeError, zero checks recorded. Even with the arity corrected the group would still have passed silently, because _check_prometheus_metric RETURNS its CheckResult rather than recording it and the value was discarded. Both halves are fixed by adopting the fan-out validate_metrics already uses: one shared deadline, a concurrency semaphore, gather, then report.add per result. The same call now records 55 checks where it previously recorded none. With the gate live, the inventory it guards had to be made honest. Two metrics could never have passed it. unl_fetch_total is emitted only from ValidatorSite::reportFetchOutcome, which indexes sites_[siteIdx]; sites_ comes from [validator_list_sites], and the harness writes a static [validators] file with no list site anywhere, so no fetch outcome is ever reported. handshake_negotiation_fail_total needs a rejected handshake, and no reject path was found to be reachable between identical localhost nodes. Both move to not_asserted.metrics_excluded, which is where the file's own description says workload-gated names belong. Eleven further conditional metrics -- the acquire, replay, disconnect, serve, jump and sweep counters -- were documented only inside free-text notes; they move to the same map. That matters beyond tidiness: _accounted_metric_names harvests metrics_excluded keys, so a name recorded only in prose is reported as unaccounted, and a prose note cannot be linted at all. Two metrics were wrongly excluded. rotation_state's callback gates only on dynamic_cast<DatabaseRotating*>, and online_delete=256 is set by both the cfg template and run-full-validation.sh, so SHAMapStoreImp builds a DatabaseRotatingImp, the cast succeeds, and both sub-series are observed on every collection tick. The note claiming the harness could not produce them conflated "no rotation runs" with "no series published"; the first is true and bounds the values, the second is false. Both are now asserted at value 0, where absence rather than the zero is the regression, and the note is corrected. The four new histograms listed only _bucket, or _bucket and _count. Each now lists _sum as well, matching the rpc_method_us and job_queued_us convention, so an exporter regression that drops one series cannot pass. On the span side, ledger.validate and ledger.store are the two ends of the per_ledger trace-join group, and the join is computed by hashing ledger_hash -- yet neither required it. Both spans take it unconditionally from makeLedgerTraceSpan, so requiring it is free, and without it a lost join key surfaces only as "spans landed in separate traces", naming the consequence instead of the cause. Deliberately unchanged: ledger.serve stays required and peer.dial keeps its current required attributes, though both look unsafe -- ledger.serve can only fire if an optional span fires first, and peer.dial's destructor exit sets neither outcome nor duration_ms. Those weaken assertions rather than add coverage, so they are reported rather than changed here. Verification: TypeError reproduced before the fix and absent after, with 55 checks recorded; both JSON files parse; no name is both asserted and excluded and none is duplicated; the declared span counters remain consistent at 48 and 74, proven by injecting an extra attribute and watching the check fail; check_otel_naming.py exits 0, and Rule K was proven to read these entries by injecting a bogus name in an owned family and observing exit 1; pre-commit passes on all three files; the levelization baseline is unchanged. NOT compiled -- no C++ changed. |
||
|
|
3836078a78 |
fix(telemetry): stop gating ledger.validate p95 and p99, which vary too much
The regression gate has been red on runs with no code change. Only two of the 25 gated keys ever tripped, both on the same span and never together: run 32862589645 failed p99 at 25.8750 ms against a 1.0600 ms baseline (+2341%), run 32867433073 failed p95 at 0.7500 ms against 0.2404 ms (+212%), and in each run the other quantile sat well inside its own bound. A real slowdown would move both. This is variance, not a defect. Measured across four CI runs: span.ledger.validate.p50 0.0484 to 0.0778 ms 1.6x spread kept span.ledger.validate.p95 0.1281 to 0.7500 ms 5.9x spread excluded span.ledger.validate.p99 0.3875 to 25.8750 ms 66.8x spread excluded Both excluded quantiles reach past their trip point on a healthy run. The mechanism is arrival timing, not slow code: the span opens only once a quorum-completing validation arrives (LedgerMaster.cpp:987, inside checkAccept, past the early return) and wraps the promotion work that follows, so one slow consensus round dominates the tail of a 3m rate window and which round that is differs every run. Widening is not available and must not be attempted later: tolerating 25.8750 ms against a 1.0600 ms baseline needs a bound of about 24.8 ms, which gates nothing. A bound admitting every healthy run's worst case admits every regression too. p50 stays gated; it is stable. THE GENERAL RULE, recorded so this does not recur: an absolute bound derived as hi_next minus baseline comes from the histogram ladder, so it budgets for quantization noise and for nothing else. It knows nothing about how far a metric moves between runs on identical code. Before gating any key, check its observed maximum across several runs against its trip point and gate it only with margin. Spread alone proves nothing: tx.apply.p50 swings 364x and never fires, because its 5 ms trip point absorbs the range. Of the 23 keys still gated the worst reaches 0.67 of its trip point. Mechanism: spans.names lists span names while _quantiles is shared, so dropping two quantiles of one span cannot be expressed by deleting a name. regression-metrics.json gains an excluded_keys map from a flat key to the reason it is not gated, subtracted by both prom_queries.py (so the key is never queried) and check_regression_bounds.py rule A. A per-name quantile override was rejected: a typo there leaves the key gating, whereas a typo in an exclusion subtracts nothing and new rule F rejects it, along with an empty reason, a leftover threshold override and a leftover baseline value. Derived figures recomputed from the committed baseline: 25 gated keys to 23, detection floor 2.02x-9.43x to 2.02x-9.42x, weakly guarded keys ten to nine, bound over baseline 102%-843% to 102%-842%. The baseline edit is a deletion of two entries only, with no value rewritten. Verified: both previously failing runs replay to zero regressions and exit 0; a tenfold increase injected into each of the 23 remaining keys in turn is still caught in all 23 cases; rule F was confirmed load-bearing by stubbing it out, which lets a stale exclusion pass. |
||
|
|
c65cb0e2a8 |
feat(telemetry): gate log-trace correlation in CI with per-leg diagnostics
The two log-correlation checks have never executed in CI: the workflow
hardcoded --skip-loki, so validate_telemetry.py never constructed
log.trace_id_present or log.trace_id_cross_reference. A green Telemetry
Validation therefore carried no evidence that a log line reaches Loki with
trace context. Drop the flag so both checks run and can fail the job.
Correlation spans four independent legs and a failed check names none of
them, so run-full-validation.sh now prints a per-leg diagnostic after the
suite whenever the checks are enabled:
node per-node debug.log line count, the count matching the injected
trace_id/span_id shape, one sample line, and the severity mix,
so "no log at all", "log level too high" and "no active sampled
span" are distinguishable
mount the container-side listing of /var/log/xrpld, taken with the
collector's own mounts and uid. That image is built from
scratch and carries no shell, so the listing runs in a
throwaway container with --volumes-from, not via docker exec
collector the receiver's watched files, logs-pipeline warnings, and the
internal log-record counters, read from inside the container's
network namespace because that endpoint binds to the
container's own localhost and its port is not published
loki the exact query used, the label inventory, and entry counts for
the stream selector with and without the line filter, so "Loki
has nothing" and "Loki has lines but none carry a trace id" are
distinguishable
The diagnostics are non-fatal by construction: every leg runs in its own
subshell with errexit off, each docker and curl call is guarded, and the
coordinator always returns success. Verified with no containers and no Loki
reachable, with an emptied PATH, and with a leg forced to exit non-zero.
validate_telemetry.py gains a matching diagnostic beside the checks,
following _log_prometheus_metric_names: warnings only, never a check
result. Its stream selector and line filter move into module constants
that the shell diagnostic reads back, so the two cannot drift into
describing different queries.
No check was widened or auto-passed, and LOG_QUERY_WINDOW_SECONDS stays at
four hours; a wider window would let a check pass on a previous run's logs.
|
||
|
|
59a0595a6e |
fix(telemetry): stop the workload harness issuing refused path-finding RPC
Every node the harness starts is a validator, and validators disable pathfinding: Config.cpp:725-726 zeroes pathSearchMax whenever a [validation_seed] or [validator_token] section is present, and run-full-validation.sh writes [validation_seed] into every generated node cfg (:308) with no [path_search] section to put the default back. So doRipplePathFind refused every call at RipplePathFind.cpp:48-49 and the 3% ripple_path_find weight bought no coverage at all. It was not free either. The pathfind.request guard is constructed at RipplePathFind.cpp:35, above that refusal, so each refused call still exported a span, and the enclosing rpc.command.ripple_path_find span carried rpc_status=error. That put a steady 3% error floor into span_calls_total for STATUS_CODE_ERROR: any error-rate threshold derived from harness data before this change was measuring the harness rather than xrpld, and needs re-deriving. Removing the load makes pathfind.request unreachable, so it moves from required to optional in expected_spans.json; without that the span check would fail on every run. Three notes in that file and three in expected_metrics.json made claims that are now false, two of them citing line numbers this commit deletes; all six are corrected. The runbook required/optional count moves 26/15 to 25/16. Two facts a future reader needs. First, the weights previously summed to 103, not 100, so every percentage the docstring stated was wrong: health checks were really 38.8%, not 40%. Dropping the 3 makes the sum exactly 100 and every stated percentage correct for the first time. expected_spans.json also carried live arithmetic off the old total, "25/103 ... roughly 43%", now 25/100 and 42%. Second, baselines/baseline-timings.json was captured WITH this load. Only span.rpc.ws_message p50/p95/p99 of the 25 gated keys sees the RPC mix, and their trip points sit 3.1x to 5.9x above baseline, so the gate will not fire. But a timing baseline is workload-specific and its profile field still reads full-validation, so nothing will flag the drift: refresh it from the next CI run's timings artifact. Pathfinding now has no coverage in this harness at all. The workload README section "Pathfinding is not exercised" records that cost, the manual verification route, and a four-step restore recipe in which steps 1 and 2 alone only reinstate the error floor. |
||
|
|
9eaa94c2f0 | Merge branch 'pratik/otel-phase9-metric-gap-fill' into pratik/otel-phase10-workload-validation | ||
|
|
879ad9fbe0 | Merge branch 'pratik/otel-phase8-log-correlation' into pratik/otel-phase9-metric-gap-fill | ||
|
|
b3a5b2b8e9 | Merge branch 'pratik/otel-phase7-native-metrics' into pratik/otel-phase8-log-correlation | ||
|
|
7eea169f0e |
Merge branch 'pratik/otel-phase6-statsd' into pratik/otel-phase7-native-metrics
Resolved .codecov.yml: kept this branch's OTelCollector ignore entries and the incoming corrected comment for the span-header globs. |
||
|
|
808a3a3cfb |
Merge branch 'pratik/otel-phase5-docs-deployment' into pratik/otel-phase6-statsd
Resolved .codecov.yml: kept the incoming telemetry ignore block, which now originates on phase-1b. It is a superset of the block this branch carried (adds *SpanLabels.h) and states the correct reason telemetry files record no coverage. |
||
|
|
115a2cc890 | Merge branch 'pratik/otel-phase4-consensus-tracing' into pratik/otel-phase5-docs-deployment | ||
|
|
d4bc355ec7 | Merge branch 'pratik/otel-phase3-tx-tracing' into pratik/otel-phase4-consensus-tracing | ||
|
|
fbc10fd953 | Merge branch 'pratik/otel-phase2-rpc-tracing' into pratik/otel-phase3-tx-tracing | ||
|
|
da171527ef | Merge branch 'pratik/otel-phase1c-rpc-integration' into pratik/otel-phase2-rpc-tracing | ||
|
|
b9527c0c96 | Merge branch 'pratik/otel-phase1b-telemetry-infra' into pratik/otel-phase1c-rpc-integration | ||
|
|
2250bdfe52 |
ci(codecov): move the telemetry ignore block to the branch that adds telemetry
The block was added on the StatsD branch, but telemetry sources start here. Codecov config only flows child-ward, so every branch between this one and that one kept reporting telemetry files as uncovered patch lines. Also corrects the rationale. The old comment said telemetry is "not enabled in coverage builds"; it is enabled — conanfile.py and CMakeLists.txt both default it ON, and the coverage matrix leg does not turn it off. The reason these files record no coverage is that the unit-test suite never starts an exporter. Adds *SpanLabels.h alongside *SpanNames.h: same compile-time-constant category, and the existing glob did not match it. This narrows the patch gap but does not close it — instrumentation added to consensus and overlay files is still counted, so codecov/patch stays red on the early phases. |
||
|
|
c863b83a1c |
feat(telemetry): warn on telemetry the harness contract does not account for
validate_metrics and validate_spans only ever run one direction: read the contract, ask the backend whether each listed name exists. Nothing looked the other way, so a metric family or span name the contract omitted was invisible by construction. Both emitted inventories were already being fetched for the CI log and neither was compared back, which is how a 345 family metric gap and 7 unknown span names went unnoticed. Add two reverse checks, metric.reverse_coverage and span.reverse_coverage. Each names every emitted family the contract never mentions, sorted, one per line, with counts in the report details. Warn only, by design. passed is hardcoded True in a single shared builder, so an unaccounted name cannot turn CI red: downstream branches legitimately add telemetry an upstream contract has not seen yet, and a hard failure would redden all of them for doing the right thing. Bulk families are accounted for declaratively. A new top level accounted_patterns list in expected_metrics.json holds anchored regexes with a written reason each, covering the 105 per job type queue gauges, the 70 per job type histogram families, the 228 overlay per category traffic families, and the Prometheus scrape plumbing that is not xrpld telemetry. Job type shapes are reduced structurally because every job type name lowercases to letters only; traffic categories are enumerated instead, because they contain underscores and a structural pattern there would swallow unrelated names. Anything outside these shapes still surfaces. Exporter shapes are folded before matching, so a histogram triple is accounted for by an entry written for its base family and is never reported as three separate gaps. Spans need no pattern list: the reverse check reuses the same matcher the forward check uses, so a glob such as rpc.command.* covers every command it expands to, and an optional entry still counts as known. Also fix the diagnostic these checks feed on: both emitted lists were logged as a single Python list repr, about 15 kB on one line for 422 families, unreadable and impossible to compare between runs. Both now print one name per line. _metric_check_targets now selects groups by testing that the value is an object, rather than by excluding two key names, so a non group top level key cannot break it. Output is byte identical: 79 metric plus 5 label checks, same names in the same order. |