mirror of
https://github.com/XRPLF/rippled.git
synced 2026-08-21 22:30:57 +00:00
Second pass on the 54-panel audit. The first commit handled the count panels and the description drift; these are the query and threshold defects. Panels 4 and 6 (DNS Resolve / Outbound Dial p95) read NaN for the whole run. Two faults compounded: both hard-coded [5m] instead of $__rate_interval, so they ignored the dashboard time range entirely, and both wrapped the histogram buckets in rate() even though DNS resolution and outbound dialling only happen during startup -- a windowed rate over a series that stopped moving is 0/0. Dropping rate() and reading the cumulative buckets gives the real distribution: p95 900 ms for DNS, 1750 ms for dialling. Panel 52 kept its rate() (consensus rounds are ongoing) but its hard-coded [5m] became $__rate_interval. Panel 15 plotted 105,892,534 "ledgers behind" during the flagship window. The underlying cause is in NetworkOPs.cpp -- getLedgersBehindNetwork() subtracts the validated sequence from a networkTarget of 0 before any peer has reported -- and that still needs a code fix. Meanwhile one sentinel spike flattened the real 0-20 backlog for the rest of the window, so the query now clamps at 1e6, far above any true backlog. Reads 4 where it read 105 million. Panel 17's sum by (from, to) dropped node identity, so with All nodes selected every node's transitions summed into one bar. It now carries service_instance_id, xrpl_branch and xrpl_work_item like every other panel. Panel 38 divided by clamp_min(op_rate, 1), which turns "no operations in this interval" into "one operation", reporting the whole duration total as if a single op had consumed it. Replaced with a `> 0` gate so an idle interval draws a gap instead of a fabricated latency. Panels 13 and 45 had inverted thresholds: green began at 1 second, so every sub-second time-to-full and time-to-first-validated rendered red -- the healthy case was the alarming colour. Now green by default, yellow past 10 minutes, red past 30. Panel 48's p95 had no outcome filter, mixing abandoned and timed-out spans (which sit at the retry ceiling by construction) into what reads as completion latency. Restricted to outcome="complete". Verified against Grafana Cloud: 66 queries, 0 parse errors, 56 with data, 9 legitimately empty (fault-only counters plus the write-timing series that pre-dates both recorded nodes). The single remaining NaN is on panel 47 and is an artifact of my verification harness forcing a hard 5m window; the panel's own $__rate_interval returns 492 ms, so it needs no change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>