mirror of
https://github.com/XRPLF/rippled.git
synced 2026-08-21 14:20:56 +00:00
Four defects found by a query-correctness audit of all 16 dashboards, all scoped to panels this branch owns. job-queue "Current Job Latency (p99 Gauge)": the histogram by-clause dropped service_instance_id, collapsing every node into one fleet-wide p99 and hiding a slow node. Measured: gauge read 1325us while the worst node was 1868us. The panel's displayName already referenced service_instance_id, so it was also rendering an empty label. Now groups by service_instance_id and xrpl_work_item, matching sibling panels 5, 6 and 7. ledger-data-sync "NodeStore Read Latency (Bottleneck Discriminator)": replaced clamp_min(<denominator>, 1) with (<denominator> > 0). clamp_min clamps the value, not just the zero case, so any node reading below 1/s got a fabricated denominator. Demonstrated with a zero-rate denominator: clamp_min invents 40.2/40.3/11.2/9.2 where the > 0 guard correctly returns no data. This panel is the bottleneck discriminator, read during a stall, which is exactly when the read rate collapses and the clamp is most wrong. Panels 21 and 23 carry the same defect but originate on phase-7 and are fixed there. node-health thresholds: percentunit fields are compared against the raw value, so a step of 80 needed 8000% and could never fire. Rescaled panels 74, 81 and 85 to 0.8. Panel 81 is a found-ratio where high is healthy, so its bands were also inverted. Note these three panels use palette-classic with thresholdsStyle off, so the steps are currently dormant rather than visibly wrong. node-health panels 81 and 85 descriptions: both described the multi-series panels they were split from. Panel 81 carried a byte-identical copy of panel 80's text, promising three plotted rate lines where it draws a single bounded ratio; panel 85's text described read-thread gauges absent from its expression. Rewritten to match the actual queries.