mirror of
https://github.com/XRPLF/rippled.git
synced 2026-07-29 18:10:34 +00:00
kSubMillisecondBoundaries existed but nothing used it, so per-fetch read latency never reached Grafana -- only the coarse read_mean_us gauge did, which cannot separate "every read took 9us" from "most took 2 and a few took 900". Add a nodestore_read_us histogram, register its view against the sub- millisecond ladder rather than kMicrosecondBoundaries (whose first edge is 100us, above the entire range a warm read occupies), and record into it from NodeStoreScheduler::onFetch using FetchReport::elapsed, which a previous change widened to microseconds for exactly this purpose. The name and its labels live in a new include/xrpl/telemetry header because the view registration (xrpld.telemetry) and the record site (xrpld.app) sit in different levelization modules; a copy-pasted literal would let them drift and silently drop the bucket override. Same reason and same placement as GetObjectMetricNames.h. No new levelization edge: xrpld.app > xrpl.telemetry already exists. NodeStoreScheduler had no registry access, so it now takes a ServiceRegistry and resolves the registry per call. It is constructed in Application's initializer list, long before metricsRegistry_ is assigned in setup() and started in startTelemetry(), so capturing a pointer at construction would capture nullptr forever; the metric macros null-check the registry, the meter and the instrument, so early fetches are simply not recorded. Labels are fetch_type and found, both already carried on the report -- 4 series, fixed at compile time. A slow async read delays prefetch while a slow sync read blocks a caller, and a miss can cost a read of every backend, so neither dimension can be collapsed. Negative elapsed times are skipped: the SDK rejects them and logs a warning on every call, which on a per-fetch path is a log flood. Zero is still recorded, since a page-cache-served read genuinely rounds to it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>