Files
rippled/include/xrpl/telemetry/HistogramBuckets.h
Pratik Mankawde 6e2b2da772 fix(telemetry): resolve microsecond latencies below 100us
The microsecond ladder's first edge was 100us, which sat ABOVE the mass of
every instrument using it. Measured on devnet: 99.3% of job_queued_us
samples, 92.5% of job_running_us and 90.4% of getobject_lookup_us fell in
that first bucket. histogram_quantile then interpolated inside bucket 0 and
returned `quantile / fraction_in_bucket_0 x first_edge` -- p75/p95/p99 of
job_queued_us read 75.52/95.66/99.69us against a prediction of
75.53/95.67/99.70. Three-decimal agreement: those panels were reporting
arithmetic on the bucket edge, not latency.

The fix was already half-written. kSubMillisecondBoundaries had been parked
in MetricsRegistry.cpp as [[maybe_unused]] with a comment noting exactly this
problem for nodestore reads. Its edges are now folded into kMicrosecondBuckets
rather than deleted, so the parked intent is carried forward: 1..1000us
resolution where the mass is, upper edges unchanged so multi-second stalls
stay measurable.

Also moves the GetObject count and charge ladders into HistogramBuckets.h, so
all five ladders have one owner and one set of invariant tests (29 now).

Adds check_bucket_parity.py, wired into the existing OTel naming workflow.
The C++ millisecond ladder and the collector's spanmetrics ladder are
specified to agree over their shared range; they were identical when shipped,
then the collector side alone was extended and nothing noticed for eleven
phases. The check asserts containment rather than equality, because jobs
outlive spans -- jobq_updatepaths averages ~60s, which no span approaches, so
demanding equality would force a ceiling that censors it. Verified it rejects
a missing collector edge, a bogus in-range edge, and a return to the 5s
ceiling.

ledger-data-sync's "Job Queue Wait p95 By Type" moves off the beast
jobq_*_q_milliseconds pair onto job_queued_us filtered by job_type. Those
beast metrics are ms-quantised at the source (Event rounds up to a whole
millisecond), so 94-100% of their samples sat in the first bucket and no
ladder change could fix them. Note the label values are camelCase
(job_type="ledgerData"), not the lowercase metric-name fragments.

Both histogram-fed alert thresholds re-validated and left unchanged, with the
measured basis recorded so neither gets tuned against the old artefact: only
0.0022% of job_queued_us samples exceed the 1s threshold, and every edge
bracketing the 1000ms ios_latency threshold survived the ladder change.

Docs: the rpc_size "known issue -- tracked separately" notes in the runbook
and 09-data-collection-reference are now resolved notes, the stale 10-edge
span_duration bucket list is corrected to the collector's real 20, and the
runbook gains a "Reading A Histogram Percentile" section covering both
saturation traps and the expected discontinuity after a ladder change.
2026-08-21 12:46:56 +01:00

245 lines
9.2 KiB
C++

#pragma once
#include <array>
#include <cstddef>
#include <span>
#include <vector>
namespace xrpl::telemetry::buckets {
/**
* @file HistogramBuckets.h
* @brief Explicit histogram bucket edges for xrpld's OTel instruments.
*
* One header owns every ladder so a reviewer sees all of them at once and a
* test can assert their invariants. Before this existed the edges lived as
* file-local `namespace {}` constants, unreachable from any test, and they
* drifted apart.
*
* Why a ladder is worth this much care: when a quantile falls in the `+Inf`
* bucket, Prometheus returns the *second-highest* edge, not `+Inf`. A
* saturated histogram therefore reports a believable constant instead of an
* obvious error. The same trap exists at the bottom -- if nearly every
* sample lands in bucket 0, `histogram_quantile` interpolates inside it and
* invents a value. A ladder is correct only when its floor sits below the
* mass of the distribution and its ceiling above the tail.
*
* sample --> [ SDK lower_bound over edges ] --> per-bucket counter
* | |
* edges come from v
* THIS header OTLP export
* |
* v
* histogram_quantile() in Grafana
*
* Ladders are `std::array<double, N>` so they are constant-initialised and
* usable in a `static_assert`. The OTel SDK wants `std::vector<double>` in
* its aggregation config, so call toVector() at the registration site
* rather than storing vectors here.
*
* Example -- register a view with the millisecond ladder:
* @code
* auto config = std::make_shared<HistogramAggregationConfig>();
* config->boundaries_ = buckets::toVector(buckets::kMillisecondBuckets);
* @endcode
*
* Example -- the edge case that motivated a second ladder. An Event whose
* samples are sizes rather than durations must not borrow a latency ladder,
* or a quarter of its samples land in `+Inf` and every quantile reads back
* as the top edge:
* @code
* config->boundaries_ = buckets::toVector(buckets::kByteBuckets);
* @endcode
*
* @note Thread safety: every member is `constexpr` and immutable, so
* reading them from any thread is safe. toVector() allocates and is
* meant for start-up registration paths, never for a record path.
* @note Limitation: changing a ladder changes the exported series count and
* ends bucket comparability across the change -- existing series keep
* their old `le` values, so panels show a break at restart. Grafana
* Cloud bills per series, so re-measure the series count after any
* edit here.
*/
/**
* Bucket edges, in milliseconds, for whole-millisecond `beast::insight`
* Events: job queue wait and run times, io latency, RPC time, pathfinding.
*
* **This list must contain every representable edge of the collector's
* spanmetrics ladder, and may extend above it.** Agreement over the shared
* range is deliberate: it lets a span-derived latency panel and a native
* histogram panel be read on the same scale. It was specified that way
* originally, then silently broken when the collector ladder alone was
* extended, which left this side capped at 5 s while spans reached 30 s and
* censored every quantile above 5 s. `check_bucket_parity.py` now enforces
* the containment -- add a collector edge, add it here too.
*
* The sub-millisecond edges the collector carries (0.01 to 0.5 ms) are
* deliberately absent. `beast::insight::Event` rounds every duration up to
* a whole millisecond before it reaches the histogram, so those edges would
* collect nothing. Metrics that genuinely need finer resolution belong on
* the microsecond ladder, on the OTel-native path.
*
* The 60 s and 120 s edges exceed the collector's 30 s top on purpose,
* because jobs outlive spans: the updatepaths job type was measured
* averaging about 60 s, so a 30 s ceiling would censor its quantiles just
* as 5 s censors them today. All these Events share one ladder, so its
* ceiling has to cover the slowest member rather than the typical one.
*
* The 2, 3 and 4 s edges resolve second-scale work that previously had to
* interpolate across a single four-second-wide bucket.
*/
inline constexpr std::array kMillisecondBuckets{
1.0,
5.0,
10.0,
25.0,
50.0,
100.0,
250.0,
500.0,
1'000.0,
2'000.0,
3'000.0,
4'000.0,
5'000.0,
10'000.0,
30'000.0,
60'000.0,
120'000.0};
/**
* Bucket edges, in bytes, for `beast::insight` Events whose samples are
* sizes rather than durations. Currently only the RPC response size.
*
* Placed from the measured distribution rather than from a guess about how
* large a response could theoretically be. Measured over 24 h: mean 2131 B,
* half of all responses under 1 kB, three quarters under 5 kB. The tail
* above 5 kB has a mean of at most 7538 B, which bounds p99 near 80 kB and
* p99.75 below 256 kB.
*
* So the resolution belongs between 512 B and 64 kB, where the
* distribution actually turns, and two further edges are ample headroom.
* Spending edges at the megabyte scale would cost cardinality on a range
* nothing measured occupies. If a genuinely multi-megabyte response ever
* shows up in the top bucket, extend this -- but extend it on evidence.
*/
inline constexpr std::array kByteBuckets{
512.0,
1'024.0,
2'048.0,
4'096.0,
8'192.0,
16'384.0,
32'768.0,
65'536.0,
262'144.0,
1'048'576.0};
/**
* Bucket edges, in microseconds, for the OTel-native duration instruments
* created directly on MetricsRegistry: job queue wait and run times, RPC
* method latency, and GetObject lookup latency.
*
* The edges from 1 to 1000 us are the ones that matter most. An earlier
* version of this ladder started at 100 us, which sat ABOVE the mass of every
* instrument using it: 99.3% of job_queued_us samples, 92.5% of
* job_running_us and 90.4% of getobject_lookup_us fell in that first bucket.
* `histogram_quantile` then interpolated inside bucket 0 and returned the
* boundary scaled by the requested quantile -- p75/p95/p99 of job_queued_us
* read 75.5/95.7/99.7 us, which is arithmetic on the bucket edge, not a
* latency. A warm nodestore read is around 1.5 us, so single-microsecond
* resolution is not excessive here.
*
* The upper edges reach a minute so multi-second stalls stay measurable. The
* SDK's own default ladder stops at 10,000, which every one of these
* instruments exceeds during catch-up.
*/
inline constexpr std::array kMicrosecondBuckets{
1.0,
2.0,
5.0,
10.0,
25.0,
50.0,
100.0,
250.0,
500.0,
1'000.0,
5'000.0,
25'000.0,
100'000.0,
500'000.0,
1'000'000.0,
5'000'000.0,
10'000'000.0,
30'000'000.0,
60'000'000.0};
/**
* Bucket edges for the GetObject request object count.
*
* Counts run from 1 to the hard reply cap (kHardMaxReplyNodes, 12288). The
* honest sync path asks for at most 8 objects, so the low edges are
* fine-grained; the upper ones follow the charge size bands up to the cap.
* Because the top edge IS the hard cap, this ladder cannot saturate.
*/
inline constexpr std::array
kObjectCountBuckets{1.0, 2.0, 4.0, 8.0, 16.0, 64.0, 256.0, 1'024.0, 4'096.0, 12'288.0};
/**
* Bucket edges for the GetObject resource charge.
*
* Charges span 0 (the free tier) to roughly 99k for a full-size all-miss
* request. The edges bracket the two thresholds that decide a peer's fate --
* the warning threshold at 5000 and the drop threshold at 25000 -- so a
* dashboard can show how close charges run to each.
*/
inline constexpr std::array
kChargeBuckets{0.0, 100.0, 500.0, 1'000.0, 5'000.0, 10'000.0, 25'000.0, 50'000.0, 100'000.0};
/**
* @brief Check that a ladder is strictly ascending and non-negative.
*
* The SDK places a sample with `std::lower_bound` over the edges, which
* silently misbuckets when edges repeat or descend. Checking at compile
* time makes that class of typo impossible to ship.
*
* @param ladder Bucket upper bounds to check.
* @return true when the ladder is non-empty, starts at or above zero, and
* every later edge is strictly greater than its predecessor.
*/
constexpr bool
isAscendingNonNegative(std::span<double const> ladder) noexcept
{
if (ladder.empty() || ladder.front() < 0.0)
return false;
for (std::size_t i = 1; i < ladder.size(); ++i)
{
if (!(ladder[i] > ladder[i - 1]))
return false;
}
return true;
}
static_assert(isAscendingNonNegative(kMillisecondBuckets));
static_assert(isAscendingNonNegative(kByteBuckets));
static_assert(isAscendingNonNegative(kMicrosecondBuckets));
static_assert(isAscendingNonNegative(kObjectCountBuckets));
static_assert(isAscendingNonNegative(kChargeBuckets));
/**
* @brief Copy a ladder into the `std::vector<double>` the OTel SDK wants.
*
* @param ladder Bucket upper bounds.
* @return A vector holding the same edges in the same order.
*/
inline std::vector<double>
toVector(std::span<double const> ladder)
{
return std::vector<double>(ladder.begin(), ladder.end());
}
} // namespace xrpl::telemetry::buckets