Files
rippled/src/tests/libxrpl/telemetry/HistogramBuckets.cpp
Pratik Mankawde 6e2b2da772 fix(telemetry): resolve microsecond latencies below 100us
The microsecond ladder's first edge was 100us, which sat ABOVE the mass of
every instrument using it. Measured on devnet: 99.3% of job_queued_us
samples, 92.5% of job_running_us and 90.4% of getobject_lookup_us fell in
that first bucket. histogram_quantile then interpolated inside bucket 0 and
returned `quantile / fraction_in_bucket_0 x first_edge` -- p75/p95/p99 of
job_queued_us read 75.52/95.66/99.69us against a prediction of
75.53/95.67/99.70. Three-decimal agreement: those panels were reporting
arithmetic on the bucket edge, not latency.

The fix was already half-written. kSubMillisecondBoundaries had been parked
in MetricsRegistry.cpp as [[maybe_unused]] with a comment noting exactly this
problem for nodestore reads. Its edges are now folded into kMicrosecondBuckets
rather than deleted, so the parked intent is carried forward: 1..1000us
resolution where the mass is, upper edges unchanged so multi-second stalls
stay measurable.

Also moves the GetObject count and charge ladders into HistogramBuckets.h, so
all five ladders have one owner and one set of invariant tests (29 now).

Adds check_bucket_parity.py, wired into the existing OTel naming workflow.
The C++ millisecond ladder and the collector's spanmetrics ladder are
specified to agree over their shared range; they were identical when shipped,
then the collector side alone was extended and nothing noticed for eleven
phases. The check asserts containment rather than equality, because jobs
outlive spans -- jobq_updatepaths averages ~60s, which no span approaches, so
demanding equality would force a ceiling that censors it. Verified it rejects
a missing collector edge, a bogus in-range edge, and a return to the 5s
ceiling.

ledger-data-sync's "Job Queue Wait p95 By Type" moves off the beast
jobq_*_q_milliseconds pair onto job_queued_us filtered by job_type. Those
beast metrics are ms-quantised at the source (Event rounds up to a whole
millisecond), so 94-100% of their samples sat in the first bucket and no
ladder change could fix them. Note the label values are camelCase
(job_type="ledgerData"), not the lowercase metric-name fragments.

Both histogram-fed alert thresholds re-validated and left unchanged, with the
measured basis recorded so neither gets tuned against the old artefact: only
0.0022% of job_queued_us samples exceed the 1s threshold, and every edge
bracketing the 1000ms ios_latency threshold survived the ladder change.

Docs: the rpc_size "known issue -- tracked separately" notes in the runbook
and 09-data-collection-reference are now resolved notes, the stale 10-edge
span_duration bucket list is corrected to the collector's real 20, and the
runbook gains a "Reading A Histogram Percentile" section covering both
saturation traps and the expected discontinuity after a ladder change.
2026-08-21 12:46:56 +01:00

236 lines
8.6 KiB
C++

/**
* GTest unit tests for the histogram bucket ladders.
*
* These ladders decide whether a Grafana percentile panel reports a
* measurement or an artefact, and neither failure mode is visible in the
* panel itself: a quantile that falls in the `+Inf` bucket reads back as the
* second-highest edge, and one that falls inside bucket 0 is interpolated.
* Both look like plausible numbers. So the invariants are asserted here
* rather than left to review.
*
* The ladders are `constexpr`, so most of this could be `static_assert`.
* They are runtime tests as well so that a failure names which edge is
* wrong instead of only failing the compile.
*/
#include <xrpl/telemetry/HistogramBuckets.h>
#include <gtest/gtest.h>
#include <algorithm>
#include <array>
#include <cmath>
#include <cstddef>
#include <span>
#include <vector>
namespace xrpl::telemetry::buckets {
// Every ladder must be strictly ascending and non-negative. The SDK places a
// sample with std::lower_bound over the edges, so a duplicated or
// out-of-order edge silently sends samples to the wrong bucket.
class HistogramBucketsTest : public ::testing::TestWithParam<std::span<double const>>
{
};
TEST_P(HistogramBucketsTest, isStrictlyAscending)
{
auto const ladder = GetParam();
ASSERT_FALSE(ladder.empty());
for (std::size_t i = 1; i < ladder.size(); ++i)
EXPECT_LT(ladder[i - 1], ladder[i]) << "edge index " << i << " does not ascend";
}
TEST_P(HistogramBucketsTest, isNonNegativeAndFinite)
{
for (double const edge : GetParam())
{
EXPECT_GE(edge, 0.0);
EXPECT_TRUE(std::isfinite(edge)) << "edge " << edge << " is not finite";
}
}
TEST_P(HistogramBucketsTest, passesTheCompileTimeValidator)
{
EXPECT_TRUE(isAscendingNonNegative(GetParam()));
}
INSTANTIATE_TEST_SUITE_P(
AllLadders,
HistogramBucketsTest,
::testing::Values(
std::span<double const>{kMillisecondBuckets},
std::span<double const>{kByteBuckets},
std::span<double const>{kMicrosecondBuckets},
std::span<double const>{kObjectCountBuckets},
std::span<double const>{kChargeBuckets}));
TEST(HistogramBucketsRange, microsecondFloorLandsBelowTheMeasuredMass)
{
// Measured: 99.3% of job_queued_us samples sat below the old 100 us floor,
// so p75/p95/p99 all interpolated inside bucket 0 and returned
// 75.5/95.7/99.7 us -- the boundary scaled by the requested quantile,
// not a latency. Warm nodestore reads are ~1.5 us, so the floor has to
// reach single microseconds and several edges must precede 100 us.
EXPECT_LE(kMicrosecondBuckets.front(), 1.0);
auto const belowHundred =
std::ranges::count_if(kMicrosecondBuckets, [](double edge) { return edge < 100.0; });
EXPECT_GE(belowHundred, 5) << "too little resolution below 100 us";
}
TEST(HistogramBucketsRange, microsecondCeilingStillReachesOneMinute)
{
// Job waits and RPC latencies routinely exceed the SDK default ceiling of
// 10,000; multi-second stalls must stay measurable rather than censored.
EXPECT_EQ(kMicrosecondBuckets.back(), 60'000'000.0);
}
TEST(HistogramBucketsRange, objectCountLadderCannotSaturate)
{
// GetObject counts run 1..kHardMaxReplyNodes, so the top edge IS the hard
// cap and censoring is impossible by construction.
EXPECT_EQ(kObjectCountBuckets.front(), 1.0);
EXPECT_EQ(kObjectCountBuckets.back(), 12'288.0);
}
TEST(HistogramBucketsRange, chargeLadderBracketsTheResourceThresholds)
{
// The two edges that decide a peer's fate must be present so a dashboard
// can show how close charges run to each: warning at 5000, drop at 25000.
// A leading 0 separates the free tier from everything else.
EXPECT_EQ(kChargeBuckets.front(), 0.0);
for (double const threshold : {5'000.0, 25'000.0})
{
EXPECT_NE(std::ranges::find(kChargeBuckets, threshold), kChargeBuckets.end())
<< threshold << " is a resource threshold and must be an edge";
}
}
// The validator must also REJECT. A predicate that only ever returns true
// would let every ladder above pass while proving nothing.
TEST(HistogramBucketsValidator, rejectsEmptyDescendingDuplicateAndNegative)
{
EXPECT_FALSE(isAscendingNonNegative(std::span<double const>{}));
constexpr std::array descending{5.0, 1.0};
EXPECT_FALSE(isAscendingNonNegative(descending));
constexpr std::array duplicated{1.0, 1.0, 2.0};
EXPECT_FALSE(isAscendingNonNegative(duplicated));
constexpr std::array negative{-1.0, 1.0};
EXPECT_FALSE(isAscendingNonNegative(negative));
}
TEST(HistogramBucketsValidator, acceptsASingleEdgeAndALeadingZero)
{
constexpr std::array single{1.0};
EXPECT_TRUE(isAscendingNonNegative(single));
// A leading zero is legal: the GetObject charge ladder starts at 0 to
// separate the free tier from everything else.
constexpr std::array leadingZero{0.0, 100.0};
EXPECT_TRUE(isAscendingNonNegative(leadingZero));
}
TEST(HistogramBucketsRange, millisecondFloorIsOneAndCeilingCoversTheSlowestJob)
{
// beast::insight::Event rounds durations up to whole milliseconds, so 1
// is the smallest edge that can ever collect a sample.
EXPECT_EQ(kMillisecondBuckets.front(), 1.0);
// The updatepaths job type was measured averaging 59,956 ms. A 30 s
// ceiling -- the collector's top edge -- would censor it just as the old
// 5 s ceiling does, so this ladder has to reach further.
EXPECT_GE(kMillisecondBuckets.back(), 120'000.0);
}
TEST(HistogramBucketsRange, millisecondLadderClearsTheMeasuredCensoringPoint)
{
// rpc_size had 24.9% of samples above the old 5000 ceiling and
// jobq_updatepaths had 100%. A ceiling at or below 5000 reintroduces the
// exact defect this ladder exists to fix.
EXPECT_GT(kMillisecondBuckets.back(), 5'000.0);
}
TEST(HistogramBucketsRange, millisecondLadderContainsEveryRepresentableCollectorEdge)
{
// Agreement with the collector's spanmetrics ladder over the shared
// range is the invariant; edges above its 30 s top are allowed because
// jobs outlive spans. Sub-millisecond collector edges are excluded
// because Event cannot represent them. check_bucket_parity.py enforces
// this against the YAML; this test pins it for the C++ side alone so a
// local edit fails fast.
constexpr std::array collectorEdges{
1.0,
5.0,
10.0,
25.0,
50.0,
100.0,
250.0,
500.0,
1'000.0,
2'000.0,
3'000.0,
4'000.0,
5'000.0,
10'000.0,
30'000.0};
for (double const edge : collectorEdges)
{
EXPECT_NE(std::ranges::find(kMillisecondBuckets, edge), kMillisecondBuckets.end())
<< edge << " ms is a collector spanmetrics edge and must be present";
}
}
TEST(HistogramBucketsRange, millisecondLadderResolvesTheOneToFiveSecondBand)
{
// Without these the 1 s to 5 s span was one four-second-wide bucket, so
// any quantile landing inside it was interpolated across four seconds.
for (double const edge : {2'000.0, 3'000.0, 4'000.0})
{
EXPECT_NE(std::ranges::find(kMillisecondBuckets, edge), kMillisecondBuckets.end())
<< edge << " ms edge missing";
}
}
TEST(HistogramBucketsRange, byteLadderBracketsTheMeasuredResponseDistribution)
{
// Measured: mean 2131 B, half under 1 kB, three quarters under 5 kB, and
// the tail above 5 kB has a mean of at most 7538 B -- which puts p99
// near 80 kB. The floor must sit at or below the measured median region
// and the ceiling well past the p99 bound.
EXPECT_LE(kByteBuckets.front(), 512.0);
EXPECT_GE(kByteBuckets.back(), 1'048'576.0);
// Most of the resolution belongs where the distribution actually turns.
auto const withinWorkingRange =
std::ranges::count_if(kByteBuckets, [](double e) { return e >= 512.0 && e <= 65'536.0; });
EXPECT_GE(withinWorkingRange, 6) << "too little resolution between 512 B and 64 kB";
}
TEST(HistogramBucketsRange, byteAndMillisecondLaddersAreDistinct)
{
// A single shared ladder is what put a byte count on a latency scale and
// censored a quarter of its samples.
EXPECT_NE(kByteBuckets.size(), kMillisecondBuckets.size());
EXPECT_GT(kByteBuckets.back(), kMillisecondBuckets.back());
}
TEST(HistogramBucketsConvert, toVectorPreservesOrderAndSize)
{
auto const converted = toVector(kByteBuckets);
ASSERT_EQ(converted.size(), kByteBuckets.size());
EXPECT_TRUE(std::ranges::equal(converted, kByteBuckets));
}
TEST(HistogramBucketsConvert, toVectorHandlesAnEmptyLadder)
{
EXPECT_TRUE(toVector(std::span<double const>{}).empty());
}
} // namespace xrpl::telemetry::buckets