mirror of
https://github.com/XRPLF/rippled.git
synced 2026-08-22 06:40:53 +00:00
Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
# Conflicts: # src/xrpld/telemetry/MetricsRegistry.cpp
This commit is contained in:
@@ -531,12 +531,12 @@ a destructor must not depend on still existing. A query that only groups by
|
||||
|
||||
The OTel Collector's SpanMetrics connector automatically generates RED (Rate, Errors, Duration) metrics from every span. No custom metrics code in xrpld is needed.
|
||||
|
||||
| Prometheus Metric | Type | Description |
|
||||
| ----------------------------------- | --------- | ------------------------------------------------------------------------------ |
|
||||
| `span_calls_total` | Counter | Total span invocations |
|
||||
| `span_duration_milliseconds_bucket` | Histogram | Latency distribution (buckets: 1, 5, 10, 25, 50, 100, 250, 500, 1000, 5000 ms) |
|
||||
| `span_duration_milliseconds_count` | Histogram | Observation count |
|
||||
| `span_duration_milliseconds_sum` | Histogram | Cumulative latency |
|
||||
| Prometheus Metric | Type | Description |
|
||||
| ----------------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `span_calls_total` | Counter | Total span invocations |
|
||||
| `span_duration_milliseconds_bucket` | Histogram | Latency distribution. Buckets come from the collector's spanmetrics config: 0.01, 0.05, 0.1, 0.25, 0.5, 1, 5, 10, 25, 50, 100, 250, 500 ms then 1, 2, 3, 4, 5, 10, 30 s. The sub-millisecond edges exist because most xrpld spans are far below 1 ms; without them every p95/p99 pinned to a constant 0.95 ms |
|
||||
| `span_duration_milliseconds_count` | Histogram | Observation count |
|
||||
| `span_duration_milliseconds_sum` | Histogram | Cumulative latency |
|
||||
|
||||
**Standard labels on every metric**: `span_name`, `status_code`, `service_name`, `span_kind`
|
||||
|
||||
@@ -684,13 +684,14 @@ prefix=xrpld
|
||||
|
||||
Quantiles collected: 0th, 50th, 90th, 95th, 99th, 100th percentile.
|
||||
|
||||
\* **`rpc_size` instrument mismatch (known issue):** response size in bytes is
|
||||
recorded through the millisecond-scaled event histogram (`makeEvent`), so it is
|
||||
exported as `rpc_size_milliseconds_bucket` with time-scaled boundaries that top
|
||||
out at 5000. Byte values above ~5 KB saturate in the last bucket, so the
|
||||
percentiles are not true byte sizes. The _RPC & Pathfinding_ panel is flagged
|
||||
accordingly. A dedicated byte-unit histogram is needed to fix this; tracked
|
||||
separately.
|
||||
\* **`rpc_size` now records bytes as bytes (fixed).** It used to go through the
|
||||
millisecond-scaled event histogram and export as `rpc_size_milliseconds_bucket`
|
||||
on a ladder topping out at 5000, so the 24.9% of responses larger than 5 kB all
|
||||
landed in the last bucket and every percentile read back as a flat 5000 — a
|
||||
plausible-looking constant rather than a byte size. `beast::insight::Event` now
|
||||
declares a `Unit`, so this instrument is created with unit `By` and exports as
|
||||
**`rpc_size_bytes_bucket`** on `kByteBuckets` (512 B to 1 MiB, placed from the
|
||||
measured distribution). Queries and panels must use the new name.
|
||||
|
||||
**Grafana dashboards**: _Node Health_ (`ios_latency`), _RPC & Pathfinding_ (`rpc_time`, `rpc_size`, `pathfind_*`)
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@
|
||||
- **OTelCounterImpl**: Wraps `opentelemetry::metrics::Counter<int64_t>`. `increment(amount)` calls `counter->Add(amount)`.
|
||||
- **OTelGaugeImpl**: Uses `opentelemetry::metrics::ObservableGauge<uint64_t>` with an async callback. `set(value)` stores value atomically; callback reads it during collection.
|
||||
- **OTelMeterImpl**: Wraps `opentelemetry::metrics::Counter<uint64_t>`. `increment(amount)` calls `counter->Add(amount)`. Semantically identical to Counter but unsigned.
|
||||
- **OTelEventImpl**: Wraps `opentelemetry::metrics::Histogram<double>`. `notify(duration)` calls `histogram->Record(duration.count())`. Uses explicit bucket boundaries matching SpanMetrics: [1, 5, 10, 25, 50, 100, 250, 500, 1000, 5000] ms.
|
||||
- **OTelEventImpl**: Wraps `opentelemetry::metrics::Histogram<double>`. `notify()` calls `histogram->Record(value.count())`. Declares its unit from `beast::insight::Unit`, which is what selects its bucket ladder: the histogram views in `Telemetry.cpp` match on unit, so a `ms` instrument gets the millisecond ladder and a `By` instrument the byte ladder. Bucket edges live in `include/xrpl/telemetry/HistogramBuckets.h` — do not restate them here. The millisecond ladder must contain every representable edge of the collector's spanmetrics ladder and may extend above it (jobs outlive spans); `.github/scripts/telemetry/check_bucket_parity.py` enforces that. An earlier version of this line specified `[1, 5, 10, 25, 50, 100, 250, 500, 1000, 5000] ms` as "matching SpanMetrics" — true when written, then silently false once the collector ladder was extended on its own, which capped every quantile above 5s at a flat 5000.
|
||||
- **OTelHookImpl**: Stores handler function. Called during periodic metric collection (same 1s pattern via PeriodicMetricReader).
|
||||
- **OTelCollectorImp**: Main class.
|
||||
- Creates `MeterProvider` with `PeriodicMetricReader` (1s export interval)
|
||||
|
||||
Reference in New Issue
Block a user