diff --git a/OpenTelemetryPlan/09-data-collection-reference.md b/OpenTelemetryPlan/09-data-collection-reference.md index 1f411e8b28..8be0b3a7e5 100644 --- a/OpenTelemetryPlan/09-data-collection-reference.md +++ b/OpenTelemetryPlan/09-data-collection-reference.md @@ -66,6 +66,12 @@ There are two independent telemetry pipelines entering a single **OTel Collector 1. **OpenTelemetry Traces** — Distributed spans with attributes, exported via OTLP/HTTP (:4318) to the collector's **OTLP Receiver**. The **Batch Processor** groups spans (1s timeout, batch size 100) before forwarding to trace backends. The **SpanMetrics Connector** derives RED metrics (rate, errors, duration) from every span and feeds them into the metrics pipeline. 2. **beast::insight OTel Metrics** — System-level gauges, counters, and histograms exported natively via OTLP/HTTP (:4318) to the same **OTLP Receiver**. These are batched and exported to Prometheus alongside span-derived metrics. The StatsD UDP transport has been replaced by native OTLP; `server=statsd` remains available as a fallback. +A third, narrower metrics path exists for instruments created at their call site through the +`XRPL_METRIC_*` macros. These use the OTel Metrics SDK directly and reach the collector's OTLP +receiver rather than the StatsD receiver, so their names carry no `xrpld_` prefix. See +[§2a](#2a-call-site-otel-metrics-metricsregistry). Code in `libxrpl` cannot use these macros and +always goes through `beast::insight` instead. + **Trace backend** — The collector exports traces via OTLP/gRPC to: - **Grafana Tempo** — Preferred trace backend. Supports TraceQL queries at `:3200`, S3/GCS object storage for cost-effective long-term trace retention, and integrates natively with Grafana. @@ -621,6 +627,147 @@ For each of the 45+ overlay traffic categories (defined in `TrafficCount.h`), fo **Grafana dashboards**: _Network Traffic_ (`network-traffic`), _Overlay Traffic Detail_ (`overlay-traffic-detail`), _Ledger Data & Sync_ (`ledger-data-sync`) +### 2.5 Per-Job-Type Saturation Gauges + +`JobTypeData` has always tracked three per-type counters — `waiting`, `running`, and `deferred` +(`include/xrpl/core/JobTypeData.h`) — but only the process-wide queue depth was exported. These +three gauge families export them per job type, so queue pressure can be attributed to a type +instead of only being visible as a single total. + +For each of the 35 non-special job types, three gauges are created in the `JobTypeData` +constructor and assigned in `JobQueue::collect()`. + +| Metric | Source File | Description | +| ------------------------- | ------------- | ----------------------------------------------- | +| `jobq_{jobtype}_waiting` | JobTypeData.h | Jobs enqueued for this type but not yet started | +| `jobq_{jobtype}_running` | JobTypeData.h | Jobs of this type currently executing | +| `jobq_{jobtype}_deferred` | JobTypeData.h | Jobs held back by this type's concurrency limit | + +**Why `deferred` matters most.** `JobQueue::addJob()` never rejects — when a type is at its +concurrency limit it increments `deferred` and returns `true` (`JobQueue.cpp`, +`addRefCountedJob`). Backpressure on a capped type therefore shows up only as latency, after the +delay has already happened. A non-zero `deferred` reading precedes that, which makes it the +leading indicator for a saturating type. It is structurally always zero for special types, which +is why they are excluded — the same `!info.special()` guard that already gates the event timers. + +**Name derivation.** These names are not literals anywhere in the source; the `jobq` segment is +added at runtime. `Application.cpp` constructs the JobQueue with +`collectorManager_->group("jobq")`, and `GroupImp::makeName()` (`Groups.cpp`) joins the group name +and the metric name with a `.`. `OTelCollector` then lowercases the result and maps `.` to `_`, so +`JtLedgerReq` (whose type name is `ledgerRequest`) exports as `jobq_ledgerrequest_deferred`. The +same derivation applies to the event timers in §2.3 and to the queue-depth gauge in §2.1, which is +why the latter appears as `jobq_job_count` rather than `job_count`. + +> **Sampling granularity**: these are gauges read by the collector hook once per second, not +> counters incremented at the event. A `deferred` spike shorter than the sampling interval can be +> missed. Saturation long enough to delay ledger sync persists across a sample, which is the case +> these gauges are for. + +> **Pipeline**: `JobQueue.cpp` is in `libxrpl` and reaches the exporter through +> `beast::insight::Collector`, like every other metric in this section. It cannot use the +> `xrpld`-only call-site metric macros used by +> [§2a](#2a-call-site-otel-metrics-metricsregistry). + +--- + +## 2a. Call-Site OTel Metrics (MetricsRegistry) + +> **See also**: [§2](#2-statsd-metrics-beastinsight) for the `beast::insight` metrics that use the +> StatsD transport instead. + +Every metric in [§2](#2-statsd-metrics-beastinsight) goes through `beast::insight`. The metrics +below are created directly at their call site through the `XRPL_METRIC_*` macros +(`src/xrpld/telemetry/MetricMacros.h`) and exported by the OTel Metrics SDK rather than over +StatsD. The macros create each instrument once and expand to an empty statement when telemetry is +compiled out, so a call site needs no conditional compilation. + +These names carry no `xrpld_` prefix: the prefix belongs to the StatsD sender, and this path +identifies the node through OTel resource attributes instead. + +> **Note**: this family is emitted by `MetricsRegistry`, which arrives with the native-metrics +> work in a later phase than this document's own StatsD pipeline. It is catalogued here because +> this document is the single source of truth for the whole telemetry surface. + +### 2a.1 Job Lifecycle Metrics + +Recorded from the three `PerfLog` job hooks that `JobQueue` already calls. + +| Metric | Type | Unit | Labels | Description | +| -------------------- | --------- | ---- | --------------------- | ---------------------------- | +| `job_queued_total` | Counter | — | `job_type`, `handler` | Jobs enqueued | +| `job_started_total` | Counter | — | `job_type`, `handler` | Jobs started | +| `job_finished_total` | Counter | — | `job_type`, `handler` | Jobs completed | +| `job_queued_us` | Histogram | µs | `job_type`, `handler` | Queue wait time distribution | +| `job_running_us` | Histogram | µs | `job_type`, `handler` | Execution time distribution | + +`job_queued_us` and `job_running_us` register explicit microsecond buckets spanning 100 µs to +60 s. Without that view the SDK default buckets stop at 10 ms and every quantile reads as 10 ms. + +#### The `handler` Label + +`job_type` names the queue a job ran on, not who produced it. Several producers share one type — +`RcvGetLedger` and `RcvGetObjByHash` both run as `JtLedgerReq`, and `JtLedgerData` has five +(`ProcessLData` and `GotStaleData` in `InboundLedgers.cpp`, `AcqDone` and the `InboundLedger` +`TimeoutCounter` job name in `InboundLedger.cpp`, and `GotFetchPack` in `LedgerMaster.cpp`) — so a +queue-wait spike cannot be pinned to one of them by `job_type` alone. The `handler` label carries +the name passed to `addJob`, which does distinguish them. + +The value is not the raw name. It passes through `MetricsRegistry::sanitiseHandler`, which keeps +the name only when it is non-empty and every character is an ASCII letter, and otherwise returns +the literal `other`. The letter test is an explicit ASCII range check rather than a +locale-dependent one, so it cannot admit non-ASCII bytes. + +**Why reduce the name.** Two `addJob` names embed a ledger sequence number: +`LedgerPersistence.cpp` builds `"Pub" + std::to_string(ledger->seq())`, and +`OrderBookDBImpl.cpp` builds `"OB" + std::to_string(ledger->seq() % 1000000000)`. Used raw, +either would mint a new time series for every ledger. Both always contain digits, so both fold to +`other`. + +Every other production job name is a compile-time literal, which makes the label domain a +function of the source rather than of traffic — it cannot grow at runtime. Names that are literals +but still fail the rule also fold to `other`: `GetConsL1` and `GetConsL2` contain digits, and the +coroutine names `RPC-Client`, `WS-Client`, and `gRPC-Client` contain a hyphen. A job name added +later that fails the rule degrades to `other` rather than becoming unbounded, which is a stronger +guarantee than a hand-maintained allowlist. + +### 2a.2 GetObject Request Metrics + +Emitted from the `TMGetObjectByHash` handler in `src/xrpld/overlay/detail/PeerImp.cpp`. Together +with `job_queued_us` and `job_running_us` filtered to `handler="RcvGetObjByHash"`, they split the +handler's end-to-end time into queue wait, NodeStore lookup, and everything else — so slowness is +attributable to a stage rather than merely observed. + +| Metric | Type | Unit | Labels | Description | +| --------------------------- | --------- | ---- | -------- | ---------------------------------------------- | +| `getobject_lookup_us` | Histogram | µs | — | Time in the NodeStore fetch loop | +| `getobject_request_objects` | Histogram | — | — | Objects requested per message | +| `getobject_lookups_total` | Counter | — | `result` | NodeStore lookup outcomes | +| `getobject_rejected_total` | Counter | — | `reason` | Requests refused before any NodeStore access | +| `getobject_charge` | Histogram | — | — | Dynamic resource charge applied to the request | + +| Label | Values | Meaning | +| -------- | ---------------------------------- | ------------------------------------------------------- | +| `result` | `hit`, `miss` | Whether the requested object was found in the NodeStore | +| `reason` | `oversize`, `malformed_ledgerhash` | Which admission check refused the request | + +**Aggregation.** `getobject_lookup_us` times the fetch loop once per request, and +`getobject_lookups_total` is added once per request carrying the batch hit and miss totals. The +loop is bounded by `Tuning::kHardMaxReplyNodes` (12288), so timing or counting each iteration +would cost more than the lookups it measures. `getobject_request_objects` records the requested +count, which is what the charge bands price on. `getobject_charge` records only the dynamic +component returned by `computeGetObjectByHashFee()`; the admission-time base charge is a constant. + +> **Bucket view required**: `getobject_lookup_us` needs an explicit microsecond-bucket view like +> the two job histograms above. Without one the SDK default buckets stop at 10 ms, and a +> lookup loop running toward its 12288 cap exceeds that routinely — so the metric would saturate +> exactly when it is needed. The view is registered alongside the other three in +> `MetricsRegistry.cpp`. Note that `MetricsRegistry` does not exist on this branch — it is +> introduced on phase-9, which is where the registration lives. + +> **Expected to read zero**: `getobject_rejected_total` counts non-conforming requests, which a +> healthy peer does not send. A zero series does not prove the metric works — verify it with a +> deliberately oversized request. + --- ## 3. Grafana Dashboard Reference