From c622da5a76e4cb5b3f9a286a8b6c49f25f3bc729 Mon Sep 17 00:00:00 2001 From: Pratik Mankawde <3397372+pratikmankawde@users.noreply.github.com> Date: Wed, 9 Sep 2026 15:40:11 +0100 Subject: [PATCH] refactor(telemetry): use the current name for the span metrics connector The pinned collector warns on every start that "spanmetrics" is a deprecated alias for "span_metrics". Rename the connector, its pipeline references and the prose that names it. The derived metric names are untouched. They come from the connector's `namespace` setting rather than its component name, so occurrences inside a metric name such as traces_spanmetrics_calls_total are deliberately left as they are; renaming those would break every span panel. The rename is applied only to the bare word, never where it is joined to a metric name by underscores. This branch introduces the connector, so the change belongs here. --- OpenTelemetryPlan/02-design-decisions.md | 2 +- OpenTelemetryPlan/07-observability-backends.md | 4 ++-- docker/telemetry/TESTING.md | 8 ++++---- docker/telemetry/docker-compose.yml | 2 +- docker/telemetry/integration-test.sh | 2 +- docker/telemetry/otel-collector-config.yaml | 12 ++++++------ docker/telemetry/prometheus.yml | 2 +- docs/telemetry-runbook.md | 12 ++++++------ 8 files changed, 22 insertions(+), 22 deletions(-) diff --git a/OpenTelemetryPlan/02-design-decisions.md b/OpenTelemetryPlan/02-design-decisions.md index c07eaf015d..d11a41d639 100644 --- a/OpenTelemetryPlan/02-design-decisions.md +++ b/OpenTelemetryPlan/02-design-decisions.md @@ -183,7 +183,7 @@ agree with the code. A CI check enforces this end to end. `validation_trusted` likewise). Defined once in the base `SpanNames.h` `namespace attr` block and re-exported (`using`) by each domain header. 3. **Collision qualifier** → `_`, only when a bare name would - collide with a DIFFERENT concept in the shared spanmetrics label space or with + collide with a DIFFERENT concept in the shared span_metrics label space or with the OTel-reserved `status` key (e.g. `rpc_status`, `grpc_status`, `consensus_phase`, `consensus_round`, `consensus_mode`). This disambiguates distinct concepts that share a word; it is NOT used to tag the same concept diff --git a/OpenTelemetryPlan/07-observability-backends.md b/OpenTelemetryPlan/07-observability-backends.md index ecca8320d3..9a271bc1e2 100644 --- a/OpenTelemetryPlan/07-observability-backends.md +++ b/OpenTelemetryPlan/07-observability-backends.md @@ -258,7 +258,7 @@ Grafana provisions three TraceQL-based alert rules (group `xrpld-tracing-alerts` {resource.service.name="xrpld" && name="consensus.round"} | avg(duration) > 5s ``` -- **RPC Error Rate Spike** (critical, `for: 2m`): fires when the error rate across `rpc.command.*` spans exceeds 5%. Error _rate_ is a ratio, so it must divide the error-span rate by the total-span rate — a single TraceQL `rate()` returns spans/second, not a percentage, and would fire on traffic volume alone. This uses span metrics emitted by the collector's `spanmetrics` connector (Prometheus datasource), not a TraceQL query: +- **RPC Error Rate Spike** (critical, `for: 2m`): fires when the error rate across `rpc.command.*` spans exceeds 5%. Error _rate_ is a ratio, so it must divide the error-span rate by the total-span rate — a single TraceQL `rate()` returns spans/second, not a percentage, and would fire on traffic volume alone. This uses span metrics emitted by the collector's `span_metrics` connector (Prometheus datasource), not a TraceQL query: ``` sum(rate(traces_spanmetrics_calls_total{service_name="xrpld", span_name=~"rpc.command.*", status_code="STATUS_CODE_ERROR"}[5m])) @@ -273,7 +273,7 @@ Grafana provisions three TraceQL-based alert rules (group `xrpld-tracing-alerts` {resource.service.name="xrpld" && name="tx.receive"} | rate() < 10 ``` -> **Note**: The Consensus Round Slow and Transaction Throughput Drop rules use TraceQL aggregates (`avg(duration)`, `rate()`), which require Tempo 2.3+ with TraceQL metrics enabled. Verify aggregate query support in your Tempo version before provisioning. The RPC Error Rate Spike rule instead queries Prometheus span metrics (collector `spanmetrics` connector), so it needs that connector enabled in the collector pipeline. +> **Note**: The Consensus Round Slow and Transaction Throughput Drop rules use TraceQL aggregates (`avg(duration)`, `rate()`), which require Tempo 2.3+ with TraceQL metrics enabled. Verify aggregate query support in your Tempo version before provisioning. The RPC Error Rate Spike rule instead queries Prometheus span metrics (collector `span_metrics` connector), so it needs that connector enabled in the collector pipeline. --- diff --git a/docker/telemetry/TESTING.md b/docker/telemetry/TESTING.md index 9e610b3cf2..f33668d508 100644 --- a/docker/telemetry/TESTING.md +++ b/docker/telemetry/TESTING.md @@ -163,7 +163,7 @@ Run the integration test script: bash docker/telemetry/integration-test.sh ``` -It checks prerequisites, clears the previous run, brings up the observability stack, generates six validator key pairs and their node configs, starts the nodes, waits for consensus and then for a validated ledger, exercises RPC and submits a transaction, verifies traces in Tempo and both the spanmetrics and the StatsD-derived metrics in Prometheus, then prints a summary and leaves the stack running. +It checks prerequisites, clears the previous run, brings up the observability stack, generates six validator key pairs and their node configs, starts the nodes, waits for consensus and then for a validated ledger, exercises RPC and submits a transaction, verifies traces in Tempo and both the span_metrics and the StatsD-derived metrics in Prometheus, then prints a summary and leaves the stack running. The script announces each step as it runs, so read its `Step N:` headers for the authoritative sequence — they are not restated here, because a numbered copy of them drifts as soon as a step is added. @@ -413,7 +413,7 @@ Base URL: `http://localhost:9090` ```bash PROM="http://localhost:9090" -# Span call counts (from spanmetrics connector) +# Span call counts (from span_metrics connector) curl -s "$PROM/api/v1/query?query=traces_span_metrics_calls_total" | jq '.data.result[] | {span: .metric.span_name, count: .value[1]}' @@ -492,13 +492,13 @@ Pre-configured datasources: ### Spanmetrics not appearing in Prometheus -1. Verify otel-collector config has `spanmetrics` connector +1. Verify otel-collector config has `span_metrics` connector 2. Check that the metrics pipeline is configured: ```yaml service: pipelines: metrics: - receivers: [spanmetrics] + receivers: [span_metrics] exporters: [prometheus] ``` 3. Verify Prometheus can reach collector: diff --git a/docker/telemetry/docker-compose.yml b/docker/telemetry/docker-compose.yml index 650f1f2e7b..13da6a56c9 100644 --- a/docker/telemetry/docker-compose.yml +++ b/docker/telemetry/docker-compose.yml @@ -30,7 +30,7 @@ services: # receiver has no auth, and xrpld runs on the host, so it reaches the # collector via 127.0.0.1. Do not expose on 0.0.0.0 in shared setups. - "127.0.0.1:8125:8125/udp" - - "8889:8889" # Prometheus metrics (spanmetrics + statsd) + - "8889:8889" # Prometheus metrics (span_metrics + statsd) - "13133:13133" # Health check volumes: # Mount collector pipeline config (receivers → processors → exporters) diff --git a/docker/telemetry/integration-test.sh b/docker/telemetry/integration-test.sh index af5a29339c..2c8b861fae 100755 --- a/docker/telemetry/integration-test.sh +++ b/docker/telemetry/integration-test.sh @@ -567,7 +567,7 @@ check_span "peer.proposal.receive" check_span "peer.validation.receive" # --------------------------------------------------------------------------- -# Step 10: Verify Prometheus spanmetrics +# Step 10: Verify Prometheus span_metrics # --------------------------------------------------------------------------- log "" log "--- Spanmetrics ---" diff --git a/docker/telemetry/otel-collector-config.yaml b/docker/telemetry/otel-collector-config.yaml index c551b6ac1a..b47c154140 100644 --- a/docker/telemetry/otel-collector-config.yaml +++ b/docker/telemetry/otel-collector-config.yaml @@ -1,11 +1,11 @@ # OpenTelemetry Collector configuration for xrpld development. # # Pipelines: -# traces: OTLP receiver -> batch processor -> debug + Tempo + spanmetrics -# metrics: StatsD receiver + spanmetrics connector -> Prometheus exporter +# traces: OTLP receiver -> batch processor -> debug + Tempo + span_metrics +# metrics: StatsD receiver + span_metrics connector -> Prometheus exporter # # xrpld sends traces via OTLP/HTTP to port 4318. The collector batches -# them, forwards to Tempo, and derives RED metrics via the spanmetrics +# them, forwards to Tempo, and derives RED metrics via the span_metrics # connector, which Prometheus scrapes on port 8889. # # xrpld's beast::insight framework sends StatsD UDP metrics to port 8125. @@ -80,7 +80,7 @@ processors: action: hash connectors: - spanmetrics: + span_metrics: # Resource attributes that define the resource-metrics grouping key. # All resource attributes are propagated onto the output metrics # regardless; promotion to Prometheus labels is handled by the @@ -144,8 +144,8 @@ service: traces: receivers: [otlp] processors: [resource/tier, resource/stripsdk, attributes/hash, batch] - exporters: [debug, otlp/tempo, spanmetrics] + exporters: [debug, otlp/tempo, span_metrics] metrics: - receivers: [spanmetrics, statsd] + receivers: [span_metrics, statsd] processors: [resource/tier, resource/stripsdk] exporters: [prometheus] diff --git a/docker/telemetry/prometheus.yml b/docker/telemetry/prometheus.yml index d99d919a55..7d87649c51 100644 --- a/docker/telemetry/prometheus.yml +++ b/docker/telemetry/prometheus.yml @@ -1,4 +1,4 @@ -# Prometheus configuration for scraping spanmetrics from OTel Collector. +# Prometheus configuration for scraping span_metrics from OTel Collector. global: scrape_interval: 15s evaluation_interval: 15s diff --git a/docs/telemetry-runbook.md b/docs/telemetry-runbook.md index 64077acc0c..c1234abc67 100644 --- a/docs/telemetry-runbook.md +++ b/docs/telemetry-runbook.md @@ -100,7 +100,7 @@ All spans instrumented in xrpld, grouped by subsystem: The three apply-pipeline spans (`tx.preflight`, `tx.preclaim`, `tx.transactor`) share a deterministic `trace_id` from `txID[0:16]`, so they group under one trace per transaction. The `stage` attribute (`preflight` / `preclaim` / -`apply`) drives the collector spanmetrics `stage` dimension, giving per-stage +`apply`) drives the collector span_metrics `stage` dimension, giving per-stage RED metrics on the _Transaction Overview_ dashboard. `current_ledger_seq` is the current (open/in-flight) ledger index a span acted on @@ -272,11 +272,11 @@ sum by (stage) (rate(traces_span_metrics_calls_total{span_name=~"tx.preflight|tx > stage rather than on a single aggregate so the failing stage is obvious. > **Sampling caveat**: these stage metrics are span-derived, so they count only -> the spans the collector's spanmetrics connector sees. Head sampling at the node +> the spans the collector's span_metrics connector sees. Head sampling at the node > is fixed at 1.0 and is not configurable (`Telemetry.h`), and the shipped > collector pipeline has no tail sampling, so today nothing is dropped and the > counts are absolute. Volume reduction is delegated to the collector: adding a -> tail-sampling processor to the traces pipeline puts it ahead of the spanmetrics +> tail-sampling processor to the traces pipeline puts it ahead of the span_metrics > connector, and these metrics would then undercount proportionally — treat them > as relative trends in that case. Native StatsD metrics are never sampled. @@ -493,7 +493,7 @@ all its normal attributes, it just lacks a cross-node parent link. ## Prometheus Metrics (Spanmetrics) -The OTel Collector's spanmetrics connector automatically derives RED (Rate, Errors, Duration) metrics from every span. No custom metrics code is needed in xrpld. +The OTel Collector's span_metrics connector automatically derives RED (Rate, Errors, Duration) metrics from every span. No custom metrics code is needed in xrpld. ### Generated Metric Names @@ -554,7 +554,7 @@ address=127.0.0.1:8125 prefix=xrpld ``` -The OTel Collector receives these via a `statsd` receiver on UDP port 8125 and exports them to Prometheus alongside spanmetrics. +The OTel Collector receives these via a `statsd` receiver on UDP port 8125 and exports them to Prometheus alongside span_metrics. ### Metric Reference @@ -671,7 +671,7 @@ collector settings make it work, both already enabled: - `prometheus.resource_to_telemetry_conversion: enabled: true` promotes resource attributes to metric labels on the local scrape surface. -- `spanmetrics.resource_metrics_key_attributes` lists the tier attributes so +- `span_metrics.resource_metrics_key_attributes` lists the tier attributes so span-derived series stay grouped per node and tier. Traces and logs carry resource attributes natively; Grafana Cloud ingests all