merge: bring the lock-free ValidationTracker forward from phase10-workload-validation

Only the workload README conflicted; the tracker, config and test files merged
clean, which closes the chain from phase-7.
This commit is contained in:
Pratik Mankawde
2026-09-09 17:10:15 +01:00
44 changed files with 2661 additions and 1266 deletions

View File

@@ -505,7 +505,7 @@ The following data is explicitly **excluded** from telemetry collection:
> 1. `PeerImp`'s constructor logs the peer's `remoteAddress_` — an `IP:port` —
> at `info` severity (`PeerImp.h:837-842`), and other overlay call sites log
> addresses too. These land in the ordinary `debug.log` stream.
> 2. The collector's `filelog` receiver tails exactly that file
> 2. The collector's `file_log` receiver tails exactly that file
> (`otel-collector-config.yaml:38-47`, `include: [/var/log/xrpld/*/debug.log]`)
> and the `logs` pipeline exports it to Loki (`:236-239`).
>
@@ -515,7 +515,7 @@ The following data is explicitly **excluded** from telemetry collection:
> fields — a `delete` action on an attribute key would not touch them.
>
> **The control points are therefore log-side, not trace-side:** Loki
> retention and access control on the log store; the `filelog` receiver's
> retention and access control on the log store; the `file_log` receiver's
> `include` list (dropping it disables log↔trace correlation entirely); or a
> collector-side transform on the log body. Do not describe the telemetry
> pipeline as IP-free without qualifying it to traces.
@@ -875,7 +875,7 @@ rather than calling `GetSpan()`, so the common no-span path costs no heap
allocation.
Because the IDs land in the ordinary `debug.log` stream, correlation is
end-to-end without touching PerfLog: the collector's `filelog` receiver parses
end-to-end without touching PerfLog: the collector's `file_log` receiver parses
`trace_id`/`span_id` as optional capture groups and ships the lines to Loki, and
Grafana links both directions (Tempo `tracesToLogs` → Loki, Loki derived fields
→ Tempo). Details in [05 §5.8.5](./05-configuration-reference.md).

View File

@@ -251,12 +251,12 @@ local stack and by CI. It carries **three** pipelines, not one:
| --------- | --------------------- | ---------------------------------------------------------------- | ------------------------------------ |
| `traces` | `otlp` | `resource/tier`, `resource/stripsdk`, `attributes/hash`, `batch` | `debug`, `otlp/tempo`, `spanmetrics` |
| `metrics` | `otlp`, `spanmetrics` | `resource/tier`, `resource/stripsdk`, `batch` | `prometheus` |
| `logs` | `filelog` | `resource/logs`, `resource/tier`, `resource/stripsdk`, `batch` | `otlphttp/loki` |
| `logs` | `file_log` | `resource/logs`, `resource/tier`, `resource/stripsdk`, `batch` | `otlp_http/loki` |
Component detail:
- **Receivers.** `otlp` on gRPC `0.0.0.0:4317` and HTTP `0.0.0.0:4318` (both
traces and native metrics arrive on 4318). `filelog` tails
traces and native metrics arrive on 4318). `file_log` tails
`/var/log/xrpld/*/debug.log` and runs a `regex_parser` that lifts
`timestamp`, `partition`, `severity` and the optional `trace_id`/`span_id`
emitted by the journal sink (§5.8.5).
@@ -282,7 +282,7 @@ Component detail:
dimensions are promoted to labels (`command`, `rpc_status`, `tx_type`,
`ter_result`, `stage`, `consensus_mode`, `outcome`, …).
- **Exporters.** `debug` (console, `verbosity: detailed`), `otlp/tempo`
(`tempo:4317`, `tls.insecure: true`), `otlphttp/loki`
(`tempo:4317`, `tls.insecure: true`), `otlp_http/loki`
(`http://loki:3100/otlp` — Loki 3.x native OTLP; the old `loki` exporter was
removed in collector-contrib v0.147.0), and `prometheus` on
`0.0.0.0:8889` with `resource_to_telemetry_conversion.enabled: true` so the
@@ -307,7 +307,7 @@ graph. The full delta:
| `basicauth/grafanacloud` | `:29` | Extension; instance id / API token from the container environment |
| `tail_sampling` | `:60` | One `probabilistic` policy at **0.5%**, `decision_wait: 10s` |
| `transform/cloudlabels` | `:119` | Copies three resource attrs onto datapoint labels for Cloud (OTLP) ingest |
| `otlphttp/grafanacloud` | `:236` | Single OTLP/HTTP exporter fanning all three signals to Grafana Cloud |
| `otlp_http/grafanacloud` | `:236` | Single OTLP/HTTP exporter fanning all three signals to Grafana Cloud |
| `metrics_flush_interval` | `:136` | `spanmetrics` flushes every 15s instead of the 60s default |
| Removed by the overlay | Consequence |
@@ -346,14 +346,14 @@ NetworkPolicy, peer trace-context validation) is covered in
The authoritative development stack lives in the repo at `docker/telemetry/docker-compose.yml`. It brings up **six** services on a shared `xrpld-telemetry` bridge network. All images are pinned to exact tags.
| Service | Image | Published ports | Role |
| ---------------- | ---------------------------------------------- | ---------------------- | ---------------------------------------------------------------- |
| `otel-collector` | `otel/opentelemetry-collector-contrib:0.158.0` | `4317`, `4318`, `8889` | OTLP ingest, spanmetrics, filelog tail, Prometheus scrape target |
| `tempo` | `grafana/tempo:2.9.4` | `3200` | Trace storage and TraceQL |
| `loki` | `grafana/loki:3.7.6` | `3100` | Log storage for log↔trace correlation |
| `prometheus` | `prom/prometheus:v3.13.2` | `9090` | Scrapes the collector's `:8889` |
| `grafana` | `grafana/grafana:13.1.2` | `3000` | Dashboards + provisioned datasources/alerts, anonymous admin |
| `renderer` | `grafana/grafana-image-renderer:v5.12.0` | `8081` | Panel→PNG rendering for image export and alert screenshots |
| Service | Image | Published ports | Role |
| ---------------- | ---------------------------------------------- | ---------------------- | ----------------------------------------------------------------- |
| `otel-collector` | `otel/opentelemetry-collector-contrib:0.158.0` | `4317`, `4318`, `8889` | OTLP ingest, spanmetrics, file_log tail, Prometheus scrape target |
| `tempo` | `grafana/tempo:2.9.4` | `3200` | Trace storage and TraceQL |
| `loki` | `grafana/loki:3.7.6` | `3100` | Log storage for log↔trace correlation |
| `prometheus` | `prom/prometheus:v3.13.2` | `9090` | Scrapes the collector's `:8889` |
| `grafana` | `grafana/grafana:13.1.2` | `3000` | Dashboards + provisioned datasources/alerts, anonymous admin |
| `renderer` | `grafana/grafana-image-renderer:v5.12.0` | `8081` | Panel→PNG rendering for image export and alert screenshots |
Two corrections to earlier drafts:
@@ -366,7 +366,7 @@ Two corrections to earlier drafts:
port mapping or run `docker compose exec`.
The collector also bind-mounts the xrpld log root read-only
(`${XRPLD_LOG_DIR:-./data/logs}` → `/var/log/xrpld`) for the `filelog`
(`${XRPLD_LOG_DIR:-./data/logs}` → `/var/log/xrpld`) for the `file_log`
receiver, and the `grafana` service reads Slack/email alert secrets from an
optional gitignored `.env.alerting`.
@@ -570,12 +570,12 @@ Fluentd or PerfLog change. Two pieces:
allocation on the (common) no-span path. This is the ordinary `debug.log`
stream — PerfLog is not involved, and the `setTraceId` hook described in
earlier drafts was never built.
2. **The collector ingests them.** The `filelog` receiver tails
2. **The collector ingests them.** The `file_log` receiver tails
`/var/log/xrpld/*/debug.log` and its `regex_parser` lifts `trace_id` and
`span_id` as optional capture groups (§5.5.1). `resource/logs` applies an
`upsert` of `service.name=xrpld`, which Loki promotes to the stream label
`service_name`, so the canonical selector is **`{service_name="xrpld"}`**.
Logs land in Loki via `otlphttp/loki`.
Logs land in Loki via `otlp_http/loki`.
> **Known issue — the collector's `job` upsert is ineffective for stream
> selection.** `resource/logs` also applies an `upsert` of a `job=xrpld` attribute
@@ -584,7 +584,7 @@ Fluentd or PerfLog change. Two pieces:
> only an **allow-listed** set of resource attributes to indexed stream labels
> (`service.name`, `service.namespace`, `service.instance.id`,
> `deployment.environment`, the `k8s.*`/`cloud.*` keys); `job` is not on that
> list, and this repo ships no Loki config override — `docker-compose.yml:75`
> list, and this repo ships no Loki config override — `docker-compose.yml:116`
> starts Loki with the image's built-in `/etc/loki/local-config.yaml`. `job`
> therefore lands in **structured metadata**, which cannot appear in a stream
> selector, so `{job="xrpld"}` returns an empty result rather than an error.

View File

@@ -633,14 +633,14 @@ See [Phase7_taskList.md](./Phase7_taskList.md) for detailed per-task breakdown.
### Motivation
xrpld's `beast::Journal` logs and OpenTelemetry traces are currently two disjoint observability signals. When investigating an issue, operators must manually correlate timestamps between log files and Tempo traces. Phase 8 bridges this gap by injecting trace context (`trace_id`, `span_id`) into every log line emitted within an active, sampled span, and ingesting those logs into Grafana Loki via the OTel Collector's filelog receiver.
xrpld's `beast::Journal` logs and OpenTelemetry traces are currently two disjoint observability signals. When investigating an issue, operators must manually correlate timestamps between log files and Tempo traces. Phase 8 bridges this gap by injecting trace context (`trace_id`, `span_id`) into every log line emitted within an active, sampled span, and ingesting those logs into Grafana Loki via the OTel Collector's file_log receiver.
#### Gains
1. **One-click trace-to-log navigation** — Click a trace in Tempo and immediately see the corresponding log lines in Loki, filtered by `trace_id`.
2. **Reverse lookup (log-to-trace)** — Loki derived fields make `trace_id` values clickable links back to Tempo.
3. **Unified observability** — All three pillars (traces, metrics, logs) flow through the same OTel Collector pipeline and are visible in a single Grafana instance.
4. **Zero new dependencies in xrpld** — Uses existing OTel SDK headers (`GetSpan`, `GetContext`) already linked in Phase 1.
4. **Zero new dependencies in xrpld** — Uses existing OTel SDK headers (`RuntimeContext`, `SpanContext`) already linked in Phase 1.
5. **Negligible overhead** — The implementation checks the thread-local context value directly, avoiding heap allocation on the no-span path (~15-20ns). On the active-span path, total cost is ~50ns per log call. At typical logging rates, overhead is negligible.
#### Losses / Risks
@@ -651,33 +651,54 @@ xrpld's `beast::Journal` logs and OpenTelemetry traces are currently two disjoin
#### Decision
The correlation value far outweighs the risks. The log format change is backward-compatible (fields are appended only when a sampled span is active), and the filelog receiver regex is straightforward to maintain.
The correlation value far outweighs the risks. The log format change is backward-compatible (fields are appended only when a sampled span is active), and the file_log receiver regex is straightforward to maintain.
### Architecture
Phase 8 has two independent sub-phases that can be developed in parallel:
- **Phase 8a (code change)**: Modify `Logs::format()` in `src/libxrpl/basics/Log.cpp` to append `trace_id=<hex32> span_id=<hex16>` when the current thread has an active OTel span. Guarded by `#ifdef XRPL_ENABLE_TELEMETRY`.
- **Phase 8b (infra only)**: Add Loki to the Docker Compose stack, configure the OTel Collector's `filelog` receiver to tail xrpld's log file, parse out structured fields (timestamp, partition, severity, trace_id, span_id, message), and export to Loki via OTLP. Configure Grafana Tempo↔Loki bidirectional linking.
- **Phase 8b (infra only)**: Add Loki to the Docker Compose stack, configure the OTel Collector's `file_log` receiver to tail xrpld's log file, parse out structured fields (timestamp, partition, severity, trace_id, span_id, message), and export to Loki via OTLP. Configure Grafana Tempo↔Loki bidirectional linking.
#### Trace ID Injection Flow
```mermaid
flowchart LR
subgraph xrpld["xrpld process"]
JLOG["JLOG(j.info())"]
Format["Logs::format()"]
OTelCtx["OTel Context<br/>(thread-local)"]
JLOG["`**JLOG(j.info())**
a log call on some thread`"]
Format["`**Logs::format()**
builds the log line`"]
OTelCtx["`**OTel thread-local context**
RuntimeContext::GetCurrent()
GetValue(kSpanKey)`"]
JLOG --> Format
OTelCtx -.->|"GetSpan()→GetContext()"| Format
OTelCtx -.->|"`GetContext()
if IsValid and IsSampled`"| Format
end
subgraph output["Log Output"]
LogLine["2024-01-15T10:30:45.123Z<br/>LedgerMaster:NFO<br/>trace_id=abc123...<br/>span_id=def456...<br/>Validated ledger 42"]
subgraph output["Log output"]
LogLine["`2026-Jan-15 10:30:45.123456789 UTC
LedgerMaster:NFO
trace_id=abc123... span_id=def456...
Validated ledger 42`"]
end
Format --> LogLine
subgraph legend["Reading the diagram"]
direction LR
L1["`**Solid arrow**
happens on every log call`"]
L2["`**Dotted arrow**
only adds ids when a sampled span is active on this thread`"]
L3["`**kSpanKey lookup**
reads the context value directly, so the no-span path allocates nothing`"]
end
L1 ~~~ L2 ~~~ L3
output ~~~ legend
style xrpld fill:#1a237e,stroke:#0d1642,color:#fff
style output fill:#1b5e20,stroke:#0d3d14,color:#fff
style JLOG fill:#283593,stroke:#1a237e,color:#fff
@@ -691,16 +712,35 @@ flowchart LR
```mermaid
flowchart LR
subgraph collector["OTel Collector"]
FR["filelog receiver<br/>tails debug.log"]
RP["regex_parser<br/>extracts trace_id,<br/>span_id, severity"]
BP["batch processor"]
LE["otlp/loki exporter"]
FR["`**file_log receiver**
tails debug.log`"]
RP["`**regex_parser**
extracts timestamp, partition,
severity, trace_id, span_id`"]
BP["`**batch processor**`"]
LE["`**otlp_http/loki exporter**`"]
FR --> RP --> BP --> LE
end
LogFile["xrpld<br/>debug.log"] --> FR
LE --> Loki["Grafana Loki<br/>:3100"]
Loki <-->|"derivedFields ↔<br/>tracesToLogs"| Tempo["Grafana Tempo"]
LogFile["`**xrpld**
debug.log`"] --> FR
LE --> Loki["`**Grafana Loki**
:3100`"]
Loki <-->|"`derivedFields
tracesToLogs`"| Tempo["`**Grafana Tempo**`"]
subgraph legend["Reading the diagram"]
direction LR
L1["`**Solid arrow**
the path every log line takes`"]
L2["`**Double arrow**
Grafana links the two backends both ways: a trace jumps to its logs, a trace_id in a log jumps back to the trace`"]
L3["`**otlp_http, not otlp**
Loki is reached over OTLP/HTTP; the old dedicated loki exporter was removed upstream`"]
end
L1 ~~~ L2 ~~~ L3
collector ~~~ legend
style collector fill:#e65100,stroke:#bf360c,color:#fff
style FR fill:#f57c00,stroke:#e65100,color:#fff
@@ -718,7 +758,7 @@ flowchart LR
| ---- | ---------------------------------------------- |
| 8.1 | Inject trace_id into Logs::format() |
| 8.2 | Add Loki to Docker Compose stack |
| 8.3 | Add filelog receiver to OTel Collector |
| 8.3 | Add file_log receiver to OTel Collector |
| 8.4 | Configure Grafana trace-to-log correlation |
| 8.5 | Update integration tests |
| 8.6 | Update documentation (runbook, reference docs) |
@@ -732,15 +772,20 @@ flowchart LR
- [x] Log lines outside spans have no trace context (no empty fields) — the
block reads the thread-local span key and appends nothing when it is
absent or the context is invalid (`Log.cpp:310-318`)
- [x] Loki ingests xrpld logs via OTel Collector filelog receiver —
`otel-collector-config.yaml:38` (`filelog`); `loki` service in
`docker-compose.yml:71`
- [x] Loki ingests xrpld logs via OTel Collector file_log receiver —
`otel-collector-config.yaml:38` (`file_log`); `loki` service in
`docker-compose.yml:112`
- [x] Grafana Tempo → Loki one-click correlation works —
`provisioning/datasources/tempo.yaml:32` (`tracesToLogs`)
- [x] Grafana Loki → Tempo reverse lookup works via derived field —
`provisioning/datasources/loki.yaml:16` (`derivedFields`)
- [ ] Integration test verifies trace_id presence in logs — implemented in the
Phase 10 harness, but CI runs it with `--skip-loki`, so it is not gated
- [ ] Integration test verifies trace_id presence in logs — CI gates this
through the Phase 10 harness's `validate_telemetry.py`, whose
`log.trace_id_present` and `log.trace_id_cross_reference` checks run
because the workflow passes no `--skip-loki`. That harness and
`.github/workflows/telemetry-validation.yml` live on the Phase 10 branch,
not here. `docker/telemetry/integration-test.sh:79-126` carries a separate
trace_id-in-logs check that no workflow under `.github/workflows/` runs
- [ ] No performance regression from trace_id injection (< 0.1% overhead) —
needs the Phase 10 benchmark suite
@@ -919,7 +964,7 @@ Alert Rules from External Dashboard**.
## 6.8.3 Phase 10: Synthetic Workload Generation & Telemetry Validation (Weeks 16-17)
> **Status**: Implemented on this branch — `docker/telemetry/workload/` (24
> **Status**: Implemented on this branch — `docker/telemetry/workload/` (25
> files) and `.github/workflows/telemetry-validation.yml` are present here.
> Upstream branches do not carry them, so the exit criteria below only hold from
> `pratik/otel-phase10-workload-validation` onward.
@@ -1003,7 +1048,7 @@ flowchart LR
- **Transaction submitter and RPC load generator** both use xrpld's native WebSocket command format (`{"command": ...}`) — not JSON-RPC format. Response data lives inside `"result"` with `"status"` at the top level.
- **Node config** requires `[signing_support] true` for server-side signing, and `[ips]` (not `[ips_fixed]`) to ensure peer connections count in `peer_finder_active_*` metrics.
- **Metric validation** uses the Prometheus `/api/v1/series` endpoint (not instant queries) to avoid false negatives from stale StatsD gauges. Every metric in `expected_metrics.json` must have > 0 series.
- **Metric validation** uses the Prometheus `/api/v1/series` endpoint (not instant queries) which polls for late-populating series and ignores Prometheus's staleness horizon. Every metric in `expected_metrics.json` must have > 0 series.
- **Gauge visibility**: the harness sets `[insight] server=otel` (`run-full-validation.sh`), so `beast::insight` gauges become OTel observable gauges whose callback is invoked on every collection cycle. A gauge that sits at 0 and never changes (e.g. `jobq_job_count`) therefore still reports, and `/api/v1/series` sees it.
- **I/O latency fix**: `io_latency_sampler` emits unconditionally on first sample, then applies the 10 ms threshold. This ensures `ios_latency` is registered in Prometheus even in low-load CI environments.
- **tx.receive span**: attribute keys are bare, not dotted — `suppressed` and `tx_status` (`TxSpanNames.h:71,75`). `suppressed` is set on both outcomes (`false` on the accepted path, `true` when the HashRouter suppresses), but `tx_status` is set **only** on the reject/known-bad/dropped paths, so it is absent on a successful receive. Assert on the attribute, not on span status.
@@ -1070,15 +1115,16 @@ See [Phase10_taskList.md](./Phase10_taskList.md) for the per-task breakdown.
### CI Deliverable (Task 10.6)
The Phase 10 CI entry point is `.github/workflows/telemetry-validation.yml`
(367 lines, on the Phase 10 branch). It runs three jobs — `linux-image-tag`,
(on the Phase 10 branch). It runs three jobs — `linux-image-tag`,
`build-xrpld`, `validate-telemetry` — and is triggered by `workflow_dispatch`
plus `push` on `pratik/otel-phase*`, `feature/otel-*` and
`feature/telemetry-*`. **There is no cron schedule**, so nothing runs this
workflow on a timer.
plus any `push` that touches one of the `paths` globs below. **There is no
branch filter**: GitHub ANDs `branches` with `paths`, so a branch glob would
decide validation by what a branch is called rather than by what it changed.
**There is no cron schedule**, so nothing runs this workflow on a timer.
> **Fixed — the `push` trigger's `paths` filter now covers the C++ telemetry
> sources.** The branch filter is only half the trigger; `push` also carries a
> `paths` filter, and it previously read:
> sources.** The `push` trigger carries a `paths` filter, and it previously
> read:
>
> ```yaml
> paths:
@@ -1092,14 +1138,14 @@ workflow on a timer.
> `include/xrpl/basics/Telemetry*.h` nor `src/xrpld/app/misc/Telemetry*` exists.
> The telemetry code lives in `src/xrpld/telemetry/**` (9 files, including
> `MetricsRegistry.cpp`), `src/libxrpl/telemetry/**` (7 files) and
> `include/xrpl/telemetry/**` (10 files), none of which were listed.
> `include/xrpl/telemetry/**` (13 files), none of which were listed.
> Consequence at the time: a pure C++ telemetry change — new instrument,
> renamed metric, changed span attribute — never triggered this workflow on
> push; only edits under `docker/telemetry/**` or to the workflow file itself
> did.
>
> The two dead globs have been replaced with the three real module directories,
> so the filter now reads:
> The two dead globs have been replaced with the real module directories, the
> name-constant headers and the checkers, so the filter now reads:
>
> ```yaml
> paths:
@@ -1107,15 +1153,22 @@ workflow on a timer.
> - "docker/telemetry/**"
> - "include/xrpl/telemetry/**"
> - "src/libxrpl/telemetry/**"
> - "src/libxrpl/beast/insight/**"
> - "src/xrpld/telemetry/**"
> - "include/xrpl/beast/insight/**"
> - "src/libxrpl/beast/insight/**"
> - "**/*SpanNames.h"
> - "**/*MetricNames.h"
> - "src/tests/libxrpl/telemetry/**"
> - ".github/scripts/otel-naming/**"
> - ".github/scripts/telemetry/**"
> ```
>
> `src/libxrpl/beast/insight/**` is included because it holds `OTelCollector.cpp`,
> the `beast::insight` OTLP export path the harness depends on. Residual gap: the
> instrumented call sites scattered through `src/xrpld/app/` are not listed, so a
> change that only adds or moves a span at a call site does not trigger the
> workflow on push. Those are reachable by manual dispatch.
> the `beast::insight` OTLP export path the harness depends on. The `*SpanNames.h`
> and `*MetricNames.h` globs cover the name constants wherever they sit, including
> under `src/xrpld/app/`. Residual gap: an instrumented call site that adds or
> moves a span without touching a name header does not trigger the workflow on
> push. Those are reachable by manual dispatch.
> **Caveat — four inert inputs (documented, not wired).** The workflow declares
> five `workflow_dispatch` inputs, but only `run_benchmark` changes behaviour.

View File

@@ -410,7 +410,7 @@ How to correlate OpenTelemetry traces with existing xrpld observability.
There is **one** collection agent, not three. Earlier drafts of this diagram
routed logs through "Promtail/Fluentd" and metrics through a "StatsD Exporter";
neither exists in this stack. Logs are read by the OTel Collector's own
`filelog` receiver, and `beast::insight` metrics arrive at the same collector
`file_log` receiver, and `beast::insight` metrics arrive at the same collector
over OTLP (`[insight] server=otel`). The single-agent shape is the point: one
process, one config file, one place to add redaction or tier tagging.
@@ -422,7 +422,7 @@ flowchart TB
insight["Beast Insight + XRPL_METRIC_*<br/>native OTLP metrics"]
end
otelc["OTel Collector<br/>receivers: otlp, filelog<br/>connector: spanmetrics<br/>3 pipelines"]
otelc["OTel Collector<br/>receivers: otlp, file_log<br/>connector: spanmetrics<br/>3 pipelines"]
subgraph storage["Storage"]
tempo[("Tempo")]
@@ -433,11 +433,11 @@ flowchart TB
dashboards["Grafana<br/>Tempo to Loki via tracesToLogs<br/>Loki to Tempo via derived fields"]
otel -->|"OTLP/HTTP :4318"| otelc
journal -->|"filelog tails<br/>/var/log/xrpld"| otelc
journal -->|"file_log tails<br/>/var/log/xrpld"| otelc
insight -->|"OTLP/HTTP :4318"| otelc
otelc -->|"otlp/tempo"| tempo
otelc -->|"otlphttp/loki"| loki
otelc -->|"otlp_http/loki"| loki
otelc -->|"prometheus :8889"| prom
tempo --> dashboards
@@ -459,7 +459,7 @@ flowchart TB
**Reading the diagram:**
- **xrpld Node (three signals, one transport)**: spans and metrics both leave over OTLP/HTTP on port 4318. Logs do not leave the node at all — the node just writes `debug.log`, and the journal sink prefixes `trace_id=`/`span_id=` whenever a span is active (`Log.cpp:304-338`).
- **OTel Collector (single agent)**: an `otlp` receiver takes spans and metrics; a `filelog` receiver tails `/var/log/xrpld/*/debug.log` and regex-parses the trace/span IDs out of each line. A `spanmetrics` connector derives RED metrics from the trace stream and feeds them into the metrics pipeline. Three pipelines, three exporters — see [05 §5.5.1](./05-configuration-reference.md).
- **OTel Collector (single agent)**: an `otlp` receiver takes spans and metrics; a `file_log` receiver tails `/var/log/xrpld/*/debug.log` and regex-parses the trace/span IDs out of each line. A `spanmetrics` connector derives RED metrics from the trace stream and feeds them into the metrics pipeline. Three pipelines, three exporters — see [05 §5.5.1](./05-configuration-reference.md).
- **PerfLog is not in this picture.** It still writes `perf.log`, but nothing collects it and it carries no trace ID; the `setTraceId` hook once planned for it was never built ([02 §2.6.5](./02-design-decisions.md)).
- **StatsD is not in this picture either.** It remains a supported `[insight] server=` choice, but selecting it takes metrics _out_ of this pipeline and requires a StatsD receiver you would have to add yourself — the compose file's StatsD port mapping is commented out.
- **Grafana**: correlation is bidirectional and configured in the datasources, not in a bespoke panel — Tempo's `tracesToLogs` (`filterByTraceID: true`) jumps trace → logs, and `loki.yaml`'s derived fields jump log → trace.
@@ -471,7 +471,7 @@ flowchart TB
| **Trace** | `trace_id` | Logs | **Live.** Tempo `tracesToLogs`, `filterByTraceID: true` |
| **Trace** | `tx_hash` | — | Live as a span attribute for search; **not** used as a cross-signal join key (`tags: []`) |
| **Trace** | `ledger_seq` | — | Live as a span attribute; not a join key |
| **Journal log** | `trace_id`, `span_id` | Traces | **Live.** Emitted by `Log.cpp:304-338` into `debug.log`, parsed by the collector's `filelog` receiver, jumped via `loki.yaml` derived fields |
| **Journal log** | `trace_id`, `span_id` | Traces | **Live.** Emitted by `Log.cpp:304-338` into `debug.log`, parsed by the collector's `file_log` receiver, jumped via `loki.yaml` derived fields |
| **PerfLog** | `trace_id` | Traces | **Not implemented.** PerfLog output has no trace ID; the planned `setTraceId` hook was never built. Use the journal log instead |
| **Insight** | `exemplar.trace_id` | Traces | **Not implemented.** No exemplar configuration exists anywhere in the code or collector config — no `exemplar_filter` on the SDK side, no `exemplarTraceIdDestinations` on the Prometheus datasource. Metric spike → trace jumps must be done by time range today |
@@ -507,7 +507,7 @@ These are journal (`debug.log`) lines, not PerfLog lines — see §7.7.2.
> **allow-listed** set of resource attributes to indexed stream labels
> (`service.name`, `service.namespace`, `service.instance.id`,
> `deployment.environment`, `k8s.*`, `cloud.*`), and `job` is not on it. This
> repo mounts no Loki config override (`docker-compose.yml:75` uses the image's
> repo mounts no Loki config override (`docker-compose.yml:116` uses the image's
> built-in `local-config.yaml`), so `job` lands in **structured metadata** —
> queryable only with a `|` filter after a selector, never as the selector
> itself. A `{job="xrpld"}` query returns empty with no error, which is why this

View File

@@ -996,7 +996,7 @@ Phase 8 injects OTel trace context into xrpld's `Logs::format()` output, enablin
Example:
```
2024-Jan-15 10:30:45.123456 UTC LedgerMaster:NFO trace_id=abc123def456789012345678abcdef01 span_id=0123456789abcdef Validated ledger 42
2024-Jan-15 10:30:45.123456789 UTC LedgerMaster:NFO trace_id=abc123def456789012345678abcdef01 span_id=0123456789abcdef Validated ledger 42
```
- **`trace_id=<hex32>`** — 32-character lowercase hex trace identifier. Links to the distributed trace in Tempo.
@@ -1010,10 +1010,10 @@ The trace context injection is implemented in `Logs::format()` (`src/libxrpl/bas
### Log Ingestion Pipeline
```
xrpld debug.log -> OTel Collector filelog receiver -> regex_parser -> Loki exporter -> Grafana Loki
xrpld debug.log -> OTel Collector file_log receiver -> regex_parser -> Loki exporter -> Grafana Loki
```
The OTel Collector's `filelog` receiver tails `debug.log` files and uses a `regex_parser` operator to extract structured fields:
The OTel Collector's `file_log` receiver tails `debug.log` files and uses a `regex_parser` operator to extract structured fields:
| Field | Type | Description |
| ----------- | -------- | -------------------------------------------------------- |
@@ -1033,7 +1033,7 @@ Bidirectional linking between logs and traces is configured via Grafana datasour
### Loki Backend
Grafana Loki (v3.7.6) serves as the log storage backend. It receives log entries from the OTel Collector's `otlphttp/loki` exporter via the native OTLP endpoint at `http://loki:3100/otlp`.
Grafana Loki (v3.7.6) serves as the log storage backend. It receives log entries from the OTel Collector's `otlp_http/loki` exporter via the native OTLP endpoint at `http://loki:3100/otlp`.
### LogQL Query Examples

View File

@@ -152,7 +152,7 @@ Before Phases 1-9 can be considered production-ready, we need proof that:
- HTTP: `rpc.http_request` -> `rpc.process` -> `rpc.command.*`
- WebSocket: `rpc.ws_message` -> `rpc.command.*` — **there is no
`rpc.process` on the WS path**. `rpc.process` is created only in
`ServerHandler::processRequest()` (`ServerHandler.cpp:705`), reached from
`ServerHandler::processRequest()` (`ServerHandler.cpp:718`), reached from
`processSession(Session, coro)`, i.e. HTTP only. Under WS-only load
`rpc.process` never appears, and `rpc.command.*` parents directly to
`rpc.ws_message`.
@@ -278,8 +278,9 @@ Before Phases 1-9 can be considered production-ready, we need proof that:
- [x] Validation suite confirms the full span / attribute / metric inventory
(totals computed dynamically from `expected_spans.json` /
`expected_metrics.json`)
- [x] Log-trace correlation validated end-to-end (Loki <-> Tempo) — implemented
and passing locally, but CI runs with `--skip-loki`, so it is not gated
- [x] Log-trace correlation validated end-to-end (Loki <-> Tempo) — implemented,
and gated in CI: the workflow passes no `--skip-loki`, so
`validate_telemetry.py` builds and runs both log-correlation checks
- [x] All 14 harness-asserted Grafana dashboards render data (no empty panels);
15 on disk, `log-derived-insights` unasserted
- [x] Overhead benchmark (`benchmark.sh`) measures telemetry-off vs telemetry-on