docs(telemetry): correct the readiness check, build steps and span triggers

The collector readiness note claimed docker-compose.yml publishes only 4317,
4318 and 8889 and that 13133 comes from a workload stack. It publishes 13133,
and that stack is not part of this branch. Probe health_check on 13133 and drop
the note; the troubleshooting entry now points at the same check instead of
carrying a second, weaker copy.

Stop restating BUILD.md. The hardcoded conan and cmake lines had drifted from
it, -Dtelemetry=ON is redundant because the Conan toolchain carries it, and the
conan-release preset resolves only from the repo root, builds into
.build/build/Release rather than .build, and sets no -Dxrpld=ON. Defer to
BUILD.md and docs/build/telemetry.md.

Test 2's keygen step reused the Devnet config with -a --start, which wrote a
genesis chain into the Devnet store, took RPC port 5005 from node 1, and was
followed by an rm -rf that also destroyed the mainnet node's store and every
log. Give it its own config under the test's temp root, as the script does.
The manual path also needs XRPLD_LOG_DIR, or the collector tails the wrong root
and Test 3 finds nothing without erroring.

Neither the template nor the script set [network_id], so a local cluster
stamped xrpl.network.type=mainnet and shared dashboard series with real mainnet
data. Set a private id in both, and say which label it produces. Also drop a
duplicate metrics_endpoint from the generated config.

Split the consensus trigger row: six families fire on a standalone
ledger_accept, and the remaining seven need the establish phase, a validator
key, or a peer. ledger.validate needs peers too, because checkAccept is
unreachable in standalone. Correct the trace-id note to 16 bytes, and name the
strategy it depends on.

The pathfinding bullet said raw account values reach Grafana Cloud. Both
accounts are already tokens when they leave the node; what differs is that the
base config hashes them a second time, so one account carries two tokens across
configs and traces must not be joined across them.

Also: the Loki allow-list is a fixed 18 keys on the pinned image with k8s and
cloud enumerated rather than wildcarded, the runbook documents 9 of 15
dashboards, and the spanmetrics block now uses one spelling with a note that
the cloud config uses the other.
This commit is contained in:
Pratik Mankawde
2026-09-11 11:31:58 +01:00
parent eb76645f69
commit 178fc58c34
3 changed files with 156 additions and 80 deletions

View File

@@ -10,17 +10,14 @@ pipeline end-to-end, from span generation through the observability stack
### Build xrpld with telemetry
Follow [BUILD.md](../../BUILD.md) with `-o telemetry=True` added. From a build directory (`.build/`):
Build as [BUILD.md](../../BUILD.md) **§ Steps** describes, adding
`-o telemetry=True` to the `conan install` line. That is the only change:
Conan carries `telemetry=ON` into the generated CMake toolchain, so no extra
CMake flag is needed. For the full telemetry build, including how to turn it
off, see [`docs/build/telemetry.md`](../../docs/build/telemetry.md).
```bash
conan install .. --output-folder . --build missing -o telemetry=True --settings build_type=Release
cmake -DCMAKE_TOOLCHAIN_FILE:FILEPATH=build/generators/conan_toolchain.cmake -DCMAKE_BUILD_TYPE=Release -Dxrpld=ON -Dtelemetry=ON ..
cmake --build . --target xrpld
```
Conan also writes a `conan-release` preset, so `cmake --preset conan-release -Dtelemetry=ON` works too. There is no preset named `default`.
The binary is at `.build/xrpld`.
This document assumes the `.build/` layout, so the binary is at `.build/xrpld`
and every command below runs from the repo root.
### Required tools
@@ -61,21 +58,14 @@ XRPLD_UID=$(id -u) XRPLD_GID=$(id -g) \
Wait for services to be ready:
```bash
# otel-collector readiness: any HTTP response on the OTLP/HTTP port means the
# receiver is listening. Do NOT use `curl -sf` here — a GET of / returns 404,
# which -f treats as failure even when the collector is healthy.
[ "$(curl -so /dev/null -w '%{http_code}' http://localhost:4318/)" != "000" ] &&
echo "collector ready"
# otel-collector readiness: the health_check extension answers on 13133, which
# docker-compose.yml publishes.
curl -sf http://localhost:13133/ >/dev/null && echo "collector ready"
# Tempo readiness
curl -sf http://localhost:3200/ready >/dev/null && echo "tempo ready"
```
> The collector's `health_check` extension listens on **13133**, but
> `docker-compose.yml` publishes only 4317, 4318 and 8889 — so 13133 is not
> reachable from the host with the base stack. It is published only by the
> workload validation stack (`docker-compose.workload.yaml`).
### Step 2: Start xrpld in standalone mode
```bash
@@ -197,34 +187,82 @@ If you prefer to run the steps manually:
#### Step 1: Start observability stack
```bash
docker compose -f docker/telemetry/docker-compose.yml up -d
XRPLD_LOG_DIR=/tmp/xrpld-integration \
docker compose -f docker/telemetry/docker-compose.yml up -d
```
The override is required here. The collector's log mount defaults to the
repo-relative `docker/telemetry/data/logs`, but this test writes its logs under
`/tmp/xrpld-integration`, so without it the `file_log` receiver tails the wrong
root, no log line reaches Loki, and Test 3 Step 3 finds nothing with no error.
#### Step 2: Generate validator keys
Start a temporary standalone xrpld:
Give the throwaway node a config of its own, under the same temp root the rest
of this test uses:
```bash
.build/xrpld --conf docker/telemetry/xrpld-telemetry.cfg -a --start &
mkdir -p /tmp/xrpld-integration/temp-keygen
cat >/tmp/xrpld-integration/temp-keygen/xrpld.cfg <<'EOCFG'
[server]
port_rpc_temp
[port_rpc_temp]
port = 5099
ip = 127.0.0.1
admin = 127.0.0.1
protocol = http
[node_db]
type=NuDB
path=/tmp/xrpld-integration/temp-keygen/nudb
online_delete=256
[database_path]
/tmp/xrpld-integration/temp-keygen/db
[debug_logfile]
/tmp/xrpld-integration/temp-keygen/debug.log
[ssl_verify]
0
EOCFG
```
Do not point this node at `docker/telemetry/xrpld-telemetry.cfg`. That is a
Devnet config whose `[node_db]`, `[database_path]` and `[debug_logfile]` all
resolve under `docker/telemetry/data`, so `--start` (a fresh-genesis start)
would write a genesis chain into the Devnet store, and deleting that directory
afterwards would also destroy the sibling mainnet node's store and every log
under `data/logs/`. Its RPC port is 5005, which is node 1's port later in this
test.
Start it and wait for RPC before asking for keys:
```bash
.build/xrpld --conf /tmp/xrpld-integration/temp-keygen/xrpld.cfg -a --start &
TEMP_PID=$!
sleep 5
until curl -sf http://localhost:5099 -d '{"method":"server_info"}' >/dev/null; do
sleep 1
done
```
Generate 6 key pairs:
```bash
for i in $(seq 1 6); do
curl -s http://localhost:5005 \
curl -s http://localhost:5099 \
-d '{"method":"validation_create"}' | jq '.result'
done
```
Record the `validation_seed` and `validation_public_key` for each.
Kill the temporary node:
Stop the temporary node and remove only its own directory:
```bash
kill $TEMP_PID
rm -rf docker/telemetry/data/
wait $TEMP_PID 2>/dev/null
rm -rf /tmp/xrpld-integration/temp-keygen
```
#### Step 3: Create node configs
@@ -247,6 +285,9 @@ port = {51234 + node_number}
ip = 0.0.0.0
protocol = peer
[network_id]
1025
[node_db]
type=NuDB
path=/tmp/xrpld-integration/node{N}/nudb
@@ -297,6 +338,14 @@ endpoint=http://localhost:4318/v1/metrics
0
```
`[network_id]` has to be a private id (anything other than 0, 1 or 2), because
the config default is id 0 and the telemetry resource maps that to `mainnet` —
without the stanza every span and metric this local cluster emits is stamped
`xrpl.network.type=mainnet` and lands on the same dashboard series as real
mainnet data. Only 0, 1 and 2 have names, so a private id is stamped
`xrpl.network.type=unknown`. That is the value to select in the dashboards'
Network Type filter when looking at this cluster.
#### Step 4: Create validators.txt
```ini
@@ -387,45 +436,51 @@ One hole worth knowing: the runbook's Span Reference tables have no row for
`method`, `grpc_role` and `grpc_status`, emitted from `GRPCServer.cpp` with the
key constants in `src/xrpld/app/main/GrpcSpanNames.h`.
If you find an older inline span inventory in this file or elsewhere, do not
trust it — the copy that used to live here had drifted badly (18 rows under a
"16 spans" heading, whole families missing, and pre-rename dotted `xrpl.*`
attribute keys the code no longer emits). The code and the runbook are the source
of truth.
### Span → How to Trigger
"Test" is the section of this file that exercises the family. `T1` = Test 1
(standalone), `T2` = Test 2 (6-node network).
| Span family (count) | Config toggle | How to trigger | Test |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- |
| **RPC** (5 total, 3 here): `rpc.http_request`, `rpc.process`, `rpc.command.<name>` | `trace_rpc=1` | Any HTTP JSON-RPC call: `curl -s http://localhost:5005 -d '{"method":"server_info"}'`. `rpc.command.<name>` is one family — the command name is part of the span name. | T1 |
| **RPC** (cont.): `rpc.ws_message`, `rpc.ws_upgrade` | `trace_rpc=1` | Needs a WebSocket client against `[port_ws_public]` (**6005**) or `[port_ws_admin_local]` (6006). `rpc.ws_upgrade` covers the handshake — force a failure to see its error path. `curl` alone will not do it. | — |
| **gRPC** (1): `grpc.<MethodName>` | `trace_rpc=1` | Call a gRPC method (`GetLedger`, `GetLedgerData`, …). **Requires a `[port_grpc]` stanza — the shipped `xrpld-telemetry*.cfg` files define none**, so add one first. | — |
| **Transaction** (6 total, 4 here): `tx.process`, `tx.preflight`, `tx.preclaim`, `tx.transactor` | `trace_transactions` | Submit any transaction (T1 Step 4). The three apply-stage spans share the tx's deterministic trace id; the `stage` attribute says where a failing tx stopped. | T1 |
| **Transaction** (cont.): `tx.receive` | `trace_transactions` | A **peer** relays a transaction. Never appears in standalone — submit on one node of the cluster and look on another. | T2 |
| **Transaction** (cont.): `tx.apply` | `trace_transactions` | Ledger close with a non-empty transaction set: submit, then `ledger_accept` (T1) or wait for consensus (T2). | T1 / T2 |
| **TxQ** (6): `txq.enqueue`, `txq.apply_direct`, `txq.batch_clear`, `txq.accept`, `txq.accept_tx`, `txq.cleanup` | `trace_transactions` | `txq.enqueue`/`apply_direct` on every submission; `txq.accept`/`accept_tx`/`cleanup` on every ledger close. To force real queueing, submit faster than ledgers close or with a fee below the required fee level. | T1 |
| **Consensus** (13): `consensus.round`, `.phase.open`, `.establish`, `.update_positions`, `.check`, `.proposal.send`, `.ledger_close`, `.accept`, `.accept.apply`, `.validation.send`, `.mode_change`, `.proposal.receive`, `.validation.receive` | `trace_consensus=1` | Requires real consensus — **standalone emits none of these**. Bring up T2 and wait for nodes to reach `proposing`; one `consensus.round` per close. `.mode_change` needs an actual mode transition (stop/start a node). | T2 |
| **Ledger** (4 total, 3 here): `ledger.build`, `ledger.validate`, `ledger.store` | `trace_ledger=1` | Any ledger close: `ledger_accept` in standalone, or consensus in T2. | T1 / T2 |
| **Ledger** (cont.): `ledger.acquire` | `trace_ledger=1` | Node fetches a **missing** ledger from peers. Start a node with no history against a running cluster, or restart one node after the others have advanced. | T2 |
| **Peer** (2): `peer.proposal.receive`, `peer.validation.receive` | `trace_peer=1` | Inbound consensus messages from peers; fresh trace roots. T2 only, and high volume. | T2 |
| **PathFind** (4): `pathfind.request`, `pathfind.compute`, `pathfind.discover`, `pathfind.update_all` | `trace_rpc=1` | `curl -s http://localhost:5005 -d '{"method":"ripple_path_find","params":[{"source_account":"…","destination_account":"…","destination_amount":"100"}]}'`. `pathfind.update_all` fires on ledger close while a request is active. | T1 |
| Span family (count) | Config toggle | How to trigger | Test |
| ------------------------------------------------------------------------------------------------------------------------------- | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------- |
| **RPC** (5 total, 3 here): `rpc.http_request`, `rpc.process`, `rpc.command.<name>` | `trace_rpc=1` | Any HTTP JSON-RPC call: `curl -s http://localhost:5005 -d '{"method":"server_info"}'`. `rpc.command.<name>` is one family — the command name is part of the span name. | T1 |
| **RPC** (cont.): `rpc.ws_message`, `rpc.ws_upgrade` | `trace_rpc=1` | Needs a WebSocket client against `[port_ws_public]` (**6005**) or `[port_ws_admin_local]` (6006). `rpc.ws_upgrade` covers the handshake — force a failure to see its error path. `curl` alone will not do it. | — |
| **gRPC** (1): `grpc.<MethodName>` | `trace_rpc=1` | Call a gRPC method (`GetLedger`, `GetLedgerData`, …). **Requires a `[port_grpc]` stanza — the shipped `xrpld-telemetry*.cfg` files define none**, so add one first. | — |
| **Transaction** (6 total, 4 here): `tx.process`, `tx.preflight`, `tx.preclaim`, `tx.transactor` | `trace_transactions` | Submit any transaction (T1 Step 4). The three apply-stage spans share the tx's deterministic trace id; the `stage` attribute says where a failing tx stopped. | T1 |
| **Transaction** (cont.): `tx.receive` | `trace_transactions` | A **peer** relays a transaction. Never appears in standalone — submit on one node of the cluster and look on another. | T2 |
| **Transaction** (cont.): `tx.apply` | `trace_transactions` | Ledger close with a non-empty transaction set: submit, then `ledger_accept` (T1) or wait for consensus (T2). | T1 / T2 |
| **TxQ** (6): `txq.enqueue`, `txq.apply_direct`, `txq.batch_clear`, `txq.accept`, `txq.accept_tx`, `txq.cleanup` | `trace_transactions` | `txq.enqueue`/`apply_direct` on every submission; `txq.accept`/`accept_tx`/`cleanup` on every ledger close. To force real queueing, submit faster than ledgers close or with a fee below the required fee level. | T1 |
| **Consensus** (13 total, 6 here): `consensus.round`, `.phase.open`, `.mode_change`, `.ledger_close`, `.accept`, `.accept.apply` | `trace_consensus=1` | A standalone `ledger_accept` drives a whole simulated round, so these six fire in T1 as well as on every real close in T2. Note `consensus.round` is ended by the **next** round's start, so a single `ledger_accept` leaves it open and Tempo will not return it. | T1 / T2 |
| **Consensus** (cont., 3): `.establish`, `.update_positions`, `.check` | `trace_consensus=1` | Need the establish phase, which the simulated round skips by jumping straight to `Accepted`. Bring up T2 and wait for a timer-driven round. | T2 |
| **Consensus** (cont., 2): `.proposal.send`, `.validation.send` | `trace_consensus=1` | Need the node to propose, which needs a **validator key** — not peers. The shipped `xrpld-telemetry*.cfg` set no `[validation_seed]`/`[validator_token]`, so a standalone node only observes. | T2 |
| **Consensus** (cont., 2): `.proposal.receive`, `.validation.receive` | `trace_consensus=1` | A peer's consensus message arriving. T2 only. | T2 |
| **Ledger** (4 total, 2 here): `ledger.build`, `ledger.store` | `trace_ledger=1` | Any ledger close: `ledger_accept` in standalone, or consensus in T2. | T1 / T2 |
| **Ledger** (cont.): `ledger.validate` | `trace_ledger=1` | Belongs to `LedgerMaster::checkAccept`, which standalone never reaches — `consensusBuilt` returns early and `switchLCL` takes its standalone branch instead. Needs peers or an inbound validation. | T2 |
| **Ledger** (cont.): `ledger.acquire` | `trace_ledger=1` | Node fetches a **missing** ledger from peers. Start a node with no history against a running cluster, or restart one node after the others have advanced. | T2 |
| **Peer** (2): `peer.proposal.receive`, `peer.validation.receive` | `trace_peer=1` | Inbound consensus messages from peers; fresh trace roots. T2 only, and high volume. | T2 |
| **PathFind** (4): `pathfind.request`, `pathfind.compute`, `pathfind.discover`, `pathfind.update_all` | `trace_rpc=1` | `curl -s http://localhost:5005 -d '{"method":"ripple_path_find","params":[{"source_account":"…","destination_account":"…","destination_amount":"100"}]}'`. `pathfind.update_all` fires on ledger close while a request is active. | T1 |
Notes that matter when a span you expect is missing:
- **Toggles are per-subsystem and all default to on** (`trace_rpc`,
`trace_transactions`, `trace_consensus`, `trace_peer`, `trace_ledger`), but
`[telemetry] enabled` defaults to **0** — nothing is emitted until it is `1`.
- **`consensus.*` and `peer.*` cannot be produced in standalone mode.** If Test 1
shows none, that is correct behaviour, not a regression — see "Expected spans
(standalone mode)" above.
- **`peer.*` cannot be produced in standalone mode.** Both peer spans are created
in inbound message handlers, and `-a` turns peerfinder's `autoConnect` off, so
the node opens no outbound peer connections and receives nothing. If Test 1
shows none, that is correct behaviour, not a regression.
- **`consensus.*` is only partly absent in standalone.** `consensus.round`,
`.phase.open`, `.mode_change`, `.ledger_close`, `.accept` and `.accept.apply`
all fire on a `ledger_accept`; the other seven need the establish phase, a
validator key, or a peer — see the Consensus rows above.
- **`rpc.ws_*` and `grpc.*` need a client and a port the quick tests do not
use.** Absence in T1/T2 is expected.
- Trace ids are deterministic for transactions (`txID[0:16]`) and consensus
rounds (`prevLedgerHash[0:16]`), so you can compute the id you expect rather
than searching for it.
- Trace ids are deterministic for transactions (from `txID`) and consensus rounds
(from `prevLedgerHash`): the trace id is the hash's first **16 bytes**, so from
a hex-printed hash take the first **32 characters**. This holds under the
default `consensus_trace_strategy=deterministic`; set it to `random` and each
node gives its round a random trace id instead, joinable only by the
`consensus_ledger_id` attribute.
---
@@ -499,8 +554,11 @@ registration.
For what each dashboard covers, see
[`docs/telemetry-runbook.md`](../../docs/telemetry-runbook.md) **§ Grafana
Dashboards** — the per-dashboard reference. Listing them here would be a second
copy that rots (this section previously named 5 of the 15 provisioned).
Dashboards**. That reference is partial: 9 of the 15 provisioned dashboards have
a section there, and six — `fee-market`, `job-queue`, `ledger-data-sync`,
`overlay-traffic-detail`, `peer-quality` and `validator-health` — do not. For
those, open a panel's info icon in Grafana; the panel descriptions carry the same
reference format.
Pre-configured datasources:
@@ -566,11 +624,17 @@ Consequences worth knowing before you debug against the cloud stack:
pipeline, so `span_*` rates stay exact while only ~1 trace in 200 is
retrievable by trace ID. A trace you can see in a metric may not exist in
Tempo.
- **Pathfinding account hashing does not happen on the cloud export.** The base
config's `attributes/hash` processor hashes `pathfind_source_account` and
`pathfind_dest_account`. It is absent from every cloud pipeline, so those two
attributes leave for Grafana Cloud (and, on that config, for Tempo) with their
raw account values.
- **The same account carries a different token on each config.** No raw account
address leaves the node: the path-finding handlers under
`src/xrpld/rpc/handlers/orderbook/` pass both accounts through
`redactAccount()` first, which is a prefix of the address's SHA-512Half digest
(contract in `include/xrpl/telemetry/Redaction.h`). The base config's
`attributes/hash` processor then hashes that token a second time; no cloud
pipeline has it. The token is deterministic, so one account stays correlatable
across nodes and restarts — but only within one config. A trace stored while
the collector ran the base config must not be joined against a trace stored
under the cloud config, because the same account appears under two different
tokens.
### Step 4: Verify data reaches Grafana Cloud
@@ -579,7 +643,7 @@ Cloud instance and confirm:
- **Traces**: Explore → hosted Tempo datasource → search `{resource.service.name="xrpld"}`
- **Metrics**: Explore → hosted Prometheus/Mimir → query `span_calls_total`
- **Logs**: Explore → hosted Loki → query `{service_name="xrpld"}` (requires `warning`+ file logging). **Not `{job="xrpld"}`** — see the note under Test 3 Step 3.
- **Logs**: Explore → hosted Loki → query `{service_name="xrpld"}` (requires file logging, at a level low enough to keep the correlated lines — the shipped devnet config's `debug` does, the mainnet config's `warning` suppresses them). **Not `{job="xrpld"}`** — see the note under Test 3 Step 3.
If nothing appears, check the collector logs for auth/export errors:
@@ -603,10 +667,13 @@ end-to-end log-trace correlation pipeline.
### Step 1: Verify trace_id in log output
After running Test 1 or Test 2 (which generate RPC spans), check the
xrpld debug.log for trace context:
xrpld debug.log for trace context. A Test 1 run writes
`docker/telemetry/data/logs/xrpld-devnet/debug.log`; the mainnet config writes
`docker/telemetry/data/logs/mainnet/debug.log` instead.
```bash
grep 'trace_id=[a-f0-9]\{32\} span_id=[a-f0-9]\{16\}' /path/to/debug.log
grep 'trace_id=[a-f0-9]\{32\} span_id=[a-f0-9]\{16\}' \
docker/telemetry/data/logs/xrpld-devnet/debug.log
```
Expected: log lines with `trace_id=<32hex> span_id=<16hex>` between the
@@ -624,7 +691,8 @@ NOT have trace context — this is expected.
Extract a `trace_id` from the log and verify it exists in Tempo:
```bash
TRACE_ID=$(grep -m1 -o 'trace_id=[a-f0-9]\{32\}' /path/to/debug.log | cut -d= -f2)
TRACE_ID=$(grep -m1 -o 'trace_id=[a-f0-9]\{32\}' \
docker/telemetry/data/logs/xrpld-devnet/debug.log | cut -d= -f2)
echo "Checking trace: $TRACE_ID"
curl -s "http://localhost:3200/api/traces/$TRACE_ID" | jq '.batches | length'
```
@@ -668,9 +736,11 @@ Counting `.data.result | length` would count streams, not log lines.
> variant also sets `job=xrpld`, in `otel-collector-config.grafanacloud.yaml`.
> Either way `{job="xrpld"}` does not work as a selector: on OTLP ingest Loki
> promotes only an allow-listed set of resource attributes to indexed stream
> labels (`service.name` → `service_name`, plus `service.namespace`,
> `service.instance.id`, `deployment.environment`, `k8s.*`, `cloud.*`), and `job`
> is not on the list. This repo mounts no Loki config override — the `loki`
> labels. On the pinned `grafana/loki:3.7.6` that list is a fixed 18 keys,
> including `service.name` → `service_name`, `service.namespace`,
> `service.instance.id`, `deployment.environment` and `container.name`. `k8s.*`
> and `cloud.*` are enumerated key lists (ten and two entries), not wildcards.
> `job` is not on the list. This repo mounts no Loki config override — the `loki`
> service runs the image's built-in `/etc/loki/local-config.yaml`
> named in `docker-compose.yml` — so `job` lands in **structured metadata**, which
> cannot be a stream selector. `{job="xrpld"}` therefore returns **zero results
@@ -717,11 +787,9 @@ Counting `.data.result | length` would count streams, not log lines.
docker compose -f docker/telemetry/docker-compose.yml logs otel-collector
```
2. Verify xrpld telemetry config has `enabled=1` and correct endpoint
3. Check that otel-collector port 4318 is accessible (`-f` would fail on the
receiver's 404 for `GET /`, so test for any HTTP status instead):
```bash
curl -so /dev/null -w '%{http_code}\n' http://localhost:4318/
```
3. Check the collector is up — the readiness check in Test 1 Step 1. Probe
`health_check` on 13133, not the OTLP/HTTP port 4318, which answers 404 to a
`GET /`
4. Increase `batch_delay_ms` or decrease `batch_size` in xrpld config
### Nodes not reaching "proposing" state
@@ -802,11 +870,13 @@ Counting `.data.result | length` would count streams, not log lines.
processors: [resource/tier, resource/stripsdk, batch]
exporters: [prometheus]
```
Both receivers are required. `spanmetrics` carries the span-derived
Both receivers are required. `span_metrics` carries the span-derived
`span_*` series; `otlp` carries the node's native `beast::insight` /
MetricsRegistry metrics, which arrive on the same OTLP port. Dropping
`otlp` silently removes every native metric while the `span_*` ones keep
working — so the dashboards only half-break.
working — so the dashboards only half-break. (The cloud config,
`otel-collector-config.grafanacloud.yaml`, spells the same connector
`spanmetrics`; both are valid ids for it.)
3. Verify Prometheus can reach collector:
```bash
curl -s http://localhost:9090/api/v1/targets | jq '.data.activeTargets'

View File

@@ -444,6 +444,12 @@ port = $PEER_PORT
ip = 0.0.0.0
protocol = peer
# A private id, so telemetry stamps xrpl.network.type=unknown. The config
# default is id 0, which maps to "mainnet" -- this cluster's spans and metrics
# would then share dashboard series with real mainnet data.
[network_id]
1025
[node_db]
type=NuDB
path=$NODE_DIR/nudb
@@ -479,7 +485,6 @@ trace_transactions=1
trace_consensus=1
trace_peer=1
trace_ledger=1
metrics_endpoint=http://localhost:4318/v1/metrics
[insight]
# server=otel is the only load-bearing key here -- it selects OTelCollector so

View File

@@ -92,16 +92,17 @@ processors:
# and arrives as the label `service_name`, which is what the LogQL
# examples in the runbook and TESTING.md select on.
#
# A custom `job` attribute is NOT on that list. Verified against
# grafana/loki:3.4.2 with the default config: after ingesting through
# this pipeline, /loki/api/v1/labels returned only `service_name` and
# `deployment_environment`, `{job="xrpld"}` matched 0 streams, and
# A custom `job` attribute is NOT on that list: the pinned
# grafana/loki:3.7.6 prints an 18-key default list under
# limits_config.otlp_config, and `job` is not one of them. Ingesting
# through this pipeline, /loki/api/v1/labels returned only `service_name`
# and `deployment_environment`, `{job="xrpld"}` matched 0 streams, and
# `job` appeared as structured metadata instead — which a `{...}`
# stream selector cannot match. Promoting it would mean mounting a Loki
# config and adding it to limits_config.otlp_config.resource_attributes
# (additive to Loki's defaults unless ignore_defaults is set), which is
# not worth a constant value — especially as Loki caps index labels at
# 15 and already promotes ~17 by default. Select on `service_name`.
# 15 and already promotes 18 by default. Select on `service_name`.
- key: service.name
value: xrpld
action: upsert