diff --git a/docker/telemetry/TESTING.md b/docker/telemetry/TESTING.md index e18b48551d..8c7ae4c5d6 100644 --- a/docker/telemetry/TESTING.md +++ b/docker/telemetry/TESTING.md @@ -10,17 +10,14 @@ pipeline end-to-end, from span generation through the observability stack ### Build xrpld with telemetry -Follow [BUILD.md](../../BUILD.md) with `-o telemetry=True` added. From a build directory (`.build/`): +Build as [BUILD.md](../../BUILD.md) **§ Steps** describes, adding +`-o telemetry=True` to the `conan install` line. That is the only change: +Conan carries `telemetry=ON` into the generated CMake toolchain, so no extra +CMake flag is needed. For the full telemetry build, including how to turn it +off, see [`docs/build/telemetry.md`](../../docs/build/telemetry.md). -```bash -conan install .. --output-folder . --build missing -o telemetry=True --settings build_type=Release -cmake -DCMAKE_TOOLCHAIN_FILE:FILEPATH=build/generators/conan_toolchain.cmake -DCMAKE_BUILD_TYPE=Release -Dxrpld=ON -Dtelemetry=ON .. -cmake --build . --target xrpld -``` - -Conan also writes a `conan-release` preset, so `cmake --preset conan-release -Dtelemetry=ON` works too. There is no preset named `default`. - -The binary is at `.build/xrpld`. +This document assumes the `.build/` layout, so the binary is at `.build/xrpld` +and every command below runs from the repo root. ### Required tools @@ -61,21 +58,14 @@ XRPLD_UID=$(id -u) XRPLD_GID=$(id -g) \ Wait for services to be ready: ```bash -# otel-collector readiness: any HTTP response on the OTLP/HTTP port means the -# receiver is listening. Do NOT use `curl -sf` here — a GET of / returns 404, -# which -f treats as failure even when the collector is healthy. -[ "$(curl -so /dev/null -w '%{http_code}' http://localhost:4318/)" != "000" ] && - echo "collector ready" +# otel-collector readiness: the health_check extension answers on 13133, which +# docker-compose.yml publishes. +curl -sf http://localhost:13133/ >/dev/null && echo "collector ready" # Tempo readiness curl -sf http://localhost:3200/ready >/dev/null && echo "tempo ready" ``` -> The collector's `health_check` extension listens on **13133**, but -> `docker-compose.yml` publishes only 4317, 4318 and 8889 — so 13133 is not -> reachable from the host with the base stack. It is published only by the -> workload validation stack (`docker-compose.workload.yaml`). - ### Step 2: Start xrpld in standalone mode ```bash @@ -197,34 +187,82 @@ If you prefer to run the steps manually: #### Step 1: Start observability stack ```bash -docker compose -f docker/telemetry/docker-compose.yml up -d +XRPLD_LOG_DIR=/tmp/xrpld-integration \ + docker compose -f docker/telemetry/docker-compose.yml up -d ``` +The override is required here. The collector's log mount defaults to the +repo-relative `docker/telemetry/data/logs`, but this test writes its logs under +`/tmp/xrpld-integration`, so without it the `file_log` receiver tails the wrong +root, no log line reaches Loki, and Test 3 Step 3 finds nothing with no error. + #### Step 2: Generate validator keys -Start a temporary standalone xrpld: +Give the throwaway node a config of its own, under the same temp root the rest +of this test uses: ```bash -.build/xrpld --conf docker/telemetry/xrpld-telemetry.cfg -a --start & +mkdir -p /tmp/xrpld-integration/temp-keygen +cat >/tmp/xrpld-integration/temp-keygen/xrpld.cfg <<'EOCFG' +[server] +port_rpc_temp + +[port_rpc_temp] +port = 5099 +ip = 127.0.0.1 +admin = 127.0.0.1 +protocol = http + +[node_db] +type=NuDB +path=/tmp/xrpld-integration/temp-keygen/nudb +online_delete=256 + +[database_path] +/tmp/xrpld-integration/temp-keygen/db + +[debug_logfile] +/tmp/xrpld-integration/temp-keygen/debug.log + +[ssl_verify] +0 +EOCFG +``` + +Do not point this node at `docker/telemetry/xrpld-telemetry.cfg`. That is a +Devnet config whose `[node_db]`, `[database_path]` and `[debug_logfile]` all +resolve under `docker/telemetry/data`, so `--start` (a fresh-genesis start) +would write a genesis chain into the Devnet store, and deleting that directory +afterwards would also destroy the sibling mainnet node's store and every log +under `data/logs/`. Its RPC port is 5005, which is node 1's port later in this +test. + +Start it and wait for RPC before asking for keys: + +```bash +.build/xrpld --conf /tmp/xrpld-integration/temp-keygen/xrpld.cfg -a --start & TEMP_PID=$! -sleep 5 +until curl -sf http://localhost:5099 -d '{"method":"server_info"}' >/dev/null; do + sleep 1 +done ``` Generate 6 key pairs: ```bash for i in $(seq 1 6); do - curl -s http://localhost:5005 \ + curl -s http://localhost:5099 \ -d '{"method":"validation_create"}' | jq '.result' done ``` Record the `validation_seed` and `validation_public_key` for each. -Kill the temporary node: +Stop the temporary node and remove only its own directory: ```bash kill $TEMP_PID -rm -rf docker/telemetry/data/ +wait $TEMP_PID 2>/dev/null +rm -rf /tmp/xrpld-integration/temp-keygen ``` #### Step 3: Create node configs @@ -247,6 +285,9 @@ port = {51234 + node_number} ip = 0.0.0.0 protocol = peer +[network_id] +1025 + [node_db] type=NuDB path=/tmp/xrpld-integration/node{N}/nudb @@ -297,6 +338,14 @@ endpoint=http://localhost:4318/v1/metrics 0 ``` +`[network_id]` has to be a private id (anything other than 0, 1 or 2), because +the config default is id 0 and the telemetry resource maps that to `mainnet` — +without the stanza every span and metric this local cluster emits is stamped +`xrpl.network.type=mainnet` and lands on the same dashboard series as real +mainnet data. Only 0, 1 and 2 have names, so a private id is stamped +`xrpl.network.type=unknown`. That is the value to select in the dashboards' +Network Type filter when looking at this cluster. + #### Step 4: Create validators.txt ```ini @@ -387,45 +436,51 @@ One hole worth knowing: the runbook's Span Reference tables have no row for `method`, `grpc_role` and `grpc_status`, emitted from `GRPCServer.cpp` with the key constants in `src/xrpld/app/main/GrpcSpanNames.h`. -If you find an older inline span inventory in this file or elsewhere, do not -trust it — the copy that used to live here had drifted badly (18 rows under a -"16 spans" heading, whole families missing, and pre-rename dotted `xrpl.*` -attribute keys the code no longer emits). The code and the runbook are the source -of truth. - ### Span → How to Trigger "Test" is the section of this file that exercises the family. `T1` = Test 1 (standalone), `T2` = Test 2 (6-node network). -| Span family (count) | Config toggle | How to trigger | Test | -| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | -| **RPC** (5 total, 3 here): `rpc.http_request`, `rpc.process`, `rpc.command.` | `trace_rpc=1` | Any HTTP JSON-RPC call: `curl -s http://localhost:5005 -d '{"method":"server_info"}'`. `rpc.command.` is one family — the command name is part of the span name. | T1 | -| **RPC** (cont.): `rpc.ws_message`, `rpc.ws_upgrade` | `trace_rpc=1` | Needs a WebSocket client against `[port_ws_public]` (**6005**) or `[port_ws_admin_local]` (6006). `rpc.ws_upgrade` covers the handshake — force a failure to see its error path. `curl` alone will not do it. | — | -| **gRPC** (1): `grpc.` | `trace_rpc=1` | Call a gRPC method (`GetLedger`, `GetLedgerData`, …). **Requires a `[port_grpc]` stanza — the shipped `xrpld-telemetry*.cfg` files define none**, so add one first. | — | -| **Transaction** (6 total, 4 here): `tx.process`, `tx.preflight`, `tx.preclaim`, `tx.transactor` | `trace_transactions` | Submit any transaction (T1 Step 4). The three apply-stage spans share the tx's deterministic trace id; the `stage` attribute says where a failing tx stopped. | T1 | -| **Transaction** (cont.): `tx.receive` | `trace_transactions` | A **peer** relays a transaction. Never appears in standalone — submit on one node of the cluster and look on another. | T2 | -| **Transaction** (cont.): `tx.apply` | `trace_transactions` | Ledger close with a non-empty transaction set: submit, then `ledger_accept` (T1) or wait for consensus (T2). | T1 / T2 | -| **TxQ** (6): `txq.enqueue`, `txq.apply_direct`, `txq.batch_clear`, `txq.accept`, `txq.accept_tx`, `txq.cleanup` | `trace_transactions` | `txq.enqueue`/`apply_direct` on every submission; `txq.accept`/`accept_tx`/`cleanup` on every ledger close. To force real queueing, submit faster than ledgers close or with a fee below the required fee level. | T1 | -| **Consensus** (13): `consensus.round`, `.phase.open`, `.establish`, `.update_positions`, `.check`, `.proposal.send`, `.ledger_close`, `.accept`, `.accept.apply`, `.validation.send`, `.mode_change`, `.proposal.receive`, `.validation.receive` | `trace_consensus=1` | Requires real consensus — **standalone emits none of these**. Bring up T2 and wait for nodes to reach `proposing`; one `consensus.round` per close. `.mode_change` needs an actual mode transition (stop/start a node). | T2 | -| **Ledger** (4 total, 3 here): `ledger.build`, `ledger.validate`, `ledger.store` | `trace_ledger=1` | Any ledger close: `ledger_accept` in standalone, or consensus in T2. | T1 / T2 | -| **Ledger** (cont.): `ledger.acquire` | `trace_ledger=1` | Node fetches a **missing** ledger from peers. Start a node with no history against a running cluster, or restart one node after the others have advanced. | T2 | -| **Peer** (2): `peer.proposal.receive`, `peer.validation.receive` | `trace_peer=1` | Inbound consensus messages from peers; fresh trace roots. T2 only, and high volume. | T2 | -| **PathFind** (4): `pathfind.request`, `pathfind.compute`, `pathfind.discover`, `pathfind.update_all` | `trace_rpc=1` | `curl -s http://localhost:5005 -d '{"method":"ripple_path_find","params":[{"source_account":"…","destination_account":"…","destination_amount":"100"}]}'`. `pathfind.update_all` fires on ledger close while a request is active. | T1 | +| Span family (count) | Config toggle | How to trigger | Test | +| ------------------------------------------------------------------------------------------------------------------------------- | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------- | +| **RPC** (5 total, 3 here): `rpc.http_request`, `rpc.process`, `rpc.command.` | `trace_rpc=1` | Any HTTP JSON-RPC call: `curl -s http://localhost:5005 -d '{"method":"server_info"}'`. `rpc.command.` is one family — the command name is part of the span name. | T1 | +| **RPC** (cont.): `rpc.ws_message`, `rpc.ws_upgrade` | `trace_rpc=1` | Needs a WebSocket client against `[port_ws_public]` (**6005**) or `[port_ws_admin_local]` (6006). `rpc.ws_upgrade` covers the handshake — force a failure to see its error path. `curl` alone will not do it. | — | +| **gRPC** (1): `grpc.` | `trace_rpc=1` | Call a gRPC method (`GetLedger`, `GetLedgerData`, …). **Requires a `[port_grpc]` stanza — the shipped `xrpld-telemetry*.cfg` files define none**, so add one first. | — | +| **Transaction** (6 total, 4 here): `tx.process`, `tx.preflight`, `tx.preclaim`, `tx.transactor` | `trace_transactions` | Submit any transaction (T1 Step 4). The three apply-stage spans share the tx's deterministic trace id; the `stage` attribute says where a failing tx stopped. | T1 | +| **Transaction** (cont.): `tx.receive` | `trace_transactions` | A **peer** relays a transaction. Never appears in standalone — submit on one node of the cluster and look on another. | T2 | +| **Transaction** (cont.): `tx.apply` | `trace_transactions` | Ledger close with a non-empty transaction set: submit, then `ledger_accept` (T1) or wait for consensus (T2). | T1 / T2 | +| **TxQ** (6): `txq.enqueue`, `txq.apply_direct`, `txq.batch_clear`, `txq.accept`, `txq.accept_tx`, `txq.cleanup` | `trace_transactions` | `txq.enqueue`/`apply_direct` on every submission; `txq.accept`/`accept_tx`/`cleanup` on every ledger close. To force real queueing, submit faster than ledgers close or with a fee below the required fee level. | T1 | +| **Consensus** (13 total, 6 here): `consensus.round`, `.phase.open`, `.mode_change`, `.ledger_close`, `.accept`, `.accept.apply` | `trace_consensus=1` | A standalone `ledger_accept` drives a whole simulated round, so these six fire in T1 as well as on every real close in T2. Note `consensus.round` is ended by the **next** round's start, so a single `ledger_accept` leaves it open and Tempo will not return it. | T1 / T2 | +| **Consensus** (cont., 3): `.establish`, `.update_positions`, `.check` | `trace_consensus=1` | Need the establish phase, which the simulated round skips by jumping straight to `Accepted`. Bring up T2 and wait for a timer-driven round. | T2 | +| **Consensus** (cont., 2): `.proposal.send`, `.validation.send` | `trace_consensus=1` | Need the node to propose, which needs a **validator key** — not peers. The shipped `xrpld-telemetry*.cfg` set no `[validation_seed]`/`[validator_token]`, so a standalone node only observes. | T2 | +| **Consensus** (cont., 2): `.proposal.receive`, `.validation.receive` | `trace_consensus=1` | A peer's consensus message arriving. T2 only. | T2 | +| **Ledger** (4 total, 2 here): `ledger.build`, `ledger.store` | `trace_ledger=1` | Any ledger close: `ledger_accept` in standalone, or consensus in T2. | T1 / T2 | +| **Ledger** (cont.): `ledger.validate` | `trace_ledger=1` | Belongs to `LedgerMaster::checkAccept`, which standalone never reaches — `consensusBuilt` returns early and `switchLCL` takes its standalone branch instead. Needs peers or an inbound validation. | T2 | +| **Ledger** (cont.): `ledger.acquire` | `trace_ledger=1` | Node fetches a **missing** ledger from peers. Start a node with no history against a running cluster, or restart one node after the others have advanced. | T2 | +| **Peer** (2): `peer.proposal.receive`, `peer.validation.receive` | `trace_peer=1` | Inbound consensus messages from peers; fresh trace roots. T2 only, and high volume. | T2 | +| **PathFind** (4): `pathfind.request`, `pathfind.compute`, `pathfind.discover`, `pathfind.update_all` | `trace_rpc=1` | `curl -s http://localhost:5005 -d '{"method":"ripple_path_find","params":[{"source_account":"…","destination_account":"…","destination_amount":"100"}]}'`. `pathfind.update_all` fires on ledger close while a request is active. | T1 | Notes that matter when a span you expect is missing: - **Toggles are per-subsystem and all default to on** (`trace_rpc`, `trace_transactions`, `trace_consensus`, `trace_peer`, `trace_ledger`), but `[telemetry] enabled` defaults to **0** — nothing is emitted until it is `1`. -- **`consensus.*` and `peer.*` cannot be produced in standalone mode.** If Test 1 - shows none, that is correct behaviour, not a regression — see "Expected spans - (standalone mode)" above. +- **`peer.*` cannot be produced in standalone mode.** Both peer spans are created + in inbound message handlers, and `-a` turns peerfinder's `autoConnect` off, so + the node opens no outbound peer connections and receives nothing. If Test 1 + shows none, that is correct behaviour, not a regression. +- **`consensus.*` is only partly absent in standalone.** `consensus.round`, + `.phase.open`, `.mode_change`, `.ledger_close`, `.accept` and `.accept.apply` + all fire on a `ledger_accept`; the other seven need the establish phase, a + validator key, or a peer — see the Consensus rows above. - **`rpc.ws_*` and `grpc.*` need a client and a port the quick tests do not use.** Absence in T1/T2 is expected. -- Trace ids are deterministic for transactions (`txID[0:16]`) and consensus - rounds (`prevLedgerHash[0:16]`), so you can compute the id you expect rather - than searching for it. +- Trace ids are deterministic for transactions (from `txID`) and consensus rounds + (from `prevLedgerHash`): the trace id is the hash's first **16 bytes**, so from + a hex-printed hash take the first **32 characters**. This holds under the + default `consensus_trace_strategy=deterministic`; set it to `random` and each + node gives its round a random trace id instead, joinable only by the + `consensus_ledger_id` attribute. --- @@ -499,8 +554,11 @@ registration. For what each dashboard covers, see [`docs/telemetry-runbook.md`](../../docs/telemetry-runbook.md) **§ Grafana -Dashboards** — the per-dashboard reference. Listing them here would be a second -copy that rots (this section previously named 5 of the 15 provisioned). +Dashboards**. That reference is partial: 9 of the 15 provisioned dashboards have +a section there, and six — `fee-market`, `job-queue`, `ledger-data-sync`, +`overlay-traffic-detail`, `peer-quality` and `validator-health` — do not. For +those, open a panel's info icon in Grafana; the panel descriptions carry the same +reference format. Pre-configured datasources: @@ -566,11 +624,17 @@ Consequences worth knowing before you debug against the cloud stack: pipeline, so `span_*` rates stay exact while only ~1 trace in 200 is retrievable by trace ID. A trace you can see in a metric may not exist in Tempo. -- **Pathfinding account hashing does not happen on the cloud export.** The base - config's `attributes/hash` processor hashes `pathfind_source_account` and - `pathfind_dest_account`. It is absent from every cloud pipeline, so those two - attributes leave for Grafana Cloud (and, on that config, for Tempo) with their - raw account values. +- **The same account carries a different token on each config.** No raw account + address leaves the node: the path-finding handlers under + `src/xrpld/rpc/handlers/orderbook/` pass both accounts through + `redactAccount()` first, which is a prefix of the address's SHA-512Half digest + (contract in `include/xrpl/telemetry/Redaction.h`). The base config's + `attributes/hash` processor then hashes that token a second time; no cloud + pipeline has it. The token is deterministic, so one account stays correlatable + across nodes and restarts — but only within one config. A trace stored while + the collector ran the base config must not be joined against a trace stored + under the cloud config, because the same account appears under two different + tokens. ### Step 4: Verify data reaches Grafana Cloud @@ -579,7 +643,7 @@ Cloud instance and confirm: - **Traces**: Explore → hosted Tempo datasource → search `{resource.service.name="xrpld"}` - **Metrics**: Explore → hosted Prometheus/Mimir → query `span_calls_total` -- **Logs**: Explore → hosted Loki → query `{service_name="xrpld"}` (requires `warning`+ file logging). **Not `{job="xrpld"}`** — see the note under Test 3 Step 3. +- **Logs**: Explore → hosted Loki → query `{service_name="xrpld"}` (requires file logging, at a level low enough to keep the correlated lines — the shipped devnet config's `debug` does, the mainnet config's `warning` suppresses them). **Not `{job="xrpld"}`** — see the note under Test 3 Step 3. If nothing appears, check the collector logs for auth/export errors: @@ -603,10 +667,13 @@ end-to-end log-trace correlation pipeline. ### Step 1: Verify trace_id in log output After running Test 1 or Test 2 (which generate RPC spans), check the -xrpld debug.log for trace context: +xrpld debug.log for trace context. A Test 1 run writes +`docker/telemetry/data/logs/xrpld-devnet/debug.log`; the mainnet config writes +`docker/telemetry/data/logs/mainnet/debug.log` instead. ```bash -grep 'trace_id=[a-f0-9]\{32\} span_id=[a-f0-9]\{16\}' /path/to/debug.log +grep 'trace_id=[a-f0-9]\{32\} span_id=[a-f0-9]\{16\}' \ + docker/telemetry/data/logs/xrpld-devnet/debug.log ``` Expected: log lines with `trace_id=<32hex> span_id=<16hex>` between the @@ -624,7 +691,8 @@ NOT have trace context — this is expected. Extract a `trace_id` from the log and verify it exists in Tempo: ```bash -TRACE_ID=$(grep -m1 -o 'trace_id=[a-f0-9]\{32\}' /path/to/debug.log | cut -d= -f2) +TRACE_ID=$(grep -m1 -o 'trace_id=[a-f0-9]\{32\}' \ + docker/telemetry/data/logs/xrpld-devnet/debug.log | cut -d= -f2) echo "Checking trace: $TRACE_ID" curl -s "http://localhost:3200/api/traces/$TRACE_ID" | jq '.batches | length' ``` @@ -668,9 +736,11 @@ Counting `.data.result | length` would count streams, not log lines. > variant also sets `job=xrpld`, in `otel-collector-config.grafanacloud.yaml`. > Either way `{job="xrpld"}` does not work as a selector: on OTLP ingest Loki > promotes only an allow-listed set of resource attributes to indexed stream -> labels (`service.name` → `service_name`, plus `service.namespace`, -> `service.instance.id`, `deployment.environment`, `k8s.*`, `cloud.*`), and `job` -> is not on the list. This repo mounts no Loki config override — the `loki` +> labels. On the pinned `grafana/loki:3.7.6` that list is a fixed 18 keys, +> including `service.name` → `service_name`, `service.namespace`, +> `service.instance.id`, `deployment.environment` and `container.name`. `k8s.*` +> and `cloud.*` are enumerated key lists (ten and two entries), not wildcards. +> `job` is not on the list. This repo mounts no Loki config override — the `loki` > service runs the image's built-in `/etc/loki/local-config.yaml` > named in `docker-compose.yml` — so `job` lands in **structured metadata**, which > cannot be a stream selector. `{job="xrpld"}` therefore returns **zero results @@ -717,11 +787,9 @@ Counting `.data.result | length` would count streams, not log lines. docker compose -f docker/telemetry/docker-compose.yml logs otel-collector ``` 2. Verify xrpld telemetry config has `enabled=1` and correct endpoint -3. Check that otel-collector port 4318 is accessible (`-f` would fail on the - receiver's 404 for `GET /`, so test for any HTTP status instead): - ```bash - curl -so /dev/null -w '%{http_code}\n' http://localhost:4318/ - ``` +3. Check the collector is up — the readiness check in Test 1 Step 1. Probe + `health_check` on 13133, not the OTLP/HTTP port 4318, which answers 404 to a + `GET /` 4. Increase `batch_delay_ms` or decrease `batch_size` in xrpld config ### Nodes not reaching "proposing" state @@ -802,11 +870,13 @@ Counting `.data.result | length` would count streams, not log lines. processors: [resource/tier, resource/stripsdk, batch] exporters: [prometheus] ``` - Both receivers are required. `spanmetrics` carries the span-derived + Both receivers are required. `span_metrics` carries the span-derived `span_*` series; `otlp` carries the node's native `beast::insight` / MetricsRegistry metrics, which arrive on the same OTLP port. Dropping `otlp` silently removes every native metric while the `span_*` ones keep - working — so the dashboards only half-break. + working — so the dashboards only half-break. (The cloud config, + `otel-collector-config.grafanacloud.yaml`, spells the same connector + `spanmetrics`; both are valid ids for it.) 3. Verify Prometheus can reach collector: ```bash curl -s http://localhost:9090/api/v1/targets | jq '.data.activeTargets' diff --git a/docker/telemetry/integration-test.sh b/docker/telemetry/integration-test.sh index 71747301bb..8efe4388b3 100755 --- a/docker/telemetry/integration-test.sh +++ b/docker/telemetry/integration-test.sh @@ -444,6 +444,12 @@ port = $PEER_PORT ip = 0.0.0.0 protocol = peer +# A private id, so telemetry stamps xrpl.network.type=unknown. The config +# default is id 0, which maps to "mainnet" -- this cluster's spans and metrics +# would then share dashboard series with real mainnet data. +[network_id] +1025 + [node_db] type=NuDB path=$NODE_DIR/nudb @@ -479,7 +485,6 @@ trace_transactions=1 trace_consensus=1 trace_peer=1 trace_ledger=1 -metrics_endpoint=http://localhost:4318/v1/metrics [insight] # server=otel is the only load-bearing key here -- it selects OTelCollector so diff --git a/docker/telemetry/otel-collector-config.yaml b/docker/telemetry/otel-collector-config.yaml index 5d238cdae7..6a81eb443d 100644 --- a/docker/telemetry/otel-collector-config.yaml +++ b/docker/telemetry/otel-collector-config.yaml @@ -92,16 +92,17 @@ processors: # and arrives as the label `service_name`, which is what the LogQL # examples in the runbook and TESTING.md select on. # - # A custom `job` attribute is NOT on that list. Verified against - # grafana/loki:3.4.2 with the default config: after ingesting through - # this pipeline, /loki/api/v1/labels returned only `service_name` and - # `deployment_environment`, `{job="xrpld"}` matched 0 streams, and + # A custom `job` attribute is NOT on that list: the pinned + # grafana/loki:3.7.6 prints an 18-key default list under + # limits_config.otlp_config, and `job` is not one of them. Ingesting + # through this pipeline, /loki/api/v1/labels returned only `service_name` + # and `deployment_environment`, `{job="xrpld"}` matched 0 streams, and # `job` appeared as structured metadata instead — which a `{...}` # stream selector cannot match. Promoting it would mean mounting a Loki # config and adding it to limits_config.otlp_config.resource_attributes # (additive to Loki's defaults unless ignore_defaults is set), which is # not worth a constant value — especially as Loki caps index labels at - # 15 and already promotes ~17 by default. Select on `service_name`. + # 15 and already promotes 18 by default. Select on `service_name`. - key: service.name value: xrpld action: upsert