Four conflicts, all where this branch's replacement of the StatsD path with native OTLP met the upstream spanmetrics -> span_metrics rename. This branch's design wins in every case; the rename is carried into its text rather than reverting it, so the connector, its pipeline references, the header comment, the TESTING.md summary and the runbook all use span_metrics while keeping the native-OTLP wording. One addition beyond a straight take-a-side: publish the collector's health check port. This branch restored the health_check extension and its own TESTING.md polls http://localhost:13133/ to decide the collector is ready, but the port was never published on this side of docker-compose.yml, so that check could not pass from the host. Verified the merged config loads with no deprecation warnings and that 13133 is published exactly once. No metric name changed: traces_span_metrics_* already read that way before the rename, which only ever touched component names and prose.
16 KiB
OpenTelemetry Integration Testing Guide
This document describes how to verify the xrpld OpenTelemetry telemetry pipeline end-to-end, from span generation through the observability stack (otel-collector, Tempo, Prometheus, Grafana).
Prerequisites
Build xrpld with telemetry
conan install . --build=missing -o telemetry=True
cmake --preset default -Dtelemetry=ON
cmake --build --preset default --target xrpld
The binary is at .build/xrpld.
Required tools
- Docker with
docker compose(v2) - curl
- jq (JSON processor)
Verify binary
.build/xrpld --version
Test 1: Single-Node Standalone (Quick Verification)
This test verifies RPC and transaction spans in standalone mode, plus the consensus spans that a simulated round still produces. The proposal, voting and peer-facing consensus spans do not fire — see the expected-spans table at the end of this test for which do and which do not.
Step 1: Start the observability stack
docker compose -f docker/telemetry/docker-compose.yml up -d
Wait for services to be ready:
# otel-collector health
curl -sf http://localhost:13133/ && echo "collector ready"
# Tempo readiness
curl -sf http://localhost:3200/ready >/dev/null && echo "tempo ready"
Step 2: Start xrpld in standalone mode
.build/xrpld --conf docker/telemetry/xrpld-telemetry.cfg -a --start
Wait a few seconds for the node to initialize.
Step 3: Exercise RPC spans
# server_info
curl -s http://localhost:5005 \
-d '{"method":"server_info"}' | jq .result.info.server_state
# server_state
curl -s http://localhost:5005 \
-d '{"method":"server_state"}' | jq .result.state.server_state
# ledger
curl -s http://localhost:5005 \
-d '{"method":"ledger","params":[{"ledger_index":"current"}]}' |
jq .result.ledger_current_index
Step 4: Submit a transaction
Close the ledger first (required in standalone mode):
curl -s http://localhost:5005 -d '{"method":"ledger_accept"}'
Submit a Payment from the genesis account:
curl -s http://localhost:5005 -d '{
"method": "submit",
"params": [{
"secret": "snoPBrXtMeMyMHUVTgbuqAfg1SUTb",
"tx_json": {
"TransactionType": "Payment",
"Account": "rHb9CJAWyB4rj91VRWn96DkukG4bwdtyTh",
"Destination": "rPMh7Pi9ct699iZUTWzJaUMR1o42VEfGqF",
"Amount": "10000000"
}
}]
}' | jq .result.engine_result
Expected result: "tesSUCCESS".
Close the ledger again to finalize:
curl -s http://localhost:5005 -d '{"method":"ledger_accept"}'
Step 5: Verify traces in Tempo
Wait 5 seconds for the batch export, then see the "Verification Queries" section below. Its span loop is a superset of what standalone mode produces, so compare its output against the "Expected spans (standalone mode)" table above rather than running a second, narrower set of queries here.
Or open Grafana Explore with Tempo datasource: http://localhost:3000
Step 6: Teardown
# Kill xrpld (Ctrl+C or)
kill $(pgrep -f 'xrpld.*xrpld-telemetry')
# Stop observability stack
docker compose -f docker/telemetry/docker-compose.yml down
# Clean xrpld data
rm -rf docker/telemetry/data/
Expected spans (standalone mode)
| Span Name | Expected | Notes |
|---|---|---|
rpc.http_request |
Yes | Every HTTP RPC call |
rpc.process |
Yes | Every RPC processing |
rpc.command.server_info |
Yes | server_info RPC |
rpc.command.server_state |
Yes | server_state RPC |
rpc.command.ledger |
Yes | ledger RPC |
rpc.command.submit |
Yes | submit RPC |
rpc.command.ledger_accept |
Yes | ledger_accept RPC |
tx.process |
Yes | Transaction submission |
tx.receive |
No | No peers in standalone |
consensus.round, .phase.open, .ledger_close, .accept, .accept.apply |
Yes | ledger_accept drives a simulated round |
consensus.establish, .update_positions, .check, .proposal.*, .validation.receive, .mode_change |
No | simulate jumps straight to Accepted; no peers |
Test 2: 6-Node Consensus Network (Full Verification)
This test verifies ALL span categories including consensus and peer transaction relay, using a 6-node validator network.
Automated
Run the integration test script:
bash docker/telemetry/integration-test.sh
It checks prerequisites, clears the previous run, brings up the observability stack, generates six validator key pairs and their node configs, starts the nodes, waits for consensus and then for a validated ledger, exercises RPC and submits a transaction, verifies traces in Tempo and both the span_metrics and the native beast::insight metrics that arrive over OTLP in Prometheus, checks that no StatsD listener is needed, then prints a summary and leaves the stack running.
The script announces each step as it runs, so read its Step N: headers for the authoritative sequence — they are not restated here, because a numbered copy of them drifts as soon as a step is added.
Its Tempo checks cover the RPC, transaction, consensus, ledger and peer span categories from a fixed list, which is narrower than the loop in the "Verification Queries" section below.
Manual
If you prefer to run the steps manually:
Step 1: Start observability stack
docker compose -f docker/telemetry/docker-compose.yml up -d
Step 2: Generate validator keys
Start a temporary standalone xrpld:
.build/xrpld --conf docker/telemetry/xrpld-telemetry.cfg -a --start &
TEMP_PID=$!
sleep 5
Generate 6 key pairs:
for i in $(seq 1 6); do
curl -s http://localhost:5005 \
-d '{"method":"validation_create"}' | jq '.result'
done
Record the validation_seed and validation_public_key for each.
Kill the temporary node:
kill $TEMP_PID
rm -rf docker/telemetry/data/
Step 3: Create node configs
For each node (1-6), create a config file. Template:
[server]
port_rpc
port_peer
[port_rpc]
port = {5004 + node_number}
ip = 127.0.0.1
admin = 127.0.0.1
protocol = http
[port_peer]
port = {51234 + node_number}
ip = 0.0.0.0
protocol = peer
[node_db]
type=NuDB
path=/tmp/xrpld-integration/node{N}/nudb
online_delete=256
[database_path]
/tmp/xrpld-integration/node{N}/db
[debug_logfile]
/tmp/xrpld-integration/node{N}/debug.log
[validation_seed]
{seed from step 2}
[validators_file]
/tmp/xrpld-integration/validators.txt
[ips_fixed]
{one "127.0.0.1 <port>" line for each port in 51235-51240 except this node's
own 51234 + node_number — a node must not list itself as a fixed peer, so
each config carries five lines, not six}
[peer_private]
1
[telemetry]
enabled=1
service_instance_id=Node-{N}
traces_endpoint=http://localhost:4318/v1/traces
metrics_endpoint=http://localhost:4318/v1/metrics
batch_size=512
batch_delay_ms=2000
max_queue_size=2048
trace_rpc=1
trace_transactions=1
trace_consensus=1
trace_peer=1
trace_ledger=1
[insight]
server=otel
endpoint=http://localhost:4318/v1/metrics
[rpc_startup]
{ "command": "log_level", "severity": "warning" }
[ssl_verify]
0
Step 4: Create validators.txt
[validators]
{public_key_1}
{public_key_2}
{public_key_3}
{public_key_4}
{public_key_5}
{public_key_6}
Step 5: Start all 6 nodes
for i in $(seq 1 6); do
.build/xrpld --conf /tmp/xrpld-integration/node$i/xrpld.cfg --start &
echo $! >/tmp/xrpld-integration/node$i/xrpld.pid
done
Step 6: Wait for consensus
Poll each node until server_state = "proposing":
for port in 5005 5006 5007 5008 5009 5010; do
while true; do
state=$(curl -s http://localhost:$port \
-d '{"method":"server_info"}' |
jq -r '.result.info.server_state')
echo "Port $port: $state"
[ "$state" = "proposing" ] && break
sleep 5
done
done
Step 7: Exercise RPC and submit transaction
# RPC calls
curl -s http://localhost:5005 -d '{"method":"server_info"}'
curl -s http://localhost:5005 -d '{"method":"server_state"}'
curl -s http://localhost:5005 -d '{"method":"ledger","params":[{"ledger_index":"current"}]}'
# Submit transaction
curl -s http://localhost:5005 -d '{
"method": "submit",
"params": [{
"secret": "snoPBrXtMeMyMHUVTgbuqAfg1SUTb",
"tx_json": {
"TransactionType": "Payment",
"Account": "rHb9CJAWyB4rj91VRWn96DkukG4bwdtyTh",
"Destination": "rPMh7Pi9ct699iZUTWzJaUMR1o42VEfGqF",
"Amount": "10000000"
}
}]
}' | jq .result.engine_result
Expected result: "tesSUCCESS", the same as Test 1 Step 4.
Wait 15 seconds for consensus and batch export.
Step 8: Verify in Tempo
See the "Verification Queries" section below.
Expected Span Catalog
A smoke-test subset: the spans a short local run reliably produces, and how to trigger each. This is not the full catalog — the authoritative span list, with the attributes each span carries, is docs/telemetry-runbook.md § Span Reference.
Attributes are deliberately not repeated here. Keeping a second copy is how this table came to list attribute keys that no longer exist anywhere in the code.
| Span Name | Source File | How to Trigger |
|---|---|---|
rpc.http_request |
ServerHandler.cpp | Any HTTP RPC call |
rpc.ws_upgrade |
ServerHandler.cpp | WebSocket upgrade |
rpc.ws_message |
ServerHandler.cpp | WebSocket RPC message |
rpc.process |
ServerHandler.cpp | RPC processing |
rpc.command.<name> |
RPCHandler.cpp | Any RPC command |
tx.process |
NetworkOPs.cpp | Submit transaction |
tx.receive |
PeerImp.cpp | Peer relays transaction |
consensus.proposal.send |
RCLConsensus.cpp | Consensus proposing phase |
consensus.ledger_close |
RCLConsensus.cpp | Ledger close event |
consensus.accept |
RCLConsensus.cpp | Ledger accepted |
consensus.validation.send |
RCLConsensus.cpp | Validation sent |
consensus.accept.apply |
RCLConsensus.cpp | Ledger apply + close time |
tx.apply |
BuildLedger.cpp | Ledger close (tx set) |
ledger.build |
BuildLedger.cpp | Ledger build |
ledger.validate |
LedgerMaster.cpp | Ledger validated |
ledger.store |
LedgerMaster.cpp | Ledger stored |
peer.proposal.receive |
PeerImp.cpp | Peer sends proposal |
peer.validation.receive |
PeerImp.cpp | Peer sends validation |
Verification Queries
Tempo API
Base URL: http://localhost:3200
TEMPO="http://localhost:3200"
# List all services
curl -s "$TEMPO/api/v2/search/tag/resource.service.name/values" | jq '.tagValues[].value'
# Query traces by operation
for op in "rpc.http_request" "rpc.ws_upgrade" "rpc.ws_message" "rpc.process" \
"rpc.command.server_info" "rpc.command.server_state" "rpc.command.ledger" \
"tx.process" "tx.receive" "tx.apply" \
"consensus.proposal.send" "consensus.ledger_close" \
"consensus.accept" "consensus.accept.apply" \
"consensus.validation.send" \
"ledger.build" "ledger.validate" "ledger.store" \
"peer.proposal.receive" "peer.validation.receive"; do
count=$(curl -s "$TEMPO/api/search" \
--data-urlencode "q={resource.service.name=\"xrpld\" && name=\"$op\"}" \
--data-urlencode "limit=5" |
jq '.traces | length')
printf "%-35s %s traces\n" "$op" "$count"
done
Prometheus API
Base URL: http://localhost:9090
PROM="http://localhost:9090"
# Span call counts (from the span_metrics connector). The span_ prefix is the
# connector's `namespace: "span"` in otel-collector-config.yaml; drop that
# setting and these become traces_span_metrics_*.
curl -s "$PROM/api/v1/query?query=span_calls_total" |
jq '.data.result[] | {span: .metric.span_name, count: .value[1]}'
# Latency histogram
curl -s "$PROM/api/v1/query?query=span_duration_milliseconds_count" |
jq '.data.result[] | {span: .metric.span_name, count: .value[1]}'
# RPC calls by command
curl -s "$PROM/api/v1/query?query=span_calls_total{span_name=~\"rpc.command.*\"}" |
jq '.data.result[] | {command: .metric.command, count: .value[1]}'
# Deployment-tier labels present on metrics (set by the collector's
# resource/tier processor and promoted via resource_to_telemetry_conversion).
# Expect deployment_environment and xrpl_network_type on each series.
curl -s "$PROM/api/v1/query?query=span_calls_total" |
jq '.data.result[0].metric | {deployment_environment, xrpl_network_type, service_name}'
Grafana
Open http://localhost:3000 (anonymous admin access enabled).
Pre-configured dashboards:
- RPC Performance: Request rates, latency percentiles by command, top commands, WebSocket rate
- Transaction Overview: Transaction processing rates, apply duration, peer relay, failed tx rate
- Consensus Health: Consensus round duration, proposer counts, mode tracking, accept heatmap
- Ledger Operations: Build/validate/store rates and durations, TX apply metrics
- Peer Network: Proposal/validation receive rates, trusted vs untrusted breakdown (requires
trace_peer=1)
Pre-configured datasources:
- Tempo: Trace data at
http://tempo:3200 - Prometheus: Metrics at
http://prometheus:9090
Troubleshooting
No traces in Tempo
- Check otel-collector logs:
docker compose -f docker/telemetry/docker-compose.yml logs otel-collector - Verify xrpld telemetry config has
enabled=1and correct endpoint - Check that otel-collector port 4318 is accessible:
curl -sf http://localhost:4318 && echo "reachable" - Increase
batch_delay_msor decreasebatch_sizein xrpld config
Nodes not reaching "proposing" state
- Check that all peer ports (51235-51240) are not in use:
for p in 51235 51236 51237 51238 51239 51240; do ss -tlnp | grep ":$p " && echo "port $p in use" done - Verify
[ips_fixed]lists the 5 other peer ports, and not the node's own - Verify
validators.txthas all 6 public keys - Check node debug logs:
tail -50 /tmp/xrpld-integration/node1/debug.log - Ensure
[peer_private]is set to1(prevents reaching out to public network)
Transaction not processing
- Verify genesis account exists:
curl -s http://localhost:5005 \ -d '{"method":"account_info","params":[{"account":"rHb9CJAWyB4rj91VRWn96DkukG4bwdtyTh"}]}' | jq .result.account_data.Balance - Check submit response for error codes
- In standalone mode, remember to call
ledger_acceptafter submitting
Spanmetrics not appearing in Prometheus
- Verify otel-collector config has
span_metricsconnector - Check that the metrics pipeline is configured:
service: pipelines: metrics: receivers: [span_metrics] exporters: [prometheus] - Verify Prometheus can reach collector:
curl -s http://localhost:9090/api/v1/targets | jq '.data.activeTargets'