fix(telemetry): correct job_count/network-traffic queries, 10s refresh, dashboard layouts

Dashboard audit against the live datasource surfaced two broken panel queries
and applied UX polish across all 14 dashboards.

- node-health "Job Queue Depth": job_count -> jobq_job_count. The JobQueue
  collector is wrapped in group("jobq") (Application.cpp), so the registered
  job_count gauge is emitted with the jobq_ prefix; the panel queried the
  unprefixed name and returned nothing.
- network-traffic "Overlay Traffic by Category" + "All Traffic Categories":
  topk(N, rate({__name__=~".*_bytes_in"}[...])) errors on Mimir ("vector
  cannot contain metrics with the same labelset") because rate() drops
  __name__ and the many counters collapse. Replaced with an enumerated
  label_replace form that re-attaches __name__ per metric, preserving the
  {{__name__}} legend and per-series display-name overrides.
- All 14 dashboards: refresh set to 10s.
- peer-quality: each panel full screen width.
- validator-health: at most two panels per row (row headers preserved).
- docs: telemetry-runbook and 06-implementation-phases updated for the
  jobq_ prefix and the network-traffic query pattern.

Verified end-to-end against a local mainnet xrpld node feeding the local
stack: jobq_job_count returns data (old job_count empty), both network-traffic
exprs execute (old form reproduces the labelset error), and panels render
through the Grafana proxy. All 14 pass validate_dashboards.py.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Pratik Mankawde
2026-07-11 19:08:12 +01:00
parent 78ce739840
commit 7231d450a4
16 changed files with 132 additions and 118 deletions

View File

@@ -336,21 +336,21 @@ xrpld has a mature metrics framework (`beast::insight`) that emits StatsD-format
### Metric Inventory
| Category | Group | Type | Count | Key Metrics |
| --------------- | ------------------ | ------------- | ---------- | ------------------------------------------------------ |
| Node State | `State_Accounting` | Gauge | 10 | `*_duration`, `*_transitions` per operating mode |
| Ledger | `LedgerMaster` | Gauge | 2 | `Validated_Ledger_Age`, `Published_Ledger_Age` |
| Ledger Fetch | — | Counter | 1 | `ledger_fetches` |
| Ledger History | `ledger.history` | Counter | 1 | `mismatch` |
| RPC | `rpc` | Counter+Event | 3 | `requests`, `time` (histogram), `size` (histogram) |
| Job Queue | — | Gauge+Event | 1 + 2×N | `job_count`, per-job `{name}` and `{name}_q` |
| Peer Finder | `Peer_Finder` | Gauge | 2 | `Active_Inbound_Peers`, `Active_Outbound_Peers` |
| Overlay | `Overlay` | Gauge | 1 | `Peer_Disconnects` |
| Overlay Traffic | per-category | Gauge | 4×57 = 228 | `Bytes_In/Out`, `Messages_In/Out` per traffic category |
| Pathfinding | — | Event | 2 | `pathfind_fast`, `pathfind_full` (histograms) |
| I/O | — | Event | 1 | `ios_latency` (histogram) |
| Resource Mgr | — | Meter | 2 | `warn`, `drop` (rate counters) |
| Caches | per-cache | Gauge | 2×N | `{cache}.size`, `{cache}.hit_rate` |
| Category | Group | Type | Count | Key Metrics |
| --------------- | ------------------ | ------------- | ---------- | ----------------------------------------------------------------------------------------------------------- |
| Node State | `State_Accounting` | Gauge | 10 | `*_duration`, `*_transitions` per operating mode |
| Ledger | `LedgerMaster` | Gauge | 2 | `Validated_Ledger_Age`, `Published_Ledger_Age` |
| Ledger Fetch | — | Counter | 1 | `ledger_fetches` |
| Ledger History | `ledger.history` | Counter | 1 | `mismatch` |
| RPC | `rpc` | Counter+Event | 3 | `requests`, `time` (histogram), `size` (histogram) |
| Job Queue | `jobq` | Gauge+Event | 1 + 2×N | `job_count`, per-job `{name}` and `{name}_q` (emitted with the `jobq_` group prefix, e.g. `jobq_job_count`) |
| Peer Finder | `Peer_Finder` | Gauge | 2 | `Active_Inbound_Peers`, `Active_Outbound_Peers` |
| Overlay | `Overlay` | Gauge | 1 | `Peer_Disconnects` |
| Overlay Traffic | per-category | Gauge | 4×57 = 228 | `Bytes_In/Out`, `Messages_In/Out` per traffic category |
| Pathfinding | — | Event | 2 | `pathfind_fast`, `pathfind_full` (histograms) |
| I/O | — | Event | 1 | `ios_latency` (histogram) |
| Resource Mgr | — | Meter | 2 | `warn`, `drop` (rate counters) |
| Caches | per-cache | Gauge | 2×N | `{cache}.size`, `{cache}.hit_rate` |
**Total**: ~255+ unique metrics (plus dynamic job-type and cache metrics)

View File

@@ -1279,5 +1279,6 @@
},
"title": "Consensus Health",
"uid": "consensus-health",
"description": "What this shows: Consensus health for XRPL nodes: how reliably and quickly the network agrees each ledger, and where agreement breaks down.\nUse it to: Spot stalled or slow consensus rounds and pinpoint the phase where agreement is failing."
"description": "What this shows: Consensus health for XRPL nodes: how reliably and quickly the network agrees each ledger, and where agreement breaks down.\nUse it to: Spot stalled or slow consensus rounds and pinpoint the phase where agreement is failing.",
"refresh": "10s"
}

View File

@@ -508,5 +508,5 @@
"title": "Fee Market & TxQ",
"uid": "fee-market",
"version": 1,
"refresh": "5s"
"refresh": "10s"
}

View File

@@ -594,5 +594,5 @@
"title": "Job Queue Analysis",
"uid": "job-queue",
"version": 1,
"refresh": "5s"
"refresh": "10s"
}

View File

@@ -635,5 +635,6 @@
"to": "now"
},
"title": "Ledger Data & Sync",
"uid": "ledger-data-sync"
"uid": "ledger-data-sync",
"refresh": "10s"
}

View File

@@ -418,6 +418,6 @@
},
"title": "Ledger Operations",
"uid": "ledger-operations",
"refresh": "5s",
"refresh": "10s",
"description": "What this shows: Ledger construction, validation, and storage activity and timing for this node.\nUse it to: Confirm ledgers are being built, validated, and stored on schedule and find the slow stage when they are not."
}

File diff suppressed because one or more lines are too long

View File

@@ -317,7 +317,7 @@
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"expr": "job_count{service_instance_id=~\"$node\", deployment_environment=~\"$deployment_environment\", xrpl_network_type=~\"$xrpl_network_type\", service_name=~\"$service_name\"}",
"expr": "jobq_job_count{service_instance_id=~\"$node\", deployment_environment=~\"$deployment_environment\", xrpl_network_type=~\"$xrpl_network_type\", service_name=~\"$service_name\"}",
"legendFormat": "Job Queue Depth [{{service_instance_id}}]"
}
],
@@ -2592,5 +2592,5 @@
},
"title": "Node Health",
"uid": "node-health",
"refresh": "5s"
"refresh": "10s"
}

View File

@@ -692,5 +692,6 @@
"to": "now"
},
"title": "Overlay Traffic Detail",
"uid": "overlay-traffic-detail"
"uid": "overlay-traffic-detail",
"refresh": "10s"
}

View File

@@ -397,5 +397,5 @@
},
"title": "Peer Network",
"uid": "peer-network",
"refresh": "5s"
"refresh": "10s"
}

View File

@@ -15,7 +15,7 @@
"type": "timeseries",
"gridPos": {
"h": 8,
"w": 12,
"w": 24,
"x": 0,
"y": 0
},
@@ -77,9 +77,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"x": 12,
"y": 0
"w": 24,
"x": 0,
"y": 8
},
"options": {
"tooltip": {
@@ -127,9 +127,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"x": 18,
"y": 0
"w": 24,
"x": 0,
"y": 16
},
"options": {
"tooltip": {
@@ -179,9 +179,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"w": 24,
"x": 0,
"y": 8
"y": 24
},
"options": {
"tooltip": {
@@ -245,9 +245,9 @@
"type": "timeseries",
"gridPos": {
"h": 8,
"w": 10,
"x": 6,
"y": 8
"w": 24,
"x": 0,
"y": 32
},
"options": {
"tooltip": {
@@ -291,9 +291,9 @@
"type": "bargauge",
"gridPos": {
"h": 8,
"w": 8,
"x": 16,
"y": 8
"w": 24,
"x": 0,
"y": 40
},
"options": {
"orientation": "horizontal",
@@ -474,5 +474,5 @@
},
"title": "Peer Quality",
"uid": "peer-quality",
"refresh": "5s"
"refresh": "10s"
}

View File

@@ -690,5 +690,5 @@
},
"title": "RPC & Pathfinding",
"uid": "rpc-pathfinding",
"refresh": "5s"
"refresh": "10s"
}

View File

@@ -1011,6 +1011,6 @@
},
"title": "RPC Performance",
"uid": "rpc-performance",
"refresh": "5s",
"refresh": "10s",
"description": "What this shows: Per-command and per-method RPC performance: call rates, error rates, and latency distributions.\nUse it to: Identify slow or failing RPC commands and track client-facing request latency."
}

View File

@@ -1006,6 +1006,6 @@
},
"title": "Transaction Overview",
"uid": "transaction-overview",
"refresh": "5s",
"refresh": "10s",
"description": "What this shows: Transaction flow through this node: receipt, processing, results, per-stage timing, and queue behavior.\nUse it to: Trace transactions from arrival to ledger, and locate stalls in processing or the queue."
}

View File

@@ -27,7 +27,7 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"w": 12,
"x": 0,
"y": 1
},
@@ -79,8 +79,8 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"x": 6,
"w": 12,
"x": 12,
"y": 1
},
"options": {
@@ -131,9 +131,9 @@
"type": "bargauge",
"gridPos": {
"h": 8,
"w": 6,
"x": 12,
"y": 1
"w": 12,
"x": 0,
"y": 9
},
"options": {
"orientation": "horizontal",
@@ -199,9 +199,9 @@
"type": "bargauge",
"gridPos": {
"h": 8,
"w": 6,
"x": 18,
"y": 1
"w": 12,
"x": 12,
"y": 9
},
"options": {
"orientation": "horizontal",
@@ -268,7 +268,7 @@
"h": 1,
"w": 24,
"x": 0,
"y": 9
"y": 17
},
"collapsed": false,
"panels": []
@@ -279,9 +279,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"w": 12,
"x": 0,
"y": 10
"y": 18
},
"options": {
"tooltip": {
@@ -329,9 +329,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"x": 6,
"y": 10
"w": 12,
"x": 12,
"y": 18
},
"options": {
"tooltip": {
@@ -363,9 +363,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"x": 12,
"y": 10
"w": 12,
"x": 0,
"y": 26
},
"options": {
"tooltip": {
@@ -429,9 +429,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"x": 18,
"y": 10
"w": 12,
"x": 12,
"y": 26
},
"options": {
"tooltip": {
@@ -479,9 +479,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"w": 12,
"x": 0,
"y": 18
"y": 34
},
"options": {
"tooltip": {
@@ -545,9 +545,9 @@
"type": "timeseries",
"gridPos": {
"h": 8,
"w": 18,
"x": 6,
"y": 18
"w": 12,
"x": 12,
"y": 34
},
"options": {
"tooltip": {
@@ -616,7 +616,7 @@
"h": 1,
"w": 24,
"x": 0,
"y": 26
"y": 42
},
"collapsed": false,
"panels": []
@@ -627,9 +627,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 6,
"w": 12,
"x": 0,
"y": 27
"y": 43
},
"options": {
"tooltip": {
@@ -661,9 +661,9 @@
"type": "timeseries",
"gridPos": {
"h": 8,
"w": 18,
"x": 6,
"y": 27
"w": 12,
"x": 12,
"y": 43
},
"options": {
"tooltip": {
@@ -707,9 +707,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 8,
"w": 12,
"x": 0,
"y": 35
"y": 51
},
"options": {
"tooltip": {
@@ -741,9 +741,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 8,
"x": 8,
"y": 35
"w": 12,
"x": 12,
"y": 51
},
"options": {
"tooltip": {
@@ -791,9 +791,9 @@
"type": "stat",
"gridPos": {
"h": 8,
"w": 8,
"x": 16,
"y": 35
"w": 12,
"x": 0,
"y": 59
},
"options": {
"tooltip": {
@@ -842,8 +842,8 @@
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 43
"x": 12,
"y": 59
},
"options": {
"tooltip": {
@@ -893,9 +893,9 @@
"type": "timeseries",
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 43
"w": 24,
"x": 0,
"y": 67
},
"options": {
"tooltip": {
@@ -1041,5 +1041,5 @@
},
"title": "Validator Health",
"uid": "validator-health",
"refresh": "5s"
"refresh": "10s"
}

View File

@@ -738,18 +738,18 @@ The `OTelCollector` implementation exports metrics via OTLP/HTTP to the same OTe
#### Gauges
| Prometheus Metric | Source | Description |
| ------------------------------------- | ------------------------- | -------------------------------------------------------------------------- |
| `ledgermaster_validated_ledger_age` | LedgerMaster.h:373 | Age of validated ledger (seconds) |
| `ledgermaster_published_ledger_age` | LedgerMaster.h:374 | Age of published ledger (seconds) |
| `state_accounting_{mode}_duration` | NetworkOPs.cpp:774 | Time in each operating mode (Disconnected/Connected/Syncing/Tracking/Full) |
| `state_accounting_{mode}_transitions` | NetworkOPs.cpp:780 | Transition count per mode |
| `peer_finder_active_inbound_peers` | PeerfinderManager.cpp:214 | Active inbound peer connections |
| `peer_finder_active_outbound_peers` | PeerfinderManager.cpp:215 | Active outbound peer connections |
| `overlay_peer_disconnects` | OverlayImpl.h:557 | Peer disconnect count |
| `job_count` | JobQueue.cpp:26 | Current job queue depth |
| `{category}_bytes_in/out` | OverlayImpl.h:535 | Overlay traffic bytes per category (57 categories) |
| `{category}_messages_in/out` | OverlayImpl.h:535 | Overlay traffic messages per category |
| Prometheus Metric | Source | Description |
| ------------------------------------- | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ledgermaster_validated_ledger_age` | LedgerMaster.h:373 | Age of validated ledger (seconds) |
| `ledgermaster_published_ledger_age` | LedgerMaster.h:374 | Age of published ledger (seconds) |
| `state_accounting_{mode}_duration` | NetworkOPs.cpp:774 | Time in each operating mode (Disconnected/Connected/Syncing/Tracking/Full) |
| `state_accounting_{mode}_transitions` | NetworkOPs.cpp:780 | Transition count per mode |
| `peer_finder_active_inbound_peers` | PeerfinderManager.cpp:214 | Active inbound peer connections |
| `peer_finder_active_outbound_peers` | PeerfinderManager.cpp:215 | Active outbound peer connections |
| `overlay_peer_disconnects` | OverlayImpl.h:557 | Peer disconnect count |
| `jobq_job_count` | JobQueue.cpp:26 | Current job queue depth (emitted as `jobq_job_count`: the JobQueue collector is wrapped in `group("jobq")`, so the registered `job_count` gauge gains the `jobq_` prefix) |
| `{category}_bytes_in/out` | OverlayImpl.h:535 | Overlay traffic bytes per category (57 categories) |
| `{category}_messages_in/out` | OverlayImpl.h:535 | Overlay traffic messages per category |
#### OTel MetricsRegistry Gauges
@@ -960,7 +960,7 @@ Requires `trace_peer=1` in the `[telemetry]` config section.
| Operating Mode (Time Share) | timeseries | `rate(state_accounting_X_duration) / sum(rate(all modes))` | — |
| Operating Mode Transitions | timeseries | `state_accounting_*_transitions` | — |
| I/O Latency | timeseries | `histogram_quantile(0.95, ios_latency_bucket)` | — |
| Job Queue Depth | timeseries | `job_count` | — |
| Job Queue Depth | timeseries | `jobq_job_count` | — |
| Ledger Fetch Rate | stat | `rate(ledger_fetches[5m])` | — |
| Ledger History Mismatches | stat | `rate(ledger_history_mismatch[5m])` | — |
| Key Jobs Execution Time | timeseries | `acceptledger{quantile="$quantile"}` (+ 10 more key jobs) | `quantile` |
@@ -982,18 +982,28 @@ Requires `trace_peer=1` in the `[telemetry]` config section.
### Network Traffic -- System Metrics (`network-traffic`)
| Panel | Type | PromQL | Labels Used |
| ------------------------------------ | ---------- | ------------------------------------------------------ | ----------- |
| Active Peers | timeseries | `peer_finder_active_*_peers` | — |
| Peer Disconnects | timeseries | `increase(overlay_peer_disconnects[$__rate_interval])` | — |
| Total Network Bytes | timeseries | `rate(total_bytes_in/out[$__rate_interval])` | — |
| Total Network Messages | timeseries | `rate(total_messages_in/out[$__rate_interval])` | — |
| Transaction Traffic | timeseries | `rate(transactions_messages_in/out[$__rate_interval])` | — |
| Proposal Traffic | timeseries | `rate(proposals_messages_in/out[$__rate_interval])` | — |
| Validation Traffic | timeseries | `rate(validations_messages_in/out[$__rate_interval])` | — |
| Traffic by Category | bargauge | `topk(10, rate(*_bytes_in[$__rate_interval]))` | — |
| Duplicate Traffic (Wasted Bandwidth) | timeseries | `rate(*_duplicate_bytes_in/out[$__rate_interval])` | — |
| All Traffic Categories (Detail) | timeseries | `topk(15, rate(*_bytes_in[$__rate_interval]))` | — |
| Panel | Type | PromQL | Labels Used |
| ------------------------------------ | ---------- | -------------------------------------------------------------------------------------------------------------------------- | ----------- |
| Active Peers | timeseries | `peer_finder_active_*_peers` | — |
| Peer Disconnects | timeseries | `increase(overlay_peer_disconnects[$__rate_interval])` | — |
| Total Network Bytes | timeseries | `rate(total_bytes_in/out[$__rate_interval])` | — |
| Total Network Messages | timeseries | `rate(total_messages_in/out[$__rate_interval])` | — |
| Transaction Traffic | timeseries | `rate(transactions_messages_in/out[$__rate_interval])` | — |
| Proposal Traffic | timeseries | `rate(proposals_messages_in/out[$__rate_interval])` | — |
| Validation Traffic | timeseries | `rate(validations_messages_in/out[$__rate_interval])` | — |
| Traffic by Category | bargauge | `topk(10, label_replace(sum by (service_instance_id)(rate(<metric>[$__rate_interval])),"__name__","<metric>","","") or …)` | — |
| Duplicate Traffic (Wasted Bandwidth) | timeseries | `rate(*_duplicate_bytes_in/out[$__rate_interval])` | — |
| All Traffic Categories (Detail) | timeseries | `topk(15, label_replace(sum by (service_instance_id)(rate(<metric>[$__rate_interval])),"__name__","<metric>","","") or …)` | — |
> **Why the per-category panels enumerate each metric.** A bare
> `rate({__name__=~".*_bytes_in"}[…])` fails on Mimir/Cloud with _"vector
> cannot contain metrics with the same labelset"_: `rate()` drops the
> `__name__` label, so the many matched counters collapse to identical
> labelsets. Wrapping in `sum by (__name__, …)` does **not** help (the inner
> vector is rejected before the outer `sum`). The working form enumerates each
> `*_bytes_in` metric and re-attaches its name with `label_replace(...,
"__name__", "<metric>", "", "")`, so the existing `{{__name__}}` legend and
> the per-series display-name overrides keep working.
### RPC & Pathfinding -- System Metrics (`rpc-pathfinding`)
@@ -1244,7 +1254,7 @@ count_over_time({service_name="xrpld"} |= "trace_id=" [5m])
2. Verify `server=otel` in the `[insight]` config section
3. Verify the endpoint in `[insight]` points to the OTLP/HTTP port (default: `http://localhost:4318/v1/metrics`)
4. Check that the `otlp` receiver is in the metrics pipeline receivers in `otel-collector-config.yaml`
5. Query Prometheus directly: `curl 'http://localhost:9090/api/v1/query?query=job_count'`
5. Query Prometheus directly: `curl 'http://localhost:9090/api/v1/query?query=jobq_job_count'`
### Server info gauge shows server_state=0