mirror of
https://github.com/XRPLF/rippled.git
synced 2026-10-11 14:18:07 +00:00
docs(telemetry): Point panel Source links at files that exist
Some panels linked a Source file that has moved or never existed, or named a Function that does not exist. Only the Source link and Function text change.
This commit is contained in:
@@ -366,7 +366,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Spare trusted validator keys above the required quorum \u2014 the single number that says whether this node can ever validate.*\n\n###### How it's computed:\n*Trusted key count minus the required quorum, matched per node.*\n\n###### Reading it:\n*Positive is healthy. Zero or negative (red) means the trusted UNL is too small to ever satisfy quorum, so the node will stay short of a validated ledger.*\n\n###### Healthy range:\n*Positive; the exact figure depends on UNL size and the configured quorum.*\n\n###### Watch for:\n*Zero or below. Pair it with UNL Fetch Outcomes (Count By Site & Outcome): a site stuck on fetch_error or expired is the usual cause of a UNL too small to meet quorum.*\n\n###### Keywords:\n- **UNL quorum headroom** *(per node)* \u2014 trusted UNL key count minus the required quorum; at or below zero the node can never declare a ledger validated.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerUnlQuorumGauge`\n\n###### References:\n[Validation quorum on xrpl.org](https://xrpl.org/docs/concepts/consensus-protocol/negative-unl) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#unl-quorum-headroom)",
|
||||
"description": "###### What this is:\n*Spare trusted validator keys above the required quorum \u2014 the single number that says whether this node can ever validate.*\n\n###### How it's computed:\n*Trusted key count minus the required quorum, matched per node.*\n\n###### Reading it:\n*Positive is healthy. Zero or negative (red) means the trusted UNL is too small to ever satisfy quorum, so the node will stay short of a validated ledger.*\n\n###### Healthy range:\n*Positive; the exact figure depends on UNL size and the configured quorum.*\n\n###### Watch for:\n*Zero or below. Pair it with UNL Fetch Outcomes (Count By Site & Outcome): a site stuck on fetch_error or expired is the usual cause of a UNL too small to meet quorum.*\n\n###### Keywords:\n- **UNL quorum headroom** *(per node)* \u2014 trusted UNL key count minus the required quorum; at or below zero the node can never declare a ledger validated.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerUnlQuorumGauge`\n\n###### References:\n[Validation quorum on xrpl.org](https://xrpl.org/docs/concepts/consensus-protocol/negative-unl) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#unl-quorum-headroom)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -643,7 +643,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Trusted validator keys currently in effect, plotted against the number of agreeing validations a ledger needs.*\n\n###### How it's computed:\n*Two series read from the same gauge: trusted_keys (usable UNL size) and quorum (validations required to declare a ledger validated).*\n\n###### Reading it:\n*Trusted Keys must sit above Quorum. Where the lines cross, or where Trusted Keys is zero, the node cannot reach a quorum and will never validate a ledger no matter how healthy the rest of the pipeline looks.*\n\n###### Healthy range:\n*Trusted Keys comfortably above Quorum and both flat.*\n\n###### Watch for:\n*Trusted Keys at zero (no usable UNL loaded) or below Quorum. Steps in Quorum track validator-list changes; steps down in Trusted Keys mean keys were dropped.*\n\n###### Keywords:\n- **UNL quorum headroom** *(per node)* \u2014 trusted UNL key count minus the required quorum; at or below zero the node can never declare a ledger validated.\n- **UNL (Unique Node List)** *(per node)* \u2014 the list of validators a node trusts not to collude; the basis for its consensus and quorum.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerUnlQuorumGauge`\n\n###### References:\n[UNL (Unique Node List)](https://xrpl.org/docs/concepts/consensus-protocol/unl) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#unl-quorum-headroom)",
|
||||
"description": "###### What this is:\n*Trusted validator keys currently in effect, plotted against the number of agreeing validations a ledger needs.*\n\n###### How it's computed:\n*Two series read from the same gauge: trusted_keys (usable UNL size) and quorum (validations required to declare a ledger validated).*\n\n###### Reading it:\n*Trusted Keys must sit above Quorum. Where the lines cross, or where Trusted Keys is zero, the node cannot reach a quorum and will never validate a ledger no matter how healthy the rest of the pipeline looks.*\n\n###### Healthy range:\n*Trusted Keys comfortably above Quorum and both flat.*\n\n###### Watch for:\n*Trusted Keys at zero (no usable UNL loaded) or below Quorum. Steps in Quorum track validator-list changes; steps down in Trusted Keys mean keys were dropped.*\n\n###### Keywords:\n- **UNL quorum headroom** *(per node)* \u2014 trusted UNL key count minus the required quorum; at or below zero the node can never declare a ledger validated.\n- **UNL (Unique Node List)** *(per node)* \u2014 the list of validators a node trusts not to collude; the basis for its consensus and quorum.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerUnlQuorumGauge`\n\n###### References:\n[UNL (Unique Node List)](https://xrpl.org/docs/concepts/consensus-protocol/unl) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#unl-quorum-headroom)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -754,7 +754,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*How far the network's agreed close time sits from this node's own clock.*\n\n###### How it's computed:\n*Signed offset in seconds, plus its magnitude so a threshold band applies in either direction.*\n\n###### Reading it:\n*Both series should hug zero. A negative signed value means the local clock runs ahead of the network, positive means it lags. The magnitude is what matters: the threshold lines sit at 1s (suspicious) and 60s.*\n\n###### Healthy range:\n*Magnitude under 1 second.*\n\n###### Watch for:\n*A magnitude above 1 second that does not decay, which delays consensus participation. Note that server_info only surfaces close_time_offset once the magnitude reaches 60 seconds, so this panel sees skew long before the API does; a persistent offset is a local NTP fault, not a network one.*\n\n###### Keywords:\n- **Clock close offset** *(per node)* \u2014 the difference between the network's agreed close time and this node's clock; a persistent offset delays consensus participation.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerClockSkewGauge`\n\n###### References:\n[Ledger close times on xrpl.org](https://xrpl.org/docs/concepts/ledgers/ledger-close-times) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#clock-close-offset)",
|
||||
"description": "###### What this is:\n*How far the network's agreed close time sits from this node's own clock.*\n\n###### How it's computed:\n*Signed offset in seconds, plus its magnitude so a threshold band applies in either direction.*\n\n###### Reading it:\n*Both series should hug zero. A negative signed value means the local clock runs ahead of the network, positive means it lags. The magnitude is what matters: the threshold lines sit at 1s (suspicious) and 60s.*\n\n###### Healthy range:\n*Magnitude under 1 second.*\n\n###### Watch for:\n*A magnitude above 1 second that does not decay, which delays consensus participation. Note that server_info only surfaces close_time_offset once the magnitude reaches 60 seconds, so this panel sees skew long before the API does; a persistent offset is a local NTP fault, not a network one.*\n\n###### Keywords:\n- **Clock close offset** *(per node)* \u2014 the difference between the network's agreed close time and this node's clock; a persistent offset delays consensus participation.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerClockSkewGauge`\n\n###### References:\n[Ledger close times on xrpl.org](https://xrpl.org/docs/concepts/ledgers/ledger-close-times) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#clock-close-offset)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -1127,7 +1127,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Two distances that say whether the sequence this node needs is inside the window its peers can serve: how much history is still reachable behind it, and how far the network tip is ahead of it.*\n\n###### How it's computed:\n*History Headroom = this node's validated_ledger_seq minus peer_ledger_supply supply_min_seq. Tip Gap = supply_max_seq minus validated_ledger_seq. Both operands are gated on being above zero, so neither line is drawn until the node has a validated ledger and at least one peer has advertised a range.*\n\n###### Reading it:\n*Plotted as distances rather than absolute sequences on purpose. The raw sequences can sit far apart, so on one linear axis the tip movement that shows whether sync is progressing is a tiny fraction of the scale and reads as a flat line. Differencing puts the meaning on the axis: zero is the boundary in both cases. History Headroom uses the left axis, Tip Gap the right, because the two differ by orders of magnitude.*\n\n###### Healthy range:\n*History Headroom comfortably positive and roughly steady. Tip Gap at or near zero.*\n\n###### Watch for:\n*History Headroom crossing below zero \u2014 no connected peer offers this node's validated ledger or anything older. That does not stop the sync, since the node catches up by fetching the newest ledger by hash, but any ledgers between this node's and the lowest one offered can stay a hole in its history that only a peer holding them can fill. Tip Gap growing steadily means the network is closing ledgers faster than this node validates them. Pair with Peers Able to Serve Needed Sequence: that panel counts how many peers can serve the next ledger, this one says how far outside the served window the node has drifted.*\n\n###### Keywords:\n- **History headroom** *(per node)* \u2014 validated sequence minus the lowest sequence any connected peer offers; how much of the history behind this node is still reachable.\n- **Tip gap** *(per node)* \u2014 the highest sequence any connected peer offers minus this node's validated sequence; how far behind the peer-reported tip this node is.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query subtracts the two series and aggregates them.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerPeerLedgerSupplyGauge`\n\n###### References:\n[Complete ledger ranges](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#complete-ledger-ranges) \u00b7 [Back-fill / catch-up](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#back-fill-catch-up)",
|
||||
"description": "###### What this is:\n*Two distances that say whether the sequence this node needs is inside the window its peers can serve: how much history is still reachable behind it, and how far the network tip is ahead of it.*\n\n###### How it's computed:\n*History Headroom = this node's validated_ledger_seq minus peer_ledger_supply supply_min_seq. Tip Gap = supply_max_seq minus validated_ledger_seq. Both operands are gated on being above zero, so neither line is drawn until the node has a validated ledger and at least one peer has advertised a range.*\n\n###### Reading it:\n*Plotted as distances rather than absolute sequences on purpose. The raw sequences can sit far apart, so on one linear axis the tip movement that shows whether sync is progressing is a tiny fraction of the scale and reads as a flat line. Differencing puts the meaning on the axis: zero is the boundary in both cases. History Headroom uses the left axis, Tip Gap the right, because the two differ by orders of magnitude.*\n\n###### Healthy range:\n*History Headroom comfortably positive and roughly steady. Tip Gap at or near zero.*\n\n###### Watch for:\n*History Headroom crossing below zero \u2014 no connected peer offers this node's validated ledger or anything older. That does not stop the sync, since the node catches up by fetching the newest ledger by hash, but any ledgers between this node's and the lowest one offered can stay a hole in its history that only a peer holding them can fill. Tip Gap growing steadily means the network is closing ledgers faster than this node validates them. Pair with Peers Able to Serve Needed Sequence: that panel counts how many peers can serve the next ledger, this one says how far outside the served window the node has drifted.*\n\n###### Keywords:\n- **History headroom** *(per node)* \u2014 validated sequence minus the lowest sequence any connected peer offers; how much of the history behind this node is still reachable.\n- **Tip gap** *(per node)* \u2014 the highest sequence any connected peer offers minus this node's validated sequence; how far behind the peer-reported tip this node is.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query subtracts the two series and aggregates them.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerPeerLedgerSupplyGauge`\n\n###### References:\n[Complete ledger ranges](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#complete-ledger-ranges) \u00b7 [Back-fill / catch-up](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#back-fill-catch-up)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -1265,7 +1265,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*PeerFinder's slot accounting in one snapshot: outbound and inbound slots in use against their configured maxima, plus outbound attempts still in flight.*\n\n###### How it's computed:\n*peerfinder_slot_census series out_active, out_max, in_active, in_max and connecting, all read from ONE PeerFinder snapshot taken under a single lock so the five values describe the same instant.*\n\n###### Reading it:\n*out_active should climb to out_max and stay there. connecting counts outbound attempts started but not yet resolved either way, so it is the in-flight term the active counts cannot show.*\n\n###### Healthy range:\n*out_active at out_max; connecting low and transient.*\n\n###### Watch for:\n*out_active pinned below out_max while connecting stays non-zero \u2014 dials are being started and never completing, so the node is trying and failing rather than sitting idle. Note out_active and in_active are also exported separately as the legacy beast::insight gauges peer_finder_active_outbound_peers and peer_finder_active_inbound_peers. Those two are unrelated single series read at different instants with no capacity, attempt or cache terms, so no reading of them can distinguish this case. The census exists to report all nine fields from one snapshot under one lock, so they are mutually consistent and joinable on one labelset.*\n\n###### Keywords:\n- **PeerFinder slot census** *(per node)* \u2014 one consistent snapshot of PeerFinder's outbound and inbound slot use, capacities, in-flight attempts, fixed peers and address caches.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerSlotCensusGauge`\n\n###### References:\n[Overlay](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#overlay) \u00b7 [Outbound dial latency](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#outbound-dial-latency)",
|
||||
"description": "###### What this is:\n*PeerFinder's slot accounting in one snapshot: outbound and inbound slots in use against their configured maxima, plus outbound attempts still in flight.*\n\n###### How it's computed:\n*peerfinder_slot_census series out_active, out_max, in_active, in_max and connecting, all read from ONE PeerFinder snapshot taken under a single lock so the five values describe the same instant.*\n\n###### Reading it:\n*out_active should climb to out_max and stay there. connecting counts outbound attempts started but not yet resolved either way, so it is the in-flight term the active counts cannot show.*\n\n###### Healthy range:\n*out_active at out_max; connecting low and transient.*\n\n###### Watch for:\n*out_active pinned below out_max while connecting stays non-zero \u2014 dials are being started and never completing, so the node is trying and failing rather than sitting idle. Note out_active and in_active are also exported separately as the legacy beast::insight gauges peer_finder_active_outbound_peers and peer_finder_active_inbound_peers. Those two are unrelated single series read at different instants with no capacity, attempt or cache terms, so no reading of them can distinguish this case. The census exists to report all nine fields from one snapshot under one lock, so they are mutually consistent and joinable on one labelset.*\n\n###### Keywords:\n- **PeerFinder slot census** *(per node)* \u2014 one consistent snapshot of PeerFinder's outbound and inbound slot use, capacities, in-flight attempts, fixed peers and address caches.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerSlotCensusGauge`\n\n###### References:\n[Overlay](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#overlay) \u00b7 [Outbound dial latency](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#outbound-dial-latency)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -1364,7 +1364,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*The address supply PeerFinder dials from \u2014 its boot cache and live cache \u2014 next to configured versus currently connected fixed peers.*\n\n###### How it's computed:\n*peerfinder_slot_census series bootcache, livecache, fixed_configured and fixed_active, read from the same single snapshot as the slot census.*\n\n###### Reading it:\n*bootcache holds seed addresses persisted across restarts; livecache holds addresses learned from peers while running. fixed_active is how many of the fixed peers named in the config are connected right now.*\n\n###### Healthy range:\n*bootcache and livecache non-zero; fixed_active equal to fixed_configured.*\n\n###### Watch for:\n*bootcache at 0 on a fresh node means there are no seed addresses to dial at all, so no outbound connection is ever attempted and every downstream sync signal on this dashboard stays empty for a reason that has nothing to do with sync. fixed_active below fixed_configured means a configured fixed peer is unreachable.*\n\n###### Keywords:\n- **PeerFinder address cache** *(per node)* \u2014 the boot cache (seed addresses persisted across restarts) and live cache (addresses learned from peers) that supply outbound dial candidates.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerSlotCensusGauge`\n\n###### References:\n[Overlay](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#overlay) \u00b7 [DNS resolve](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#dns-resolve)",
|
||||
"description": "###### What this is:\n*The address supply PeerFinder dials from \u2014 its boot cache and live cache \u2014 next to configured versus currently connected fixed peers.*\n\n###### How it's computed:\n*peerfinder_slot_census series bootcache, livecache, fixed_configured and fixed_active, read from the same single snapshot as the slot census.*\n\n###### Reading it:\n*bootcache holds seed addresses persisted across restarts; livecache holds addresses learned from peers while running. fixed_active is how many of the fixed peers named in the config are connected right now.*\n\n###### Healthy range:\n*bootcache and livecache non-zero; fixed_active equal to fixed_configured.*\n\n###### Watch for:\n*bootcache at 0 on a fresh node means there are no seed addresses to dial at all, so no outbound connection is ever attempted and every downstream sync signal on this dashboard stays empty for a reason that has nothing to do with sync. fixed_active below fixed_configured means a configured fixed peer is unreachable.*\n\n###### Keywords:\n- **PeerFinder address cache** *(per node)* \u2014 the boot cache (seed addresses persisted across restarts) and live cache (addresses learned from peers) that supply outbound dial candidates.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerSlotCensusGauge`\n\n###### References:\n[Overlay](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#overlay) \u00b7 [DNS resolve](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#dns-resolve)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -1476,7 +1476,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*How long the node took, from process start, to reach the full server state for the first time.*\n\n###### How it's computed:\n*sync_state series initial_full_duration_us, converted from microseconds to seconds.*\n\n###### Reading it:\n*A value appears only once the node has actually synced. Zero (red) means it has never reached full \u2014 that is the signal, not missing data.*\n\n###### Healthy range:\n*Seconds to a few minutes on a warm node; longer on a fresh one that must acquire history.*\n\n###### Watch for:\n*A flat zero. The value never changes after the first full transition, so it either fills in or the node never synced.*\n\n###### Keywords:\n- **Time to first FULL** *(per node)* \u2014 elapsed time from process start until the node first reached the full server state.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerSyncStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#time-to-first-full)",
|
||||
"description": "###### What this is:\n*How long the node took, from process start, to reach the full server state for the first time.*\n\n###### How it's computed:\n*sync_state series initial_full_duration_us, converted from microseconds to seconds.*\n\n###### Reading it:\n*A value appears only once the node has actually synced. Zero (red) means it has never reached full \u2014 that is the signal, not missing data.*\n\n###### Healthy range:\n*Seconds to a few minutes on a warm node; longer on a fresh one that must acquire history.*\n\n###### Watch for:\n*A flat zero. The value never changes after the first full transition, so it either fills in or the node never synced.*\n\n###### Keywords:\n- **Time to first FULL** *(per node)* \u2014 elapsed time from process start until the node first reached the full server state.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerSyncStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#time-to-first-full)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -1547,7 +1547,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Whether the node is still waiting to see a full network ledger before it will participate.*\n\n###### How it's computed:\n*sync_state series network_ledger_gate: 1 while the gate is closed, 0 once it opens.*\n\n###### Reading it:\n*0 (green) is healthy. A persistent 1 (red) means the node has never seen a complete network ledger, so it refuses transactions and can never reach full no matter how healthy the rest of the pipeline looks.*\n\n###### Healthy range:\n*0 within the first few minutes of startup.*\n\n###### Watch for:\n*A 1 that never clears. Pair it with the Bootstrap row \u2014 no peers or no quorum is the usual cause.*\n\n###### Keywords:\n- **Network ledger gate** *(per node)* \u2014 the startup guard that holds a node back until it has seen a full ledger from the network.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerSyncStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#network-ledger-gate)",
|
||||
"description": "###### What this is:\n*Whether the node is still waiting to see a full network ledger before it will participate.*\n\n###### How it's computed:\n*sync_state series network_ledger_gate: 1 while the gate is closed, 0 once it opens.*\n\n###### Reading it:\n*0 (green) is healthy. A persistent 1 (red) means the node has never seen a complete network ledger, so it refuses transactions and can never reach full no matter how healthy the rest of the pipeline looks.*\n\n###### Healthy range:\n*0 within the first few minutes of startup.*\n\n###### Watch for:\n*A 1 that never clears. Pair it with the Bootstrap row \u2014 no peers or no quorum is the usual cause.*\n\n###### Keywords:\n- **Network ledger gate** *(per node)* \u2014 the startup guard that holds a node back until it has seen a full ledger from the network.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerSyncStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#network-ledger-gate)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -1815,7 +1815,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*How far this node's validated ledger trails the highest ledger any connected peer reports holding.*\n\n###### How it's computed:\n*sync_state series ledgers_behind: the peer-reported network tip minus this node's validated sequence, floored at zero.*\n\n###### Reading it:\n*Trending to zero is healthy convergence. Flat or rising means the node is not catching up. Zero also covers \"no peer has reported a newer ledger\", which on a node with no peers is the same thing.*\n\n###### Healthy range:\n*0 to 1 on a synced node.*\n\n###### Watch for:\n*A plateau or a climb during initial sync: the node is acquiring slower than the network advances, so it will never converge. Correlate with the acquire and job-queue panels.*\n\n###### Keywords:\n- **Ledgers behind network** *(per node)* \u2014 the gap between the peer-reported network tip and this node's own validated ledger sequence.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerSyncStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#ledgers-behind-network)",
|
||||
"description": "###### What this is:\n*How far this node's validated ledger trails the highest ledger any connected peer reports holding.*\n\n###### How it's computed:\n*sync_state series ledgers_behind: the peer-reported network tip minus this node's validated sequence, floored at zero.*\n\n###### Reading it:\n*Trending to zero is healthy convergence. Flat or rising means the node is not catching up. Zero also covers \"no peer has reported a newer ledger\", which on a node with no peers is the same thing.*\n\n###### Healthy range:\n*0 to 1 on a synced node.*\n\n###### Watch for:\n*A plateau or a climb during initial sync: the node is acquiring slower than the network advances, so it will never converge. Correlate with the acquire and job-queue panels.*\n\n###### Keywords:\n- **Ledgers behind network** *(per node)* \u2014 the gap between the peer-reported network tip and this node's own validated ledger sequence.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerSyncStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#ledgers-behind-network)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2081,7 +2081,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Outstanding SHAMap nodes the busiest in-flight ledger acquire still needs, split by tree. This is the signal that separates a sync that is merely slow from one that will never finish.*\n\n###### How it's computed:\n*sync_acquire series missing_state_nodes_max and missing_tx_nodes_max: the largest outstanding node count across all in-flight acquires, refreshed after each getMissingNodes sweep. The maximum, not the sum, so one stuck acquire stays visible instead of being averaged away.*\n\n###### Reading it:\n*Falling toward zero means the acquire is progressing. A value pinned at 256 is the sweep cap, meaning there are at least that many nodes outstanding. Zero on one tree with a value on the other means that tree is already complete.*\n\n###### Healthy range:\n*Falling to 0 within seconds per ledger.*\n\n###### Watch for:\n*A flat, non-zero value across several minutes: no peer is serving that tree, so this acquire will never complete. Pair with Acquire Stalls \u2014 No Progress (Count) \u2014 both flat and climbing together is the definitive stuck-sync signature.*\n\n###### Keywords:\n- **Missing SHAMap node** *(per node)* \u2014 a tree node this node needs to complete a ledger but does not yet hold.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerSyncAcquireGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#missing-shamap-node)",
|
||||
"description": "###### What this is:\n*Outstanding SHAMap nodes the busiest in-flight ledger acquire still needs, split by tree. This is the signal that separates a sync that is merely slow from one that will never finish.*\n\n###### How it's computed:\n*sync_acquire series missing_state_nodes_max and missing_tx_nodes_max: the largest outstanding node count across all in-flight acquires, refreshed after each getMissingNodes sweep. The maximum, not the sum, so one stuck acquire stays visible instead of being averaged away.*\n\n###### Reading it:\n*Falling toward zero means the acquire is progressing. A value pinned at 256 is the sweep cap, meaning there are at least that many nodes outstanding. Zero on one tree with a value on the other means that tree is already complete.*\n\n###### Healthy range:\n*Falling to 0 within seconds per ledger.*\n\n###### Watch for:\n*A flat, non-zero value across several minutes: no peer is serving that tree, so this acquire will never complete. Pair with Acquire Stalls \u2014 No Progress (Count) \u2014 both flat and climbing together is the definitive stuck-sync signature.*\n\n###### Keywords:\n- **Missing SHAMap node** *(per node)* \u2014 a tree node this node needs to complete a ledger but does not yet hold.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerSyncAcquireGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#missing-shamap-node)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2288,7 +2288,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Peer packets stashed across all in-flight acquires waiting to be applied, alongside how many acquires are running.*\n\n###### How it's computed:\n*sync_acquire series received_data_depth (summed across acquires) and in_flight (the acquire count). The depth mirrors the receive stash size; in_flight gives the context that makes an all-zero reading legible.*\n\n###### Reading it:\n*A depth near zero means node data is applied as fast as it arrives. in_flight at zero means the node is idle, which is why an all-zero Missing Nodes panel is not by itself a healthy reading.*\n\n###### Healthy range:\n*Depth 0 to a few; in_flight low single digits during sync.*\n\n###### Watch for:\n*A growing depth means arriving data outpaces processing, which is a job-queue or disk problem rather than a peer-supply one. Check the job-queue backlog next.*\n\n###### Keywords:\n- **Received-data stash** *(per node)* \u2014 peer packets held for later processing because the acquire cannot apply them as fast as they arrive.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerSyncAcquireGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#received-data-stash)",
|
||||
"description": "###### What this is:\n*Peer packets stashed across all in-flight acquires waiting to be applied, alongside how many acquires are running.*\n\n###### How it's computed:\n*sync_acquire series received_data_depth (summed across acquires) and in_flight (the acquire count). The depth mirrors the receive stash size; in_flight gives the context that makes an all-zero reading legible.*\n\n###### Reading it:\n*A depth near zero means node data is applied as fast as it arrives. in_flight at zero means the node is idle, which is why an all-zero Missing Nodes panel is not by itself a healthy reading.*\n\n###### Healthy range:\n*Depth 0 to a few; in_flight low single digits during sync.*\n\n###### Watch for:\n*A growing depth means arriving data outpaces processing, which is a job-queue or disk problem rather than a peer-supply one. Check the job-queue backlog next.*\n\n###### Keywords:\n- **Received-data stash** *(per node)* \u2014 peer packets held for later processing because the acquire cannot apply them as fast as they arrive.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerSyncAcquireGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#received-data-stash)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2419,7 +2419,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Fraction of SHAMap tree-node lookups served from memory instead of the node store. The layer above the NuDB cache: a miss here is what causes a node-store read.*\n\n###### How it's computed:\n*shamap_cache_hit_rate series treenode, from TaggedCache::getHitRate() on the node family's tree-node cache, normalized from 0-100 to 0.0-1.0.*\n\n###### Reading it:\n*Near 1.0 on a warm node. Low during a fresh sync while the cache fills. Distinct from the NuDB Cache Hit Ratio panel on the Ledger Data Sync dashboard, which measures the node-store layer beneath this one.*\n\n###### Healthy range:\n*> 0.9 on a warm node.*\n\n###### Watch for:\n*A persistently low rate on a node that should be warm: the working set does not fit the cache, or continuous re-acquisition is churning it, so every tree walk pays disk latency. Read with Acquire Source.*\n\n###### Keywords:\n- **SHAMap cache hit rate** *(per node)* \u2014 the share of SHAMap tree-node lookups answered from the in-memory cache rather than the node store.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerCacheHitRateDetailGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#shamap-cache-hit-rate)",
|
||||
"description": "###### What this is:\n*Fraction of SHAMap tree-node lookups served from memory instead of the node store. The layer above the NuDB cache: a miss here is what causes a node-store read.*\n\n###### How it's computed:\n*shamap_cache_hit_rate series treenode, from TaggedCache::getHitRate() on the node family's tree-node cache, normalized from 0-100 to 0.0-1.0.*\n\n###### Reading it:\n*Near 1.0 on a warm node. Low during a fresh sync while the cache fills. Distinct from the NuDB Cache Hit Ratio panel on the Ledger Data Sync dashboard, which measures the node-store layer beneath this one.*\n\n###### Healthy range:\n*> 0.9 on a warm node.*\n\n###### Watch for:\n*A persistently low rate on a node that should be warm: the working set does not fit the cache, or continuous re-acquisition is churning it, so every tree walk pays disk latency. Read with Acquire Source.*\n\n###### Keywords:\n- **SHAMap cache hit rate** *(per node)* \u2014 the share of SHAMap tree-node lookups answered from the in-memory cache rather than the node store.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerCacheHitRateDetailGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#shamap-cache-hit-rate)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2539,7 +2539,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Share of the worker-thread pool currently executing a job. This is the pool-wide view: when the pool itself is exhausted, every subsystem queued behind it looks independently slow, and this panel attributes the whole slowdown once.*\n\n###### How it's computed:\n*jobq_saturation running_tasks divided by worker_threads (the denominator is clamped to at least 1). The thread count is exported rather than hardcoded because it is derived at startup from [workers], node size and hardware concurrency.*\n\n###### Reading it:\n*Below 80% (green) means the pool has spare capacity, so a slow stage is that stage's own fault. At 100% every worker is busy \u2014 read Worker Pool Capacity & Total Backlog next: 100% with a queue is an exhausted pool, 100% with an empty queue is merely busy.*\n\n###### Healthy range:\n*< 80%.*\n\n###### Watch for:\n*A sustained 100% together with a non-zero backlog. Every job type is then starved by the pool, so fix pool capacity or the long-running jobs holding it, not the individual victim subsystems.*\n\n###### Keywords:\n- **Worker-pool saturation** *(per node)* \u2014 worker threads executing a job as a share of the threads the pool is configured to run.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerJobQueueSaturationGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#worker-pool-saturation)",
|
||||
"description": "###### What this is:\n*Share of the worker-thread pool currently executing a job. This is the pool-wide view: when the pool itself is exhausted, every subsystem queued behind it looks independently slow, and this panel attributes the whole slowdown once.*\n\n###### How it's computed:\n*jobq_saturation running_tasks divided by worker_threads (the denominator is clamped to at least 1). The thread count is exported rather than hardcoded because it is derived at startup from [workers], node size and hardware concurrency.*\n\n###### Reading it:\n*Below 80% (green) means the pool has spare capacity, so a slow stage is that stage's own fault. At 100% every worker is busy \u2014 read Worker Pool Capacity & Total Backlog next: 100% with a queue is an exhausted pool, 100% with an empty queue is merely busy.*\n\n###### Healthy range:\n*< 80%.*\n\n###### Watch for:\n*A sustained 100% together with a non-zero backlog. Every job type is then starved by the pool, so fix pool capacity or the long-running jobs holding it, not the individual victim subsystems.*\n\n###### Keywords:\n- **Worker-pool saturation** *(per node)* \u2014 worker threads executing a job as a share of the threads the pool is configured to run.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerJobQueueSaturationGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#worker-pool-saturation)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2610,7 +2610,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*The three raw worker-pool numbers behind the saturation ratio: tasks in flight, threads configured, and total jobs queued across every job type.*\n\n###### How it's computed:\n*jobq_saturation series running_tasks, worker_threads and total_waiting, all read from one JobQueue sample so the ratio and the backlog describe the same instant.*\n\n###### Reading it:\n*running_tasks tracking worker_threads means the pool is fully committed. total_waiting is what makes that legible: queued work behind a fully committed pool is exhaustion, no queued work is just a busy moment.*\n\n###### Healthy range:\n*running_tasks below worker_threads; total_waiting near 0.*\n\n###### Watch for:\n*total_waiting climbing while running_tasks is pinned at worker_threads. Distinct from the jobq_job_count depth panel on the Ledger Data Sync dashboard, which has no capacity term at all, so a depth reading there cannot say whether the pool is the cause.*\n\n###### Keywords:\n- **Worker thread pool** *(per node)* \u2014 the fixed set of threads that execute all job-queue work, sized at startup.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerJobQueueSaturationGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#worker-pool-saturation)",
|
||||
"description": "###### What this is:\n*The three raw worker-pool numbers behind the saturation ratio: tasks in flight, threads configured, and total jobs queued across every job type.*\n\n###### How it's computed:\n*jobq_saturation series running_tasks, worker_threads and total_waiting, all read from one JobQueue sample so the ratio and the backlog describe the same instant.*\n\n###### Reading it:\n*running_tasks tracking worker_threads means the pool is fully committed. total_waiting is what makes that legible: queued work behind a fully committed pool is exhaustion, no queued work is just a busy moment.*\n\n###### Healthy range:\n*running_tasks below worker_threads; total_waiting near 0.*\n\n###### Watch for:\n*total_waiting climbing while running_tasks is pinned at worker_threads. Distinct from the jobq_job_count depth panel on the Ledger Data Sync dashboard, which has no capacity term at all, so a depth reading there cannot say whether the pool is the cause.*\n\n###### Keywords:\n- **Worker thread pool** *(per node)* \u2014 the fixed set of threads that execute all job-queue work, sized at startup.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerJobQueueSaturationGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#worker-pool-saturation)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2795,7 +2795,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*How long the node took, from process start, to pass the pre-accept quorum gate for the first time.*\n\n###### How it's computed:\n*ledger_quorum_publish series time_to_first_validated_us, converted from microseconds to seconds.*\n\n###### Reading it:\n*A one-shot measurement: it fills in the moment the node first fully validates a ledger and never changes again, so it has no trend to read. Exactly two readings matter \u2014 a duration, meaning the node got there and this is how long it took, or zero (red), meaning it never has.*\n\n###### Healthy range:\n*Seconds to a few minutes; longer on a fresh node that must acquire history first.*\n\n###### Watch for:\n*A flat zero while Time to First FULL shows a value: the node reached the full server state but has still never fully validated a ledger, which points at the quorum gate rather than at acquire. The measurement is clamped to a minimum of 1 microsecond so a genuine reading can never be confused with the never-reached zero.*\n\n###### Keywords:\n- **Time to first validated ledger** *(per node)* \u2014 elapsed time from process start until the node first declared a ledger fully validated.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerLedgerQuorumPublishGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#time-to-first-validated-ledger)",
|
||||
"description": "###### What this is:\n*How long the node took, from process start, to pass the pre-accept quorum gate for the first time.*\n\n###### How it's computed:\n*ledger_quorum_publish series time_to_first_validated_us, converted from microseconds to seconds.*\n\n###### Reading it:\n*A one-shot measurement: it fills in the moment the node first fully validates a ledger and never changes again, so it has no trend to read. Exactly two readings matter \u2014 a duration, meaning the node got there and this is how long it took, or zero (red), meaning it never has.*\n\n###### Healthy range:\n*Seconds to a few minutes; longer on a fresh node that must acquire history first.*\n\n###### Watch for:\n*A flat zero while Time to First FULL shows a value: the node reached the full server state but has still never fully validated a ledger, which points at the quorum gate rather than at acquire. The measurement is clamped to a minimum of 1 microsecond so a genuine reading can never be confused with the never-reached zero.*\n\n###### Keywords:\n- **Time to first validated ledger** *(per node)* \u2014 elapsed time from process start until the node first declared a ledger fully validated.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerLedgerQuorumPublishGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#time-to-first-validated-ledger)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2866,7 +2866,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Trusted validations counted at the most recent pre-accept gate, plotted against the number that gate required.*\n\n###### How it's computed:\n*Two series from the same gauge: trusted_validation_tally (agreeing trusted validations seen for the candidate ledger, after the negative-UNL filter) and quorum_target (what the gate demanded). Both are snapshotted on every gate evaluation, whether it passed or failed, so a node that keeps failing still reports both numbers.*\n\n###### Reading it:\n*Read the shape of the tally, not any single value. A tally climbing toward the target is a slow sync that will finish, so keep waiting. A tally flat below the target is stuck: it will never reach quorum on its own, and nothing in the acquire pipeline can fix it. Expect a sawtooth on a healthy node: each series is a snapshot of the most recent gate evaluation, and the first evaluation of every round runs before peer validations arrive, so a sampled low reading between higher ones is normal. Judge it over minutes, and only the sustained floor of the tally against the target carries the signal.*\n\n###### Healthy range:\n*Tally at or above Target, both flat, on a node that is validating.*\n\n###### Watch for:\n*A tally pinned below the target \u2014 too few trusted validators are reachable, or the UNL / negative-UNL configuration excludes the ones that are. Also watch the target jumping to about 9.2e18 (signed 64-bit maximum): that is the explicit quorum-disabled sentinel, meaning too many publishers are unavailable and the trusted list switched quorum off entirely, so the node can never validate however far the tally climbs. It is reported as that maximum rather than wrapping negative precisely so it cannot be misread as a tally that already exceeds its target. Both series flat at 0 means the gate has never been evaluated \u2014 nothing has been offered for validation yet, which sends you back to the Bootstrap row.*\n\n###### Keywords:\n- **Quorum shortfall** *(per node)* \u2014 trusted validations for a candidate ledger falling short of the quorum needed to declare it validated, so the node holds the ledger and still cannot call it validated.\n- **Validation quorum** *(per node)* \u2014 the number of agreeing trusted validations a ledger needs before this node treats it as validated.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerLedgerQuorumPublishGauge`\n\n###### References:\n[Negative UNL and validation quorum on xrpl.org](https://xrpl.org/docs/concepts/consensus-protocol/negative-unl) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#quorum-shortfall)",
|
||||
"description": "###### What this is:\n*Trusted validations counted at the most recent pre-accept gate, plotted against the number that gate required.*\n\n###### How it's computed:\n*Two series from the same gauge: trusted_validation_tally (agreeing trusted validations seen for the candidate ledger, after the negative-UNL filter) and quorum_target (what the gate demanded). Both are snapshotted on every gate evaluation, whether it passed or failed, so a node that keeps failing still reports both numbers.*\n\n###### Reading it:\n*Read the shape of the tally, not any single value. A tally climbing toward the target is a slow sync that will finish, so keep waiting. A tally flat below the target is stuck: it will never reach quorum on its own, and nothing in the acquire pipeline can fix it. Expect a sawtooth on a healthy node: each series is a snapshot of the most recent gate evaluation, and the first evaluation of every round runs before peer validations arrive, so a sampled low reading between higher ones is normal. Judge it over minutes, and only the sustained floor of the tally against the target carries the signal.*\n\n###### Healthy range:\n*Tally at or above Target, both flat, on a node that is validating.*\n\n###### Watch for:\n*A tally pinned below the target \u2014 too few trusted validators are reachable, or the UNL / negative-UNL configuration excludes the ones that are. Also watch the target jumping to about 9.2e18 (signed 64-bit maximum): that is the explicit quorum-disabled sentinel, meaning too many publishers are unavailable and the trusted list switched quorum off entirely, so the node can never validate however far the tally climbs. It is reported as that maximum rather than wrapping negative precisely so it cannot be misread as a tally that already exceeds its target. Both series flat at 0 means the gate has never been evaluated \u2014 nothing has been offered for validation yet, which sends you back to the Bootstrap row.*\n\n###### Keywords:\n- **Quorum shortfall** *(per node)* \u2014 trusted validations for a candidate ledger falling short of the quorum needed to declare it validated, so the node holds the ledger and still cannot call it validated.\n- **Validation quorum** *(per node)* \u2014 the number of agreeing trusted validations a ledger needs before this node treats it as validated.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerLedgerQuorumPublishGauge`\n\n###### References:\n[Negative UNL and validation quorum on xrpl.org](https://xrpl.org/docs/concepts/consensus-protocol/negative-unl) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#quorum-shortfall)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -2973,7 +2973,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*How many ledgers this node has fully validated but not yet published to its clients and subscribers.*\n\n###### How it's computed:\n*ledger_quorum_publish series publish_lag: the validated ledger sequence minus the published ledger sequence, floored at zero. The published sequence was never exported before, so this gap was not derivable from any other series.*\n\n###### Reading it:\n*Publishing trails validation by design, so a small lag that drains each round is normal. A lag that stays positive, or grows, means validation is healthy and the publish pipeline is not \u2014 a different fault from anything the quorum or acquire panels can show.*\n\n###### Healthy range:\n*0 to 1 ledger.*\n\n###### Watch for:\n*A monotonic climb: the publish loop is falling behind a chain tip the node already holds, so clients and subscriptions see stale data while the node itself is current. Read it with Worker Pool Saturation and the per-job-type jobq_<jobtype>_deferred gauges \u2014 a starved job queue is the usual cause. A flat 0 is only healthy on a node that is validating: on one that never has, the 0 means nothing has been validated to publish, so read Trusted Validations vs Quorum Target first.*\n\n###### Keywords:\n- **Publish lag** *(per node)* \u2014 validated ledgers not yet published to clients and subscribers, i.e. the gap between the validated and the published sequence.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerLedgerQuorumPublishGauge`\n\n###### References:\n[Ledger close and publication on xrpl.org](https://xrpl.org/docs/concepts/consensus-protocol) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#publish-lag)",
|
||||
"description": "###### What this is:\n*How many ledgers this node has fully validated but not yet published to its clients and subscribers.*\n\n###### How it's computed:\n*ledger_quorum_publish series publish_lag: the validated ledger sequence minus the published ledger sequence, floored at zero. The published sequence was never exported before, so this gap was not derivable from any other series.*\n\n###### Reading it:\n*Publishing trails validation by design, so a small lag that drains each round is normal. A lag that stays positive, or grows, means validation is healthy and the publish pipeline is not \u2014 a different fault from anything the quorum or acquire panels can show.*\n\n###### Healthy range:\n*0 to 1 ledger.*\n\n###### Watch for:\n*A monotonic climb: the publish loop is falling behind a chain tip the node already holds, so clients and subscriptions see stale data while the node itself is current. Read it with Worker Pool Saturation and the per-job-type jobq_<jobtype>_deferred gauges \u2014 a starved job queue is the usual cause. A flat 0 is only healthy on a node that is validating: on one that never has, the 0 means nothing has been validated to publish, so read Trusted Validations vs Quorum Target first.*\n\n###### Keywords:\n- **Publish lag** *(per node)* \u2014 validated ledgers not yet published to clients and subscribers, i.e. the gap between the validated and the published sequence.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerLedgerQuorumPublishGauge`\n\n###### References:\n[Ledger close and publication on xrpl.org](https://xrpl.org/docs/concepts/consensus-protocol) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#publish-lag)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -3274,7 +3274,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Seconds until an unsupported amendment activates and this node stops validating for good. This is the LEADING indicator \u2014 the countdown before the block, while there is still time to upgrade.*\n\n###### How it's computed:\n*amendment_block series seconds_to_block. A value of -1 is the sentinel for \"no unsupported amendment is pending\", the same convention validator_health{metric=\"unl_expiry_days\"} already uses.*\n\n###### Reading it:\n*Because -1 is the healthy sentinel and a small positive number is the emergency, the colour scale is not monotonic \u2014 read the value, not only the colour. -1 is green and means nothing is pending. A large positive value is green above 7 days and yellow under 7 days: an unsupported amendment holds majority but there is still time to upgrade. A small positive value under 1 day is red: at 0 this node stops validating and does not resume without a software upgrade.*\n\n###### Healthy range:\n*-1.*\n\n###### Watch for:\n*Any value at or above 0. This is distinct from the Amendment Blocked stat on the Validator Health dashboard (validator_health{metric=\"amendment_blocked\"}), which reports the TERMINAL state \u2014 already blocked, too late to act. This panel is the window before that happens, so the two are read together: countdown first, terminal state as confirmation. The identity of the blocking amendment is deliberately NOT a metric label, because the network can vote on an arbitrary 256-bit amendment id and a label would be unbounded cardinality. The hash is logged instead by AmendmentTableImpl::doValidatedLedger (\"Unsupported amendment <hash> reached majority at ...\") and is available via Loki.*\n\n###### Keywords:\n- **Amendment block countdown** *(per node)* \u2014 seconds until an unsupported amendment that has reached majority activates and blocks this node; -1 when none is pending.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerAmendmentBlockGauge`\n\n###### References:\n[Amendments on xrpl.org](https://xrpl.org/docs/concepts/networks-and-servers/amendments) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#amendment-blocked)",
|
||||
"description": "###### What this is:\n*Seconds until an unsupported amendment activates and this node stops validating for good. This is the LEADING indicator \u2014 the countdown before the block, while there is still time to upgrade.*\n\n###### How it's computed:\n*amendment_block series seconds_to_block. A value of -1 is the sentinel for \"no unsupported amendment is pending\", the same convention validator_health{metric=\"unl_expiry_days\"} already uses.*\n\n###### Reading it:\n*Because -1 is the healthy sentinel and a small positive number is the emergency, the colour scale is not monotonic \u2014 read the value, not only the colour. -1 is green and means nothing is pending. A large positive value is green above 7 days and yellow under 7 days: an unsupported amendment holds majority but there is still time to upgrade. A small positive value under 1 day is red: at 0 this node stops validating and does not resume without a software upgrade.*\n\n###### Healthy range:\n*-1.*\n\n###### Watch for:\n*Any value at or above 0. This is distinct from the Amendment Blocked stat on the Validator Health dashboard (validator_health{metric=\"amendment_blocked\"}), which reports the TERMINAL state \u2014 already blocked, too late to act. This panel is the window before that happens, so the two are read together: countdown first, terminal state as confirmation. The identity of the blocking amendment is deliberately NOT a metric label, because the network can vote on an arbitrary 256-bit amendment id and a label would be unbounded cardinality. The hash is logged instead by AmendmentTableImpl::doValidatedLedger (\"Unsupported amendment <hash> reached majority at ...\") and is available via Loki.*\n\n###### Keywords:\n- **Amendment block countdown** *(per node)* \u2014 seconds until an unsupported amendment that has reached majority activates and blocks this node; -1 when none is pending.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerAmendmentBlockGauge`\n\n###### References:\n[Amendments on xrpl.org](https://xrpl.org/docs/concepts/networks-and-servers/amendments) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#amendment-blocked)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -3507,7 +3507,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Whether an amendment this build does not support has reached majority on the network.*\n\n###### How it's computed:\n*amendment_block series warned: 1 while an unsupported amendment holds majority, 0 otherwise.*\n\n###### Reading it:\n*0 is healthy. A 1 is the first warning that an upgrade is required, and it is raised before the amendment activates rather than after.*\n\n###### Healthy range:\n*0.*\n\n###### Watch for:\n*The transition from 0 to 1 \u2014 that is the moment the upgrade clock starts. Read Amendment Block Countdown next for how long is left, and the Amendment Blocked stat on the Validator Health dashboard for whether the block has already happened; that one is the terminal state, this one is the warning. Which amendment is blocking is not a label (an arbitrary 256-bit amendment id would be unbounded cardinality); the hash is logged by AmendmentTableImpl::doValidatedLedger and is available via Loki.*\n\n###### Keywords:\n- **Amendment warning** *(per node)* \u2014 an unsupported amendment has reached majority; the node still validates, but will stop when that amendment activates.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerAmendmentBlockGauge`\n\n###### References:\n[Amendments on xrpl.org](https://xrpl.org/docs/concepts/networks-and-servers/amendments) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#amendment-blocked)",
|
||||
"description": "###### What this is:\n*Whether an amendment this build does not support has reached majority on the network.*\n\n###### How it's computed:\n*amendment_block series warned: 1 while an unsupported amendment holds majority, 0 otherwise.*\n\n###### Reading it:\n*0 is healthy. A 1 is the first warning that an upgrade is required, and it is raised before the amendment activates rather than after.*\n\n###### Healthy range:\n*0.*\n\n###### Watch for:\n*The transition from 0 to 1 \u2014 that is the moment the upgrade clock starts. Read Amendment Block Countdown next for how long is left, and the Amendment Blocked stat on the Validator Health dashboard for whether the block has already happened; that one is the terminal state, this one is the warning. Which amendment is blocking is not a label (an arbitrary 256-bit amendment id would be unbounded cardinality); the hash is logged by AmendmentTableImpl::doValidatedLedger and is available via Loki.*\n\n###### Keywords:\n- **Amendment warning** *(per node)* \u2014 an unsupported amendment has reached majority; the node still validates, but will stop when that amendment activates.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerAmendmentBlockGauge`\n\n###### References:\n[Amendments on xrpl.org](https://xrpl.org/docs/concepts/networks-and-servers/amendments) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#amendment-blocked)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -3924,7 +3924,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Node-store write latency next to read latency, in microseconds per operation. The write side is the signal: a node with a large existing database back-fills slower than a fresh one, and back-fill is write-bound, so no read-side metric can show it.*\n\n###### How it's computed:\n*nodestore_state series node_writes_duration_us and node_reads_duration_us, each divided by its own count series (node_writes, node_reads_total) so the reading is the latency during the selected interval rather than the average since boot. All three concrete store paths time themselves, so the write side is live on an ordinary node.*\n\n###### Reading it:\n*Compare the two lines. Reads far above writes points at the read path or a cold cache; writes far above reads points at backend write pressure, which is the large-existing-database case.*\n\n###### Healthy range:\n*Both well under a few hundred microseconds on healthy local storage.*\n\n###### Watch for:\n*A rising write line during history back-fill: the backend cannot absorb writes fast enough and sync will stay slow no matter how many peers are available. Read with Peers Able to Serve Needed Sequence to tell a data-supply problem from a disk problem. This is a mean, not a percentile \u2014 a tail that matters will move it, but p99 is not available from this signal.*\n\n###### Keywords:\n- **Node-store write latency** *(per node)* \u2014 how long the node store takes to persist one object.\n- **Node-store read latency** *(per node)* \u2014 how long the node store takes to retrieve one object.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerNodeStoreGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#node-store-write-latency)",
|
||||
"description": "###### What this is:\n*Node-store write latency next to read latency, in microseconds per operation. The write side is the signal: a node with a large existing database back-fills slower than a fresh one, and back-fill is write-bound, so no read-side metric can show it.*\n\n###### How it's computed:\n*nodestore_state series node_writes_duration_us and node_reads_duration_us, each divided by its own count series (node_writes, node_reads_total) so the reading is the latency during the selected interval rather than the average since boot. All three concrete store paths time themselves, so the write side is live on an ordinary node.*\n\n###### Reading it:\n*Compare the two lines. Reads far above writes points at the read path or a cold cache; writes far above reads points at backend write pressure, which is the large-existing-database case.*\n\n###### Healthy range:\n*Both well under a few hundred microseconds on healthy local storage.*\n\n###### Watch for:\n*A rising write line during history back-fill: the backend cannot absorb writes fast enough and sync will stay slow no matter how many peers are available. Read with Peers Able to Serve Needed Sequence to tell a data-supply problem from a disk problem. This is a mean, not a percentile \u2014 a tail that matters will move it, but p99 is not available from this signal.*\n\n###### Keywords:\n- **Node-store write latency** *(per node)* \u2014 how long the node store takes to persist one object.\n- **Node-store read latency** *(per node)* \u2014 how long the node store takes to retrieve one object.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerNodeStoreGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#node-store-write-latency)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -4031,7 +4031,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*Node-store write and read operation rates \u2014 the denominators behind the latency panel.*\n\n###### How it's computed:\n*Rate of the nodestore_state node_writes and node_reads_total series.*\n\n###### Reading it:\n*Writes climb while a node is back-filling history and fall to near the ledger-close rate once it is caught up.*\n\n###### Healthy range:\n*Non-zero writes whenever the node is ingesting ledgers.*\n\n###### Watch for:\n*Write rate at zero while the node is still behind the network: nothing is being persisted, so the stall is upstream of the node store \u2014 check peer supply and the acquire panels rather than storage. A flat latency with a collapsing operation rate also means the latency figure above has gone stale rather than good.*\n\n###### Keywords:\n- **Node-store operation rate** *(per node)* \u2014 stores and fetches per second against the node store.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerNodeStoreGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#node-store-operation-rate)",
|
||||
"description": "###### What this is:\n*Node-store write and read operation rates \u2014 the denominators behind the latency panel.*\n\n###### How it's computed:\n*Rate of the nodestore_state node_writes and node_reads_total series.*\n\n###### Reading it:\n*Writes climb while a node is back-filling history and fall to near the ledger-close rate once it is caught up.*\n\n###### Healthy range:\n*Non-zero writes whenever the node is ingesting ledgers.*\n\n###### Watch for:\n*Write rate at zero while the node is still behind the network: nothing is being persisted, so the stall is upstream of the node store \u2014 check peer supply and the acquire panels rather than storage. A flat latency with a collapsing operation rate also means the latency figure above has gone stale rather than good.*\n\n###### Keywords:\n- **Node-store operation rate** *(per node)* \u2014 stores and fetches per second against the node store.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerNodeStoreGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#node-store-operation-rate)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
@@ -4246,7 +4246,7 @@
|
||||
"type": "prometheus",
|
||||
"uid": "${DS_PROMETHEUS}"
|
||||
},
|
||||
"description": "###### What this is:\n*When an online-delete rotation is running, and the rate of the extra writes it forces. A rotation rewrites into the new backend any node body the doomed archive serves, which is I/O an ordinary fetch would never perform \u2014 and it exists only on a populated, already-rotated database, which is why it never appears on a fresh node.*\n\n###### How it's computed:\n*rotation_state with metric=in_flight plotted raw (it is a 0/1 state flag), and metric=copy_forward \u2014 a cumulative write total \u2014 plotted as a rate. Both are read from the node store on each collection tick. The copy-forward count existed before as a log-only per-rotation tally that reset on every swap; the total behind this panel never resets, so it can be rated.*\n\n###### Reading it:\n*The two must move together: copy-forward writes should only appear while the window flag is 1. Read the write rate against the node-store write latency panel above \u2014 that is what tells extra rotation writes from a slow backend.*\n\n###### Healthy range:\n*Flag at 0 most of the time, rising to 1 briefly once per delete interval, with the write rate non-zero only inside those windows.*\n\n###### Watch for:\n*A copy-forward rate that is large enough to move node-store write latency: rotation is competing with sync I/O, which is the whole hypothesis this panel tests. Copy-forward writes while the flag reads 0 would mean the window flag leaked, not that rotation is cheap. NO SERIES AT ALL on either query means online_delete is not configured on this node, which is different from a rotation that costs nothing.*\n\n###### Keywords:\n- **Rotation window** *(per node)* \u2014 the interval during which an online-delete backend swap is in progress.\n- **Copy-forward write** *(per node)* \u2014 rewriting a node body from the backend about to be deleted into the one replacing it.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`registerRotationStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#rotation-window)",
|
||||
"description": "###### What this is:\n*When an online-delete rotation is running, and the rate of the extra writes it forces. A rotation rewrites into the new backend any node body the doomed archive serves, which is I/O an ordinary fetch would never perform \u2014 and it exists only on a populated, already-rotated database, which is why it never appears on a fresh node.*\n\n###### How it's computed:\n*rotation_state with metric=in_flight plotted raw (it is a 0/1 state flag), and metric=copy_forward \u2014 a cumulative write total \u2014 plotted as a rate. Both are read from the node store on each collection tick. The copy-forward count existed before as a log-only per-rotation tally that reset on every swap; the total behind this panel never resets, so it can be rated.*\n\n###### Reading it:\n*The two must move together: copy-forward writes should only appear while the window flag is 1. Read the write rate against the node-store write latency panel above \u2014 that is what tells extra rotation writes from a slow backend.*\n\n###### Healthy range:\n*Flag at 0 most of the time, rising to 1 briefly once per delete interval, with the write rate non-zero only inside those windows.*\n\n###### Watch for:\n*A copy-forward rate that is large enough to move node-store write latency: rotation is competing with sync I/O, which is the whole hypothesis this panel tests. Copy-forward writes while the flag reads 0 would mean the window flag leaked, not that rotation is cheap. NO SERIES AT ALL on either query means online_delete is not configured on this node, which is different from a rotation that costs nothing.*\n\n###### Keywords:\n- **Rotation window** *(per node)* \u2014 the interval during which an online-delete backend swap is in progress.\n- **Copy-forward write** *(per node)* \u2014 rewriting a node body from the backend about to be deleted into the one replacing it.\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[AppMetricGauges.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/telemetry/AppMetricGauges.cpp)\n\n###### Function:\n`registerRotationStateGauge`\n\n###### References:\n[Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#rotation-window)",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
|
||||
@@ -806,7 +806,7 @@
|
||||
},
|
||||
{
|
||||
"title": "State Changes Rate [$xrpl_network_type]",
|
||||
"description": "###### What this is:\n*Rate of server operating-state changes per hour.*\n\n###### How it's computed:\n*Per-second rate of the state-change counter over the dashboard's rate interval, scaled to per hour.*\n\n###### Reading it:\n*Near zero is healthy; each increment is one state transition.*\n\n###### Healthy range:\n*Near 0 changes per hour.*\n\n###### Watch for:\n*Frequent transitions, which point to network instability or configuration problems.*\n\n###### Keywords:\n- **Operating mode / server state** *(per node)* \u2014 the node's sync level: Disconnected, Connected, Syncing, Tracking, Full (and Validating/Proposing).\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[MetricsRegistry.cpp](https://github.com/XRPLF/rippled/blob/develop/src/libxrpl/telemetry/MetricsRegistry.cpp)\n\n###### Function:\n`incrementStateChanges (caller NetworkOPs.cpp)`\n\n###### References:\n[Operating mode / server state](https://xrpl.org/docs/references/http-websocket-apis/api-conventions/xrpld-server-states) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#operating-mode-server-state)",
|
||||
"description": "###### What this is:\n*Rate of server operating-state changes per hour.*\n\n###### How it's computed:\n*Per-second rate of the state-change counter over the dashboard's rate interval, scaled to per hour.*\n\n###### Reading it:\n*Near zero is healthy; each increment is one state transition.*\n\n###### Healthy range:\n*Near 0 changes per hour.*\n\n###### Watch for:\n*Frequent transitions, which point to network instability or configuration problems.*\n\n###### Keywords:\n- **Operating mode / server state** *(per node)* \u2014 the node's sync level: Disconnected, Connected, Syncing, Tracking, Full (and Validating/Proposing).\n\n###### Computation boundary:\n*Result: Per node \u2014 each series is one server's own value.*\n*Computed in xrpld code (MetricsRegistry, OpenTelemetry SDK) and exported as a metric; the collector only forwards it; the Grafana query selects and aggregates it.*\n\n###### Source:\n[NetworkOPs.cpp](https://github.com/XRPLF/rippled/blob/develop/src/xrpld/app/misc/NetworkOPs.cpp)\n\n###### Function:\n`NetworkOPsImp::setMode`\n\n###### References:\n[Operating mode / server state](https://xrpl.org/docs/references/http-websocket-apis/api-conventions/xrpld-server-states) \u00b7 [Telemetry glossary](https://github.com/XRPLF/rippled/blob/develop/docs/telemetry-glossary.md#operating-mode-server-state)",
|
||||
"type": "stat",
|
||||
"gridPos": {
|
||||
"h": 10,
|
||||
|
||||
Reference in New Issue
Block a user