mirror of
https://github.com/XRPLF/rippled.git
synced 2026-10-01 01:08:03 +00:00
Dashboards and alert rules reference 186 metrics; the harness asserted 57. Excluding the 107 per-category overlay-traffic expansions, the meaningful gap was 20 names. This closes it under the contract file's own doctrine: assert only what the workload guarantees, and record the rest with a precise reason. Asserted 18, taking the metric checks from 61 to 79 and the whole metric phase from 66 to 84. No pre-existing check name or position changes. statsd_gauges gains the nine state_accounting_* siblings of the one member already asserted, plus the two NodeFamily full-below-cache gauges and overlay_peer_disconnects. All twelve rest on one mechanism the group description now spells out: on the OTel path a beast gauge is an Int64ObservableGauge, every instance self-registers in its constructor, onCollectionReady arms all of them unconditionally, and the armed callback Observes on every export cycle whether or not set was ever called, so the series exist at 0. The state_accounting family is set in one unconditional block in NetworkOPsImp::collectMetrics, and full_transitions is the input to the NodeStateFlapping alert rule, so the alert's own signal had been going unverified. A new job_queue_per_type_gauges group asserts the six per-job-type gauges that a panel or a rule names literally, jobq_manifest_waiting among them as the ManifestJobQueueConvoy rule's input. The description records why those six and not all 105: the guarantee is identical for every non-special job type, so the discriminator is consumer coverage, and the remaining names are only reached through topk queries over the family that do not depend on any single type being present. Recorded four more in not_asserted rather than asserting them. pathfind_fast_milliseconds is unreachable for this workload, not merely rare: reportFast fires only from the doCreate fast pass, which is guarded by !hasCompletion(), and both ripple_path_find entry points construct the request with a completion function. Only the path_find subscription reaches it, and the generator does not use it. pathfind_full_milliseconds is reachable but only one ledger close after the request, through PathRequestManager::updateAll, and nothing in the harness arranges or checks that, so the guarantee is probabilistic. warn_total and drop_total are resource-manager meters gated on a consumer crossing the warn or drop threshold; their rpc-pathfinding panels are correct and render empty only because the condition has not occurred, which is worth stating because both were briefly mis-read as phantoms. Runbook check counts updated to match.
241 lines
29 KiB
JSON
241 lines
29 KiB
JSON
{
|
|
"description": "Expected metric inventory for xrpld telemetry validation. Metric names have no prefix (the xrpld_ prefix was removed). beast::insight metrics are lowercased by formatName. Every name here was verified against its declaration in MetricsRegistry.cpp or include/xrpl/telemetry/GetObjectMetricNames.h and against a panel query under docker/telemetry/grafana/dashboards/. IMPORTANT: validate_telemetry.py has no notion of an optional metric — validate_metrics() iterates every group that has a \"metrics\" key and hard-fails any name with 0 Prometheus series after a 45 s poll. A metric is therefore listed only when the harness workload guarantees it will appear: observable gauges/counters whose callbacks Observe unconditionally (series exist at value 0), or push counters/histograms on a path every run exercises. Workload-gated and defect-gated names are recorded in the \"not_asserted\" group, which intentionally has no \"metrics\" key so the validator skips it. Only series existence is checked, never a value, except for the four bounds checks hardcoded in PARITY_VALUE_SANITY. A group may additionally declare \"required_labels\": each label there becomes one check that at least one of the group's series carries it with a non-empty value, matched as <label>!=\"\" because Prometheus cannot tell an absent label from an empty one. The same guarantee rule applies — list a label only where the workload guarantees it.",
|
|
"spanmetrics": {
|
|
"description": "SpanMetrics-derived RED metrics from the OTel Collector spanmetrics connector.",
|
|
"metrics": [
|
|
"span_calls_total",
|
|
"span_duration_milliseconds_bucket",
|
|
"span_duration_milliseconds_count",
|
|
"span_duration_milliseconds_sum"
|
|
],
|
|
"required_labels": [
|
|
"span_name",
|
|
"status_code",
|
|
"service_name",
|
|
"span_kind"
|
|
],
|
|
"dimension_labels": [
|
|
"command",
|
|
"rpc_status",
|
|
"consensus_mode",
|
|
"local",
|
|
"proposal_trusted",
|
|
"validation_trusted",
|
|
"tx_type",
|
|
"ter_result",
|
|
"stage",
|
|
"txq_status",
|
|
"close_time_correct",
|
|
"consensus_state",
|
|
"suppressed"
|
|
],
|
|
"_dimension_labels_note": "Bare label names as configured in otel-collector-config.yaml spanmetrics dimensions. Informational only (not asserted by the validator)."
|
|
},
|
|
"statsd_gauges": {
|
|
"description": "beast::insight gauges exported via OTLP/HTTP to the collector (server=otel). What guarantees these is the export mechanism, not the workload. On the OTel path a beast gauge is an Int64ObservableGauge (OTelCollector.cpp:703); every OTelGaugeImpl registers itself with the collector in its own constructor (:690); onCollectionReady() arms every registered gauge with no per-metric condition (:941-979); and the armed callback Observes currentValue() on each export cycle regardless of the value and regardless of whether set() was ever called (:719-727). Series therefore exist from the first export, at 0 if nothing has happened. The 'gauges only mark dirty on value changes' caveat in _poll_series_count is a property of the StatsD backend and does not apply here. The consequence is that a beast gauge whose object is constructed on an unconditional startup path is as safe to assert as an observable gauge from MetricsRegistry — which is why the state_accounting family below is asserted whole rather than one member deep. The ten state_accounting_* names are one family with one guarantee. NetworkOPsImp::Stats creates all ten in the NetworkOPsImp constructor (NetworkOPs.cpp:1069-1080; makeGauge(\"State_Accounting\", \"Full_duration\") joins on '.' at Collector.h:151-155 and formatName lowercases it, giving state_accounting_full_duration), and NetworkOPsImp::collectMetrics() set()s all ten unconditionally in a single block (:5204-5231). No state has to be entered for its pair to exist: a node that never disconnects still publishes state_accounting_disconnected_duration and _transitions at 0. All ten appear on the node-health dashboard and its grafanacloud copy, and state_accounting_full_transitions is also the input to the NodeStateFlapping alert rule (provisioning/alerting/rules.yaml, uid xrpld-node-state-flapping). Asserting only full_duration left the other nine unchecked, the alert's own input among them, even though a regression could not plausibly drop one sibling without dropping full_duration too. node_family_full_below_cache_hit_rate and _size are the TaggedCache Stats pair (TaggedCache.h:282-283), created in the cache constructor (TaggedCache.ipp:63-66) and published by collectMetrics(), which reports a hit_rate of 0 rather than skipping the gauge when hits+misses is 0 (TaggedCache.ipp:751-765). The instance is NodeFamily's full-below cache, constructed with the real collector in NodeFamily's constructor (NodeFamily.cpp:28-35), which every node builds. Both are on node-health. overlay_peer_disconnects is the first member of OverlayImpl::Stats (OverlayImpl.h:603) and OverlayImpl::collectMetrics() assigns it getPeerDisconnect() unconditionally (:646), in the same hook that publishes the overlay_traffic group's gauges. It is on network-traffic and log-derived-insights. It is listed here rather than in overlay_traffic because that group's contract is an explicit subset of the per-category traffic family, and a disconnect count is not one of those categories.",
|
|
"metrics": [
|
|
"ledgermaster_validated_ledger_age",
|
|
"ledgermaster_published_ledger_age",
|
|
"state_accounting_full_duration",
|
|
"peer_finder_active_inbound_peers",
|
|
"peer_finder_active_outbound_peers",
|
|
"jobq_job_count",
|
|
"state_accounting_connected_duration",
|
|
"state_accounting_connected_transitions",
|
|
"state_accounting_disconnected_duration",
|
|
"state_accounting_disconnected_transitions",
|
|
"state_accounting_full_transitions",
|
|
"state_accounting_syncing_duration",
|
|
"state_accounting_syncing_transitions",
|
|
"state_accounting_tracking_duration",
|
|
"state_accounting_tracking_transitions",
|
|
"node_family_full_below_cache_hit_rate",
|
|
"node_family_full_below_cache_size",
|
|
"overlay_peer_disconnects"
|
|
]
|
|
},
|
|
"statsd_counters": {
|
|
"description": "beast::insight counters exported via OTLP/HTTP. The OTel Prometheus exporter appends _total to monotonic counters.",
|
|
"metrics": ["rpc_requests_total", "ledger_fetches_total"]
|
|
},
|
|
"io_latency": {
|
|
"description": "io_context scheduling-latency histogram — a beast::insight Event exported over OTLP (Application.cpp:515, makeEvent('ios_latency') with the default millisecond unit). It is the one beast Event the node itself guarantees, which is why it is asserted while rpc_time_milliseconds and the jobq_* pairs sit in not_asserted: the sampler starts on the unconditional startup path (Application.cpp:1697, outside the 'if (withTimers)' guard), and its handler always emits the first sample whatever its value (Application.cpp:185-189, 'firstSample_.exchange(false) || lastSample >= 10ms', with a comment stating the point is to register the metric downstream). The OTel histogram is cumulative, so that one notify creates series that persist for the rest of the run. Named with the _bucket/_count/_sum suffixes the Prometheus exporter emits for a histogram — no bare ios_latency_milliseconds series exists, the same convention rpc_method_us and span_duration_milliseconds follow, and every consumer queries ios_latency_milliseconds_bucket: 2 distinct panels (ledger-data-sync 'I/O Scheduler Latency p95' and node-health 'I/O Latency', each mirrored in a grafanacloud copy) plus one alert rule (grafana/provisioning/alerting/rules.yaml:586). Only existence is asserted: past the first sample the handler reports only latencies >= 10 ms, so neither the sample count nor the value is predictable.",
|
|
"metrics": [
|
|
"ios_latency_milliseconds_bucket",
|
|
"ios_latency_milliseconds_count",
|
|
"ios_latency_milliseconds_sum"
|
|
]
|
|
},
|
|
"overlay_traffic": {
|
|
"description": "Overlay traffic metrics (subset — full list has 45+ categories).",
|
|
"metrics": [
|
|
"total_bytes_in",
|
|
"total_bytes_out",
|
|
"total_messages_in",
|
|
"total_messages_out"
|
|
]
|
|
},
|
|
"nodestore_io": {
|
|
"description": "NodeStore I/O observable gauge (MetricsRegistry via OTLP). Single metric with 'metric' label distinguishing sub-metrics.",
|
|
"metrics": ["nodestore_state"]
|
|
},
|
|
"cache_hit_rates": {
|
|
"description": "Cache hit rate observable gauge (MetricsRegistry via OTLP). Single metric with 'metric' label.",
|
|
"metrics": ["cache_metrics"]
|
|
},
|
|
"transaction_queue": {
|
|
"description": "Transaction queue observable gauge (MetricsRegistry via OTLP). Single metric with 'metric' label.",
|
|
"metrics": ["txq_metrics"]
|
|
},
|
|
"rpc_method_detail": {
|
|
"description": "Per-RPC-method counters and duration histogram (MetricsRegistry.cpp:351-357). rpc_method_errored_total is deliberately absent — see not_asserted below. rpc_method_us is a Histogram, so the Prometheus exporter emits only the _bucket/_count/_sum triple and there is no bare rpc_method_us series to match — same convention as span_duration_milliseconds in the spanmetrics group above.",
|
|
"metrics": [
|
|
"rpc_method_started_total",
|
|
"rpc_method_finished_total",
|
|
"rpc_method_us_bucket",
|
|
"rpc_method_us_count",
|
|
"rpc_method_us_sum"
|
|
]
|
|
},
|
|
"job_queue": {
|
|
"description": "Job-queue counters and latency histograms (MetricsRegistry.cpp:360-366). Every xrpld job passes through these, so they populate under any workload. Both histograms are recorded in the same function bodies as job_started_total / job_finished_total, under the same guard and with the same labels, so their presence is equally guaranteed. They are named with the _bucket/_count/_sum suffixes the Prometheus exporter emits: regression-metrics.json and the job-queue dashboard both query job_queued_us_bucket / job_running_us_bucket, and no bare series exists. This group also carries the xrpl_node_id required_labels assertion. 070d29b465 added the xrpl.node.id resource attribute so Grafana Cloud trace ingest would stop folding distinct nodes into one — ingest groups ResourceSpans but ignores service.instance.id, so before the attribute existed the boards reported one node's ledger.build five times over and saw 2 of 9 nodes. Nothing asserted it, so a regression would silently restore the folding. It is asserted here rather than on a statsd_* group because these metrics are MetricsRegistry-backed and guaranteed by any workload, whereas the beast::insight meter provider is constructed before the wallet DB exists and so has no node public key to stamp: those metrics legitimately lack the label and 070d29b465 omits it rather than writing it blank.",
|
|
"required_labels": ["xrpl_node_id"],
|
|
"metrics": [
|
|
"job_queued_total",
|
|
"job_started_total",
|
|
"job_finished_total",
|
|
"job_queued_us_bucket",
|
|
"job_queued_us_count",
|
|
"job_queued_us_sum",
|
|
"job_running_us_bucket",
|
|
"job_running_us_count",
|
|
"job_running_us_sum"
|
|
]
|
|
},
|
|
"job_queue_per_type_gauges": {
|
|
"description": "Per-job-type queue-depth gauges from the beast::insight \"jobq\" group, for the six job types the ledger-sync diagnostics single out. JobTypeData's constructor creates waiting/running/deferred for every non-special job type (JobTypeData.h:96-103), the JobQueue constructor eagerly constructs a JobTypeData for every type in JobTypes (JobQueue.cpp:45-53), and JobQueue::collect() assigns all three from one snapshot for every type on every hook cycle, with no per-type condition (JobQueue.cpp:104-106). The collector is the \"jobq\" group, so GroupImp::makeName() prefixes \"jobq.\" and formatName lowercases the job type: JtTxnData's \"fetchTxnData\" + \"_running\" exports as jobq_fetchtxndata_running (JobTypeData.h:29-37). Existence rests on the same ObservableGauge mechanism documented on statsd_gauges, so these series exist at 0 on an idle node. All six were present in the CI validation run's own emitted-metric list, which is the strongest evidence behind any entry in this file. All six job types are non-special: makeFetchPack (limit 1), manifest (maxLimit), ledgerRequest (3), ledgerData (3), updatePaths (1) and fetchTxnData (5) all declare a non-zero limit in JobTypes.h, so none is skipped by the special() guard in JobTypeData's constructor. The 11 special types (limit 0) get no gauges at all, which is why no jobq_pathfind_* or jobq_peercommand_* name can be asserted. Why these six and not all 105. The guarantee is identical for every one of the 35 non-special types x 3 states, so the discriminator here is consumer coverage rather than emission: these six are the names a dashboard or an alert queries literally, so losing one blanks a specific panel or silently disarms a rule. jobq_manifest_waiting is the input to the ManifestJobQueueConvoy alert rule (provisioning/alerting/rules.yaml, uid xrpld-manifest-job-convoy); the other five are per-type saturation panels on node-health and its grafanacloud copy. The remaining names are only ever reached through the __name__=~\"jobq_.*_waiting\" / \"jobq_.*_deferred\" topk queries on ledger-data-sync, which keep working as long as the family exists and do not depend on any single job type being present. Asserting all 105 would add 99 checks that share one failure mode with these six while costing poll bandwidth in the shared 45 s window, so it would buy no signal.",
|
|
"metrics": [
|
|
"jobq_fetchtxndata_running",
|
|
"jobq_ledgerdata_running",
|
|
"jobq_ledgerrequest_running",
|
|
"jobq_makefetchpack_running",
|
|
"jobq_manifest_waiting",
|
|
"jobq_updatepaths_running"
|
|
]
|
|
},
|
|
"rpc_in_flight": {
|
|
"description": "In-flight RPC gauge via the XRPL_METRIC_UPDOWN_ADD call-site macro (PerfLogImp.cpp, +1 rpcStart / -1 rpcEnd). UpDownCounter: no _total suffix.",
|
|
"metrics": ["rpc_in_flight_requests"]
|
|
},
|
|
"object_counts": {
|
|
"description": "Counted object instances observable gauge (MetricsRegistry via OTLP).",
|
|
"metrics": ["object_count"]
|
|
},
|
|
"load_factors": {
|
|
"description": "Fee escalation and load factor observable gauge (MetricsRegistry via OTLP).",
|
|
"metrics": ["load_factor_metrics"]
|
|
},
|
|
"parity_validation_agreement": {
|
|
"description": "External dashboard parity: validation agreement percentages (MetricsRegistry).",
|
|
"metrics": [
|
|
"validation_agreement{metric=\"agreement_pct_1h\"}",
|
|
"validation_agreement{metric=\"agreement_pct_24h\"}"
|
|
]
|
|
},
|
|
"parity_validator_health": {
|
|
"description": "External dashboard parity: validator health indicators (MetricsRegistry).",
|
|
"metrics": [
|
|
"validator_health{metric=\"amendment_blocked\"}",
|
|
"validator_health{metric=\"unl_expiry_days\"}"
|
|
]
|
|
},
|
|
"parity_peer_quality": {
|
|
"description": "External dashboard parity: peer quality metrics (MetricsRegistry).",
|
|
"metrics": [
|
|
"peer_quality{metric=\"peer_latency_p90_ms\"}",
|
|
"peer_quality{metric=\"peers_insane_count\"}"
|
|
]
|
|
},
|
|
"parity_ledger_economy": {
|
|
"description": "External dashboard parity: ledger economy metrics (MetricsRegistry.cpp:1401). transaction_rate is observed on every export, in both branches of the ledger-age test (MetricsRegistry.cpp:1444-1451). base_fee_xrp is observed only inside the 'if (ledger)' guard on getValidatedLedger() (MetricsRegistry.cpp:1418-1423), and that returns validLedger_ (LedgerMaster.cpp:1569-1572), which stays null until a ledger validates — the same precondition complete_ledgers has. Both are asserted because run-full-validation.sh waits for a validated ledger before running the workload. base_fee_xrp absent while transaction_rate is present is the signature of a cluster that never validated, not of a missing metric.",
|
|
"metrics": [
|
|
"ledger_economy{metric=\"base_fee_xrp\"}",
|
|
"ledger_economy{metric=\"transaction_rate\"}"
|
|
]
|
|
},
|
|
"parity_state_tracking": {
|
|
"description": "External dashboard parity: server state tracking (MetricsRegistry).",
|
|
"metrics": ["state_tracking{metric=\"state_value\"}"]
|
|
},
|
|
"parity_counters": {
|
|
"description": "External dashboard parity: monotonic counters (MetricsRegistry). validations_checked_total is incremented unconditionally at the top of NetworkOPsImp::recvValidation (NetworkOPs.cpp:2681), and run-full-validation.sh brings up a 5-node validator cluster, so inbound validations are guaranteed.",
|
|
"metrics": [
|
|
"ledgers_closed_total",
|
|
"validations_sent_total",
|
|
"validations_checked_total",
|
|
"state_changes_total"
|
|
]
|
|
},
|
|
"parity_storage": {
|
|
"description": "External dashboard parity: storage detail metrics (MetricsRegistry).",
|
|
"metrics": ["storage_detail{metric=\"stored_object_bytes\"}"]
|
|
},
|
|
"node_health_gauges": {
|
|
"description": "Node-health observable gauges (MetricsRegistry.cpp:997, :1081, :1102, :1161). server_info, build_info and db_metrics Observe unconditionally on every periodic export (build_info observes a literal 1; server_info and db_metrics read live services), so their series exist regardless of workload shape. complete_ledgers is the exception and is asserted on a narrower guarantee: its callback returns without observing when the range is empty (MetricsRegistry.cpp:1113-1114) and skips any segment that carries no '-' (:1122-1127), and a one-sequence range renders with no '-' (RangeSet.h:70-71), so it needs a complete range spanning at least two sequences. completeLedgers_ is filled by setFullLedger (LedgerMaster.cpp:862-863), which on a peered node is reached only from the publish path in doAdvance (LedgerMaster.cpp:1972) — closing a ledger is not enough, it has to validate. run-full-validation.sh waits for that before the workload starts, so on a healthy cluster the series always exists — a 5-node run yields 10 series, one start and one end per node. If this check ever fails, read the Step 3 output first: a run that logged 'No validated ledger' cannot produce this series and the cluster, not the exporter, is what broke.",
|
|
"metrics": ["server_info", "build_info", "complete_ledgers", "db_metrics"]
|
|
},
|
|
"overlay_reduce_relay": {
|
|
"description": "Transaction reduce-relay efficiency gauge (MetricsRegistry.cpp:1354, peer-network dashboard). Backed by Overlay::txMetrics(); TxMetrics::json() emits txr_selected_cnt / txr_suppressed_cnt / txr_not_enabled_cnt unconditionally (TxMetrics.cpp:121-127), so the gauge always reports at least the selected_peers series.",
|
|
"metrics": ["reduce_relay_metrics"]
|
|
},
|
|
"overlay_overflow": {
|
|
"description": "Job-queue transaction overflow total (MetricsRegistry.cpp:609, job-queue dashboard). An ObservableCounter that reads Overlay::getJqTransOverflow() and Observes unconditionally, so the series exists at value 0 even when no overflow occurs.",
|
|
"metrics": ["jq_trans_overflow_total"]
|
|
},
|
|
"validation_lifetime_counters": {
|
|
"description": "Lifetime validation agreement/miss ObservableCounters (MetricsRegistry.cpp:1636, :1658, validator-health dashboard). Both callbacks reconcile the tracker and Observe unconditionally, so the series exist even on a node that has not yet agreed or missed (value 0). Only existence is asserted, never the value — validation_missed_total legitimately dominates on a non-validating node.",
|
|
"metrics": ["validation_agreements_total", "validation_missed_total"]
|
|
},
|
|
"not_asserted": {
|
|
"description": "Emitted-and-dashboarded metrics deliberately left unasserted because they are workload-gated or defect-gated: the harness workload cannot guarantee they appear, and a check that fails on a healthy run is worse than no check. This group has no \"metrics\" key, so validate_telemetry.py skips it (validate_metrics iterates category_data.get(\"metrics\", [])). Promote an entry into an asserted group only after the workload is changed to guarantee it.",
|
|
"metrics_excluded": {
|
|
"rpc_method_errored_total": "MetricsRegistry.cpp:354, push counter — needs an RPC that returns an error. rpc_load_generator.py issues only well-formed server_info / fee / ledger / ripple_path_find calls, so no series may ever be created.",
|
|
"ledger_history_mismatch_total": "MetricsRegistry.cpp:377, incremented only from LedgerHistory.cpp:332 on a built-vs-validated ledger mismatch. On a healthy run it never fires — asserting it would mean asserting a defect.",
|
|
"txq_expired_total": "MetricsRegistry.cpp:379, incremented only at TxQ.cpp:1428 when a queued tx expires past its LastLedgerSequence. CI does run a txq-burst phase (workload-profiles.json:41, 30 s of single-type Payment at 60 TPS), but that does not guarantee sustained fee escalation followed by expiry: a run in which every other check passed still exposed only txq_metrics and no txq_expired_total.",
|
|
"txq_dropped_total": "MetricsRegistry.cpp:381, incremented only at TxQ.cpp:1302 / :1347 on queue-full admission refusal. Same reason as txq_expired_total.",
|
|
"getobject_rejected_total": "GetObjectMetricNames.h:81, emitted from PeerImp.cpp:2725/:2743 only for a TMGetObjectByHash message refused as oversize or malformed_ledgerhash. A cooperating cluster never sends one.",
|
|
"getobject_request_objects": "GetObjectMetricNames.h:86, emitted from PeerImp.cpp:2926 only while serving an inbound TMGetObjectByHash. The XRPL_METRIC_* macros create their instrument lazily on first use (MetricMacros.h:174-285), so no series exists until a peer actually requests objects by hash — which a 5-node cluster started at genesis and already in sync may never do.",
|
|
"getobject_lookup_us": "GetObjectMetricNames.h:95, PeerImp.cpp:2929. Same lazy-creation and same inbound-request gate as getobject_request_objects.",
|
|
"getobject_lookups_total": "GetObjectMetricNames.h:100, PeerImp.cpp:2949/:2956. Same gate.",
|
|
"getobject_charge": "GetObjectMetricNames.h:105, PeerImp.cpp:2931. Same gate.",
|
|
"rpc_size_bytes": "ServerHandler.cpp:191, group('rpc')->makeEvent('size', Unit::Bytes). The OTLP Prometheus exporter derives the metric-name suffix from the declared unit, so a byte unit yields rpc_size_bytes. The Unit::Bytes declaration itself landed earlier, in 76c9051203; what 24094e427b changed was the exporter finally consuming it, replacing a hardcoded CreateDoubleHistogram(name, 'Duration in ms', 'ms') with otelUnitDescription(unit)/otelUnitCode(unit), and that is what renamed the series off rpc_size_milliseconds and the millisecond bucket ladder. Neither name was ever recorded here, so the harness could confirm neither the rename nor a regression back onto that ladder. Notified from ServerHandler::processRequest:1133, the HTTP JSON-RPC path — it computes an HTTP status and appends a trailing newline — and the load generators are WebSocket-only, the same gate regression-metrics.json:4 records for rpc.process, so only the harness's handful of HTTP health polls reach it. Real coverage needs an HTTP JSON-RPC phase in rpc_load_generator.py; that is a workload change rather than a harness correction, and is deliberately out of scope here.",
|
|
"rpc_time_milliseconds": "ServerHandler.cpp:192, group('rpc')->makeEvent('time') with the default millisecond unit. Notified from ServerHandler::processRequest:1129, the same HTTP JSON-RPC call site as rpc_size_bytes and behind the same WebSocket-only gate.",
|
|
"pathfind_fast_milliseconds": "PathRequestManager.h:35, makeEvent('pathfind_fast') with the default millisecond unit, so the exported form is the pathfind_fast_milliseconds_bucket/_count/_sum triple and there is no bare series — the same convention io_latency and rpc_method_us follow, and rpc-pathfinding queries the _bucket. Notified from PathRequestManager::reportFast (:81-84), whose only caller is PathRequest.cpp:852, inside the 'if (fast && quickReply_ == {})' branch of PathRequest::doUpdate. The only doUpdate call that passes fast=true is in PathRequest::doCreate (PathRequest.cpp:259), and it is guarded by '!hasCompletion()'. That guard is what makes this unreachable for the harness: ripple_path_find has two entry points and both construct the PathRequest with a completion function, so hasCompletion() (:161-164) is true and the fast pass is skipped. doRipplePathFind with no ledger specified goes to makeLegacyPathRequest, which passes the coroutine-post lambda (RipplePathFind.cpp:140-160); with a ledger specified it goes to doLegacyPathRequest, which passes an empty-body but non-null lambda (PathRequestManager.cpp:317). Only the path_find streaming subscription reaches reportFast: makePathRequest builds the request from a subscriber with no completion (PathRequestManager.cpp:261), so hasCompletion() is false. The load generator does not use path_find and says so at rpc_load_generator.py:72-73, so the 3% ripple_path_find weight cannot produce this metric at all — not rarely, never. That matches the Grafana Cloud reading of zero series in 180 days on devnet nodes serving live traffic. Covering it needs a path_find subscription phase in the generator, which is a workload change and out of scope here.",
|
|
"pathfind_full_milliseconds": "PathRequestManager.h:36, makeEvent('pathfind_full'); same histogram naming as pathfind_fast_milliseconds above. Notified from reportFull (:87-90) via PathRequest.cpp:857, the 'else if (!fast && fullReply_ == {})' branch. Unlike the fast event this one is reachable from the harness workload, but not on the request's own turn, which is why it is recorded here rather than asserted. The generator sends no ledger/ledger_index/ledger_hash (rpc_load_generator.py:343-350) and CI is not standalone, so doRipplePathFind takes the makeLegacyPathRequest branch, which only enqueues the request; the fast=false update happens later, when the next ledger close drives PathRequestManager::updateAll into its one-shot branch (PathRequestManager.cpp:160-166). Emission is therefore one ledger close behind the request and conditional on the request still being alive and still passing isValid() at that point (PathRequest.cpp:772-773) — none of which the harness arranges or checks, and a phase that ends before the next close records nothing. The synchronous route that would make this deterministic, doLegacyPathRequest calling doUpdate(cache, false) directly (PathRequestManager.cpp:321), needs an explicit ledger parameter the generator does not send. A probabilistic path is exactly what the doctrine at the top of this file excludes. Grafana Cloud shows zero series in 180 days on the devnet nodes, consistent with those nodes receiving no pathfinding RPC rather than with a broken exporter.",
|
|
"warn_total": "include/xrpl/resource/detail/Logic.h:41, makeMeter('warn'). makeMeter maps to CreateUInt64Counter (OTelCollector.cpp:878-881 -> :773-777), so the Prometheus exporter appends _total; the meter is created on the bare collector with no group, hence the unprefixed name. Incremented only at Logic.h:481, inside the 'if (notify)' branch reached when a consumer's balance crosses kWarningThreshold. A cooperating 5-node cluster plus a rate-limited load generator never charges a consumer that far, and Grafana Cloud confirms zero series in 180 days. Recorded explicitly because this was briefly mis-diagnosed as a phantom metric: the rpc-pathfinding panel that queries it is correct, and renders empty only because the condition has not occurred.",
|
|
"drop_total": "include/xrpl/resource/detail/Logic.h:42, makeMeter('drop'); same CreateUInt64Counter mapping and same _total suffix as warn_total. Incremented only at Logic.h:505, when a consumer's balance is at or above kDropThreshold and the connection is dropped. Grafana Cloud shows 2 live series, so unlike warn_total this one does fire in the wild — but only on a genuinely abusive consumer, which the harness deliberately does not create, so it is condition-gated all the same. Its rpc-pathfinding panel is likewise correct rather than phantom.",
|
|
"jobq_*_milliseconds, jobq_*_q_milliseconds": "This key is a pattern rather than a literal metric name — unlike every other entry in this map it stands for a whole family, one pair per job type. Created per job type in JobTypeData.h:97-98 from info.name() and info.name() + kSuffixQueued ('_q'), so the exported names are jobq_<jobtype>_milliseconds and jobq_<jobtype>_q_milliseconds with the job type lowercased by formatName. Which job types appear depends on which jobs a run happens to schedule, so no individual name is guaranteed. They are also rounded up to a whole millisecond at source (Event.h:47-51 applies ceil to a millisecond value type), which is why 6e2b2da772 moved the ledger-data-sync q-wait panels off jobq_<jobtype>_q_milliseconds_bucket onto job_queued_us_bucket — they are poor assertion targets regardless."
|
|
}
|
|
},
|
|
"grafana_dashboards": {
|
|
"description": "All 15 Grafana dashboards provisioned on disk under docker/telemetry/grafana/dashboards/ (UID == file stem for every one). validate_dashboards() checks that each UID resolves via GET /api/dashboards/uid/<uid> and reports its panel count — it verifies provisioning and loadability, not panel data. log-derived-insights is included on that basis even though its panels are Loki-backed and CI runs with --skip-loki: the dashboard itself must still provision cleanly. Its panel data is not asserted anywhere.",
|
|
"uids": [
|
|
"rpc-performance",
|
|
"transaction-overview",
|
|
"consensus-health",
|
|
"ledger-operations",
|
|
"peer-network",
|
|
"peer-quality",
|
|
"fee-market",
|
|
"job-queue",
|
|
"validator-health",
|
|
"node-health",
|
|
"network-traffic",
|
|
"rpc-pathfinding",
|
|
"overlay-traffic-detail",
|
|
"ledger-data-sync",
|
|
"log-derived-insights"
|
|
]
|
|
}
|
|
}
|