mirror of
https://github.com/XRPLF/rippled.git
synced 2026-08-21 22:30:57 +00:00
The harness manifests asserted things the code cannot produce and missed most of what it does. Two assertions were failing every run, and the metric set covered 16 of the ~41 emitted names. expected_spans.json: rpc.process was required with rpc.ws_message as its parent, but it is created only in ServerHandler::processRequest() on the HTTP path, so a WebSocket-only workload never produces it -- it is now optional and parented to rpc.http_request, and the rpc.process -> rpc.command.* edge is skipped with the real reason instead of a coroutine-context-loss diagnosis that was never the cause. Adds the missing rpc.ws_upgrade span, corrects four parents (consensus.mode_change, pathfind.request, and update_positions/check, which are children of consensus.establish rather than consensus.round), and demotes conditionally-set attributes out of required_attributes so a healthy run stops failing. Counts recomputed from the file: 41 span types, 62 unique required attributes. expected_metrics.json: 16 -> 52 asserted entries across the job-queue, RPC method, reduce-relay, overflow and validation families, plus the fifteenth dashboard uid. Metrics the harness workload cannot exercise -- erroring RPC, ledger-mismatch, TxQ overflow, and the lazily-created getobject_* instruments -- are listed in a not_asserted group the validator skips, rather than as assertions that would fail on a healthy node. The workflow's push trigger listed two globs matching nothing (include/xrpl/basics/Telemetry*.h, src/xrpld/app/misc/Telemetry*), so no C++ telemetry change ever triggered validation. Replaced with the paths the code actually lives in, including src/libxrpl/beast/insight/** for the insight export path the harness depends on. The four inert workflow_dispatch inputs are now labelled UNUSED rather than looking like working knobs. Docs: the workload README described a StatsD dirty-flag mechanism under a member name that does not exist, on a code path the harness never uses -- it sets [insight] server=otel, so gauges export through an observable-gauge callback every cycle. Adds the missing txq-burst phase, reconciles three different dashboard counts, and drops "posts summary to PR", which the workflow has no permission to do. The runbook's phase-10 section loses the last sampling_ratio reference (not a config key), gains a Regression Gate and CI subsection covering the gate that can fail CI, and its compose-logs command now names the workload compose file. cmake --preset default is left for a separate change: no CMakePresets.json is tracked, so it is wrong everywhere it appears. Also drops the dead exporter=otlp_http key the harness wrote into every node config, and stops capture_timings.py defaulting --profile to a profile that does not exist.
105 lines
5.1 KiB
Markdown
105 lines
5.1 KiB
Markdown
# Performance Baselines
|
|
|
|
This directory holds the committed baseline file used by the OTel-driven regression gate.
|
|
|
|
## How the gate works
|
|
|
|
After the validation suite runs, `capture_timings.py` queries Prometheus for the timings
|
|
declared in [`../regression-metrics.json`](../regression-metrics.json) and writes a
|
|
`timings.json`. Then `compare_to_baseline.py` reads [`baseline-timings.json`](./baseline-timings.json),
|
|
[`../regression-thresholds.json`](../regression-thresholds.json), and the captured
|
|
`timings.json`. The comparator picks one of two modes automatically:
|
|
|
|
- **Placeholder baseline** (`"placeholder": true` or empty `metrics`): the comparator
|
|
prints the captured timings JSON in exactly the format expected for this file, then
|
|
exits 0 without gating. This is how we bootstrap the baseline.
|
|
- **Populated baseline**: the comparator diffs per-metric, enforces the thresholds
|
|
(regression = current exceeds baseline on BOTH the percentage AND absolute bound),
|
|
and exits non-zero on any regression.
|
|
|
|
The regression gate runs against whatever workload profile `run-full-validation.sh`
|
|
was invoked with. Capture and comparison are profile-agnostic — they only read
|
|
Prometheus — so all existing profiles (`full-validation`, `quick-smoke`, `stress`)
|
|
continue to work unchanged.
|
|
|
|
## Bootstrapping the baseline
|
|
|
|
1. Merge a CI run with a `"placeholder": true` baseline. The telemetry-validation
|
|
workflow runs, fails no gate, and prints the captured timings block to the workflow
|
|
Step Summary under the heading `### Paste into baselines/baseline-timings.json`.
|
|
2. Open a new PR. Copy the full JSON block from the Step Summary (or download the
|
|
`timings.json` artifact) into this file, replacing the placeholder contents. The
|
|
JSON is emitted in the exact byte-for-byte format this file expects — sorted keys,
|
|
2-space indent, trailing newline.
|
|
3. The committed baseline PR needs reviewer approval just like any other code change.
|
|
This is the primary audit point for "who moved the performance bar."
|
|
|
|
## Refreshing the baseline
|
|
|
|
Refresh when a legitimate performance change lands on `develop` (for example, a
|
|
deliberate rewrite that changes a span's structure). The process is identical to
|
|
bootstrapping: run CI with the current baseline, inspect the delta, and if the
|
|
new numbers should become the norm, open a PR pasting the fresh timings into
|
|
`baseline-timings.json`. The reviewer decides whether the new baseline is acceptable.
|
|
|
|
Do **not** edit `baseline-timings.json` by hand outside of this process — every entry
|
|
should trace back to a real CI run so variance characteristics are preserved.
|
|
|
|
## Schema
|
|
|
|
```json
|
|
{
|
|
"schema_version": 1,
|
|
"captured_at": "2026-04-24T17:30:00Z",
|
|
"window": "3m",
|
|
"git_sha": "<SHA of the commit that produced these numbers>",
|
|
"profile": "<workload profile used>",
|
|
"metrics": {
|
|
"span.tx.process.p99": { "value": 12.4, "unit": "ms" },
|
|
"job.transaction.queued.p95": { "value": 1500.0, "unit": "us" }
|
|
}
|
|
}
|
|
```
|
|
|
|
Keys follow `{category}.{name}.p{quantile}`. Only two categories are actually
|
|
produced today — `span.*` and `job.*` — because `build_query_plan()` in
|
|
`prom_queries.py` reads the `spans` and `job_queue` groups of
|
|
`regression-metrics.json`, and that file defines only those two.
|
|
|
|
Placeholder baselines additionally include `"placeholder": true`. The comparator
|
|
detects this field (or an empty `metrics` object) to switch into "populate" mode
|
|
instead of enforcing thresholds. Remove the `placeholder` key when pasting real
|
|
captured timings.
|
|
|
|
Missing metrics (value `null`) in a captured run do not count as regressions. In
|
|
`regression-report.json`, `summary.missing_in_current` is a **count** only; the
|
|
identities are in the `metrics[]` array, as the entries whose `note` is
|
|
`"not captured in current run"`. Filter for those to see which keys went missing:
|
|
|
|
```bash
|
|
jq -r '.metrics[] | select(.note == "not captured in current run") | .key' \
|
|
/tmp/xrpld-validation/reports/regression-report.json
|
|
```
|
|
|
|
This keeps the gate robust when a profile doesn't exercise every span on every run.
|
|
|
|
## Known gap: no `rpc.*` metric can gate (FU-4)
|
|
|
|
Per-RPC-method timings are **not** gated, and would not gate even if they were
|
|
captured. Two independent blockers:
|
|
|
|
1. **Nothing emits an `rpc.*` key.** `build_query_plan()` in `prom_queries.py`
|
|
builds `rpc.*` entries from `cfg.get("rpc_methods", {})`, and
|
|
`regression-metrics.json` has no `rpc_methods` block — so the group resolves
|
|
to empty and no `rpc.*` key ever reaches `timings.json` or this baseline.
|
|
2. **Even a captured `rpc.*` key would silently not gate.** `resolve_thresholds()`
|
|
in `compare_to_baseline.py` maps the `rpc` category to the threshold group
|
|
`rpc_method`, but `regression-thresholds.json` defines only
|
|
`defaults.span` and `defaults.job_queue`. With no `rpc_method` block the
|
|
lookup returns `(None, None)`, which the comparator treats as "no threshold
|
|
configured" — the metric is reported but can never fail the build.
|
|
|
|
Closing this needs **both** an `rpc_methods` group in `regression-metrics.json`
|
|
and a `defaults.rpc_method` block in `regression-thresholds.json`. Adding only
|
|
the first produces metrics that look gated in the report but are not.
|