mirror of
https://github.com/XRPLF/rippled.git
synced 2026-09-27 07:26:51 +00:00
The harness killed and probed node directories named `node<N>`, but the
directories it creates are `validator-<N>` in run-full-validation.sh and
`bench-node-<N>` in benchmark.sh. Verified with pgrep against processes whose
command lines mimic the real ones: the pattern matched nothing either script
produces. Three consequences, all live:
- `--cleanup` deleted the workdir and left the xrpld processes running. They
are host processes, so the compose teardown does not reach them.
- The pre-run cleanup could not free the previous run's RPC, WS and peer
ports, which surfaces much later as a cluster that never reaches consensus.
- The startup crash fast-fail read a pid path that never exists, so its
`stopped > 0` branch was unreachable and a dead node waited out the full
120-attempt window.
Rather than patch four literals, derive every node path, kill pattern and log
glob from one NODE_PREFIX per script. The directory name is also the node's
identity: the collector's file_log receiver lifts that segment into
service.instance.id, so the directory and the [telemetry] service_instance_id
must agree. Deriving both from one value is what stops them drifting again.
Also in the same files, each confirmed by test rather than inspection:
- The collector readiness probe could never fail. curl -w '%{http_code}'
prints 000 on a refused connection and then exits non-zero, so the
`|| echo 000` inside the substitution appended a second 000 and the
"not ready" comparison never matched. Move the fallback outside.
- The generated config wrote [ips], the starter-list section. A loopback mesh
that must reach quorum is the [ips_fixed] case, which is what the variable,
the comment and the sibling cfg template already said.
- benchmark.sh returned exit 1 for a row it could not measure, though the
exit-code table reserves 1 for "every metric was measured and one breached".
Report 2 there instead.
- Five bc computations fell back to 0, which clears every threshold. The
guards beside them already fall back to the inconclusive token; these now
do too.
- A comment claimed a `|| guard` after a heredoc lands in the heredoc, and
that claim had removed a real guard from the config write. It does not: the
guard runs, and fires when cat fails.
- The EXIT trap was installed 88 lines before stop_workload was defined. If it
fired in that window, errexit aborted the handler on "command not found" and
the cluster reap never ran. Install it below both handlers.
- jq exits 5 on malformed JSON, outside this script's documented codes, so
read_metric now routes that through cannot_measure.
- --nodes and --duration were unvalidated, and --nodes 0 made the pid-count
guard compare 0 with 0 and pass, handing the sampler no pids at all.
- --cleanup now passes -v so the named tempo-data volume goes with it.
Otherwise the next run's Tempo still serves the previous run's traces and a
span assertion can be satisfied by them.
- Five messages reported an attempt count as seconds, though each attempt is
a sleep plus every node's probe.
The baselines README and the two regression JSON files had gone stale when the
baseline was refreshed to a three-run median: they described 20 gated keys and
five exclusions, against an actual 19 and six, and cited the superseded run,
date and commit. Re-derive every affected figure from the committed files. The
detection floors are recomputed (2.00x to 7.41x, so a 10x regression is now
caught on all 19 keys), the newly excluded span.ledger.build.p99 is documented,
and figures that no committed artifact can verify are either replaced with
derivable ones or labelled with their numerator.
No baseline value, threshold bound or derivation entry changes.
735 lines
28 KiB
Bash
Executable File
735 lines
28 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# benchmark.sh — Performance benchmark for rippled telemetry overhead.
|
|
#
|
|
# Runs the same client workload twice against a rippled cluster:
|
|
# 1. Baseline: telemetry disabled ([telemetry] enabled=0)
|
|
# 2. Telemetry: full telemetry enabled (traces + native OTel metrics)
|
|
#
|
|
# Both arms drive rpc_load_generator.py and tx_submitter.py at one fixed rate for
|
|
# the whole sample window, so the delta is attributable to telemetry rather than
|
|
# to a difference in offered load. The workload is not optional: with only the
|
|
# sampler's own ~1 request/sec of server_info, the tx.*, txq.*, transactor-stage
|
|
# and non-server_info rpc.command.* spans are never entered, and those are where
|
|
# a per-operation span cost shows up. A pass on an idle cluster says nothing.
|
|
#
|
|
# Compares CPU, memory, RPC latency, TPS, and mean consensus round time.
|
|
# Outputs a Markdown table with pass/fail against configured thresholds.
|
|
#
|
|
# Usage:
|
|
# ./benchmark.sh --xrpld /path/to/xrpld --duration 300
|
|
#
|
|
# Thresholds (configurable via environment variables):
|
|
# BENCH_CPU_OVERHEAD_PCT=3 CPU overhead < 3%
|
|
# BENCH_MEM_OVERHEAD_MB=5 Memory overhead < 5MB
|
|
# BENCH_RPC_LATENCY_IMPACT_MS=2 RPC p99 latency impact < 2ms
|
|
# BENCH_TPS_IMPACT_PCT=5 Throughput impact < 5%
|
|
# BENCH_CONSENSUS_IMPACT_PCT=1 Consensus round time impact < 1%
|
|
#
|
|
# Exit codes:
|
|
# 0 Every overhead metric was measured and is within its threshold.
|
|
# 1 Every overhead metric was measured and at least one exceeded its
|
|
# threshold. This is the only "telemetry is too expensive" signal.
|
|
# Also returned for a command-line usage error, which the pipeline
|
|
# cannot trigger (run-full-validation.sh passes a fixed flag list).
|
|
# 2 The overhead could not be measured at all — a missing prerequisite, a
|
|
# cluster that never reached consensus, or an incomplete metric
|
|
# collection. Nothing was compared, so nothing was breached.
|
|
#
|
|
# run-full-validation.sh depends on that split: it folds 1 into its own
|
|
# "checks failed" exit and 2 into its "infrastructure error" exit. Reporting a
|
|
# run that measured nothing as a threshold breach would be a false regression.
|
|
|
|
set -euo pipefail
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Colored output helpers
|
|
# ---------------------------------------------------------------------------
|
|
log() { printf "\033[1;34m[BENCH]\033[0m %s\n" "$*"; }
|
|
ok() { printf "\033[1;32m[BENCH]\033[0m %s\n" "$*"; }
|
|
warn() { printf "\033[1;33m[BENCH]\033[0m %s\n" "$*"; }
|
|
fail() { printf "\033[1;31m[BENCH]\033[0m %s\n" "$*"; }
|
|
|
|
# Usage error. Exit 1 by shell convention; see the exit-code block above.
|
|
die() {
|
|
fail "$*" >&2
|
|
exit 1
|
|
}
|
|
|
|
# Fatal, and no overhead figure was produced. Exit 2 keeps exit 1 exclusively
|
|
# for a measured threshold breach, so the caller never grades a run that
|
|
# measured nothing as a performance regression.
|
|
cannot_measure() {
|
|
fail "$*" >&2
|
|
exit 2
|
|
}
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Defaults and thresholds
|
|
# ---------------------------------------------------------------------------
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
REPO_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
|
|
|
# Configurable thresholds via environment variables.
|
|
CPU_THRESHOLD="${BENCH_CPU_OVERHEAD_PCT:-3}"
|
|
MEM_THRESHOLD="${BENCH_MEM_OVERHEAD_MB:-5}"
|
|
RPC_THRESHOLD="${BENCH_RPC_LATENCY_IMPACT_MS:-2}"
|
|
TPS_THRESHOLD="${BENCH_TPS_IMPACT_PCT:-5}"
|
|
CONSENSUS_THRESHOLD="${BENCH_CONSENSUS_IMPACT_PCT:-1}"
|
|
|
|
XRPLD="${BENCH_XRPLD:-$REPO_ROOT/.build/xrpld}"
|
|
DURATION=300
|
|
NUM_NODES=3
|
|
WORKDIR="/tmp/xrpld-benchmark"
|
|
RESULTS_DIR="$SCRIPT_DIR/benchmark-results"
|
|
RPC_PORT_BASE=5020
|
|
PEER_PORT_BASE=51250
|
|
# Above run-full-validation.sh's 6006.. so both harnesses can share a box.
|
|
WS_PORT_BASE=6020
|
|
|
|
# Name stem of each node's directory: node i lives in $WORKDIR/$NODE_PREFIX-$i.
|
|
# Every node path and kill pattern below derives from it, and so does the
|
|
# service_instance_id the telemetry arm reports. The telemetry arm exports to the
|
|
# same endpoint the workload harness uses, so a stem that differs from that
|
|
# harness's is what keeps the two clusters apart in the backend. Deriving both
|
|
# from one value is what stops the directory and the reported id drifting.
|
|
NODE_PREFIX="bench-node"
|
|
|
|
# Head start the generators get before the sampler opens its window.
|
|
# tx_submitter.py creates and funds eight accounts from genesis and then waits
|
|
# for those payments to validate, so without a lead the first seconds of every
|
|
# window carry no transaction load. Identical in both arms, so it cancels out.
|
|
WORKLOAD_LEAD_SEC=20
|
|
|
|
# One flat offered rate, not a workload-profiles.json profile: both arms must
|
|
# issue the same work for the delta to mean anything, and a profile's phase
|
|
# shaping only adds variance. Payment-only for the same reason -- a rejected
|
|
# transaction costs a different amount of work than an applied one.
|
|
WORKLOAD_RPC_RATE="${BENCH_RPC_RATE:-30}"
|
|
WORKLOAD_TX_TPS="${BENCH_TX_TPS:-3}"
|
|
|
|
# This arm's generator pids, reaped by wait_workload and killed by the trap.
|
|
WORKLOAD_PIDS=()
|
|
|
|
# Hard ceiling on every RPC probe below. curl applies no overall timeout of its
|
|
# own, so a node that accepts the connection and then stops answering parks the
|
|
# poll loop for the rest of the run. The loops here count attempts, not seconds,
|
|
# so without this their stated timeouts are not bounds at all.
|
|
CURL_MAX_TIME="${CURL_MAX_TIME:-5}"
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Argument parsing
|
|
# ---------------------------------------------------------------------------
|
|
usage() {
|
|
echo "Usage: $0 [OPTIONS]"
|
|
echo ""
|
|
echo "Options:"
|
|
echo " --xrpld PATH Path to xrpld binary (default: \$REPO_ROOT/.build/xrpld)"
|
|
echo " --duration SECS Benchmark duration per run (default: 300)"
|
|
echo " --nodes NUM Number of validator nodes (default: 3)"
|
|
echo " --output DIR Results output directory"
|
|
echo " -h, --help Show this help"
|
|
exit 0
|
|
}
|
|
|
|
while [ $# -gt 0 ]; do
|
|
case "$1" in
|
|
--xrpld)
|
|
XRPLD="$2"
|
|
shift 2
|
|
;;
|
|
--duration)
|
|
DURATION="$2"
|
|
shift 2
|
|
;;
|
|
--nodes)
|
|
NUM_NODES="$2"
|
|
shift 2
|
|
;;
|
|
--output)
|
|
RESULTS_DIR="$2"
|
|
shift 2
|
|
;;
|
|
-h | --help) usage ;;
|
|
*) die "Unknown option: $1" ;;
|
|
esac
|
|
done
|
|
|
|
# Validate prerequisites. A missing binary or tool means no measurement can be
|
|
# taken, which is "cannot measure", not "too slow".
|
|
# Validate the two numeric options before anything derives ports, loop counts or
|
|
# pid-count guards from them. --nodes 0 is the dangerous one: node_pids_csv then
|
|
# yields an empty list, awk reports 0 fields, and the "found N of NUM_NODES pids"
|
|
# guard compares 0 with 0 and passes, handing the sampler no pids at all.
|
|
case "$NUM_NODES" in
|
|
'' | *[!0-9]*) cannot_measure "--nodes must be a positive integer, got '$NUM_NODES'" ;;
|
|
esac
|
|
[ "$NUM_NODES" -ge 1 ] || cannot_measure "--nodes must be at least 1, got '$NUM_NODES'"
|
|
case "$DURATION" in
|
|
'' | *[!0-9]*) cannot_measure "--duration must be a positive integer, got '$DURATION'" ;;
|
|
esac
|
|
[ "$DURATION" -ge 1 ] || cannot_measure "--duration must be at least 1, got '$DURATION'"
|
|
|
|
[ -x "$XRPLD" ] || cannot_measure "xrpld not found at $XRPLD"
|
|
command -v jq >/dev/null 2>&1 || cannot_measure "jq not found"
|
|
command -v bc >/dev/null 2>&1 || cannot_measure "bc not found"
|
|
command -v curl >/dev/null 2>&1 || cannot_measure "curl not found"
|
|
command -v python3 >/dev/null 2>&1 ||
|
|
cannot_measure "python3 not found (the load generators need it)"
|
|
python3 -c 'import websockets' 2>/dev/null ||
|
|
cannot_measure "python3 'websockets' package not found -- pip install -r $SCRIPT_DIR/requirements.txt"
|
|
|
|
mkdir -p "$RESULTS_DIR" || cannot_measure "Could not create the results directory $RESULTS_DIR"
|
|
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Node cluster management
|
|
# ---------------------------------------------------------------------------
|
|
|
|
# True while xrpld children spawned by start_cluster may still be alive.
|
|
# Read by stop_cluster so it can be called any number of times.
|
|
CLUSTER_RUNNING=false
|
|
|
|
start_cluster() {
|
|
local telemetry_enabled="$1"
|
|
local label="$2"
|
|
|
|
log "Starting $NUM_NODES-node cluster ($label, telemetry=$telemetry_enabled)..."
|
|
|
|
rm -rf "$WORKDIR" || cannot_measure "Could not clear the workdir $WORKDIR"
|
|
mkdir -p "$WORKDIR" || cannot_measure "Could not create the workdir $WORKDIR"
|
|
|
|
# Generate keys using a temporary standalone node. The helper fails through
|
|
# its own die(), which exits 1 -- the code this script reserves for a
|
|
# measured breach. Remap it, or a keygen failure reads as "too expensive".
|
|
bash "$SCRIPT_DIR/generate-validator-keys.sh" "$XRPLD" "$NUM_NODES" "$WORKDIR" ||
|
|
cannot_measure "generate-validator-keys.sh failed; no keys for the $NUM_NODES-node cluster"
|
|
|
|
# Set before the spawn loop so a failure part-way through it still gets
|
|
# cleaned up by the EXIT trap.
|
|
CLUSTER_RUNNING=true
|
|
|
|
# Build per-node configs.
|
|
for i in $(seq 1 "$NUM_NODES"); do
|
|
local node_dir="$WORKDIR/$NODE_PREFIX-$i"
|
|
mkdir -p "$node_dir/nudb" "$node_dir/db" ||
|
|
cannot_measure "Could not create node$i directories under $node_dir"
|
|
|
|
local rpc_port
|
|
rpc_port=$((RPC_PORT_BASE + i - 1))
|
|
local peer_port
|
|
peer_port=$((PEER_PORT_BASE + i - 1))
|
|
local ws_port
|
|
ws_port=$((WS_PORT_BASE + i - 1))
|
|
# Split from the declaration on purpose: `local seed=$(...)` reports
|
|
# local's status, not jq's, so the guard below would never fire.
|
|
local seed
|
|
seed=$(jq -r ".[$((i - 1))].seed" "$WORKDIR/validator-keys.json") ||
|
|
cannot_measure "Could not read node$i's seed from $WORKDIR/validator-keys.json"
|
|
# jq prints "null" and exits 0 when the array is short, so the exit
|
|
# status alone does not catch a truncated key file.
|
|
case "$seed" in
|
|
"" | null)
|
|
cannot_measure "node$i has no seed in $WORKDIR/validator-keys.json"
|
|
;;
|
|
esac
|
|
|
|
# Build ips_fixed list.
|
|
local ips_fixed=""
|
|
for j in $(seq 1 "$NUM_NODES"); do
|
|
if [ "$j" -ne "$i" ]; then
|
|
ips_fixed="${ips_fixed}127.0.0.1 $((PEER_PORT_BASE + j - 1))
|
|
"
|
|
fi
|
|
done
|
|
|
|
# Build telemetry section.
|
|
local telemetry_section=""
|
|
if [ "$telemetry_enabled" = "1" ]; then
|
|
telemetry_section="
|
|
[telemetry]
|
|
enabled=1
|
|
service_instance_id=$NODE_PREFIX-${i}
|
|
endpoint=http://localhost:4318/v1/traces
|
|
exporter=otlp_http
|
|
batch_size=512
|
|
batch_delay_ms=2000
|
|
max_queue_size=2048
|
|
trace_rpc=1
|
|
trace_transactions=1
|
|
trace_consensus=1
|
|
trace_peer=1
|
|
trace_ledger=1
|
|
|
|
[insight]
|
|
# No prefix is set on purpose: OTelCollector applies no prefix to instrument
|
|
# names, so setting one here would be inert. Same reasoning as
|
|
# run-full-validation.sh.
|
|
server=otel
|
|
endpoint=http://localhost:4318/v1/metrics"
|
|
else
|
|
telemetry_section="
|
|
[telemetry]
|
|
enabled=0"
|
|
fi
|
|
|
|
# Guarded like every other fallible write. An unwritable node_dir would
|
|
# otherwise surface as cat's status 1, the code this script reserves for a
|
|
# measured threshold breach.
|
|
cat >"$node_dir/xrpld.cfg" <<EOCFG || cannot_measure "Could not write $node_dir/xrpld.cfg"
|
|
[server]
|
|
port_rpc
|
|
port_ws
|
|
port_peer
|
|
|
|
[port_rpc]
|
|
port = $rpc_port
|
|
ip = 127.0.0.1
|
|
admin = 127.0.0.1
|
|
protocol = http
|
|
|
|
# The generators speak WebSocket only. admin is required rather than cosmetic:
|
|
# tx_submitter.py calls wallet_propose and submits with "secret". Declared in
|
|
# BOTH arms, so the listener itself is not part of the measured delta.
|
|
[port_ws]
|
|
port = $ws_port
|
|
ip = 127.0.0.1
|
|
admin = 127.0.0.1
|
|
protocol = ws
|
|
|
|
[port_peer]
|
|
port = $peer_port
|
|
ip = 0.0.0.0
|
|
protocol = peer
|
|
|
|
[node_db]
|
|
type=NuDB
|
|
path=$node_dir/nudb
|
|
online_delete=256
|
|
|
|
[database_path]
|
|
$node_dir/db
|
|
|
|
[debug_logfile]
|
|
$node_dir/debug.log
|
|
|
|
[validation_seed]
|
|
$seed
|
|
|
|
[validators_file]
|
|
$WORKDIR/validators.txt
|
|
|
|
[ips_fixed]
|
|
${ips_fixed}
|
|
[peer_private]
|
|
1
|
|
${telemetry_section}
|
|
|
|
[rpc_startup]
|
|
# The log level deliberately diverges from run-full-validation.sh, which uses
|
|
# info. This script measures telemetry OVERHEAD as the delta
|
|
# between a telemetry-off and a telemetry-on arm, so it needs a quiet, identical
|
|
# baseline in both arms rather than good observability. Logging is synchronous, so
|
|
# admitting more log lines would add I/O to both arms and inflate the CPU,
|
|
# memory and latency figures the thresholds gate. Do not "align" this with the
|
|
# workload harness -- that would silently degrade the measurement.
|
|
{ "command": "log_level", "severity": "warning" }
|
|
|
|
[ssl_verify]
|
|
0
|
|
EOCFG
|
|
|
|
"$XRPLD" --conf "$node_dir/xrpld.cfg" --start >"$node_dir/stdout.log" 2>&1 &
|
|
echo $! >"$node_dir/xrpld.pid"
|
|
done
|
|
|
|
# Wait for consensus. Reaching the limit is fatal: numbers taken from a
|
|
# cluster that never got to "proposing" would silently corrupt the
|
|
# baseline-vs-telemetry comparison.
|
|
local max_wait=120
|
|
log "Waiting for consensus (up to ${max_wait}s)..."
|
|
for attempt in $(seq 1 "$max_wait"); do
|
|
local ready=0
|
|
for i in $(seq 1 "$NUM_NODES"); do
|
|
local port
|
|
port=$((RPC_PORT_BASE + i - 1))
|
|
local state
|
|
state=$(curl -sf --max-time "$CURL_MAX_TIME" "http://localhost:$port" \
|
|
-d '{"method":"server_info"}' 2>/dev/null |
|
|
jq -r '.result.info.server_state' 2>/dev/null || echo "")
|
|
if [ "$state" = "proposing" ]; then
|
|
ready=$((ready + 1))
|
|
fi
|
|
done
|
|
if [ "$ready" -ge "$NUM_NODES" ]; then
|
|
ok "All $NUM_NODES nodes proposing (attempt $attempt)"
|
|
break
|
|
fi
|
|
if [ "$attempt" -eq "$max_wait" ]; then
|
|
cannot_measure "Consensus timeout — only $ready/$NUM_NODES nodes proposing after ${max_wait}s"
|
|
fi
|
|
sleep 1
|
|
done
|
|
|
|
# Let the cluster stabilize.
|
|
sleep 5
|
|
}
|
|
|
|
stop_cluster() {
|
|
# Idempotent. The happy path calls this directly and the EXIT trap calls
|
|
# it again, so a second call must not re-kill or log a misleading message.
|
|
[ "$CLUSTER_RUNNING" = true ] || return 0
|
|
CLUSTER_RUNNING=false
|
|
|
|
log "Stopping cluster..."
|
|
for i in $(seq 1 "$NUM_NODES"); do
|
|
local pidfile="$WORKDIR/$NODE_PREFIX-$i/xrpld.pid"
|
|
if [ -f "$pidfile" ]; then
|
|
kill "$(cat "$pidfile")" 2>/dev/null || true
|
|
fi
|
|
done
|
|
# Belt and braces for a node whose pidfile is missing or stale. Matched on the
|
|
# per-node config path start_cluster launches with, not the workdir alone: a
|
|
# workdir-only pattern would also match anything that merely mentions the
|
|
# tree, such as a developer's `tail -f` on a node's debug.log.
|
|
pkill -f "$WORKDIR/$NODE_PREFIX-[0-9]+/xrpld\.cfg" 2>/dev/null || true
|
|
|
|
# Guarded on purpose. This runs as the EXIT trap, where any unguarded
|
|
# failure makes `set -e` exit with that command's status and discard the
|
|
# status the script meant to report — a threshold breach would surface as
|
|
# a plain 1 and "cannot measure" would lose its 2. With every command here
|
|
# guarded, the explicit `return 0` below is reachable and authoritative.
|
|
sleep 3 || true
|
|
|
|
return 0
|
|
}
|
|
|
|
# Build RPC ports CSV string.
|
|
rpc_ports_csv() {
|
|
local ports=""
|
|
for i in $(seq 1 "$NUM_NODES"); do
|
|
[ -n "$ports" ] && ports="$ports,"
|
|
ports="$ports$((RPC_PORT_BASE + i - 1))"
|
|
done
|
|
echo "$ports"
|
|
}
|
|
|
|
# Collects one leg of the benchmark.
|
|
#
|
|
# The collector exits non-zero when it cannot run (1) or when a measurement
|
|
# source came back empty (3). An all-zero or partial sample set clears every
|
|
# threshold, so an incomplete leg aborts with "cannot measure" instead of being
|
|
# compared and passed.
|
|
# Echoes one ws:// endpoint per node, space separated.
|
|
ws_endpoints() {
|
|
local i out=""
|
|
for i in $(seq 1 "$NUM_NODES"); do
|
|
out="$out ws://localhost:$((WS_PORT_BASE + i - 1))"
|
|
done
|
|
printf '%s' "${out# }"
|
|
}
|
|
|
|
# This cluster's xrpld pids, comma separated, for the sampler's process filter.
|
|
# Without it the sampler matches every xrpld on the host: run-full-validation.sh
|
|
# leaves five validation nodes running while these three start, so both arms
|
|
# average eight processes and the delta is diluted away.
|
|
node_pids_csv() {
|
|
local i out="" pid
|
|
for i in $(seq 1 "$NUM_NODES"); do
|
|
pid=$(cat "$WORKDIR/$NODE_PREFIX-$i/xrpld.pid" 2>/dev/null) || continue
|
|
[ -n "$pid" ] && out="$out,$pid"
|
|
done
|
|
printf '%s' "${out#,}"
|
|
}
|
|
|
|
# Starts this arm's generators, then waits out the funding lead so the sampler's
|
|
# whole window is under load. Logs and JSON summaries go to RESULTS_DIR, not
|
|
# WORKDIR: the next arm's start_cluster rm -rf's WORKDIR.
|
|
start_workload() {
|
|
local label="$1"
|
|
local gen_duration=$((DURATION + WORKLOAD_LEAD_SEC))
|
|
local logdir="$RESULTS_DIR/workload-${TIMESTAMP}"
|
|
mkdir -p "$logdir" || cannot_measure "Could not create the workload log dir $logdir"
|
|
|
|
log "Starting workload ($label): ${WORKLOAD_RPC_RATE} rpc/s + ${WORKLOAD_TX_TPS} tps for ${gen_duration}s..."
|
|
|
|
# shellcheck disable=SC2046 # --endpoints takes a list; splitting is intended
|
|
python3 "$SCRIPT_DIR/rpc_load_generator.py" \
|
|
--endpoints $(ws_endpoints) \
|
|
--rate "$WORKLOAD_RPC_RATE" \
|
|
--duration "$gen_duration" \
|
|
--output "$logdir/$label-rpc.json" \
|
|
>"$logdir/$label-rpc.log" 2>&1 &
|
|
WORKLOAD_PIDS+=("$!")
|
|
|
|
python3 "$SCRIPT_DIR/tx_submitter.py" \
|
|
--endpoint "ws://localhost:$WS_PORT_BASE" \
|
|
--tps "$WORKLOAD_TX_TPS" \
|
|
--duration "$gen_duration" \
|
|
--weights '{"Payment": 100}' \
|
|
--output "$logdir/$label-tx.json" \
|
|
>"$logdir/$label-tx.log" 2>&1 &
|
|
WORKLOAD_PIDS+=("$!")
|
|
|
|
sleep "$WORKLOAD_LEAD_SEC"
|
|
}
|
|
|
|
# Reaps this arm's generators. A non-zero generator means the arms did not do
|
|
# the same work, so nothing is attributable: "cannot measure", never "too slow".
|
|
wait_workload() {
|
|
local label="$1"
|
|
local pid status=0
|
|
for pid in ${WORKLOAD_PIDS[@]+"${WORKLOAD_PIDS[@]}"}; do
|
|
wait "$pid" || status=$?
|
|
done
|
|
WORKLOAD_PIDS=()
|
|
[ "$status" -eq 0 ] ||
|
|
cannot_measure "$label workload generator exited $status; the arms did not do identical work -- see $RESULTS_DIR/workload-${TIMESTAMP}/"
|
|
}
|
|
|
|
# Kills any generator still running, so an aborted arm leaves no python3 holding
|
|
# WebSocket connections. Guarded like stop_cluster's commands: this runs from the
|
|
# EXIT trap, where an unguarded failure would discard the real exit status.
|
|
stop_workload() {
|
|
local pid
|
|
for pid in ${WORKLOAD_PIDS[@]+"${WORKLOAD_PIDS[@]}"}; do
|
|
kill "$pid" 2>/dev/null || true
|
|
done
|
|
WORKLOAD_PIDS=()
|
|
return 0
|
|
}
|
|
|
|
# Reap the workload and the cluster on every exit path. Installed below both
|
|
# handlers, not after argument parsing: if the trap fires while either name is
|
|
# still undefined, errexit aborts the handler on "command not found" and the
|
|
# rest of it -- the cluster reap -- never runs. Without the trap, any failure
|
|
# between start_cluster and stop_cluster leaks the xrpld children along with
|
|
# their RPC ports (5020+) and peer ports (51250+).
|
|
trap 'stop_workload; stop_cluster' EXIT
|
|
|
|
collect_metrics() {
|
|
local label="$1"
|
|
local out_file="$2"
|
|
|
|
local status=0
|
|
# An empty or short list would silently fall back to host-wide sampling,
|
|
# which is the outcome the pid argument exists to prevent.
|
|
local pids
|
|
pids=$(node_pids_csv)
|
|
local n_pids
|
|
n_pids=$(printf '%s' "$pids" | awk -F, '{print NF}')
|
|
[ "${n_pids:-0}" -eq "$NUM_NODES" ] ||
|
|
cannot_measure "$label: found $n_pids of $NUM_NODES node pids, so the sampler cannot be scoped to this cluster"
|
|
|
|
bash "$SCRIPT_DIR/collect_system_metrics.sh" \
|
|
"$(rpc_ports_csv)" "$DURATION" "$out_file" "$pids" || status=$?
|
|
[ "$status" -eq 0 ] ||
|
|
cannot_measure "$label metric collection failed (exit $status) — refusing to compare an incomplete run"
|
|
|
|
# Only an explicit false counts. The flag is absent from older artifacts,
|
|
# and jq's "//" operator would turn a real false into the default.
|
|
local complete
|
|
complete=$(jq -r '.metrics_complete' "$out_file" 2>/dev/null || echo "null")
|
|
[ "$complete" != "false" ] ||
|
|
cannot_measure "$label metrics are flagged incomplete — refusing to compare an incomplete run"
|
|
}
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Run benchmark
|
|
# ---------------------------------------------------------------------------
|
|
log "="
|
|
log " rippled Telemetry Performance Benchmark"
|
|
log " Nodes: $NUM_NODES | Duration: ${DURATION}s | Binary: $XRPLD"
|
|
log "="
|
|
|
|
# --- Baseline run ---
|
|
BASELINE_FILE="$RESULTS_DIR/baseline-${TIMESTAMP}.json"
|
|
start_cluster "0" "baseline"
|
|
start_workload "baseline"
|
|
collect_metrics "baseline" "$BASELINE_FILE"
|
|
wait_workload "baseline"
|
|
stop_cluster
|
|
|
|
# --- Telemetry run ---
|
|
TELEMETRY_FILE="$RESULTS_DIR/telemetry-${TIMESTAMP}.json"
|
|
start_cluster "1" "telemetry"
|
|
start_workload "telemetry"
|
|
collect_metrics "telemetry" "$TELEMETRY_FILE"
|
|
wait_workload "telemetry"
|
|
stop_cluster
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Compare results
|
|
# ---------------------------------------------------------------------------
|
|
log "Comparing results..."
|
|
|
|
# Written into an impact variable when its baseline could not be used.
|
|
# check_threshold turns this into an INCONCLUSIVE verdict.
|
|
INCONCLUSIVE="n/a"
|
|
|
|
read_metric() {
|
|
local file="$1"
|
|
local key="$2"
|
|
# jq exits 5 on malformed JSON and 2 on a missing file. Neither is one of
|
|
# this script's three documented codes, and a bare assignment would let it
|
|
# escape through errexit, so route both through cannot_measure.
|
|
jq -r ".$key // 0" "$file" 2>/dev/null ||
|
|
cannot_measure "$file is not readable JSON — refusing to compare an unreadable run"
|
|
}
|
|
|
|
BASE_CPU=$(read_metric "$BASELINE_FILE" "cpu_pct_avg")
|
|
TELE_CPU=$(read_metric "$TELEMETRY_FILE" "cpu_pct_avg")
|
|
CPU_DELTA=$(echo "scale=2; $TELE_CPU - $BASE_CPU" | bc 2>/dev/null || echo "$INCONCLUSIVE")
|
|
|
|
BASE_MEM=$(read_metric "$BASELINE_FILE" "memory_rss_mb_peak")
|
|
TELE_MEM=$(read_metric "$TELEMETRY_FILE" "memory_rss_mb_peak")
|
|
MEM_DELTA=$(echo "scale=2; $TELE_MEM - $BASE_MEM" | bc 2>/dev/null || echo "$INCONCLUSIVE")
|
|
|
|
BASE_RPC=$(read_metric "$BASELINE_FILE" "rpc_p99_ms")
|
|
TELE_RPC=$(read_metric "$TELEMETRY_FILE" "rpc_p99_ms")
|
|
RPC_DELTA=$(echo "scale=2; $TELE_RPC - $BASE_RPC" | bc 2>/dev/null || echo "$INCONCLUSIVE")
|
|
|
|
# Both impacts below are ratios of the baseline, so a non-positive baseline
|
|
# leaves them undefined. Reporting that as "0% impact" would clear the
|
|
# threshold and hide a failed baseline run. The reachable case is not the
|
|
# collector's tps=0 placeholder -- that marks the run incomplete and stops it
|
|
# earlier -- but quantization: tps is printed to two decimals, so a barely
|
|
# advancing cluster yields "0.00" with the metrics still flagged complete.
|
|
#
|
|
# Both expressions scale by 100 before dividing. bc truncates at "scale" after
|
|
# every operation, so dividing first would floor the ratio to 2 decimals and
|
|
# then multiply the lost precision by 100 — a real 1.25% consensus impact came
|
|
# out as exactly 1.00 and passed the 1% threshold.
|
|
BASE_TPS=$(read_metric "$BASELINE_FILE" "tps")
|
|
TELE_TPS=$(read_metric "$TELEMETRY_FILE" "tps")
|
|
if [[ "$(echo "$BASE_TPS > 0" | bc 2>/dev/null)" = "1" ]]; then
|
|
TPS_IMPACT=$(echo "scale=2; ($BASE_TPS - $TELE_TPS) * 100 / $BASE_TPS" | bc 2>/dev/null || echo "$INCONCLUSIVE")
|
|
else
|
|
TPS_IMPACT="$INCONCLUSIVE"
|
|
fi
|
|
|
|
BASE_CONS=$(read_metric "$BASELINE_FILE" "consensus_round_mean_ms")
|
|
TELE_CONS=$(read_metric "$TELEMETRY_FILE" "consensus_round_mean_ms")
|
|
if [[ "$(echo "$BASE_CONS > 0" | bc 2>/dev/null)" = "1" ]]; then
|
|
CONS_IMPACT=$(echo "scale=2; ($TELE_CONS - $BASE_CONS) * 100 / $BASE_CONS" | bc 2>/dev/null || echo "$INCONCLUSIVE")
|
|
else
|
|
CONS_IMPACT="$INCONCLUSIVE"
|
|
fi
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Pass/fail checks
|
|
# ---------------------------------------------------------------------------
|
|
PASS_COUNT=0
|
|
FAIL_COUNT=0
|
|
INCONCLUSIVE_COUNT=0
|
|
|
|
# Records the verdict for one row of the report.
|
|
#
|
|
# Arguments: metric name, measured value, threshold, unit, and the name of the
|
|
# variable to write the bare verdict into.
|
|
#
|
|
# The verdict travels through that named variable and every diagnostic goes to
|
|
# stderr. Calling this through a command substitution would run it in a
|
|
# subshell, which drops the counter updates and captures the colored log line
|
|
# into the caller's variable.
|
|
check_threshold() {
|
|
local name="$1"
|
|
local actual="$2"
|
|
local threshold="$3"
|
|
local unit="$4"
|
|
local result_var="$5"
|
|
|
|
# Unusable measurement. Counted as a failure so the exit gate fires: an
|
|
# undefined result must never read as a pass.
|
|
if [ "$actual" = "$INCONCLUSIVE" ]; then
|
|
fail "$name: INCONCLUSIVE — the baseline was zero or missing, or the arithmetic failed" >&2
|
|
FAIL_COUNT=$((FAIL_COUNT + 1))
|
|
INCONCLUSIVE_COUNT=$((INCONCLUSIVE_COUNT + 1))
|
|
printf -v "$result_var" 'INCONCLUSIVE'
|
|
return
|
|
fi
|
|
|
|
# Compare: actual <= threshold
|
|
if [[ "$(echo "$actual <= $threshold" | bc 2>/dev/null)" = "1" ]]; then
|
|
ok "$name: ${actual}${unit} <= ${threshold}${unit} PASS" >&2
|
|
PASS_COUNT=$((PASS_COUNT + 1))
|
|
printf -v "$result_var" 'PASS'
|
|
else
|
|
fail "$name: ${actual}${unit} > ${threshold}${unit} FAIL" >&2
|
|
FAIL_COUNT=$((FAIL_COUNT + 1))
|
|
printf -v "$result_var" 'FAIL'
|
|
fi
|
|
}
|
|
|
|
# Formats a delta for the report table: appends the unit to a real number and
|
|
# leaves the INCONCLUSIVE placeholder bare.
|
|
fmt_delta() {
|
|
local value="$1"
|
|
local unit="$2"
|
|
if [ "$value" = "$INCONCLUSIVE" ]; then
|
|
printf '%s' "$value"
|
|
else
|
|
printf '%s%s' "$value" "$unit"
|
|
fi
|
|
}
|
|
|
|
check_threshold "CPU overhead" "$CPU_DELTA" "$CPU_THRESHOLD" "%" CPU_RESULT
|
|
check_threshold "Memory overhead" "$MEM_DELTA" "$MEM_THRESHOLD" "MB" MEM_RESULT
|
|
check_threshold "RPC p99 impact" "$RPC_DELTA" "$RPC_THRESHOLD" "ms" RPC_RESULT
|
|
check_threshold "TPS impact" "$TPS_IMPACT" "$TPS_THRESHOLD" "%" TPS_RESULT
|
|
check_threshold "Consensus impact" "$CONS_IMPACT" "$CONSENSUS_THRESHOLD" "%" CONS_RESULT
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Output Markdown table
|
|
# ---------------------------------------------------------------------------
|
|
REPORT_FILE="$RESULTS_DIR/benchmark-report-${TIMESTAMP}.md"
|
|
|
|
cat >"$REPORT_FILE" <<EOMD
|
|
# Telemetry Performance Benchmark Report
|
|
|
|
**Date**: $(date -u +"%Y-%m-%d %H:%M:%S UTC")
|
|
**Nodes**: $NUM_NODES | **Duration**: ${DURATION}s per run
|
|
**Binary**: $XRPLD
|
|
|
|
## Results
|
|
|
|
| Metric | Baseline | Telemetry | Delta | Threshold | Result |
|
|
|--------|----------|-----------|-------|-----------|--------|
|
|
| CPU (avg %) | ${BASE_CPU}% | ${TELE_CPU}% | ${CPU_DELTA}% | < ${CPU_THRESHOLD}% | ${CPU_RESULT} |
|
|
| Memory RSS (peak MB) | ${BASE_MEM} MB | ${TELE_MEM} MB | ${MEM_DELTA} MB | < ${MEM_THRESHOLD} MB | ${MEM_RESULT} |
|
|
| RPC p99 Latency (ms) | ${BASE_RPC} ms | ${TELE_RPC} ms | ${RPC_DELTA} ms | < ${RPC_THRESHOLD} ms | ${RPC_RESULT} |
|
|
| Throughput (TPS) | ${BASE_TPS} | ${TELE_TPS} | $(fmt_delta "$TPS_IMPACT" "%") | < ${TPS_THRESHOLD}% | ${TPS_RESULT} |
|
|
| Consensus Round Mean (ms) | ${BASE_CONS} ms | ${TELE_CONS} ms | $(fmt_delta "$CONS_IMPACT" "%") | < ${CONSENSUS_THRESHOLD}% | ${CONS_RESULT} |
|
|
|
|
\`INCONCLUSIVE\` means the baseline for that row was zero or missing, so the
|
|
impact could not be computed. Such rows count as failures.
|
|
|
|
\`Consensus Round Mean\` is the mean inter-ledger interval derived from the
|
|
collector's 2 s ledger-sequence samples, not a percentile.
|
|
|
|
## Summary
|
|
|
|
- **Passed**: $PASS_COUNT / $((PASS_COUNT + FAIL_COUNT))
|
|
- **Failed**: $FAIL_COUNT / $((PASS_COUNT + FAIL_COUNT))
|
|
- **Inconclusive**: $INCONCLUSIVE_COUNT (included in Failed)
|
|
|
|
## Raw Data
|
|
|
|
- Baseline: \`$(basename "$BASELINE_FILE")\`
|
|
- Telemetry: \`$(basename "$TELEMETRY_FILE")\`
|
|
EOMD
|
|
|
|
ok "Benchmark report written to $REPORT_FILE"
|
|
cat "$REPORT_FILE"
|
|
|
|
# Exit per the code table at the top of this file. A breach and an unmeasurable
|
|
# row both have to fail the gate, but they are different codes: exit 1 promises
|
|
# every metric was measured, so a run carrying an INCONCLUSIVE row reports 2
|
|
# instead. FAIL_COUNT includes the inconclusive rows, so subtract them to see
|
|
# whether any real threshold was exceeded.
|
|
if [ "$((FAIL_COUNT - INCONCLUSIVE_COUNT))" -gt 0 ]; then
|
|
exit 1
|
|
elif [ "$INCONCLUSIVE_COUNT" -gt 0 ]; then
|
|
fail "$INCONCLUSIVE_COUNT overhead row(s) could not be measured — reporting an infrastructure error, not a breach" >&2
|
|
exit 2
|
|
fi
|