From 85fc1c5399f1c79b896f129f0eeae6180a9a258b Mon Sep 17 00:00:00 2001 From: Pratik Mankawde <3397372+pratikmankawde@users.noreply.github.com> Date: Thu, 17 Sep 2026 17:01:29 +0100 Subject: [PATCH] fix(telemetry): keep the workload cluster on [ips], not [ips_fixed] The previous commit switched the generated node config from [ips] to [ips_fixed] on the grounds that the variable, the comment and the sibling cfg template all named ips_fixed, and that ips_fixed is the section whose documented meaning fits a private cluster. Both of those are still true. The switch is reverted anyway, because it is a workload change rather than a naming fix. Measured on CI, parent commit against this branch's previous tip, one functional config line apart: span.consensus.ledger_close.p95 0.57 ms -> 6.43 ms (tripped the gate) span.consensus.ledger_close.p99 0.94 ms -> 9.50 ms span.consensus.accept.p50 0.97 ms -> 2.63 ms span.tx.process.p50 0.36 ms -> 0.18 ms (faster) job.acceptLedger.running.p95 21157 us -> 10938 us (faster) Every consensus-path span rose and every transaction-path metric fell, which is the shape a denser always-connected mesh produces and not the shape of run-to-run variance. [ips_fixed] holds connections open to all four peers instead of treating the list as a discovery hint, so each node processes proposals and validations from the full mesh every round. Nothing else in that commit touches the consensus path: the emitted config differed in exactly three lines, of which one is a die message and one expands to an identical string. The committed baseline describes the [ips] topology. Adopting [ips_fixed] therefore needs a refreshed baseline and re-derived bounds, which is the process baselines/README.md already documents for a workload change. Left as its own work item rather than smuggled in behind a section rename, and the reason is now recorded beside the line so it is not repeated. This also falsified a claim the previous commit had written into baselines/README.md and regression-thresholds.json: that none of the six weakly-guarded keys fires on any observed run. Corrected in both, and the measurement above is cited in place of the absolute. --- docker/telemetry/workload/baselines/README.md | 14 ++++++++++++-- .../telemetry/workload/regression-thresholds.json | 2 +- docker/telemetry/workload/run-full-validation.sh | 11 +++++++++-- 3 files changed, 22 insertions(+), 5 deletions(-) diff --git a/docker/telemetry/workload/baselines/README.md b/docker/telemetry/workload/baselines/README.md index 29f627deee..02acb37445 100644 --- a/docker/telemetry/workload/baselines/README.md +++ b/docker/telemetry/workload/baselines/README.md @@ -136,8 +136,18 @@ The guarantee costs sensitivity where the ladder is coarse: the detection floor | `span.tx.process.p99` | 0.9940 ms | 5 ms | 5.03x | 1 ms → 5 ms | Four of the six are limited by the same `1 ms → 5 ms` step, which is where this ladder is coarsest -relative to how the spans actually behave. None of the six fires on any observed run, so all six -stay gated; the weak floor is recorded here so it is visible rather than surprising. +relative to how the spans actually behave. All six stay gated; the weak floor is recorded here so it +is visible rather than surprising. + +None of the six fires on an observed run **of this workload** — but the qualifier is load-bearing, +and there is now a measurement behind it. Changing one line of the generated node config from +`[ips]` to `[ips_fixed]`, which holds peer connections open instead of treating the list as a +discovery hint, moved `span.consensus.ledger_close.p95` from 0.57 ms to 6.43 ms and tripped this +gate, while every transaction-path metric fell. Nothing else in that commit touched the consensus +path. So a weak floor is not the only way one of these keys reddens: a change to the cluster's +topology is enough on its own, which is exactly why +[Refreshing the baseline](#refreshing-the-baseline) treats a workload change as requiring a new +baseline. The fix is a 2 ms edge (ideally 3 ms as well) in the collector's spanmetrics `buckets` list plus the matching entries in `kMillisecondBuckets`, and 2000 us plus 50000 us edges in `kMicrosecondBuckets`. diff --git a/docker/telemetry/workload/regression-thresholds.json b/docker/telemetry/workload/regression-thresholds.json index 973b099831..876dec9814 100644 --- a/docker/telemetry/workload/regression-thresholds.json +++ b/docker/telemetry/workload/regression-thresholds.json @@ -1,5 +1,5 @@ { - "_absolute_bound_derivation": "HOW EVERY max_abs_increase_* NUMBER BELOW WAS OBTAINED. Rule: locate the baseline value in the half-open bucket (lo, hi] of its own ladder, take hi_next = the next edge above hi, and set the bound to (hi_next - baseline). The trip point is therefore exactly hi_next: the gate fires only when the reported value EXCEEDS the top of the bucket above the baseline's own bucket. WHY THAT AND NOT A MULTIPLE OF THE BUCKET WIDTH: histogram_quantile returns a value interpolated inside whichever bucket the true quantile falls in, so a reading taken while the true quantile sits anywhere in the baseline's bucket OR anywhere in the one immediately above is at most hi_next and cannot fire. Firing requires the true quantile to have moved at least two buckets up. A multiple of the ENCLOSING width cannot deliver that, because once the quantile crosses hi the interpolation happens across the NEXT bucket, which on this ladder is up to 8x wider \u2014 (0.5,1] has width 0.5 and (1,5] has width 4 \u2014 so the reading's excursion is not bounded by any multiple of the enclosing width. Worked example: span.tx.process.p99 has baseline 0.9940ms in bucket (0.5, 1], hi_next = 5, so its bound is 4.0060ms and the gate fires only above 5ms. Bounds are stored as exact doubles rather than rounded figures so that rounding cannot break the guarantee and so check_regression_bounds.py can assert each one against the ladder to within a 1e-12 relative tolerance -- tight enough that a bound rounded for readability, such as 4.0060 for 4.006030232040599, is rejected; _derivation_table below shows the arithmetic for each one. Measured over the committed 2026-09-10 baseline this rule yields a detection floor of 2.00x to 7.41x of baseline, per key. WHAT THIS RULE DOES NOT COVER, AND THE ONE CHECK TO RUN BEFORE GATING ANY KEY: hi_next - baseline is derived from the LADDER, so it budgets for QUANTIZATION noise -- one bucket of interpolation headroom -- and for nothing else. It knows nothing about how much the metric itself moves between runs on identical code. Where run-to-run workload variance is the larger term the bound is simply the wrong size, and the gate reddens on a healthy run. So before adding a key here, capture it over several runs and check its OBSERVED MAXIMUM against its trip point (baseline + bound); gate it only if the observed maximum stays below that trip point with margin. Spread on its own proves nothing -- it is spread RELATIVE TO THE TRIP POINT that decides, and a baseline that lands at the LOW end of a metric's own range shrinks that trip point even though nothing about the metric changed. THREE KEYS FAILED THIS TEST ON THE 2026-08-26 BASELINE AND ARE EXCLUDED, all of them p50 (six keys are excluded in total; the other three are span.ledger.validate.p95, span.ledger.validate.p99 and span.ledger.build.p99): span.tx.apply.p50 (bound 0.0440ms, trips at 0.05ms, observed max 2.3378ms = 46.76x its trip point), span.ledger.build.p50 (bound 0.3849ms, trips at 0.5ms, observed max 2.3826ms = 4.77x) and span.consensus.ledger_close.p50 (bound 0.0613ms, trips at 0.1ms, observed max 0.2377ms = 2.38x). Their spreads across three runs are 391.8x, 20.7x and 6.1x. This is the general rule above being APPLIED, not a new exception: a key is gateable only when its run-to-run spread fits inside its bound, and these three do not. The evidence that settles it is span.tx.apply.p50's own history -- it read 0.7917ms in the 2026-08-24 baseline and 0.00597ms in the 2026-08-26 one, a 132x difference between two runs of the SAME workload. At the old value the identical rule produced a 4.21ms bound whose 5ms trip point absorbed the whole range; at the new one it produces 0.0440ms and cannot. Whether the gate functioned was therefore decided by where in its distribution the captured run happened to land, which is not a threshold needing tuning but a key that cannot be gated from a single-run baseline at all. Before the exclusion, replaying the two preceding CI runs 32862589645 and 32867433073 against the 2026-08-26 baseline reported exactly those three and nothing else on BOTH runs, and 32867433073 carries the same post-path-finding-removal workload as the baseline itself -- so the movement was metric variance, not a workload difference. After it, both runs replay clean. On the current 2026-09-10 baseline the 19 gated keys sit between 0.14 and 0.50 of baseline over trip point, the tightest being span.tx.process.p95 at 0.50. That ratio is derivable from this file and baseline-timings.json. A headroom figure against each key's OBSERVED MAXIMUM is not derivable here, because no per-run timings.json is committed -- so do not restate one without citing the run it came from. See _excluded_shape in regression-metrics.json for what the excluded keys have in common and for the multi-run-baseline work that would let them be gated again. A key that fails this test is not fixed by widening its bound: see excluded_keys in regression-metrics.json. WHAT THIS REPLACED, IN TWO GENERATIONS: (1) a single flat pair of bounds (10ms for span p50/p95, 15ms for span p99, 20000us for job_queue p95) justified as 'roughly two bucket widths in the 5-25ms band where most span quantiles actually sit'. The 2026-08-24 capture falsifies that premise \u2014 18 of the 28 quantiles gated at that time sat below 1ms \u2014 so the absolute bound sat 1.15x to 2000x above the metric it guarded and, because the rule is an AND, the percentage bound could never carry a regression on its own; a 10x regression injected into each key in turn was caught on only 5 of 28, and a 100x regression injected into span.ledger.store.p95 produced 0 regressions and exit 0. (2) a first correction to 2 \u00d7 the ENCLOSING bucket width, which caught 10x on 28 of 28 but placed the trip point INSIDE the adjacent bucket -- and so left a single-crossing false positive reachable -- on 21 of the 25 keys gated at the time, 4 of them tripping on a tail-mass shift under 1.5% of samples. That is the assumption this rule removes. RE-DERIVE THESE NUMBERS whenever baseline-timings.json is refreshed or either ladder changes: a refreshed baseline can land in a different bucket, which changes hi_next. .github/scripts/telemetry/check_regression_bounds.py enforces the rule in CI so a stale bound cannot survive a baseline refresh. LIMITATION \u2014 WHICH KEYS ARE ONLY WEAKLY GUARDED: the guarantee costs sensitivity wherever the ladder is coarse, and the detection floor is hi_next/baseline, so a baseline sitting just above an edge is guarded loosely. On the current 2026-09-10 baseline the six weakest keys are span.consensus.ledger_close.p95 (baseline 0.6750ms, fires at 5ms, 7.41x), span.consensus.accept.p50 (1.4364ms, 10ms, 6.96x), job.acceptLedger.running.p95 (15967.74us, 100000us, 6.26x), span.rpc.ws_message.p95 (0.8122ms, 5ms, 6.16x), span.rpc.ws_message.p99 (0.9873ms, 5ms, 5.06x) and span.tx.process.p99 (0.9940ms, 5ms, 5.03x). Four of the six are limited by the 1ms\u21925ms step; the other two by 5ms\u219210ms (span.consensus.accept.p50) and 25000us\u2192100000us (job.acceptLedger.running.p95). None of the six fires on any observed run, so all six stay gated, but the weak floors are recorded here so they are visible rather than surprising. Because every floor is now under 10x, a 10x regression is caught on all 19 gated keys -- that is derived from the floors, not sampled. On the 2026-08-26 baseline job.acceptLedger.running.p95 had a 16.28x floor and was the one key 10x missed; its baseline rose 6142.86us to 15967.74us while hi_next stayed at 100000us, which pulled its floor to 6.26x. Sensitivity therefore moves with each refresh even when no code changes, so re-derive these floors whenever the baseline is refreshed. The fix is a 2ms edge (and ideally 3ms) in the collector's spanmetrics ladder plus the matching edges in kMillisecondBuckets, and 2000us plus 50000us edges in kMicrosecondBuckets \u2014 that work belongs to the branch that owns the ladders, not here. Until then do not read these keys as guarded. span.ledger.store is absent from the overrides below because it is excluded from the gated surface entirely: its quantiles are the ladder floor times the quantile, so no bound can gate it. See _excluded_ledger_store in regression-metrics.json. REFRESHED 2026-09-10 from the median of CI runs 34495527952, 34505215266 and 34507425933, the first three runs with the account-funding race fixed. The 2026-08-26 baseline predated that fix, so the phases that lost their traffic captured artificially low ledger and transaction timings. Applying the observed-maximum test to the refreshed numbers leaves 19 of 20 keys between 0.17 and 0.76 of their trip points, and disqualifies span.ledger.build.p99 -- see excluded_keys in regression-metrics.json. span.tx.process.p95 is the tightest survivor at 0.76 and is the key to re-measure first if the gate reddens again.", + "_absolute_bound_derivation": "HOW EVERY max_abs_increase_* NUMBER BELOW WAS OBTAINED. Rule: locate the baseline value in the half-open bucket (lo, hi] of its own ladder, take hi_next = the next edge above hi, and set the bound to (hi_next - baseline). The trip point is therefore exactly hi_next: the gate fires only when the reported value EXCEEDS the top of the bucket above the baseline's own bucket. WHY THAT AND NOT A MULTIPLE OF THE BUCKET WIDTH: histogram_quantile returns a value interpolated inside whichever bucket the true quantile falls in, so a reading taken while the true quantile sits anywhere in the baseline's bucket OR anywhere in the one immediately above is at most hi_next and cannot fire. Firing requires the true quantile to have moved at least two buckets up. A multiple of the ENCLOSING width cannot deliver that, because once the quantile crosses hi the interpolation happens across the NEXT bucket, which on this ladder is up to 8x wider \u2014 (0.5,1] has width 0.5 and (1,5] has width 4 \u2014 so the reading's excursion is not bounded by any multiple of the enclosing width. Worked example: span.tx.process.p99 has baseline 0.9940ms in bucket (0.5, 1], hi_next = 5, so its bound is 4.0060ms and the gate fires only above 5ms. Bounds are stored as exact doubles rather than rounded figures so that rounding cannot break the guarantee and so check_regression_bounds.py can assert each one against the ladder to within a 1e-12 relative tolerance -- tight enough that a bound rounded for readability, such as 4.0060 for 4.006030232040599, is rejected; _derivation_table below shows the arithmetic for each one. Measured over the committed 2026-09-10 baseline this rule yields a detection floor of 2.00x to 7.41x of baseline, per key. WHAT THIS RULE DOES NOT COVER, AND THE ONE CHECK TO RUN BEFORE GATING ANY KEY: hi_next - baseline is derived from the LADDER, so it budgets for QUANTIZATION noise -- one bucket of interpolation headroom -- and for nothing else. It knows nothing about how much the metric itself moves between runs on identical code. Where run-to-run workload variance is the larger term the bound is simply the wrong size, and the gate reddens on a healthy run. So before adding a key here, capture it over several runs and check its OBSERVED MAXIMUM against its trip point (baseline + bound); gate it only if the observed maximum stays below that trip point with margin. Spread on its own proves nothing -- it is spread RELATIVE TO THE TRIP POINT that decides, and a baseline that lands at the LOW end of a metric's own range shrinks that trip point even though nothing about the metric changed. THREE KEYS FAILED THIS TEST ON THE 2026-08-26 BASELINE AND ARE EXCLUDED, all of them p50 (six keys are excluded in total; the other three are span.ledger.validate.p95, span.ledger.validate.p99 and span.ledger.build.p99): span.tx.apply.p50 (bound 0.0440ms, trips at 0.05ms, observed max 2.3378ms = 46.76x its trip point), span.ledger.build.p50 (bound 0.3849ms, trips at 0.5ms, observed max 2.3826ms = 4.77x) and span.consensus.ledger_close.p50 (bound 0.0613ms, trips at 0.1ms, observed max 0.2377ms = 2.38x). Their spreads across three runs are 391.8x, 20.7x and 6.1x. This is the general rule above being APPLIED, not a new exception: a key is gateable only when its run-to-run spread fits inside its bound, and these three do not. The evidence that settles it is span.tx.apply.p50's own history -- it read 0.7917ms in the 2026-08-24 baseline and 0.00597ms in the 2026-08-26 one, a 132x difference between two runs of the SAME workload. At the old value the identical rule produced a 4.21ms bound whose 5ms trip point absorbed the whole range; at the new one it produces 0.0440ms and cannot. Whether the gate functioned was therefore decided by where in its distribution the captured run happened to land, which is not a threshold needing tuning but a key that cannot be gated from a single-run baseline at all. Before the exclusion, replaying the two preceding CI runs 32862589645 and 32867433073 against the 2026-08-26 baseline reported exactly those three and nothing else on BOTH runs, and 32867433073 carries the same post-path-finding-removal workload as the baseline itself -- so the movement was metric variance, not a workload difference. After it, both runs replay clean. On the current 2026-09-10 baseline the 19 gated keys sit between 0.14 and 0.50 of baseline over trip point, the tightest being span.tx.process.p95 at 0.50. That ratio is derivable from this file and baseline-timings.json. A headroom figure against each key's OBSERVED MAXIMUM is not derivable here, because no per-run timings.json is committed -- so do not restate one without citing the run it came from. See _excluded_shape in regression-metrics.json for what the excluded keys have in common and for the multi-run-baseline work that would let them be gated again. A key that fails this test is not fixed by widening its bound: see excluded_keys in regression-metrics.json. WHAT THIS REPLACED, IN TWO GENERATIONS: (1) a single flat pair of bounds (10ms for span p50/p95, 15ms for span p99, 20000us for job_queue p95) justified as 'roughly two bucket widths in the 5-25ms band where most span quantiles actually sit'. The 2026-08-24 capture falsifies that premise \u2014 18 of the 28 quantiles gated at that time sat below 1ms \u2014 so the absolute bound sat 1.15x to 2000x above the metric it guarded and, because the rule is an AND, the percentage bound could never carry a regression on its own; a 10x regression injected into each key in turn was caught on only 5 of 28, and a 100x regression injected into span.ledger.store.p95 produced 0 regressions and exit 0. (2) a first correction to 2 \u00d7 the ENCLOSING bucket width, which caught 10x on 28 of 28 but placed the trip point INSIDE the adjacent bucket -- and so left a single-crossing false positive reachable -- on 21 of the 25 keys gated at the time, 4 of them tripping on a tail-mass shift under 1.5% of samples. That is the assumption this rule removes. RE-DERIVE THESE NUMBERS whenever baseline-timings.json is refreshed or either ladder changes: a refreshed baseline can land in a different bucket, which changes hi_next. .github/scripts/telemetry/check_regression_bounds.py enforces the rule in CI so a stale bound cannot survive a baseline refresh. LIMITATION \u2014 WHICH KEYS ARE ONLY WEAKLY GUARDED: the guarantee costs sensitivity wherever the ladder is coarse, and the detection floor is hi_next/baseline, so a baseline sitting just above an edge is guarded loosely. On the current 2026-09-10 baseline the six weakest keys are span.consensus.ledger_close.p95 (baseline 0.6750ms, fires at 5ms, 7.41x), span.consensus.accept.p50 (1.4364ms, 10ms, 6.96x), job.acceptLedger.running.p95 (15967.74us, 100000us, 6.26x), span.rpc.ws_message.p95 (0.8122ms, 5ms, 6.16x), span.rpc.ws_message.p99 (0.9873ms, 5ms, 5.06x) and span.tx.process.p99 (0.9940ms, 5ms, 5.03x). Four of the six are limited by the 1ms\u21925ms step; the other two by 5ms\u219210ms (span.consensus.accept.p50) and 25000us\u2192100000us (job.acceptLedger.running.p95). None of the six fires on an observed run of this workload, so all six stay gated, but the weak floors are recorded here so they are visible rather than surprising. The qualifier is measured, not hedging: switching the generated node config from [ips] to [ips_fixed] moved span.consensus.ledger_close.p95 from 0.57ms to 6.43ms and tripped this gate on an otherwise clean run, so a topology change alone can redden one of these keys. Because every floor is now under 10x, a 10x regression is caught on all 19 gated keys -- that is derived from the floors, not sampled. On the 2026-08-26 baseline job.acceptLedger.running.p95 had a 16.28x floor and was the one key 10x missed; its baseline rose 6142.86us to 15967.74us while hi_next stayed at 100000us, which pulled its floor to 6.26x. Sensitivity therefore moves with each refresh even when no code changes, so re-derive these floors whenever the baseline is refreshed. The fix is a 2ms edge (and ideally 3ms) in the collector's spanmetrics ladder plus the matching edges in kMillisecondBuckets, and 2000us plus 50000us edges in kMicrosecondBuckets \u2014 that work belongs to the branch that owns the ladders, not here. Until then do not read these keys as guarded. span.ledger.store is absent from the overrides below because it is excluded from the gated surface entirely: its quantiles are the ladder floor times the quantile, so no bound can gate it. See _excluded_ledger_store in regression-metrics.json. REFRESHED 2026-09-10 from the median of CI runs 34495527952, 34505215266 and 34507425933, the first three runs with the account-funding race fixed. The 2026-08-26 baseline predated that fix, so the phases that lost their traffic captured artificially low ledger and transaction timings. Applying the observed-maximum test to the refreshed numbers leaves 19 of 20 keys between 0.17 and 0.76 of their trip points, and disqualifies span.ledger.build.p99 -- see excluded_keys in regression-metrics.json. span.tx.process.p95 is the tightest survivor at 0.76 and is the key to re-measure first if the gate reddens again.", "_bucket_note": "SpanMetrics latency histograms use explicit buckets [0.01,0.05,0.1,0.25,0.5,1,5,10,25,50,100,250,500]ms then [1,2,3,4,5,10,30]s (20 edges; docker/telemetry/otel-collector-config.yaml is the authoritative list). Second-scale consensus spans have 2s/3s/4s boundaries, so their quantiles quantize to ~1s widths there \u2014 the ladder is NOT uniformly 2x-or-coarser, which matters for _percentage_bound_note. The native job_queue histograms are microsecond-valued on the ladder [1,2,5,10,25,50,100,250,500,1000,5000,25000,100000,500000]us then [1,5,10,30,60]s (19 edges; include/xrpl/telemetry/HistogramBuckets.h is authoritative). NOTE: BOTH ladders were re-cut, and a baseline captured before its own ladder changed is an interpolation artefact, not a latency. The job_queue floor moved 100us \u2192 1us. The span floor is 0.01ms; a span baseline captured against a 1ms floor is void below 1ms \u2014 a p95 reading 0.95ms there is 0.95 \u00d7 that 1ms first edge, not a measurement. Do not assume a surviving span baseline is unaffected by ladder work: every span quantile below 1ms is affected. Only the band from 1ms to 1s is safe: those edges are byte-identical across the two ladders. The re-cut also ADDED edges above 1s (2s/3s/4s/10s/30s), so a span whose quantiles land in the second-scale range \u2014 consensus.round ~3.9s, consensus.establish ~1.9s, the ledger.acquire tail \u2014 is distorted just as much, and any pre-2026-08-04 baseline for it is equally void. Do not read this note as licensing a stale second-scale baseline.", "_defaults_note": "A MISSING OVERRIDE IS DETECTED BY CI, NOT BY THESE DEFAULTS. .github/scripts/telemetry/check_regression_bounds.py fails the build at lint time, naming the key and the exact value its bound should have, before the workload ever runs. That is the mechanism; the defaults below are only a runtime backstop for the case where that check is bypassed. The defaults carry the FLOOR of each ladder as their absolute bound \u2014 0.01ms for spans, 1us for job_queue \u2014 deliberately too small to bind for any real metric, which leaves max_pct_increase (50%) as the operative bound on this path. Measured: a metric with no override and a baseline of 3900ms passes at +49% and fires at +51%; a job metric with a baseline of 5000us behaves the same. The backstop is honestly imperfect and should not be oversold. At 50% relative it CAN false-fire: a metric whose baseline is 1.06ms inside the 4ms-wide (1,5] bucket fires on a single-bucket-width move (measured: 1.06 \u2192 5.06ms, +377%, regressed). That false fire is NOT to be read as 'the intended signal that the override is missing' \u2014 CI prints REGRESSION and a reader cannot tell it from a real one, and rejecting a tighter alternative for exactly that cries-wolf risk while shipping it here would be inconsistent. The check is what makes the signal legible. The backstop is kept only because a metric silently not gated at all is the worse of the two failures.", "_derivation_table": { diff --git a/docker/telemetry/workload/run-full-validation.sh b/docker/telemetry/workload/run-full-validation.sh index 7a356fa85b..604c4436bf 100755 --- a/docker/telemetry/workload/run-full-validation.sh +++ b/docker/telemetry/workload/run-full-validation.sh @@ -351,7 +351,14 @@ for i in $(seq 1 "$NUM_NODES"); do "" | null) die "$NODE_PREFIX-$i has no seed in $WORKDIR/validator-keys.json — the file holds fewer than $NUM_NODES entries, or entry $((i - 1)) carries no seed" ;; esac - # Build ips_fixed. + # Peer list for the loopback mesh. Emitted as [ips], NOT [ips_fixed], + # even though [ips_fixed] is the section whose documented meaning fits a + # private cluster. [ips_fixed] holds the connections open to all peers, and + # measured against this same commit that moved consensus.ledger_close.p95 + # from 0.57 ms to 6.43 ms and tripped the regression gate, while every + # transaction-path metric fell. The committed baseline describes the [ips] + # topology, so switching sections is a deliberate workload change that has + # to arrive with a refreshed baseline and re-derived bounds. IPS_FIXED="" for j in $(seq 1 "$NUM_NODES"); do if [ "$j" -ne "$i" ]; then @@ -400,7 +407,7 @@ $SEED [validators_file] $WORKDIR/validators.txt -[ips_fixed] +[ips] ${IPS_FIXED} [telemetry]