fix(telemetry): keep the workload cluster on [ips], not [ips_fixed]

The previous commit switched the generated node config from [ips] to
[ips_fixed] on the grounds that the variable, the comment and the sibling cfg
template all named ips_fixed, and that ips_fixed is the section whose
documented meaning fits a private cluster. Both of those are still true. The
switch is reverted anyway, because it is a workload change rather than a
naming fix.

Measured on CI, parent commit against this branch's previous tip, one
functional config line apart:

  span.consensus.ledger_close.p95    0.57 ms -> 6.43 ms   (tripped the gate)
  span.consensus.ledger_close.p99    0.94 ms -> 9.50 ms
  span.consensus.accept.p50          0.97 ms -> 2.63 ms
  span.tx.process.p50                0.36 ms -> 0.18 ms   (faster)
  job.acceptLedger.running.p95      21157 us -> 10938 us  (faster)

Every consensus-path span rose and every transaction-path metric fell, which
is the shape a denser always-connected mesh produces and not the shape of
run-to-run variance. [ips_fixed] holds connections open to all four peers
instead of treating the list as a discovery hint, so each node processes
proposals and validations from the full mesh every round. Nothing else in that
commit touches the consensus path: the emitted config differed in exactly
three lines, of which one is a die message and one expands to an identical
string.

The committed baseline describes the [ips] topology. Adopting [ips_fixed]
therefore needs a refreshed baseline and re-derived bounds, which is the
process baselines/README.md already documents for a workload change. Left as
its own work item rather than smuggled in behind a section rename, and the
reason is now recorded beside the line so it is not repeated.

This also falsified a claim the previous commit had written into
baselines/README.md and regression-thresholds.json: that none of the six
weakly-guarded keys fires on any observed run. Corrected in both, and the
measurement above is cited in place of the absolute.
This commit is contained in:
Pratik Mankawde
2026-09-17 17:01:29 +01:00
parent 3c05bdfdd8
commit 85fc1c5399
3 changed files with 22 additions and 5 deletions

View File

@@ -136,8 +136,18 @@ The guarantee costs sensitivity where the ladder is coarse: the detection floor
| `span.tx.process.p99` | 0.9940 ms | 5 ms | 5.03x | 1 ms → 5 ms |
Four of the six are limited by the same `1 ms → 5 ms` step, which is where this ladder is coarsest
relative to how the spans actually behave. None of the six fires on any observed run, so all six
stay gated; the weak floor is recorded here so it is visible rather than surprising.
relative to how the spans actually behave. All six stay gated; the weak floor is recorded here so it
is visible rather than surprising.
None of the six fires on an observed run **of this workload** — but the qualifier is load-bearing,
and there is now a measurement behind it. Changing one line of the generated node config from
`[ips]` to `[ips_fixed]`, which holds peer connections open instead of treating the list as a
discovery hint, moved `span.consensus.ledger_close.p95` from 0.57 ms to 6.43 ms and tripped this
gate, while every transaction-path metric fell. Nothing else in that commit touched the consensus
path. So a weak floor is not the only way one of these keys reddens: a change to the cluster's
topology is enough on its own, which is exactly why
[Refreshing the baseline](#refreshing-the-baseline) treats a workload change as requiring a new
baseline.
The fix is a 2 ms edge (ideally 3 ms as well) in the collector's spanmetrics `buckets` list plus the
matching entries in `kMillisecondBuckets`, and 2000 us plus 50000 us edges in `kMicrosecondBuckets`.

File diff suppressed because one or more lines are too long

View File

@@ -351,7 +351,14 @@ for i in $(seq 1 "$NUM_NODES"); do
"" | null) die "$NODE_PREFIX-$i has no seed in $WORKDIR/validator-keys.json — the file holds fewer than $NUM_NODES entries, or entry $((i - 1)) carries no seed" ;;
esac
# Build ips_fixed.
# Peer list for the loopback mesh. Emitted as [ips], NOT [ips_fixed],
# even though [ips_fixed] is the section whose documented meaning fits a
# private cluster. [ips_fixed] holds the connections open to all peers, and
# measured against this same commit that moved consensus.ledger_close.p95
# from 0.57 ms to 6.43 ms and tripped the regression gate, while every
# transaction-path metric fell. The committed baseline describes the [ips]
# topology, so switching sections is a deliberate workload change that has
# to arrive with a refreshed baseline and re-derived bounds.
IPS_FIXED=""
for j in $(seq 1 "$NUM_NODES"); do
if [ "$j" -ne "$i" ]; then
@@ -400,7 +407,7 @@ $SEED
[validators_file]
$WORKDIR/validators.txt
[ips_fixed]
[ips]
${IPS_FIXED}
[telemetry]