Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics

This commit is contained in:
Pratik Mankawde
2026-09-15 14:27:16 +01:00
2 changed files with 4 additions and 5 deletions

View File

@@ -133,8 +133,8 @@ log "Collecting metrics for ${DURATION}s (${SAMPLES} samples, ${#RPC_PORTS[@]} n
# exit 0 with an all-zero JSON — a silent false pass.
#
# The clock has to be cheap as well as precise, because the latency it
# measures is compared against a 2 ms threshold. Measured on a dev box: `date
# +%s%N` costs ~1.2 ms per call, forking python3 for the same value ~13 ms.
# measures is compared against a 2 ms threshold. Measured on one Linux host:
# `date +%s%N` costs ~1.2 ms per call, forking python3 for the same value ~13 ms.
# Two calls bracket every request, so a python3 fallback would add ~26 ms of
# its own overhead to a 2 ms budget and make the number meaningless. There is
# no cheap alternative worth having, so probe once and refuse to run without

View File

@@ -3390,9 +3390,8 @@ increase(nodestore_state{metric="acquire_ledger_timeouts", service_instance_id=~
#### Measured reference points
**Provenance.** The two columns below are our own measurements: node2 on the AWS
dev box, build `e3c2f8279a`, 2026-07-27/28, same host and same binary for both
runs, differing only in the state of the store. Use them as the shape to compare
**Provenance.** The two columns below are our own measurements: one mainnet node,
same host and same binary for both runs, differing only in the state of the store. Use them as the shape to compare
against, not as thresholds. The read figures below come from the `read_mean_us`
gauge, the only read-latency signal exported; the "highest sample" row is the
largest value that gauge reached over the run, not a read-latency percentile. The third dataset in this section — the 25-minute devnet stall and its