From 866ab77ececf6dfd4f6b17f4e111c525e1edba77 Mon Sep 17 00:00:00 2001 From: Pratik Mankawde <3397372+pratikmankawde@users.noreply.github.com> Date: Tue, 15 Sep 2026 14:26:17 +0100 Subject: [PATCH 1/2] docs(telemetry): describe rotation measurements without naming the host The runbook provenance paragraph named the internal AWS dev box and a build hash and dates. State what was measured (one mainnet node, same host and binary, differing only in store state) without the deployment detail, which belongs in an internal runbook, not the public repo. --- docs/telemetry-runbook.md | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/docs/telemetry-runbook.md b/docs/telemetry-runbook.md index e1098a30b4..1274ee51df 100644 --- a/docs/telemetry-runbook.md +++ b/docs/telemetry-runbook.md @@ -3308,9 +3308,8 @@ increase(nodestore_state{metric="acquire_ledger_timeouts", service_instance_id=~ #### Measured reference points -**Provenance.** The two columns below are our own measurements: node2 on the AWS -dev box, build `e3c2f8279a`, 2026-07-27/28, same host and same binary for both -runs, differing only in the state of the store. Use them as the shape to compare +**Provenance.** The two columns below are our own measurements: one mainnet node, +same host and same binary for both runs, differing only in the state of the store. Use them as the shape to compare against, not as thresholds. The read figures below come from the `read_mean_us` gauge, the only read-latency signal exported; the "highest sample" row is the largest value that gauge reached over the run, not a read-latency percentile. The third dataset in this section — the 25-minute devnet stall and its From 50eff17dd458030f7b0c98215c688551f5a6cf10 Mon Sep 17 00:00:00 2001 From: Pratik Mankawde <3397372+pratikmankawde@users.noreply.github.com> Date: Tue, 15 Sep 2026 14:26:23 +0100 Subject: [PATCH 2/2] docs(telemetry): drop the host name from the sampling-clock comment The comment measured date +%s%N cost 'on a dev box'; say 'on one Linux host' instead. The number is the point, not where it was taken. --- docker/telemetry/workload/collect_system_metrics.sh | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docker/telemetry/workload/collect_system_metrics.sh b/docker/telemetry/workload/collect_system_metrics.sh index 931b0ef6ed..87cc707bab 100755 --- a/docker/telemetry/workload/collect_system_metrics.sh +++ b/docker/telemetry/workload/collect_system_metrics.sh @@ -133,8 +133,8 @@ log "Collecting metrics for ${DURATION}s (${SAMPLES} samples, ${#RPC_PORTS[@]} n # exit 0 with an all-zero JSON — a silent false pass. # # The clock has to be cheap as well as precise, because the latency it -# measures is compared against a 2 ms threshold. Measured on a dev box: `date -# +%s%N` costs ~1.2 ms per call, forking python3 for the same value ~13 ms. +# measures is compared against a 2 ms threshold. Measured on one Linux host: +# `date +%s%N` costs ~1.2 ms per call, forking python3 for the same value ~13 ms. # Two calls bracket every request, so a python3 fallback would add ~26 ms of # its own overhead to a 2 ms budget and make the number meaningless. There is # no cheap alternative worth having, so probe once and refuse to run without