fix(telemetry): Accept every legitimate unl_expiry_days reading in the validator

The sanity bound required the value to be strictly positive, which rejected
three of the four readings the gauge can legitimately produce: a negative
count once the validator list has expired, the -1 sentinel for "no published
list fetched", and +inf for a config-listed list that never expires.

The negative case is the one that matters. It is the signal that expiry has
already passed, so the gauge deliberately does not clamp at zero, and a gate
that rejects it would fail exactly when an operator most needs the reading.
The -1 sentinel already violated the bound and had simply never been hit,
because the validation cluster always fetches a published list.

The floor is now a century, which still catches a broken clock. Detecting the
unsigned wrap this bound used to hide moves to the MetricsRegistry::daysUntil
unit tests, which are deterministic and do not need a running cluster.
This commit is contained in:
Pratik Mankawde
2026-09-22 19:04:48 +01:00
parent 406e38275a
commit cbed889f4b

View File

@@ -2151,9 +2151,14 @@ PARITY_VALUE_SANITY: list[dict[str, Any]] = [
{
"name": "unl_expiry_days",
"query": 'validator_health{metric="unl_expiry_days"}',
"lo": 0,
# Four reading classes are all legitimate: days remaining, a negative
# count once the list has expired, -1 when no published list has been
# fetched, and +inf for a config-listed list that never expires. Only a
# reading past a century is nonsense, and that means a broken clock.
# The wrap this floor used to hide is caught deterministically by the
# MetricsRegistry::daysUntil unit tests instead.
"lo": -36500,
"hi": None,
"exclusive_lo": True,
},
{
"name": "peer_latency_p90_ms",