Files
rippled/.github/scripts/otel-naming
Pratik Mankawde fddf78567d Merge branch 'pratik/otel-phase10-workload-validation' into pratik/otel-sync-diagnostics
Two conflicts, both additive-vs-additive; each resolution keeps both sides.

check_otel_naming.py -- phase-10 taught the L6 label extractor to match the
label MAP first and to resolve a key hoisted into a `k...Label` constant,
scanning headers as well as sources. Our side had added the two-regex
first/subsequent literal scan and the `metric_constants(root)[1]` union that
covers the `namespace label` header style.

Kept phase-10's mechanism whole: METRIC_LABEL_MAP + the `(?:^|\{)` key regex
already subsumes what METRIC_LABEL_NEXT did, since matching inside the map body
makes every pair after the first open with a single `{`. So METRIC_LABEL_NEXT is
dropped as genuinely redundant rather than kept as a duplicate scan, and the
reason it existed is folded into METRIC_LABEL's comment. Re-added our
`metric_constants(root)[1]` union on top: LABEL_CONST_DEF only matches
`k`-prefixed identifiers, so it cannot see MetricNames.h's `label::jobType`
style, and without that union Rule D would reject dashboards querying labels
Rule I forced into constants. The two derivations are complementary and both
are now documented as such.

MetricsRegistry.cpp -- both sides added a new sibling view-registration helper
next to addMicrosecondHistogramView, and both added a registration call in
initExporterAndProvider(). Kept all four helpers
(addHistogramView/Microsecond/RoundDuration/SubMillisecond) and every
registration: phase-10's addSubMillisecondHistogramView + kNodeStoreReadUs
alongside our addRoundDurationHistogramView, sweepMallocTrimUs and the two
millisecond dial/resolve ladders.

phase-10's nodestore_read_us histogram does not duplicate our work. The
nodestore_latency gauge that would have overlapped it was retired in c4e434d520
before this merge, and the surviving nodestore_state gauge is complementary
rather than duplicative: both read the same fetch measurement, but the gauge
publishes only a since-boot mean via scaledMean() and cannot yield a
percentile -- the consequence observeNodeStoreTotals' own docs state plainly --
while the histogram buckets each fetch and can. The histogram also splits by
fetch_type and found, which the gauge cannot. phase-10 registered its
explicit-bucket View, so it does not inherit the SDK default ladder.

Each file keeps its own existing naming style: phase-10's k-prefixed constants
are left as-is, ours stay namespaced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 14:09:05 +01:00
..

OTel naming-consistency check

check_otel_naming.py enforces the OpenTelemetry span-attribute naming convention documented in CONTRIBUTING.md across every layer of the telemetry pipeline. The *SpanNames.h constants are the single source of truth (L1); every other layer must agree with them.

Running locally

python .github/scripts/otel-naming/check_otel_naming.py

It takes no arguments, can be run from any directory inside the repo, and uses only the Python standard library (no pip install, matching the levelization check). A non-zero exit code means a violation was found; the output lists each violation as RULE | location | token | expected.

What it checks

The valid key set is derived dynamically from the OTel code — there is no hardcoded allowlist:

  • L1 keys come from the namespace attr { ... } blocks of every *SpanNames.h, resolving the makeStr("x") / join(seg::a, seg::b) DSL (cross-file, so join(seg::rpc, ...) resolves seg::rpc from the base SpanNames.h). Each constant is resolved against its own header, so two headers that define a same-named constant (e.g. a base attr::ledgerHash and a domain attr::ledgerHash) each contribute their real wire key — a later header cannot clobber an earlier one's value in a flat table.
  • Legitimate dotted keys = ONLY the keys the code actually sets as resource attributes, i.e. the entries inside Telemetry.cpp's Resource::Create({...}) call: the semconv::service::* keys (service.*) plus any attr::<name> constants passed there (xrpl.network.*). A dotted key that is declared in a header but never set as a resource attr is a span attribute in resource clothing — a Rule-A violation, even if it lives in the base SpanNames.h.
  • L1-metrics — instrument names, label keys and bounded label values come from the namespace metric / namespace label / namespace lval blocks of every *MetricNames.h, read as inline constexpr char NAME[] = "wire";. These headers deliberately do not use the makeStr/StaticStr DSL the span headers use: the OTel C++ API takes nostd::string_view, which constructs from char const* but has no constructor from std::string_view, so a StaticStr will not compile in an instrument-name or label-key position.

Rules (each fails the build, when its inputs are present)

Rule Check
A No stray dotted span-attribute key (only the derived resource keys may be dotted).
G Attribute keys are lower_snake_case (^[a-z][a-z0-9_]*$ per dot-segment) — no camelCase, UPPERCASE, or spaces.
F No string literals as attribute keys or span-name arguments in setAttribute/addEvent/span/rootSpan/childSpan (rootSpan shares span's (cat, prefix, name) signature). Attribute values are exempt (runtime data); *SpanNames.h definitions and test files are exempt.
B Every collector spanmetrics.dimensions name exists in the L1 key set.
C Every Tempo span-filter tag exists in the L1 key set.
D Every dashboard label resolves to an L1 span attribute, a native-metric label (L6, emitted by MetricsRegistry), or a Prometheus/Grafana builtin. TraceQL scope prefixes (span./resource./…) are stripped before the L1 lookup.
E No dotted xrpl.<domain>.<field> attribute key in the runbook (only the L1 resource attrs xrpl.network.* may be dotted). Span names, filenames, OTel-standard keys, and metric labels are not flagged.
I No string literals as metric instrument names or label keys — the mirror of Rule F. Applies to the name passed to an XRPL_METRIC_* macro or a meter->Create* factory and to the label keys in its label set. Label values, descriptions, *MetricNames.h, MetricMacros.h and test files are exempt. Scoped by metric family (first underscore segment): declaring a constant opts that family in, so the metric surface can be converted subsystem by subsystem. Unconverted families warn as Rule L.
J Metric instrument names follow the suffix conventions: lower_snake_case, no xrpld_/xrpl_ prefix (the exporter adds it), a counter ends _total, a histogram ends _us/_ms/_seconds, a gauge does not end _total. The instrument kind is read from the emit site, never guessed from words in the name — so a multi-series gauge carrying units in its label values (e.g. nodestore_state observing write_mean_us) is not a violation.
K Every metric named in docker/telemetry/workload/expected_metrics.json resolves to a declared constant, so a rename in code cannot leave the workload validator asserting a name nothing emits. PromQL selectors (m{label="v"}) and exporter-appended histogram suffixes (_bucket/_count/_sum) are normalized away first; groups fed by another emit path (statsd_gauges, statsd_counters, spanmetrics) are out of scope by design.

Rule F runs unconditionally (it is a purely syntactic check on the call-sites and needs no *SpanNames.h), so a code path that calls SpanGuard::span/setAttribute directly without ever defining a header is still caught.

Warnings (printed, never fail the build)

Rule Check
H A namespace-qualified constant (e.g. foo::bar::myKey) used at a telemetry call-site is not defined in any *SpanNames.h. The constant should live in the proper header; defining it in-place bypasses rules A/G/F. Warns rather than fails — the argument may be a legitimately dynamic value, and the header may live on a later branch. Bare locals and std:: names are not warned.
L A literal metric name in a family that has no *MetricNames.h constants yet. Rule I's ratchet defers these instead of failing the build on the whole pre-existing metric surface at once; the warning keeps the outstanding conversion work visible rather than silently accepted.

Presence-gated

Every rule runs only when the source files it needs are present in the tree and is otherwise skipped (printed as SKIP: <rule> — <reason>), never failed. This keeps the check correct no matter how telemetry work is split across PRs — a stacked chain, one large PR, or independent per-stage PRs where (for example) the collector config lands before the dashboards. The collector/Tempo/dashboard/ runbook layers are introduced in later phases; on a branch without them, only the L1-intrinsic rules (A, G, F) run.