Two suspects from the 3.3.0 slowdown investigation had no signal. Both were already computing the numbers and throwing them away, so this exposes them rather than adding measurement. Per-sweep heap trim. The trim runs after every cache sweep, and its cost scales with resident heap, so it is the leading explanation for a node with a populated database syncing slower than a fresh one. The report already carried duration, fault deltas and reclaimed pages, but the whole measurement sat behind a debug-journal check, so an ordinary node measured nothing, and the call site discarded the result. The measurement now always runs and only the log line stays gated. Records trim duration, minor faults and reclaimed kilobytes. Measured cost of the always-on path is about six microseconds per sweep against a trim costing milliseconds, at a cadence of ten to a hundred and twenty seconds. Honest limit, stated in the runbook: the fault delta spans only the trim call, so it shows the trim itself faulting but not the faults that follow as caches refill. The duration is the signal to correlate against sweep-job queueing. Rotation writes. Rotation copies archive-served reads forward and re-stores nodes missing from both backends, both of which compete with sync I/O and only happen on a populated online_delete database. The copy-forward count existed but was reset by the rotation's own log line, so a metric reading it would drop to zero on every swap; a never-reset total sits beside it now. The re-store count was not measured at all. Rotation duration is deliberately not recorded: the health throttle sleeps at eight points inside the sequence and dominates exactly when the node is unhealthy, so the number would conflate work with waiting. Nothing added for the other two suspects. Get-object serving is already covered by the handler label, the lookup histogram and the deferred and saturation gauges; peer churn by the disconnect-reason counter. Also replaces nine per-file cspell ignores with one ignoreRegExpList entry for the telemetry macro names, and picks up the levelization baseline for the consensus span-name test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Levelization
Levelization is the term used to describe efforts to prevent xrpld from having or creating cyclic dependencies.
xrpld code is organized into directories under src/xrpld, src/libxrpl (and
src/test) representing modules. The modules are intended to be
organized into "tiers" or "levels" such that a module from one level can
only include code from lower levels. Additionally, a module
in one level should never include code in an impl or detail folder of any level
other than its own.
The codebase is split into two main areas:
- libxrpl (
src/libxrpl,include/xrpl): Reusable library modules with public interfaces - xrpld (
src/xrpld): Application-specific implementation code
Unfortunately, over time, enforcement of levelization has been inconsistent, so the current state of the code doesn't necessarily reflect these rules. Whenever possible, developers should refactor any levelization violations they find (by moving files or individual classes). At the very least, don't make things worse.
The table below summarizes the desired division of modules, based on the current state of the xrpld code. The levels are numbered from the bottom up with the lower level, lower numbered, more independent modules listed first, and the higher level, higher numbered modules with more dependencies listed later.
tl;dr: The modules listed first are more independent than the modules listed later.
libxrpl Modules (Reusable Libraries)
| Level / Tier | Module(s) |
|---|---|
| 01 | xrpl/beast |
| 02 | xrpl/basics |
| 03 | xrpl/json xrpl/crypto |
| 04 | xrpl/protocol |
| 05 | xrpl/core xrpl/resource xrpl/server |
| 06 | xrpl/ledger xrpl/nodestore xrpl/net |
| 07 | xrpl/shamap |
xrpld Modules (Application Implementation)
| Level / Tier | Module(s) |
|---|---|
| 05 | xrpld/conditions xrpld/consensus |
| 06 | xrpld/core xrpld/peerfinder |
| 07 | xrpld/shamap xrpld/overlay |
| 08 | xrpld/app |
| 09 | xrpld/rpc |
| 10 | xrpld/perflog |
Test Modules
| Level / Tier | Module(s) |
|---|---|
| 11 | test/jtx test/beast test/csf |
| 12 | test/unit_test |
| 13 | test/crypto test/conditions test/json test/resource test/shamap test/peerfinder test/basics test/overlay |
| 14 | test |
| 15 | test/net test/protocol test/ledger test/consensus test/core test/server test/nodestore |
| 16 | test/rpc test/app |
(Note that test levelization is much less important and much less
strictly enforced than xrpl/xrpld levelization, other than the requirement
that test code should never be included in xrpl or xrpld code.)
Validation
The levelization script takes no parameters,
reads no environment variables, and can be run from any directory,
as long as it is in the expected location in the xrpld repo.
It can be run at any time from within a checked out repo, and will
do an analysis of all the #includes in
the xrpld source. The only caveat is that it runs much slower
under Windows than in Linux. It hasn't yet been tested under MacOS.
It generates many files of results:
rawincludes.txt: The raw dump of the#includespaths.txt: A second dump grouping the source module to the destination module, de-duped, and with frequency counts.includes/: A directory where each file represents a module and contains a list of modules and counts that the module includes.included_by/: Similar toincludes/, but the other way around. Each file represents a module and contains a list of modules and counts that include the module.loops.txt: A list of direct loops detected between modules as they actually exist, as opposed to how they are desired as described above. In a perfect repo, this file will be empty. This file is committed to the repo, and is used by the levelization Github workflow to validate that nothing changed.ordering.txt: A list showing relationships between modules where there are no loops as they actually exist, as opposed to how they are desired as described above. This file is committed to the repo, and is used by the levelization Github workflow to validate that nothing changed.levelization.ymlGithub Actions workflow to test that levelization loops haven't changed. Unfortunately, if changes are detected, it can't tell if they are improvements or not, so if you have resolved any issues or done anything else to improve levelization, rungenerate.py, and commit the updated results.
The loops.txt and ordering.txt files relate the modules
using comparison signs, which indicate the number of times each
module is included in the other.
A > Bmeans that A should probably be at a higher level than B, because B is included in A significantly more than A is included in B. These results can be included in bothloops.txtandordering.txt. Becauseordering.txtonly includes relationships where B is not included in A at all, it will only include these types of results.A ~= Bmeans that A and B are included in each other a different number of times, but the values are so close that the script can't definitively say that one should be above the other. These results will only be included inloops.txt.A == Bmeans that A and B include each other the same number of times, so the script has no clue which should be higher. These results will only be included inloops.txt.
The committed files hide the detailed values intentionally, to prevent false alarms and merging issues, and because it's easy to get those details locally.
- Run
generate.py - Grep the modules in
paths.txt.- For example, if a cycle is found
A ~= B, simplygrep -w A .github/scripts/levelization/results/paths.txt | grep -w B
- For example, if a cycle is found