Use-case path tracing: importance, relevance drift and path changes¶
Tool: scripts/usecase_paths.py (data: scripts/usecase_flows.yaml),
CI: .github/workflows/usecase-paths.yml. Report-only today.
Why¶
Line coverage says whether a test reached a line. It does not say whether any real use of abicheck does, how many use cases depend on a function, or whether a change moved a use case onto different code. A first, hand-run measurement (the real CLI over the catalog under coverage) found about 10k lines that no command or documented API reached, and several building blocks with two implementations answering one question differently. This plan turns that one-off measurement into a repeatable process with three goals:
- Importance. Code on the paths of known use cases matters more and should be tested and reviewed harder than code no use case runs.
- Relevance drift. Code that stops being reached is a signal to look, not a deletion list.
- Path changes. A change that routes a use case through different code is visible even when every test passes.
How it works¶
record runs each use-case source under coverage, one coverage context
per run, and maps every executed line to the function whose body ran.
The def line, decorators and default arguments run at import and are
never counted: counting them made every function of every imported module
read as reached (tests/test_usecase_paths.py checks attribution against
generated code that logs its own calls). A one-line def f(): body shares
its def line and is left unattributed -- 33 of about 9,300 functions,
almost all protocol stubs.
Two sources, each run tagged with a docs/contribute/usecase-registry.yaml
id:
| Source | What runs | Cost | Reaches |
|---|---|---|---|
scenarios |
the 29 automated scenarios of tests/scenarios/*.yaml |
~15 s, no compiler | policy, reporting, snapshot storage |
flows |
11 real CLI command lines x 12 catalog cases (usecase_flows.yaml) |
~4 min plus the catalog build | also binary readers, both header frontends, compat, deps |
rank places each function in a tier by distinct use cases reaching it:
shared (3 or more), use-case (1-2), unreached. --unit-coverage
joins a unit-suite coverage file and lists use-case code whose own unit
coverage is low.
diff BASE HEAD (with --git-base) reports, in this order:
- the functions the change edits, most relied-on first -- where review and test effort should go (goal 1);
- functions some use case reached on base and none reaches on head, split into still-defined (look here) and removed/renamed (goal 2);
- runs whose path changed, grouped by the change they share (goal 3);
- newly reached functions, and runs that failed.
dead RECORDING classifies every unreached function by production
reference (scripts/production_references.py): dead when every
reference to its name in abicheck/, scripts/, action/, actions/,
.github/ or pyproject.toml lies inside another dead function (a
greatest fixpoint), with documented API and ADR/plan-named functions listed
apart and treated as roots. It is the recomputed form of the hand-made list
dead-code-and-single-owner started from.
record --root measures another checkout with this tool, so CI records
the PR base and head with one recorder.
First measurements¶
On main at the time of writing, scenarios plus flows (153 runs):
| Tier | Functions |
|---|---|
| shared | 2,418 |
| use-case | 1,096 |
| unreached | 5,772 |
Recording main~5 against main showed exactly the functions those five
commits introduced entering scenario paths (six distinct changes over 29
runs), and listed the 96 functions they edited with 25 in the shared tier
(checker.compare, the snapshot encoder, the fact codec among them).
The flows source also caught a live defect on main: compat could not
read back a dump it wrote under a .dump name (fixed in PR #1448).
Status and next steps¶
Landed: the recorder, the four reports, the PR and weekly workflow, and the attribution, diff and dead-code fixpoint tests.
Open, in order:
- Widen the sources. Flows for the composite Action, PR-comment
rendering,
post_manifest,debian_symbols, the hybrid frontend and a multi-library release directory, which no recorded run reaches yet. - Windows and macOS. PE, Mach-O and PDB readers read as unreached on Linux only because no such toolchain runs there; record the flows on those runners and merge the recordings before trusting their tier.
- Make tiers act. Once the ranking is stable over a few weeks: a shared-tier function edited without a test change gets a PR notice, and the mutation lane prioritises shared-tier modules. Both stay advisory until the false-positive rate is known.
- Order, not only membership. A path is a set of functions today. If set membership proves too coarse (same functions, different order or branch), record branch arcs as well.
- Gate.
diff --fail-on dropped|changed|failedexists; turning any of them on is a decision for after the report has run on real PRs.