Skip to content

Comparison Performance

This page documents the runtime time cost of comparing real, large shared libraries. Its companion, Comparison Memory, covers what a multi-library comparison keeps resident, how to measure that in a way that attributes a peak to an owner, and the current owners.

This page documents the runtime cost of comparing real, large shared libraries, the bottlenecks that were found and fixed, and the tooling that guards against regressions in CI.

TL;DR

  • Dump scales fine. Snapshotting libonedal_core.so (~10,550 exported functions) takes ~5 s.
  • Compare used to blow up. On the same library, compare did not finish within 60 s. The cost was entirely in the post-processing detectors, not the core symbol diff. A profiling sweep found six super-linear paths — several quadratic, one effectively cubic — all now fixed (see What was fixed).
  • A synthetic scaling harness (scripts/benchmark_scaling.py) reproduces each path without a real binary, compiler, or castxml, and a slow regression test guards the realistic hot path.

What was fixed

Every fix preserves detector behaviour (the full unit suite, the FP-rate gate, and the metamorphic/oracle detector tests all stay green); they only change how the work is organised.

# Path Was Fix Result
1 Public-surface scoping (surface.classify_change_surface) Recomputed four old∪new set unions per finding → O(findings × surface). Made every comparison quadratic. Compute the unions once per pass (surface_unions) and reuse. add_remove 4000: 9.1 s → 0.32 s (linear)
2 Namespace detection (diff_namespaces) demangle_batch called one symbol at a time → one c++filt subprocess per symbol. Batch-demangle each snapshot once (_batch_demangle_public) and thread the map through. elf_namespace 4000: 5.2 s → 0.33 s (linear)
3 Variable / symbol diffing Quadratic via the same per-finding surface unions (#1). Fixed by #1. var_churn 4000: 2.1 s → 0.06 s (linear)
4 Batch-rename heuristic (diff_symbols._find_rename_pairs) O(removed × added) suffix scan. Reversed-name index + binary search (endswith → reversed prefix lookup). folded into add_remove win
5 Type-spelling fallback (diff_type_spellings) Rebuilt a set(...) inside a comprehension → O(n²). Hoist the set once. folded into add_remove win
6 Affected-symbol enrichment / ancestor closure (diff_filtering) Transitive ancestor function lists accumulated duplicates, then re-sorted per change → effectively cubic on nested type graphs. Use sets (dedup on union); sort once. nested_types n=200: >60 s → 0.16 s
7 ELF-only rename matching (binary_fingerprint, diff_symbols._plausible_rename) O(removed × added) name-similarity scan; the name predicate re-demangled both names per pair. Scan only the size-tolerance window via the existing size index; cache the per-name parse; cap the heuristic pass for mass-rename inputs. rename_churn n=1000: 13.2 s → 2.1 s, larger inputs bounded
8 Affected-symbol enrichment type↔function/field mapping (diff_filtering._build_type_to_funcs, _build_type_embed_index) any(tname in ft ...) nested inside the type loop and the function/field loop → O(types × functions × refs); quadratic when many distinct types churn (a header refactor or versioned upgrade). The original perf sweep only sampled these scenarios at n=500, so the exponent was never computed and the table mislabelled them "linear". One Aho-Corasick _SubstringMatcher over the affected type names, built once and shared; each ref/field is matched in O(len) with identical substring semantics. typedef_churn n=4000: 6.3 s → 0.73 s; union_churn 9.1 s → 1.10 s; vtable_churn 7.8 s → 1.03 s; enum_churn ≈1.8 → 1.0; opaque_filter ≈1.7 → 1.2 (all now linear)
9 Opaque-handle pointer-only / factory check (diff_filtering._is_pointer_only_type, _has_public_pointer_factory via _filter_opaque_size_changes) Each opaque candidate rescanned every public function/variable with a word-boundary regex → O(candidates × functions), a regex per pair (type_churn n=4000: ~3.2 M searches). One indexed pass per snapshot (_opaque_usage_index): an Aho-Corasick prefilter narrows each type string to the candidates present, then the same regex oracle decides — so the verdict is unchanged (verified by a fuzz test vs the per-candidate functions). type_churn n=4000: 1.13 s → 0.38 s (≈1.6 → linear)

With fixes #8 and #9 the compare pipeline has no remaining quadratic path at the tracked sizes — every scenario is linear (tail exponent ≈1.0–1.3) except the inherently deep nested_types chain. The opaque-handle pointer-only check (_is_pointer_only_type) used to be O(candidates × functions) with a word-boundary regex per pair (type_churn n=4000: ~3.2 M regex searches, ≈1.6); fix #9 replaced the per-candidate rescan with one indexed pass.

How to reproduce

No real binary, compiler, or castxml required — the harness synthesises AbiSnapshot pairs that exercise each path:

# Sweep all scenarios and print a table with a scaling exponent per scenario.
python scripts/benchmark_scaling.py

# Focus one path and emit machine-readable JSON.
python scripts/benchmark_scaling.py --scenario type_churn \
    --sizes 1000 2000 4000 --json-out reports/perf/scaling.json

Scenarios (add_remove is the linear control; the rest target a specific path). The first group exercises compare() (the original focus); the second group, added later, extends coverage beyond compare() to the suppression and reporting stages — see Coverage beyond compare():

Scenario Measures Stresses
add_remove compare() Core symbol diff + surface scoping (control)
type_churn compare() Affected-symbol enrichment, opaque filtering (structs)
enum_churn compare() Enum diffing (diff_types._diff_enums)
typedef_churn compare() Typedef base-change diffing (_diff_typedefs)
union_churn compare() Union member diffing
wide_struct compare() Per-field diffing within large records
vtable_churn compare() Vtable / virtual-layout diffing
elf_namespace compare() Namespace detection + demangling (stripped lib)
pe_churn compare() PE/COFF export diffing (diff_platform PE arm)
macho_churn compare() Mach-O export diffing (diff_platform Mach-O arm)
var_churn compare() Public-surface classification
rename_churn compare() ELF-only fingerprint rename matching — the reject path (disjoint names, no match emitted)
fuzzy_rename_churn compare() ELF-only fingerprint rename matching — the accept path (every symbol genuinely renamed → one func_likely_renamed per pair). The ICU/LLVM cost driver (P11: rename detection, not symbol count)
version_node_churn compare() Version-node migration fan-out — every export moves LIB_1.0 → LIB_2.0 → n symbol_moved_version_node findings (the LLVM 17→18 36,991-finding shape)
versioned_rename_churn compare() (collapse on) Versioned-symbol-scheme detection and collapse over 2×n churn (ICU/OpenSSL u_*_NN)
nested_types compare() Transitive type-ancestor closure
opaque_filter compare() Opaque-handle size filter (the known O(candidates × functions) residual)
suppression_audit SuppressionList.audit() Rule-vs-finding matching (O(rules × findings))
severity categorize_changes() Severity categorization of findings
serialize snapshot_to_json → from_dict Snapshot serialize/load round-trip (dump-pipeline proxy)
report_html generate_html_report() HTML document assembly
report_sarif to_sarif_str() SARIF JSON assembly
report_junit to_junit_xml() JUnit XML assembly

Peak memory

Every measurement also records the peak tracked heap (peak_mb, via tracemalloc) of the timed call. The inputs are built outside the traced window, so the figure attributes only the call's own allocations. The memory pass also runs cold: process-wide caches warmed by the timing loop (e.g. the functools.lru_cache demanglers) are cleared first, so input-scaled cache growth is counted rather than hidden behind a warm cache. A flat per-item time alongside a rising peak_mb flags an intermediate O(n²) space blow-up that a wall-clock-only gate would miss. Disable with --no-memory (timing only); gate with --max-memory-mb <budget>.

Alongside it, each point records the process peak RSS (rss_mb, via resource.getrusage). Unlike peak_mb, which only sees Python-heap allocations, RSS also counts native memory — pyelftools parse buffers and c++filt subprocess pages — which is what dominates real libraries (the field eval observed ~330 MiB RSS at LLVM scale, invisible to tracemalloc). ru_maxrss is a process high-water mark, so it is monotonic across sizes and the largest/last value is the true peak (it overstates a single call's own footprint, since inputs are built in-process); gate the peak with --max-rss-mb <budget>. RSS is unavailable on Windows (resource is Unix-only), where the column is simply absent.

Coverage beyond compare()

The original sweep (PR #331) only covered compare() post-processing. A follow-up gap analysis extended it to the two other stages that build the largest data structures from the finding set:

  • Suppression audit (suppression.py, SuppressionList.audit) tests every rule against every change — O(rules × findings). The suppression_audit scenario holds the rule count fixed (a project's ruleset is roughly fixed while its library grows) and scales findings, so it stays linear in findings; a regression that makes per-finding matching itself super-linear (e.g. recompiling a pattern per change) shows up as a rising exponent.
  • Reporting — to_markdown/to_json were already guarded by slow tests; report_html and report_sarif extend that to the HTML and SARIF renderers, which assemble the largest output documents. Both are linear.

Measured scaling (after fixes)

Most scenarios are linear at the sizes a real library reaches (per-change cost roughly flat); type_churn and enum_churn are mildly super-linear (~1.7) but bounded and tracked:

Figures are indicative local timings (absolute seconds vary with runner speed — the tail exponent is the portable signal). The first group times compare(); the second group, added in PR #336, times the suppression and reporting stages (see Coverage beyond compare()).

Scenario time @ size tail exponent
add_remove 0.32 s @ n=4000 ~0.9 (linear)
var_churn 0.06 s @ n=4000 ~1.0 (linear)
elf_namespace 0.33 s @ n=4000 ~1.1 (linear)
pe_churn / macho_churn <0.05 s @ n=500 ~1.0 (linear)
wide_struct 0.1–0.2 s @ n=500 ~1.0 (linear)
typedef_churn / union_churn / vtable_churn 0.7–1.1 s @ n=4000 ~1.0 (linear, after fix #8)
enum_churn 1.0 s @ n=4000 ~1.0 (linear, after fix #8 — was ≈1.8)
type_churn 0.38 s @ n=4000 ~1.0 (linear, after fix #9 — was ≈1.6)
opaque_filter 0.45 s @ n=1000 ~1.2 (linear at tracked sizes after fix #8)
rename_churn 2.1 s @ n=1000, capped above bounded
fuzzy_rename_churn 0.39 s @ n=4000 ~1.0 (linear)
version_node_churn 0.86 s @ n=10000 (10 k moves) ~1.0 (linear)
versioned_rename_churn 0.87 s @ n=8000 (16 k changes) ~1.1–1.2 (mild)
nested_types 0.70 s @ n=400 inherent for deep chains
suppression_audit 0.09 s @ n=2000 (fixed 40-rule set) ~1.0 (linear in findings)
severity <0.01 s @ n=1000 ~1.0 (linear)
serialize 0.12 s @ n=1000 ~1.0 (linear)
report_html / report_sarif / report_junit ≤0.04 s @ n≤2000 ~1.0 (linear)

CI integration

.github/workflows/performance.yml runs the scaling benchmark and the slow performance tests. Now that every compare() scenario is linear, the lane is gating:

  • Triggers: weekly schedule, manual workflow_dispatch (with size / budget inputs), and every PR (opened/reopened/synchronize/labeled) — but the expensive jobs (scaling, regression, header-graph-perf — schedule/dispatch only, see below — and l2-cli-perf, which also runs the header-graph PR-vs-base gate) only actually run when a classify job decides the PR touches performance-sensitive code, the classifier-job pattern (scripts/classify_perf_paths.py, tests/test_classify_perf_paths.py). This replaced an earlier pull_request.paths: trigger-level filter, for a real, not theoretical, reason: a trigger-level filter means the workflow's check run never registers at all on a non-matching PR, and the pattern list lived only in unversioned, untestable YAML glob strings — a real, perf-relevant change (the P0.2 Bazel aquery/cquery root-target scoping, the P0.3 auto-applied-L3-context-to-L2-headers change) merged with no Performance check run at all because the list hadn't been updated to cover it (see "Coverage gaps this workflow does not close" below). The classify job always runs, diffs the PR's changed files against PERF_SENSITIVE_PATTERNS — the detector core (abicheck/diff_*.py, checker.py, post_processing.py, demangle.py, binary_fingerprint.py, surface.py, ...), all of abicheck/buildsource/**, compare/dump orchestration (dry_run_estimate.py/service_input_resolution.py and the other service_*/cache modules — service_scan.py was renamed to dry_run_estimate.py and scan_engine.py was deleted outright, along with the rest of scan, in ADR-068 Phase 6), the benchmark scripts, and the perf tests — and reports a run output the four downstream jobs each gate on (if: needs.classify.outputs.run == 'true'). Adding the performance label force-runs the lane regardless of changed paths (fixed alongside the classify job — the label previously had no effect unless the changed paths also happened to match, despite an existing comment claiming otherwise); for a PR that touches neither, run it on demand with workflow_dispatch.
  • Armed budgets: the scaling step runs with --max-exponent 1.4 (the tail, largest-two-size slope) and --max-rss-mb 1024; the regression job blocks on a PR-vs-base slowdown exceeding max(15%, 100ms) per (scenario, size) point (--regress-tolerance 0.15 --regress-min-delta-seconds 0.1 — the "stable synthetic PR scenario" tier; a flat 50 % tolerance is an emergency stop, not real regression protection — several consecutive +15 % merges would otherwise nearly double the runtime before anything caught it). continue-on-error is dropped on both, so a catastrophic regression fails the lane; scripts/benchmark_scaling.py's own main() additionally fails closed if --baseline loads to zero points, or if zero (scenario, size) points end up shared between base and head — so a scenario-set/--sizes mismatch can't silently turn the comparison into a no-op that reports a clean pass. The thresholds are CLI flags so the budget lives in the workflow, not the script — loosen a threshold rather than re-adding continue-on-error if normal drift ever flakes a lane.
  • Per-scenario tuned sweeps, not one global --sizes. The PR/schedule triggers no longer pass --sizes at all, so each scenario's own tuned default sweep (SCENARIOS in benchmark_scaling.py) applies — onedal_large_surface reaches its tuned 5k-20k range, nested_types gets its full multi-point sweep (so a tail exponent is actually computable), and rename_churn/opaque_filter keep all three of their tuned points. A previous version of this workflow always passed a single 500 1000 2000 4000 ladder to every scenario regardless, silently overriding all of the above. --sizes (and the workflow_dispatch input of the same name) still works for a manual, deliberately-scoped run.
  • Median, not fastest-of-N. Each point runs one untimed warmup, then 5 timed repeats on every trigger (7 on the weekly schedule, since it isn't blocking a merge) — the reported/gated figure is the median of those repeats, with min/max/p95/coefficient-of-variation recorded alongside it (scripts/perf_measurement.py). Keeping the fastest of a couple of runs (the previous behaviour, --repeat 2 + min()) systematically favours the run least likely to have hit GC/scheduler noise — i.e. the one least representative of what a regression would actually look like — and gives no noise signal at all.
  • Base/head identity in the JSON report. --meta git_sha=... --meta side=base|head --meta os_image=... stamps the report each side was actually measured on, so a report downloaded later can be matched back to the exact commit/runner it came from.
  • The --max-exponent gate is per-scenario opt-out: nested_types is an inherently super-linear embedding chain, so it carries gate_exponent=False and is exempted (its tail slope is still printed for visibility, just not gated). Every other scenario is gated.
  • The exponent gate has a noise floor, checked on the slope's lower endpoint. A tail slope is only as precise as the smaller of its two points, so the gate applies only when both tail points take at least EXPONENT_FLOOR_SECONDS (0.2 s, scripts/perf_measurement.py). Until 2026-10 the floor was checked against the scenario's peak. A cheap scenario whose 4000 point had just crossed 0.2 s was then gated on a slope whose 2000 point sat at ~0.1 s in runner jitter, and report_sarif read 1.48 against the 1.4 budget on unchanged code (locally the same code read 1.18–1.35). The cheap gated scenarios (pe_churn, macho_churn, var_churn, the three report_*, fuzzy_rename_churn, onedal_mass_removal) now sweep one size step higher (TAIL_ABOVE_FLOOR_SIZES), so both tail points clear the floor and they stay gated. A scenario whose lower tail point later drifts under the floor prints exponent gate inactive with its tail value, never silently; raise that scenario's sizes to re-arm it.
  • Publishes the scaling table to the job summary and uploads the JSON.

slow regression guards also live in tests/test_performance.py — TestTypeChurnScaling (compare back to genuine O(n²)), TestSuppressionAuditScaling (audit stays linear in findings), and the HTML/SARIF cases in TestReporterScaling. They run in the existing slow lane with generous thresholds, so a catastrophic regression fails fast without flaking on normal drift.

The same workflow also carries a second, independent pair of jobs for the G31 Phase D header-graph attach-cost gate (scripts/check_header_graph_perf.py, see the G31 Phase D follow-up plan): header-graph-perf is report-only trend data on schedule/dispatch (on a pull request the trend point is the header-graph gate's own head measurement, uploaded under the same performance-header-graph artifact name, so a PR does not spend a second runner re-measuring the same head; no stable committed baseline number would survive a runner/toolchain change, the same reasoning check_mutation_score.py's SURVIVOR_BASELINE bootstrap avoids); The header-graph PR-vs-base gate (steps of the l2-cli-perf job since 2026-10, formerly its own header-graph-regression job on a separate runner) follows this page's own --baseline/--regress-tolerance same-runner base-vs-head pattern (see Baseline regression below) and gates from day one, since that pattern never needs a stale committed number to begin with.

Three measurement levels — which number means what

The single most common way to misread this page is to quote one level's number as another's. There are three harnesses, they measure genuinely different things, and none of them is a substitute for the others:

Level Harness What it measures Compiler? Interpreter startup?
1. Synthetic, in-process scripts/benchmark_scaling.py compare(), suppression audit, severity, serialization and the HTML/SARIF/JUnit renderers over hand-built snapshots. Scaling exponents, peak heap, process RSS. no no
2. Real L2, in-process scripts/check_header_graph_perf.py A real dumper.dump() of a real .so + synthetic header sweep, plus service._attach_header_graph's marginal cost. Gated as three separate metrics: dump_ms, attach_ms, total_ms. yes no
3. Full CLI scripts/check_l2_cli_perf.py The whole abicheck CLI as a subprocess against a real compiled C++ fixture, across the six supported L2 forms. yes (except the stored-operand forms, which must run none) yes

Boundaries worth stating explicitly, because each has been a real source of confusion:

  • Level 2's total_ms is not a CLI wall time. It is dump + attach in-process, and it excludes interpreter startup, config resolution, input resolution, serialization, comparison and rendering. It was previously named baseline_ms, which invited reading it as "the baseline cost of a run"; it is the cost of one phase.
  • Level 3's measured window is the subprocess's whole lifetime. Fixture compilation, snapshot pre-dumping for a stored-operand scenario, artifact download and dependency installation are all setup, measured separately and excluded. Every correctness check runs after the timed window closes.
  • Old binary-only numbers are not L2 numbers. The ~5 s libonedal_core.so dump quoted in this page's TL;DR is a historical binary-plus-DWARF figure. An L2 run of the same library additionally parses its public headers, and mixing the two is how a "dump got 10x slower" conclusion gets manufactured.

What the full-CLI harness asserts besides duration

A timing harness that checks only duration rewards the worst regression available to it: getting faster by doing less. Level 3 therefore gates correctness alongside time, and a scenario whose validation fails is a failure, never a fast data point. Per scenario, outside the timed window:

  • the resolved evidence depth really reached headers on every side (a binary-only fallback still accepts --depth headers, and it is faster);
  • public scoping resolved and did not fall back;
  • the expected declarations survived into the snapshot (names, not counts — a count can be matched by a degraded parse that found a different set);
  • the header call, include and type graph extractor passes all ran and none is recorded as degraded, and the include graph collected at least one edge;
  • the fixture's deliberate break produced both a removal-family and a layout-family finding — a much stronger claim than "the verdict was BREAKING", which a single unrelated finding can satisfy;
  • the unchanged control produced neither, i.e. no false positive;
  • --no-baseline read as an audit, with no manufactured compatibility verdict.

Native invocations are observed, not assumed

perf_receipt.NativeInvocationSpy prepends a directory of exec-ing shims to PATH, one per spied tool, each logging its own argv before handing off to the real binary. classify_invocation then bins each observed invocation as header extraction, include pass, or version probe — a distinction that is load-bearing rather than cosmetic: a cold L2 dump makes three castxml calls and only one of them is a parse (the others are --version and -dumpmachine), so a flat per-tool count cannot tell a cached parse from a skipped probe.

That turns three otherwise-unfalsifiable claims into measurements:

  • a stored-snapshot/stored-snapshot comparison performs zero header extractions and zero include passes — measured, not assumed;
  • a stored/live comparison performs exactly one side's worth, so the operand handed over as a snapshot is demonstrably not re-extracted;
  • a run labelled "warm cache" really was served by a cache, proven by its extraction count dropping rather than by the fact that it ran second.

The last one matters because on a small fixture a second run served by nothing is indistinguishable from a warm one by wall time alone.

Three rules about which runs those measurements cover, each of which started as a defect where one favourable observation certified a batch:

  • a forbidden contract is a zero over every invocation kind, not only the extraction bucket. "No compiler ran" is the claim, so an include pass or a version probe falsifies it exactly as a parse does.
  • a live contract additionally requires the include-graph pass to have run, at least once per live side. Header-AST extraction is not the whole of the measured L2 work, and a run that stops doing the clang -M pass is faster while still resolving depth headers and still finding the deliberate break.
  • every cold/warm repetition is checked individually, paired by index, and the reported cache service is the worst repetition. Reducing each batch with min() let one warm repetition certify a scenario whose others re-extracted in full, so the gated median could describe an uncached run under a receipt claiming a served cache.

The same "every repetition, not the lucky one" rule governs correctness: the scenario's semantic validation runs at the end of each repetition, against the outputs that repetition just wrote, and every file an invocation is declared to produce is deleted beforehand and required afterwards. Validating once at the end inspected only the final report, so an earlier repetition that emitted degraded evidence while still writing a file and exiting with an allowed code kept its faster timing in the median whenever the last repetition happened to be correct.

Cache states are three things, not two

The harness separates, and never conflates:

  • cold application cache — a fresh process with a fresh XDG_CACHE_HOME. This is explicitly not a cold disk: the OS page cache still holds the fixture and the interpreter, and no attempt is made to drop it. Dropping a CI runner's page cache needs privilege and would measure the host.
  • warm AST cache — a fresh process against the same cache root and byte-identical inputs.
  • invalidated — the same again after a transitive dependency header changes. The cache must not serve; if it does, the product is reusing stale evidence, which the harness reports as a failure rather than a speedup.

Which one actually served is read off the counters (observed_cache_service: none / partial / full), never inferred from run order.

Memory is reported as two differently-derived numbers

sampled_peak_tree_bytes is the largest simultaneous sum of RSS across the measured process and every descendant alive at one sampling instant, taken from a parent-side sampler at a recorded interval. It is deliberately not called a peak:

  • a spike shorter than the sampling interval is missed entirely;
  • pages shared between processes are counted once per process, so it can also overstate real physical usage;
  • a descendant that exits between two samples is never seen.

ru_maxrss_bytes is carried separately because it is a different thing: the kernel's own high-water mark, which never misses a spike but is a maximum over individual processes and so cannot see two live children's combined footprint. Both are reported, labelled, with the interval and the observation limits — picking one would hide the other's failure mode.

ru_maxrss_bytes comes from RUSAGE_CHILDREN, which is cumulative over every child the harness has reaped, so it is reported only for the step that raised it. A step that did not raise it gets null plus a ru_maxrss_scope naming the earlier, heavier child that holds the mark. Without that rule, one 2 GB step makes every following step report 2 GB as its own RSS — which is exactly what the first published PVXS receipt did, showing 1.9 GB against a run whose sampled tree peak was 438 MB. Read sampled_peak_tree_bytes for such a step.

Measured cost of the full-CLI lanes

All figures local (gcc 13.3.0 / castxml 0.7.0 / clang 18.1.3, 4 CPUs, Linux, --repeat 3 unless stated). Reproduce with the commands in the harness's own module docstring. These are lane costs, not per-scenario costs; per-scenario numbers are in the receipt.

Lane Wall User CPU Native invocations Fixture build (setup, excluded)
--suite pr 44–48 s 40–43 s 28 header extractions, 14 include passes, 100 probes ~0.9 s
--suite extended (--repeat 1) ~65 s — — ~4.7 s
--suite extended (--repeat 3) ~180 s — — ~5.1 s

The PR lane's cost is dominated by interpreter startup, not by analysis: the lane makes roughly 30 CLI invocations (8 gated steps plus setup and resolution steps, times 3 repeats) at ~0.6 s of startup each, so well over a third of the lane is spent before any evidence work happens.

Per-repetition validation (each repetition's own outputs checked, rather than only the final report) triples the validation work, all of it outside every timed window. It does not materially change the lane: re-measured at 43 s wall, inside the range above. That measurement was taken while the extended suite ran concurrently on the same 4-CPU host, so read it as an upper bound — which is what makes it usable here, since an upper bound inside the existing range is enough to say the range still holds. The range is deliberately left as it was rather than narrowed to a contended number.

Per-scenario gated full_cli medians on that fixture, with the coefficient of variation that sets the gate's noise floor (--repeat 3, except the two-format row, re-measured at --repeat 1 after the export grammar landed):

Scenario Step Median cv
dump_l2 dump 1.074 s 6.4%
compare_live_live compare 1.219 s 3.9%
compare_stored_live compare 1.389 s 5.3%
compare_stored_stored compare 1.272 s 15.9%
compare_no_baseline audit 1.125 s 11.6%
compare_two_formats compare_exporting_two_formats 1.140 s —
compare_live_live (unchanged) compare 1.327 s 6.1%

Those cv figures (up to ~16%) are why the lane's absolute floor is 0.5 s rather than 0: a purely relative 30% tolerance on a ~1.2 s measurement would be only ~0.36 s, inside what this fixture's own run-to-run variance already covers.

On the extended axes (--repeat 1): templates ~1.36 s, 8 headers ~2.48 s, 32 headers ~10.4 s, and five libraries ~1.15–1.50 s each — the same per-library cost whether they share one dependency header or each have their own, because the header-frontend invocation count follows the top-level headers rather than their dependencies. The shared arm resolves through one physical file at a common include root, and header_contexts is counted from the resolved paths actually built rather than from the flag that asked for them; an earlier version gave each library its own byte-identical copy, which (the AST cache keying on resolved path) made the "shared" arm a second distinct-path workload.

The multi-library set is measured as one cache lifecycle — reset once before the set, not before each member — since otherwise each library's comparison starts from an empty cache and cross-library reuse is unobservable by construction, which is the only thing separating the shared arm from the distinct one. With that in place the measurement says something it previously could not: at --repeat 2 every member of both arms performs 4 header extractions and 2 include passes, identical in the shared and distinct arms, so no cross-library reuse happens today even when five libraries resolve one physical dependency header through a common include root. That is an observation about the product, not a harness gap, and it is recorded here rather than acted on: this work deliberately changes no caching strategy (see "Scope" above). It is the cross-library half of the ⚠️ row for bundle/multi-library orchestration in the coverage table below.

One comparison, two artifacts. -o FORMAT=DESTINATION is repeatable and every export renders the one completed analysis (ADR-068 slices 7m/7n), so a JSON report and a human report come from a single compare. The harness verifies that rather than assuming it: over a stored-old/live-new pair the two-export invocation performs exactly one side's worth of header extraction, so the second renderer provably does not re-run the analysis.

Instrumentation overhead, and a worked example of why ordering matters. Measured by running the identical lane with and without --no-spy.

A single unordered pair (spy, then no-spy) gave +2.1 s wall (+4.6%) and +2.3 s CPU — a plausible-looking result, and one it would have been easy to publish. Repeating it in ABBA order (spy, no-spy, no-spy, spy), so drift across the sweep cannot be read as a configuration difference, gave medians of 45.2 s with the spy against 46.0 s without it: −0.7 s (−1.6%), i.e. the opposite sign, against a largest within-configuration spread of 2.2 s.

So the honest statement is that the spy's overhead is not resolvable above run-to-run noise on this host, bounded by roughly ±5% of the lane, and the first measurement's +4.6% was noise wearing a plausible number. Mechanically that is what one expects: ~142 extra shim invocations per lane, each a /bin/sh startup plus a printf plus an exec, against castxml parses that each cost hundreds of milliseconds.

Note also that --no-spy disables every extraction-count assertion, so it is a measurement aid, never a cheaper way to run the lane.

L2 scaling gate (headers x libraries)

scripts/check_l2_scaling_perf.py runs in the l2-cli-perf PR job after the PR-vs-base comparison. That comparison measures one small fixture, so it catches a constant-factor regression but not a change in shape. This gate sweeps two axes through the real CLI and gates how cost grows:

Axis Operation Sizes Budget (marginal exponent) Measured (2026-09, 4 CPUs)
headers compare --depth headers, one library 1 / 4 / 12 / 30 headers 1.4 0.95 (1.3 s → 2.6 s)
libraries directory compare (release fan-out), 2 headers each 1 / 3 / 6 / 10 libraries 1.7 1.39 (1.2 s → 8.6 s); 1.64 (→ 11.5 s) before the fixes below

Peak process-tree RSS is gated at 1024 MB per point (observed ≤ 365 MB). The whole gate takes about 75 s at --repeat 3.

The gated number is the marginal exponent: the slope of log(wall(n) - wall(1)) against log(n - 1). Every CLI run pays a ~1 s fixed floor (startup, imports, config, report writing), so a raw log-log slope at these sizes reads close to 0 whatever the product does. Subtracting the n = 1 floor measures the work the axis adds: a linear step reads ~1.0, and a step that redoes all previous units' work reads ~2.0. If the largest point is less than 0.25 s above the floor, the sweep fails as unfittable rather than passing.

The library axis is still super-linear, and the budget does not bless that. Every member of a directory compare is dumped against the release's union header and include set, so any per-member step that walks that set grows with the member count. The total then grows faster than linearly. The two largest such steps are now memoized:

  • contract fingerprinting's path resolution (comparability_fields), about 35% of wall time at 16 libraries;
  • the C++20 dialect scan (extract/header_scan_memo.py), which ran several times per dump over the identical set.

The first measurement blamed bundle_symbol_status → qualified_name_segments_walk. That was a profiling artifact: cProfile attributed worker threads' time wrongly, and py-spy corrected it. Before → after, 4 CPUs, cold cache: 6 libraries 5.0 s → 4.2 s, 10 libraries 11.5 s → 8.6 s, and 16 libraries 32.2 s → 19.1 s. The 1.7 budget catches regression; lower it toward ~1.1 once members stop receiving the union set. Recorded in Known gaps.

Real-integration profiles (oneDAL, SVS, PVXS)

scripts/l2_real_profiles.py pins the live integrations declaratively: each profile's revisions (and how each side's operands are obtained), the libraries and headers actually in L2 scope, the required tools and approximate build cost, and reproducible prepare commands. They are periodic/manual only — an ordinary PR must never download and build oneDAL (~120 build-minutes, ~25 GB).

SVS carries two profiles, deliberately not folded into one: svs compares the released v0.4.0 runtime distribution against PR #387's head — the comparison the integration actually gates on, and the one that exposed the real ABI change — while svs_pr_base compares the PR's merge base against its head. The latter is a useful additional smoke test and cannot substitute for the former, since it cannot expose a change that entered the branch before the merge base. The released side is consumed, not rebuilt: rebuilding a release from its tag measures the rebuilding host's toolchain rather than the artifact consumers received, and when the release and CI artifacts already exist there is no reason to rebuild at all.

The rule the module exists to enforce: an unavailable profile is reported PARTIAL/BLOCKED/NOT_RUN with a concrete reason, never silently replaced by a synthetic substitute, and a synthetic number is never published under a real project's name. Five further constraints it encodes:

  • No declarative L2 bundle capability exists today. A multi-library profile is measured as the supported set of per-library L2 operations — the set's total, each library's own cost, the number of distinct header contexts, and how much work was genuinely repeated across them. That is not a bundle scan and the module never calls it one. Adding a bundle capability is product work.
  • A library with no public API of its own is not an L2 case. oneDAL's libonedal_thread is recorded as a non-case with a stated reason, rather than inflated into a sixth L2 library by pointing it at someone else's headers.
  • A historical baseline must be historical. validate_side_headers rejects a plan that resolves both sides' headers to one root — the easy accidental substitution (check out the new revision, build both binaries, point both --header sets at the working tree) runs fine, is faster, and is not a temporal L2 comparison. PVXS must also not be measured by running the project's own script with --depth source: that is an L4/L5 measurement. It must also be the baseline the integration declares — which is why SVS's release-to-candidate and PR-base comparisons are separate profiles.
  • Readiness is not a measurement. resolve_status answers only whether a host could measure a profile, so its positive answers are READY and PARTIAL. It once returned MEASURED for any request whose prepared_root merely existed — an empty directory, neither side's library, neither side's headers, no comparison run. Every declared operand is now checked on both sides (missing_inputs), and MEASURED is reachable only through promote_to_measured, which requires a completed timed window with validated output.
  • A scenario's findings mean what that scenario says they mean. Every declared scenario carries an expectation (SCENARIO_EXPECTATIONS). "Any finding is a false positive by construction" holds for a literal self-comparison and for nothing else: two independent builds under one controlled contract are to be investigated against the recorded compiler, flags, dependencies and artifact evidence, and two intentionally different build variants differ in contract on purpose. SVS's own PR artifacts make the point — the default and public-only builds have byte-identical runtime headers and materially different exported-symbol sets — so identical header text plainly does not guarantee identical binary evidence. Treating all three alike risks reading correct detection of a build-induced ABI change as a scanner defect, or suppressing it to satisfy the wrong expectation.

oneDAL solo L2 compare across the 2026-09 perf series

User-supplied receipts for one identical command (solo run, 125 header roots, fresh cache) at four revisions. The operands are not in this repository and no CI lane reproduces this run, so these are reference expectations for the manual real-integration profile, not gated numbers:

Revision Wall Peak RSS Exit Verdict lambda at spellings
0.6.0 2:43:53 15,176,116 KB 4 BREAKING 146
main e9d820797 11:47.46 16,239,532 KB 4 BREAKING 290
main 963138528 11:04.71 16,525,332 KB 2 API_BREAK 0
main 577d856a4 7:49.88 16,525,160 KB 0 NO_CHANGE 0

Read it as three signals, not one. Wall time fell ~21x since 0.6.0 and ~33% across the last step. Peak RSS did not fall (15.2 → 16.5 GB) — the series bought time, not memory, and ~16 GB sits at the edge of a nominal 16 GB host (see memory.md). And the verdict moved BREAKING → API_BREAK → NO_CHANGE as checkout-path-dependent lambda at spellings were stripped (#1343, #1355): at this scale, a correctness regression that re-introduces path-dependent identity shows up first as a spurious verdict, which no synthetic lane below would catch. A re-measurement should record all five columns, not wall time alone.

Complexity and cost gates beyond wall-clock time

Timing exponents need several sizes, repeats and the slow lane. Most real regressions also have a cheaper, exact symptom, so these gates measure that instead and run where noted. All of them share the synthetic workloads in tests/_compare_workloads.py (signature, rename, enum, variable, nested-type and type churn, add/remove), whose every entity population grows with n.

Gate What it pins Lane
tests/test_compare_call_complexity.py No first-party function's call count grows faster than (size ratio)^1.5 in compare(), for every workload in three modes (default, --contract evaluation, pattern verdicts + surface metrics). Names the path:line(function). Oracle is the input's size ratio, not a recorded baseline. unit
tests/test_pipeline_call_complexity.py The same, for snapshot serialization round trips, every report format (JSON, Markdown, SARIF, HTML, JUnit) over real findings, and release reconciliation as the member count grows. unit
tests/test_compare_cost_budgets.py + tests/perf_call_budgets.json Exact ratchet budgets per mode: same-argument repeat calls of a reviewed list of expensive functions (graph/surface/idiom builders, demangle_batch, ...) and child processes per compare(), which must also not grow with input size. A figure above or below its budget fails; re-record with python scripts/audit_repeated_calls.py --write-budgets. unit
tests/test_extract_call_complexity.py + tests/_cpp_corpus.py Call-count complexity of dump() over a generated real C++ library (namespaces, virtuals, overloads, templates, typedef chains), and of compare() over two dumped versions at 10% and 100% churn. integration
perf-antipatterns (scripts/perf_antipatterns.py) No new list-membership, re.compile, json.loads/deepcopy, self-copying accumulation, subprocess call, loop-invariant re-sort/copy or str += inside a loop under abicheck/; existing sites are per-function counts in scripts/perf_antipatterns_baseline.json. ai-readiness
tests/test_history_scaling.py build_longitudinal_history stays linear in the number of releases (call counts, K=5 vs 20) and its time exponent over up to 50 releases stays sub-quadratic. unit + slow
tests/test_compare_scaling_shapes.py Wall-clock exponent per workload shape -- catches a linear number of calls whose per-call cost grows. slow

The first of these found a real functions x types scan in --pattern-verdicts (per-finding rebuild of every type name), and the repeated-call audit found the namespace detectors demangling each snapshot once per detector per side; both are fixed.

Investigating by hand

  • Which functions are called repeatedly with the same arguments? python scripts/audit_repeated_calls.py --workload type_churn --n 400 --top 30 (plain values compare by value, other objects by identity; generator resumptions are not counted).
  • Which call counts grow with input? In a test or REPL, call profile_call_counts (tests/_call_counts.py) at two sizes and pass both tables to superlinear_call_sites. Salt each run's names (the workloads' tag argument): demangling and canonical-spelling caches are process-wide.
  • Where does the time go? python -m cProfile -o out.prof -m abicheck compare OLD NEW, then python -m pstats out.prof (sort cumulative, stats 40); for a flame graph of a live run, py-spy record -o flame.svg -- abicheck compare OLD NEW (py-spy is not a dependency; install it ad hoc). For memory, see memory.md and ABICHECK_MEMORY_TRACE.
  • Every anti-pattern site, baseline or not: python scripts/perf_antipatterns.py --all.

Coverage gaps this workflow does not close

An external performance audit (2026-08) found that compare()/dump/scan scaling has real CI protection (this page's own subject), but two adjacent lanes had false-green failure modes — a check reporting success while actually verifying nothing. Both are fixed; recorded here so a future reader doesn't rediscover them from scratch.

The eval-suite.yml source-tier (L3/L4/L5) lane reported success while scanning zero libraries. skills-src/evaluation/field/runner.py's _dump_sources() called abicheck dump --sources ... --depth full — a rung retired from the public CLI (ADR-043 D2, collapsed into --depth source; see the "Scan-level scalability sweep" note above for the same retirement). Every source-tier scan therefore failed identically with the same click.BadParameter, and the source-tier job's blanket continue-on-error: true (deliberately set, since one library's own build/network flakiness is meant to be tolerated — see the job's own comment) meant a 0-of-N-scanned run still reported as a passing, green job, with REPORT.md's source-tier table still published looking like real coverage. Fixed two ways, matching this page's own "gate on the systemic signal, not the individual one" pattern: the --depth full → --depth source argv fix itself, and a new skills-src/evaluation/field/runner.py --fail-on-empty-source gate (source_tier_broken()) that fails only when the whole tier is broken — zero libraries scanned successfully, or every "successful" scan captured zero real L3 build evidence — while still tolerating one library's own failure exactly as before. continue-on-error is no longer set at the job level; the new gate is what now distinguishes "tolerable per-library flake" (still passes) from "the tool itself is broken" (now fails loudly).

Still open, deliberately not attempted in this pass (each needs its own scoped design, per this repo's "known gaps over risky reactive patches" convention — root AGENTS.md):

  • True interleaved base/head measurement. The regression job runs the base branch to completion, then the head branch to completion, on the same runner — not alternating scenario-by-scenario. Interleaving would average out slow drift over the run's wall-clock (thermal throttling, a noisy neighbour) that a strictly-sequential base-then-head comparison cannot distinguish from a real regression. Doing this soundly needs each benchmark_scaling.py invocation to run a single repeat and be invoked alternately from the workflow (or a driver script that shells out to both venvs in turn), then aggregates medians across rounds — a real, separate piece of orchestration, not a flag on the existing single-shot invocation.
  • No scheduled real-scale lane. Every gated lane is synthetic and small (≤ 20k functions in-process, ≤ 30 headers / 10 libraries through the CLI — see "L2 scaling gate" above). The oneDAL receipts above — minutes of wall time, ~16 GB RSS, and a verdict that depends on path-independent identity — are reproducible only by hand via scripts/l2_real_profiles.py, and scripts/bench_release_memory.py, bench_graph_materialization.py and bench_extraction_scope.py are wired into no workflow. A peak-RSS regression on a real multi-GB AST, or a scale-only identity/verdict drift, therefore has no automated guard.
  • A maintained end-to-end depth/backend matrix. The slow perf tests and the scaling/header-graph harnesses cover compare() and the L2 attach cost well; there is no equivalent maintained CI matrix over binary / headers (clang + castxml) / build (CMake + Bazel scoped/fallback) / source (seeded + unseeded) × cold/warm cache. skills-src/evaluation/field/scan_level_scaling.py sweeps the level axis but is manual-only (real clang time), and skills-src/evaluation/field/scaling.py's ABICHECK_L4_JOBS sweep is likewise manual.
  • Per-run performance receipts — now implemented for the L2 CLI path (scripts/perf_receipt.py, consumed by scripts/check_l2_cli_perf.py). A run now emits a versioned receipt carrying wall/user/sys time with its CPU accounting scope named, sampled concurrent process-tree RSS alongside a separately-labelled ru_maxrss, nested phase windows, per-kind native invocation counts (header extraction vs. include pass vs. version probe), output sizes, correctness-validation status, and the effective thresholds that gated the run. What is still open is the wider depth/backend coverage below, and an L3/L4/L5 equivalent: the receipt layer is generic, but only the L2 CLI harness feeds it today.
  • Repeated L3 collection under dump --sources/--build-info (P0.3) — open, accepted cost, not this PR's to close. Include seeding and compile-context derivation (derive_l2_include_dirs/ derive_l2_compile_context, abicheck/buildsource/l2_seed.py) each independently trigger a full L3 collection (collect_inline_pack(layers=("L3",))) for the exact same (sources, build_info) inputs — up to three L3 collection passes per side for some input shapes, counting the embed step's own separate collection. A sentinel-based evidence= sharing mechanism was prototyped in an earlier revision of this PR, then dropped: main had independently landed a parallel, far more heavily reviewed P0.3 pass in the interim (10+ review rounds on derive_l2_compile_context alone — forced-language interaction with a matched compile unit's derived -std=, MSVC /std:-vs--std= precedence, ambiguity-signature narrowing order, ...) whose own _L2SeedPackArgs docstring already considered and explicitly declined folding these two calls into one, calling the double collection "an accepted, documented cost" so each derive_l2_* function's return shape stays independently additive. Re-litigating that call from inside a merge-conflict resolution, in code this delicate and this recently stabilized, was judged the wrong place to relitigate it — see this repository's own "known gaps over risky reactive patches" convention. This is now in scope for the classify job's PERF_SENSITIVE_PATTERNS (see "CI integration" above), so a change to it will run this workflow even though no benchmark scenario targets its specific cost yet. Closing it for real needs its own dedicated pass, reviewed on its own terms rather than folded into an unrelated CI-tooling PR.
  • Per-job shard classification. The classify job (see "CI integration" above) reports one shared run output gating every downstream PR job uniformly — the exact same set of jobs a changed path used to trigger together under the old paths: filter, just moved into tested code. Splitting PERF_SENSITIVE_PATTERNS into separate shards (e.g. one gating only scaling/regression, another gating only the two header-graph-* jobs) would cut CI cost for a PR that only touches one area, but was deliberately not attempted: check_header_graph_perf.py imports abicheck.buildsource.header_graph (reached through service._attach_header_graph), so even a directory as specifically-named as abicheck/buildsource/ is relevant to the header-graph jobs, not just the compare()-scaling jobs its name suggests — a real per-job split needs a verified transitive-import trace from each script's own entry point, not a guess from a module's directory or filename. A wrong split would silently under-cover one shard, exactly the failure mode the classify job exists to close, so it's tracked here as future work rather than attempted speculatively.

Coverage gap analysis & remaining gaps

A second pass (continuation of PR #331) audited the whole pipeline for scaling risk and extended the harness to the highest-value uncovered paths plus per-call peak-memory tracking and PR-vs-base drift detection. Current status:

Path Status Notes
compare() post-processing ✅ covered Original PR #331 scenarios.
Suppression audit ✅ covered suppression_audit scenario + slow test. O(rules × findings); linear in findings for a fixed ruleset.
HTML / SARIF / JUnit reporting ✅ covered report_html / report_sarif / report_junit scenarios + slow tests; all linear. (to_markdown/to_json already guarded.)
Enum / typedef / union / wide-struct / vtable diffing ✅ covered enum_churn, typedef_churn, union_churn, wide_struct, vtable_churn. Sweeping typedef/union/vtable/enum across sizes (the original table only sampled n=500, so no exponent was ever computed) exposed a genuine ≈O(n²) in the affected-symbol enrichment — see fix #8; all four are linear after it, and opaque_filter dropped from ≈1.7 to ≈1.2 as a side effect (its cost was the enrichment, not _filter_opaque_size_changes).
PE/COFF & Mach-O diff arms ✅ covered pe_churn / macho_churn build pe=/macho= snapshots so diff_platform's PE/Mach-O detectors run.
Opaque-handle pointer-only check ✅ covered Was the O(candidates × functions) residual (_is_pointer_only_type, regex per pair, surfaced by type_churn ≈1.6); fix #9 linearized it via _opaque_usage_index (one indexed pass). Both type_churn and opaque_filter are now linear.
Versioned-symbol-scheme collapse (ICU/OpenSSL) ✅ covered versioned_rename_churn reproduces the field-eval P08 ICU 75→78 shape (16 k removed/added churn findings + the scheme-collapse pass). Profiling it surfaced a per-finding name re-tokenization in the namespace detectors (diff_namespaces._segments), now fast-pathed for plain names. ~1.1–1.2 tail exponent; the residual is the post-processing detector fan-out, not the scheme recogniser.
Severity categorization ✅ covered severity scenario over categorize_changes; linear.
Fuzzy rename matching (accept path) ✅ covered fuzzy_rename_churn — every symbol genuinely renamed → one func_likely_renamed per pair, the cost driver P11-refined identified (ICU 2134 renames = 94.5 s; rename detection, not symbol count, dominates). The pre-existing rename_churn only exercised the reject path (disjoint names, zero matches). Linear at ICU scale (≤8 k).
Version-node migration fan-out (LLVM bump) ✅ covered version_node_churn — every export moves LIB_1.0 → LIB_2.0, reproducing the LLVM 17→18 36,991-symbol_moved_version_node shape and the post-processing fan-out over it. Linear to 50 k.
Peak memory (all scenarios) ✅ covered tracemalloc peak_mb column + --max-memory-mb gate (cold-cache pass), plus process rss_mb (resource.getrusage) + --max-rss-mb gate — RSS catches native (pyelftools / c++filt) allocations tracemalloc cannot see (the ~330 MiB LLVM-scale figure).
Historical / PR-vs-base memory regression ✅ covered (gating) --regress-memory-tolerance/--regress-min-delta-mb + the memory half of the regression workflow job compare peak tracked heap against the base branch under max(20%, 4 MiB). Before it, memory was gated only against an absolute ceiling, which a doubling well under that ceiling passed silently. See Memory regression.
Historical / PR-vs-base regression ✅ covered (now gating) --baseline/--regress-tolerance + the regression workflow job measure the base branch and PR head on the same runner and flag scenarios that got slower by more than the tolerance — catching gradual drift the per-run exponent misses. continue-on-error is dropped, so it now blocks. See Baseline regression.
Dump / snapshot creation (DWARF/PE/PDB) ⚠️ partial The synthetic harness can't run the real parsers. The ELF symbol-table parse and the DWARF debug-info parse (-g build) are now guarded by tests/test_perf_dump_scaling.py (integration, gcc-only) — DWARF being the dominant real-library dump cost (ICU 18.6 MB snapshot, openblas 23 MB / 9.5 s). The serialize scenario proxies the rest of the pipeline. PE/COFF + PDB parsing remains unbenchmarked — those need a committed binary or a synthetic byte-stream generator (no Linux-only toolchain produces them).
Appcompat HTML / stack analysis / appcompat filtering ⚠️ not benchmarked stack_checker runs one compare() per dependency (inherent). Appcompat filtering uses set-membership lookups (appcompat.py — O(1) per change, likely already fine) and appcompat_html.py is linear by inspection; neither is timed.
Directory multi-library compare (release fan-out) ✅ covered (scaling) check_l2_scaling_perf.py's library axis: 1 / 3 / 6 / 10 libraries through the real CLI, marginal exponent gated at 1.7 (measured 1.39).
Bundle / environment-matrix compare ⚠️ not benchmarked Per-library cost is covered; bundle and environment-matrix orchestration is not.
  1. ~~Wire a budget gate~~ — done: the lane now runs --max-exponent 1.4 (nested_types exempt via gate_exponent=False) and --max-rss-mb 1024, and continue-on-error is dropped on both the scaling and regression jobs. The --regress-tolerance 0.5 PR-vs-base check also blocks now; loosen a threshold rather than re-adding continue-on-error if runner variance flakes a lane.
  2. Extend the dump/parse guard to PE/PDB — the ELF symbol-table and DWARF parses are now covered (tests/test_perf_dump_scaling.py, integration, gcc + -g); the PE/COFF and PDB parsers still need a committed binary or a synthetic byte-stream generator behind the integration marker (no Linux-only toolchain emits them).
  3. Benchmark bundle / environment-matrix orchestration — directory multi-library scaling is covered by the L2 scaling gate. Bundle and environment-matrix orchestration, including appcompat/stack fan-out, is still untimed.
  4. ~~Optimise the super-linear residuals~~ — done: fix #8 linearized the enrichment (typedef/union/vtable/enum/opaque), fix #9 the opaque pointer-only check (type_churn). No quadratic compare() path remains at tracked sizes.

Baseline regression

The per-run scaling exponent catches catastrophic blow-ups but not a gradual 15–20 % slowdown. To catch drift, the harness can compare against a baseline:

# On the base branch / a prior commit, capture a baseline (5 repeats -> a
# meaningful median/cv; --meta records who/what this measurement is):
python scripts/benchmark_scaling.py --repeat 5 \
    --meta git_sha=$(git rev-parse HEAD) --meta side=base \
    --json-out base.json

# On the PR head, measure and compare (fails if any shared point regresses
# by more than max(15%, 100ms) vs. its baseline):
python scripts/benchmark_scaling.py --repeat 5 \
    --meta git_sha=$(git rev-parse HEAD) --meta side=head \
    --baseline base.json \
    --regress-tolerance 0.15 --regress-min-delta-seconds 0.1

Memory regression

The rule above gates time. Peak memory has had the same treatment since the memory gate was added (originally its own memory-regression workflow job, now the second half of the regression job):

# Same two-step shape, with memory tracking left ON (no --no-memory), and
# --repeat 1 because a tracemalloc peak is an allocation count, not a
# wall-clock duration, so repeats buy far less than they do for timing.
python scripts/benchmark_scaling.py --repeat 1 --json-out base-memory.json

python scripts/benchmark_scaling.py --repeat 1 \
    --baseline base-memory.json \
    --regress-tolerance 100 \
    --regress-memory-tolerance 0.20 --regress-min-delta-mb 4

A point regresses once its peak exceeds the baseline's by more than max(--regress-memory-tolerance x baseline, --regress-min-delta-mb) — the same combined relative/absolute rule as the timing gate, computed by the same perf_measurement.combined_regression_threshold, so the two cannot drift apart. The memory defaults are tighter (20 % / 4 MiB, against 50 % / 0 s): a tracemalloc peak counts bytes the interpreter actually allocated and does not move with GC timing, scheduler preemption or a cold cache, so the noise that forces a loose timing tolerance is largely absent. Baseline peaks below an 8 MiB floor are skipped, for the same reason the timing rule ignores sub-50 ms points: at that size, fixture and import allocations dominate the figure rather than the code under test.

The memory baseline is read from the same --baseline report — a benchmark_scaling.py report already carries peak_mb on every point, so there is no second file to keep in sync. A baseline produced with --no-memory (or by a build predating this gate) carries no peak_mb at all; that is reported as an inactive memory gate rather than treated as "allocated nothing", which would otherwise flag every point in every run. The timing gate still fails closed on an empty baseline, so a wholly missing or malformed baseline is still caught.

Why this is a separate measurement run (the memory half of the regression job, not another flag on its timing run): tracing the heap is not timing-neutral. measure()'s memory pass clears every live lru_cache between sizes and runs an extra untimed cold call, which is precisely the bias the timing job's own --no-memory comment documents. Timing and memory therefore cannot be gated from one run. In the memory run memory is on for both sides, so that bias applies equally and cancels; the timing tolerance there is explicitly neutralised (--regress-tolerance 100) so a distorted timing can never fail the memory gate. The two runs need two runs, not two runners: they share one job, one checkout pair and one pair of venvs, and the memory steps start only after the timing steps have finished, so tracing never overlaps a timed measurement. (Until 2026-10 they were two jobs, which checked out and installed both sides twice on two runners for no measurement benefit.)

What it closes. peak_mb was recorded long before it was gated, and was checked only against the absolute ceilings --max-memory-mb/--max-rss-mb. An absolute ceiling catches a regression only once it crosses the ceiling, so a change doubling a scenario's allocation from 200 MiB to 400 MiB passed the then-2048 MiB ceiling in silence — exactly the gradual drift the timing side had had a base-branch comparison for since PR #768.

Each point's reported/gated figure is the median of its timed repeats (plus one untimed warmup) — not the fastest one; see scripts/perf_measurement.py for why "keep the minimum across a couple of runs" systematically hides regressions rather than catching them, and reports no run-to-run noise signal at all (min/max/p95/coefficient-of-variation are recorded alongside the median for exactly that reason).

Only scenarios present on both sides are compared (a scenario new in the PR has no baseline and is skipped), and baseline times below a 50 ms noise floor are ignored regardless of the threshold rule. A point regresses once its absolute slowdown exceeds max(tolerance x baseline, --regress-min-delta- seconds) — the combined relative/absolute rule protects a small baseline (a few ms) from flagging on ordinary noise while still catching a genuinely large percentage regression on a large baseline; --regress-min-delta-seconds 0 (the default) reduces to a pure percentage tolerance. A --baseline that loads to zero points, or that shares zero (scenario, size) points with what was actually measured, is a hard failure, not a silently-skipped comparison that reports a clean pass — this closes a real incident where the regression job's base-measurement step swallowed its own failure with || true, so a PR could merge having "compared" against a baseline that was never actually captured.

The regression workflow job automates this on PRs: it installs the base branch and the PR head into separate venvs on the same runner, runs both, and prints the regressions to the job summary. It gates (a regression past the threshold above fails the job) — loosen --regress-tolerance/--regress-min-delta-seconds rather than re-adding continue-on-error if runner variance proves noisy.

The header-graph attach-cost gate (scripts/check_header_graph_perf.py, see below) follows the identical median/warmup/combined-threshold design over its own --repeat/ --regress-tolerance/--regress-min-delta-ms flags — the two scripts share the underlying statistics via scripts/perf_measurement.py so their regression math can't independently drift.

Scan level cost model: one cliff at L4

A real scan-level sweep on two UXL libraries (oneTBB v2021.12→.13, C++; UMF v0.10→v0.11, C; raw data in skills-src/evaluation/validation/data/uxl_scan_results_2026-06.json) shows the cost has one cliff, at the L4 AST-replay boundary, and the cheap tier below it is dominated by the binary dump + always-on pattern scan, not by the source layer:

Level Reaches oneTBB (C++, 40 TUs) UMF (C, 50 TUs)
s0 diff classifier — (L0/L1 + pattern) ~29 s ~17 s
s1 compile-DB +L3 ~29 s ~17 s
s3 lexical (pattern only) ~29 s ~17 s
s4 symbol/graph index +L3 +L5 ~29 s ~17 s
s5 targeted AST +L4 (changed TUs) +L5 ~222 s ~22 s
s6 full AST +L4 (all TUs) ~215 s ~21 s

Rules of thumb:

  • The cliff height is a C++ phenomenon. L4 cost = clang per-TU AST replay; it scales with C++ template/STL instantiation depth, not .so or TU count. Heavy C++ (oneTBB) jumps ~7× (29→222 s); plain C (UMF) barely moves (~1.3×, 17→21 s). Budget L4 by how templated the source is.
  • The cheap tier (s0–s4) is one price. All four cost the same — the floor is the DWARF dump + lexical scan of the tree. Pick by coverage you need, not cost: s0≈s3 (L0/L1 + pattern only), s1 adds L3, s4 adds the L5 reachability graph without paying for L4. s4 is the structure sweet spot.
  • s5 is only cheaper than s6 with a diff seed. Without --since/ --changed-path the changed-TU set is empty and s5 replays every TU — same cost as s6. With a one-file seed, oneTBB s5 dropped from 222 s to 11.5 s (~19×) for the identical verdict. This scoping applies only to the source-changed collect mode — i.e. s5 and --mode pr. The other AST modes replay full scope regardless of any seed: --mode pr-deep resolves to graph-full, and --mode baseline/s6 to full (source_replay.CI_MODE_TO_SCOPE: source-changed→changed, graph-full→full), so pinning those in CI will not produce the scoped speedup.
  • audit costs the same as the baseline modes — the wall-clock is L4/L5 collection of the new side, not the baseline diff.

The verdict was identical across all levels on both libraries: the authoritative L0/L1 binary diff sets the gate; L3–L5 add coverage/localization, not a different pass/fail. For a CI gate, the cheap tier suffices; spend on L4 only when you want source-body semantics or PR localization for humans.

Scan-level scalability sweep

The UXL run above fixed the corpus (two real libs) and varied the level. The complementary question — how each level scales as a project's complexity grows — is swept by skills-src/evaluation/field/scan_level_scaling.py, a self-contained harness (no network/repo) that synthesises STL/template-heavy C++ trees of increasing TU count, builds them with the host compiler, and runs scan at each level against a slightly-changed baseline — recording wall time and peak child RSS (os.wait4) per (size, level).

Two results (measured back when the harness still had a separate full (s6) sweep entry alongside seedless source (s5) — see the note at the end of this section for why that entry was later removed, without invalidating either finding below):

  • The cheap tier is flat in TU count. binary/headers/build/graph (s4) cost the same at 4, 8, and 16 TUs (tail exponent ≈0) — they are priced on the binary dump + L2 header AST + L3 compile-DB parse, none of which grow with the number of .cpp files. Seedless (full-tree-replay) source-level scanning is linear in TU count (every TU is replayed). Both as expected.
  • Seedless --depth source (s5) used to hide a full-tree cost — now fixed. It cost ~2× the wall time and ~2.5× the RSS of the seeded run for the identical L4 coverage (both report L4=1/1), and the gap widened with TU count. The seed scopes both the L4 replay and the L5 clang call-graph pass to the changed TU; without a seed the L4 replay fell back to headers-only (one TU) but the call-graph pass ran over the whole compile DB — a second clang -ast-dump=json over every TU. The unseeded call-graph pass now scopes to the same compile units the L4 replay used (headers-only), so it is consistent with the L4 surface and no longer scales with the tree (~2.4× faster on a synthetic n=8 tree, identical verdict). Seeded runs are unchanged.

Why there is no separate full sweep entry anymore. ADR-043 D2 retired --depth full from the public CLI ladder entirely — it collapsed into --depth source, since replay scope (a change seed present vs. absent), not evidence depth, was the only thing distinguishing them, and scan itself now resolves that scope from whether --since/--changed-path was given (see abicheck/model/evidence_depth_levels.py's EvidenceDepth.FULL/SourceScope docstrings). The harness's own seedless "source" entry is the shape the old "full" entry measured — keeping both would just re-run the identical --depth source argv twice under two names. (Concretely, before this fix --depth full was a hard click.BadParameter — a harness bug in the same family as the skills-src/evaluation/field/runner.py one described in "Coverage gaps this workflow does not close" below, just in a manual-only harness rather than a scheduled CI lane, so it never produced a false-green.)

That whole-DB call-graph pass shells out to the same multi-GiB clang -ast-dump=json as the L4 replay, but its worker count (call_graph._call_graph_jobs) was CPU-bound only — it lacked the RAM-aware, cgroup-aware clamp the L4 replay grew (_l4_jobs → _l4_mem_cap) after the UXL oneTBB/oneDNN OOM. On a constrained host the L4 pass was protected but the unseeded call-graph pass was not. _call_graph_jobs now shares the L4 memory cap (_call_graph_mem_cap → _l4_mem_cap, same ABICHECK_L4_JOB_MEM_GIB budget); ABICHECK_CALL_GRAPH_JOBS still overrides the CPU count but memory wins over an over-eager override, exactly like _l4_jobs.

L2 header-scan deadline enforcement (pathological headers)

A real-world field report (Intel SVS) found the cheap tier's flatness above has an exception: a pathological header (deep #include/template complexity) can make the L2 clang/castxml AST dump itself run far longer than its on-disk size suggests — the report's own scan --dry-run estimate read 0.51 s for a header set whose actual parse ran over 15,000 s and 3+ GiB RSS before an external SIGKILL, because --budget was checked only once, after the whole scan had already finished, and the clang/castxml subprocess.run(timeout=120) call had no process-group isolation (a timeout only killed the direct child, orphaning any compiler-driver grandchild).

The fix (abicheck/deadline.py) threads a shrinking --budget deadline down to the L2 subprocess boundary (checked before each clang/castxml invocation, not only at the end) and runs that subprocess in its own process group so a timeout kills the whole tree. This is a bounding fix, not a speedup — a genuinely pathological header still costs whatever clang/castxml need, up to whatever --budget is given; it now fails cleanly at that boundary instead of running unbounded.

Regression/perf-tracking coverage, deliberately without needing the SVS corpus itself (see "Extract minimal synthetic fixtures" guidance):

  • tests/test_deadline.py — fast, synthetic (sh/sleep), proves the process-group kill and mid-stage budget check mechanisms directly.
  • tests/test_header_scan_deadline_integration.py — real clang, self-skips if absent. test_pathological_header_aborts_within_bounded_time_under_tiny_budget reproduces the SVS shape with a genuinely expensive (not simulated) 4-line header: a recursive template chain whose clang -ast-dump=json output grows steeply super-linearly with recursion depth (calibrated locally: depth 100 → ~40 MB/0.2 s, depth 200 → ~280 MB/0.6 s, depth 300 → ~900 MB/1.5 s — kept at depth 150 in the test to stay CI-safe), and asserts a tiny budget bounds it. The slow-marked companion test_pathological_header_natural_cost_is_tracked records that header's unbudgeted natural cost so a future regression (lost disk cache, a clang upgrade changing dump behaviour) shows up in the existing per-test duration trend (tests/conftest.py's ABICHECK_DURATIONS_JSON hook → scripts/summarize_test_durations.py → the CI run summary) — the same mechanism this page already relies on for the compare()-scaling story, rather than a new bespoke benchmark harness. scripts/benchmark_scaling.py is deliberately not the home for this: it is pure-Python by design ("no compiler/castxml" — see its module docstring) and this concern is inherently compiler-driven.

Fixed (follow-up): the L2 path (dumper._clang_header_dump, via the new dumper_clang_errors.run_clang_to_ast_file) now spills clang's AST-dump stdout straight to a temp file, mirroring the L4 per-TU replay (source_extractors/clang.py's _run_ast_to_file) instead of capturing it into a Python str. The calibration above showed a tiny header can legitimately produce hundreds of MB to multiple GB of AST-dump output, which capture_output=True would buffer on top of the parsed dict this code also builds; measured ~27% lower Python-heap peak (tracemalloc) on the depth-150 fixture (364.5 MB → 267.1 MB) with the fix. The same deadline.run_bounded treatment (shrinking --budget deadline, process-group kill on timeout) was also extended to preprocessor_facts.py's live extractor and both L4 source extractors (source_extractors/clang.py, source_extractors/castxml.py), which previously used the same fixed-timeout/no-process-group pattern the P0 fix closed for L2 — and a --budget deadline expiring during a PE/Mach-O header-scoped dump (service._try_header_scoped_dump) is no longer silently swallowed by the broad except Exception that falls back to export-table mode for a merely-unavailable header backend.

L2 acquisition: cache-key walk, AST cache write, include-map fan-out

Three costs around the L2 header parse that are not the parse itself. All three were measured locally (this host: 4 CPUs, clang 20, Python 3.13) against current-main function bodies and synthetic-but-realistically-shaped inputs -- not through a full oneDAL/SVS/PVXS dump/compare run, so read them as component figures, not an end-to-end speedup claim.

Request-local context sharing

Scalar comparison and release-member fan-out now own one request-local L2 acquisition table. A content-addressed frontend context has one in-flight clang/CastXML producer and one export-neutral SemanticIR normalization. Each binary still runs the concrete parser's legacy declaration construction against its own export set: those parsers own inclusion, fallback identity, and surface facts, so reconstructing their binding from a neutral parse was experimentally rejected after it lost real removal findings. Waiters have deadline-bounded queue time, do not cancel a producer another member still needs, and failed or input-mutated acquisitions are never retained as reusable results.

For Intel DPC++ the ordinary host context is also acquired directly with -fsycl -fsycl-host-only. This prevents the compiler from serializing a full unused device JSON document before the shared host evidence can be normalized. An explicit device request continues through the multi-context framing route: host-only and device ASTs are different parse contexts and must never share a cache entry merely because their entry headers match.

The compact result retained after each release-member comparison also carries that member's already-validated ElfMetadata and recorded native filename. Final bundle-graph assembly consumes those fields directly; it falls back to a stored-snapshot decode or live ELF parse only for a member that did not resolve earlier. Filesystem alias probing remains enabled only for genuinely live paths, so carrying metadata forward does not re-resolve a stored identity against the caller's current working directory.

A six-library local control used one real shared public header declaring six C functions and six separately compiled DSOs, each exporting a different one. Cold-cache clang acquisition changed from 6 compiler + 6 normalization calls, 4.253 s to 1 compiler + 1 neutral normalization, 3.375 s; CastXML 0.6.11 changed from 6 + 6, 1.900 s to 1 compiler + 1 neutral normalization, 1.492 s. Legacy declaration binding remains one pass per member and is not included in the eliminated-call claim. Every member retained all six header declarations and only its own export was marked binary-exported. These are small-fixture local measurements, not oneDAL results; the periodic real-profile lane remains the owner of oneDAL/SVS/PVXS claims.

  • Include-tree inventory for cache keys (extract/cache_header_scan.py). Every header-parse cache key walks each include root to fold in (path, mtime) for every header-like descendant -- dumper_ast_config. _cache_key for the AST cache, snapshot_cache for the whole-snapshot one. Path.rglob("*") materializes a Path per entry visited and then sorts those objects; carrying path strings through the traversal and building a Path only for the surviving, already-ordered entries measured 2.4x faster (/usr/include, 4188 matching entries: 47 ms -> 19 ms; a generated 15k-header tree: 151 ms -> 62 ms) with a byte-identical key. Two traversal behaviours are load-bearing and were matched deliberately rather than "cleaned up": the order is component-wise (PurePath compares normcased parts, so a/b/c.h precedes a/b.h) and symlinked directories are not descended into (which is also what makes a symlink loop terminate). A sorted(str(p) ...) "simplification" changes every cache key on disk while looking like a no-op -- tests/test_cache_header_walk.py pins both against a sorted(rglob("*")) oracle over generated trees.
  • The clang AST cache write (dumper_cache._atomic_write_json). json.dump never reaches the C encoder: JSONEncoder.iterencode selects c_make_encoder only under _one_shot, which is dumps' path. So the DPC++ host/device cache write was encoding a large AST a fragment at a time in pure Python. Descending the large containers in Python and handing each bounded subtree to the one-shot C encoder is byte-identical and measured 5x faster on a deep-template AST fixture (0.25 s -> 0.05 s for 2.1 MB of output; the ratio grows with document size). The obvious _atomic_write(path, json.dumps(obj).encode()) is not the fix -- it holds a second full-size copy of a tree that can be multiple GB, which is why the streaming write exists at all. The plain (non-DPC++) path still streams a raw file copy of clang's own stdout and pays none of this.
  • Per-header clang -M include-map fan-out (buildsource/include_graph_workers.py). The probes are independent subprocesses, so the pass was wall-clock-bound on nothing but serialization: 32 wrapper headers over a shared include took 1.71 s at one worker, 0.44 s at four (3.9x), with byte-identical depfiles; eight workers bought nothing measurable on 4 CPUs, which is why the auto default is CPU-derived rather than a fixed 8. This trades memory for wall time rather than removing compiler work -- child CPU time is unchanged -- so the pool is clamped by the same process_resources RAM probe the L4 pool uses (at a much smaller per-worker budget: clang -M is preprocess-only). One One sizing subtlety worth keeping: the shared gate is sized from the host budget alone and never replaced, while each pool's own size is that budget narrowed to min(..., unit_count). Deriving the gate from a pool's size instead is what a review round caught here -- two sides with differing header counts resolve differing sizes, and rebuilding the gate for the second left the first pool holding an orphaned semaphore, so both admitted their full quota at once (tests/test_include_graph_parallel.py:: test_sides_with_different_unit_counts_still_share_one_gate, and perf.shared_resource_gate_keyed_on_a_per_caller_value in tests/regressions/manifest.py). Two further review rounds on the same gate are worth reading before touching it: the serial path (one unit, or jobs=1) must take a slot too -- one child is not no children, and an ungated one alongside the other side's full quota is host_limit + 1 -- and the wait for a slot must be bounded by the run's aggregate and --budget deadlines, or a short-budget request blocks past its own budget behind an unrelated request's long probe and reports the overrun only afterwards. The third round is why the gate is a counter plus a condition variable rather than a BoundedSemaphore at all: a semaphore holds a size, and a size cached at construction cannot notice the host budget shrinking under a long-lived process -- every pool re-reads the budget and narrows itself while the stale gate keeps admitting the old number. Re-reading the limit on each admission removes the staleness instead of resizing around it, which would need to detect when every previous holder had drained. One thing that is not a valid shortcut here, and was ruled out with a real clang: replacing the per-header probes with a single umbrella TU. A header with an include guard that a previous header in the umbrella already pulled in disappears from the second header's dependency list entirely, so the umbrella silently loses edges the per-header probes see. The safe win is scheduling the same invocations better, never changing the compilation context.

L4 source-replay (dump-side) performance

The scaling harness above is pure-Python and times the compare pipeline. The dump-side L4 source ABI replay (clang per-TU AST extraction) is a separate cost, timed by skills-src/evaluation/field/scaling.py on real source trees (it needs clang + a built tree, so it is manual, not in CI).

Knobs and the reasoning behind them (abicheck/buildsource/source_replay.py):

  • ABICHECK_L4_JOBS — worker count for the per-TU extract pool. Auto = min(TUs, cpu_count, 8). An explicit override is clamped to max(8, 2×cpu_count) (logged when it fires) so a stray =64 can't oversubscribe a host into thrash (skills-src/evaluation/field/SCALING.md already saw jobs=8 on 4 CPUs regress). Set =1 to force serial (determinism).
  • Memory cap (auto + override). A single template-heavy C++ TU's clang -ast-dump=json output — and its in-Python parse — can reach several GiB, so the worker count is also capped by available RAM (min(…, available / ABICHECK_L4_JOB_MEM_GIB), default 3.0 GiB/worker, Linux only). "Available" is the smaller of host MemAvailable (/proc/meminfo) and the cgroup memory headroom (v2 memory.max − memory.current, or v1 memory.limit_in_bytes − memory.usage_in_bytes), so a container/pod confined to a small cgroup on a large host sizes its workers to what it is actually allowed to use rather than to host RAM. On a low-memory host this stops N concurrent giant ASTs from exhausting one process and getting the whole replay OOM-killed (the kernel SIGKILLs it → exit -9, all L4 work lost — observed on the UXL oneTBB/oneDNN s5/s6 full-target replays on a 15 GiB host). The clamp is logged. For a template-heavy tree on a constrained host, prefer a seeded/scoped scan (--since/--changed-path → a handful of TUs) over a full-target s5/s6; it sidesteps both the time and the memory cliff. ABICHECK_L4_JOB_MEM_GIB tunes the per-worker budget (lower = more workers).
  • ABICHECK_L4_EXECUTOR (thread default / process) — after clang returns, the extractor parses clang's large JSON AST dump and builds structural fingerprints: pure-Python, GIL-bound work. A thread pool parallelizes only the clang subprocess wait, so that post-processing serializes on the GIL — part of the ~60–83 % "serial fraction" in skills-src/evaluation/field/SCALING.md. process runs the extract phase in a ProcessPoolExecutor, parallelizing the AST work too (at the cost of pickling each SourceAbiTu and per-process spawn). It is opt-in pending a measured win — compare the curves with python skills-src/evaluation/field/scaling.py --jobs 1,2,4 --executor process vs thread. The driver falls back to serial if a process pool can't start (sandbox, spawn import error), so it never aborts L4.
  • Concurrent AST memory: clang's output is spilled to a temp file, not captured. A template-heavy TU's clang -ast-dump=json output can be multiple GiB. Capturing it (capture_output=True) holds the whole AST string in the heap from the moment clang finishes — and because the C json parse holds the GIL, the default thread pool serializes parsing, so all N workers sit holding their giant AST strings (≈ N × text) while queued behind the GIL. Spilling clang's stdout to a temp file keeps those payloads on disk until each worker's turn to parse, so the heap holds roughly one payload at a time instead of N. json.load still reads the file back to parse, so a single TU's parse peak is unchanged (≈ serialized text + tree) — this is a concurrency win, not a per-TU one — and it also drops the text=True decode copy (bytes parse) and frees the tree before the macro pass. The per-TU tree itself (~2–5× the AST text) is irreducible without a streaming JSON parser (a dependency the project avoids); for a template-heavy tree on a constrained host, a seeded/scoped scan is still the structural win.
  • ABICHECK_L4_CACHE_DIR — persists the per-TU cache (SourceAbiCache, content-addressed + per-included-file dependency-hash invalidation) across dump --sources runs. Previously the inline path passed no cache, so every dump re-extracted every TU; wiring this dir makes a cold run (eval E4: zstd 48.6 s) reuse the warm cache (3.4 s). Point it at a CI cache directory restored via actions/cache to start every CI run warm. The cache validation phase is serial, so the dependency digest is memoized per replay pass — a public header included by N TUs is hashed once, not N times.

S2 preprocessor pre-scan performance (scan --depth build)

scan's S2 preprocessor pre-scan (ADR-035 D2, buildsource/preprocessor_facts.py) is the conditional tier that runs once L3 build evidence is available: per-TU ABI-macro-value capture (clang -E -dM) and public-header-leak detection (clang -M). It is advisory-only (never a verdict on its own), but on a real-world build it dominated scan --depth build's wall time: a reported 4-minute-to-20-minute jump on a library with ~2000 translation units, ~920s spent in this tier alone — one serial clang -E -dM invocation per compile unit, with no cap, even though L3 ingestion itself (bazel aquery) cost only single-digit seconds.

Parallel probing, semantics-preserving for the ABI-macro-value/leak facts this tier reports: each probe is one I/O-bound clang -E/-M subprocess wait — unlike L4's AST parse, it is not GIL- or RAM-heavy — so probes run concurrently via a thread pool, one probe per compile unit / public header (unchanged from before this fix).

Deliberately NOT deduped by compile context. An earlier revision of this fix probed once per distinct (language, cwd, flags) signature and fanned the result out to every compile unit sharing it, on the premise that the curated ABI-macro list (_GLIBCXX_USE_CXX11_ABI, NDEBUG, _ITERATOR_DEBUG_LEVEL, …) is almost always driven by command-line/ predefined macros, not by a TU's own source text. Reverted (Codex review, fresh evidence): clang -E -dM reflects the TU's own source and #include chain too — two TUs sharing identical flags can legitimately resolve a curated macro differently (a debug-only TU #undef-ing NDEBUG, a conditionally-included config header) — and that is exactly the shape of bug find_macro_divergence exists to catch. Flag-based dedup would silently fan the first such TU's value out over the rest, hiding a real divergence rather than merely under-covering a rare case — a detection regression in the tier's own core purpose, worse than the perf win it bought. See this repo's "known gaps over risky reactive patches" convention (root AGENTS.md).

Knobs (abicheck/buildsource/preprocessor_facts.py), mirroring the L4 conventions above:

  • ABICHECK_PREPROCESSOR_SCAN_JOBS — worker count for the probe pool. Auto = one worker per compile unit / public header, capped by the same max(8, 2×cpu_count) ceiling L4 uses. Set =1 to force serial (determinism).
  • ABICHECK_PREPROCESSOR_SCAN_MAX_PROBES (default 512) — caps the number of compile units / public headers probed, bounding worst-case cost on a build with an unusually large number of either. Truncation is reported in the scan's diagnostics and folds the coverage row down to partial (never silent — this file's "no silent caps" convention).
  • ABICHECK_PREPROCESSOR_SCAN (default on) — set =0 to skip the S2 tier entirely while keeping L3 for the other tiers. Reported as an honest not_collected coverage row, the same as a missing compile DB or missing clang — never silently counted as clean.

The remaining, undeduped cost (~920s serial → roughly that divided by the worker count, parallel) is a real, unavoidable floor for this tier's design: every compile unit's macro-affecting #include chain can only be observed by actually running the preprocessor over it. A caller that wants the D2 coverage row without this cost has the disable knob above; there is no "fast and still divergence-complete" third option.

Why not precompiled headers (PCH) / modules?

A natural idea to cut the repeated per-TU header parse is a PCH over the public headers. It does not apply here: clang -Xclang -ast-dump=json does not re-emit declarations that came from a PCH, so loading one would silently drop the very header surface L4 exists to capture — a correctness bug, not a speedup. The right levers for repeated-parse cost are therefore the per-TU cache and the replay scope (changed/target), both already in place.

--budget mid-step preemption gap (found live, pvxs full-version-matrix scan)

Real-world evidence: scan --depth source on a real 62-TU library, --ast-frontend clang (no castxml on that host), did not complete within a 3+ minute --budget on a 4-core host — RSS climbed past 4.9 GiB before an external kill, with no indication --budget itself was actually bounding the work.

Investigation traced this to a gap in the loop driver, not the deadline mechanism itself: deadline.check() is already threaded per-TU inside each extractor (source_extractors/clang.py/castxml.py call it before/after the AST load, and bounded_timeout() shrinks the subprocess's own timeout to whatever's left of the scan-wide deadline before the clang/castxml process is even spawned) — so an already-dispatched TU always self-aborts quickly once the budget is gone. The gap was that _extract_cache_misses's serial fallback loop (used when jobs<=1, a single miss unit, or when the opt-in ProcessPoolExecutor fails to start) and _replay_cache_lookup's per-unit cache-key loop kept iterating to the next unit without checking the deadline first — each subsequent unit still self-aborted fast once dispatched, but only after paying the interpreter/extractor-startup overhead to get there. Fixed: both loops now call deadline.check() before each iteration (abicheck/buildsource/source_replay.py), so a hundred-unit miss list under an already-exhausted budget costs roughly one unit's overhead, not a hundred — regression-guarded by tests/test_source_replay.py::test_extract_cache_misses_serial_path_stops_dispatching_once_budget_is_gone (synthetic, fast — proven to fail without the fix by temporarily reverting it). The default parallel pool.map() path is unchanged: it still relies on each already-dispatched unit's own fast self-abort rather than a stop-enqueuing check, since pool.map submits its whole batch up front — a "don't submit further work once budget is gone" gate there would need a bespoke, non-pool.map dispatch loop, a larger change than this fix.

Not resolved by this fix, still an open question: whether the underlying per-TU cost on this real 62-TU pvxs binary was itself pathological (something superlinear in this specific run) or is simply what clang-frontend L4 replay genuinely costs per TU on a template-heavy real C++ codebase without castxml — the first report's equivalent pass completed in 129s, but that ran with castxml, not available on the pvxs scan's host. Distinguishing those two needs a dedicated profiling pass on a real or skills-src/evaluation/field/scan_level_scaling.py- synthesized multi-TU tree with a --budget sweep added to that harness (mirroring how the L2 pathological-header investigation above was profiled), not assumed from a single real-world data point. Until profiled, the safe recommendation for a clang-only (no castxml) CI runner on a library this size remains: skip scan --depth source in favor of compare for the L1/L2 release gate, or scope it with --since/--changed-path to just the changed files rather than the whole library.

Weekly optimization report

python scripts/perf_report.py [--corpus N] [-o FILE] ranks hot functions, costed same-argument repeats (audit_repeated_calls.py --by-cost), calls per declaration (per_decl_x10:* budgets) and anti-pattern sites. performance.yml runs it weekly into the step summary. Record what you decide about a candidate in the findings registry.