Comparison Performance¶
This page documents the runtime time cost of comparing real, large shared libraries. Its companion, Comparison Memory, covers what a multi-library comparison keeps resident, how to measure that in a way that attributes a peak to an owner, and the current owners.
This page documents the runtime cost of comparing real, large shared libraries, the bottlenecks that were found and fixed, and the tooling that guards against regressions in CI.
TL;DR¶
- Dump scales fine. Snapshotting
libonedal_core.so(~10,550 exported functions) takes ~5 s. - Compare used to blow up. On the same library,
comparedid not finish within 60 s. The cost was entirely in the post-processing detectors, not the core symbol diff. A profiling sweep found six super-linear paths — several quadratic, one effectively cubic — all now fixed (see What was fixed). - A synthetic scaling harness (
scripts/benchmark_scaling.py) reproduces each path without a real binary, compiler, or castxml, and aslowregression test guards the realistic hot path.
What was fixed¶
Every fix preserves detector behaviour (the full unit suite, the FP-rate gate, and the metamorphic/oracle detector tests all stay green); they only change how the work is organised.
| # | Path | Was | Fix | Result |
|---|---|---|---|---|
| 1 | Public-surface scoping (surface.classify_change_surface) |
Recomputed four old∪new set unions per finding → O(findings × surface). Made every comparison quadratic. | Compute the unions once per pass (surface_unions) and reuse. |
add_remove 4000: 9.1 s → 0.32 s (linear) |
| 2 | Namespace detection (diff_namespaces) |
demangle_batch called one symbol at a time → one c++filt subprocess per symbol. |
Batch-demangle each snapshot once (_batch_demangle_public) and thread the map through. |
elf_namespace 4000: 5.2 s → 0.33 s (linear) |
| 3 | Variable / symbol diffing | Quadratic via the same per-finding surface unions (#1). | Fixed by #1. | var_churn 4000: 2.1 s → 0.06 s (linear) |
| 4 | Batch-rename heuristic (diff_symbols._find_rename_pairs) |
O(removed × added) suffix scan. | Reversed-name index + binary search (endswith → reversed prefix lookup). |
folded into add_remove win |
| 5 | Type-spelling fallback (diff_type_spellings) |
Rebuilt a set(...) inside a comprehension → O(n²). |
Hoist the set once. | folded into add_remove win |
| 6 | Affected-symbol enrichment / ancestor closure (diff_filtering) |
Transitive ancestor function lists accumulated duplicates, then re-sorted per change → effectively cubic on nested type graphs. | Use sets (dedup on union); sort once. | nested_types n=200: >60 s → 0.16 s |
| 7 | ELF-only rename matching (binary_fingerprint, diff_symbols._plausible_rename) |
O(removed × added) name-similarity scan; the name predicate re-demangled both names per pair. | Scan only the size-tolerance window via the existing size index; cache the per-name parse; cap the heuristic pass for mass-rename inputs. | rename_churn n=1000: 13.2 s → 2.1 s, larger inputs bounded |
| 8 | Affected-symbol enrichment type↔function/field mapping (diff_filtering._build_type_to_funcs, _build_type_embed_index) |
any(tname in ft ...) nested inside the type loop and the function/field loop → O(types × functions × refs); quadratic when many distinct types churn (a header refactor or versioned upgrade). The original perf sweep only sampled these scenarios at n=500, so the exponent was never computed and the table mislabelled them "linear". |
One Aho-Corasick _SubstringMatcher over the affected type names, built once and shared; each ref/field is matched in O(len) with identical substring semantics. |
typedef_churn n=4000: 6.3 s → 0.73 s; union_churn 9.1 s → 1.10 s; vtable_churn 7.8 s → 1.03 s; enum_churn ≈1.8 → 1.0; opaque_filter ≈1.7 → 1.2 (all now linear) |
| 9 | Opaque-handle pointer-only / factory check (diff_filtering._is_pointer_only_type, _has_public_pointer_factory via _filter_opaque_size_changes) |
Each opaque candidate rescanned every public function/variable with a word-boundary regex → O(candidates × functions), a regex per pair (type_churn n=4000: ~3.2 M searches). |
One indexed pass per snapshot (_opaque_usage_index): an Aho-Corasick prefilter narrows each type string to the candidates present, then the same regex oracle decides — so the verdict is unchanged (verified by a fuzz test vs the per-candidate functions). |
type_churn n=4000: 1.13 s → 0.38 s (≈1.6 → linear) |
With fixes #8 and #9 the compare pipeline has no remaining quadratic path at
the tracked sizes — every scenario is linear (tail exponent ≈1.0–1.3) except the
inherently deep nested_types chain. The opaque-handle pointer-only check
(_is_pointer_only_type) used to be O(candidates × functions) with a
word-boundary regex per pair (type_churn n=4000: ~3.2 M regex searches, ≈1.6);
fix #9 replaced the per-candidate rescan with one indexed pass.
How to reproduce¶
No real binary, compiler, or castxml required — the harness synthesises
AbiSnapshot pairs that exercise each path:
# Sweep all scenarios and print a table with a scaling exponent per scenario.
python scripts/benchmark_scaling.py
# Focus one path and emit machine-readable JSON.
python scripts/benchmark_scaling.py --scenario type_churn \
--sizes 1000 2000 4000 --json-out reports/perf/scaling.json
Scenarios (add_remove is the linear control; the rest target a specific
path). The first group exercises compare() (the original focus); the second
group, added later, extends coverage beyond compare() to the suppression and
reporting stages — see Coverage beyond compare():
| Scenario | Measures | Stresses |
|---|---|---|
add_remove |
compare() |
Core symbol diff + surface scoping (control) |
type_churn |
compare() |
Affected-symbol enrichment, opaque filtering (structs) |
enum_churn |
compare() |
Enum diffing (diff_types._diff_enums) |
typedef_churn |
compare() |
Typedef base-change diffing (_diff_typedefs) |
union_churn |
compare() |
Union member diffing |
wide_struct |
compare() |
Per-field diffing within large records |
vtable_churn |
compare() |
Vtable / virtual-layout diffing |
elf_namespace |
compare() |
Namespace detection + demangling (stripped lib) |
pe_churn |
compare() |
PE/COFF export diffing (diff_platform PE arm) |
macho_churn |
compare() |
Mach-O export diffing (diff_platform Mach-O arm) |
var_churn |
compare() |
Public-surface classification |
rename_churn |
compare() |
ELF-only fingerprint rename matching — the reject path (disjoint names, no match emitted) |
fuzzy_rename_churn |
compare() |
ELF-only fingerprint rename matching — the accept path (every symbol genuinely renamed → one func_likely_renamed per pair). The ICU/LLVM cost driver (P11: rename detection, not symbol count) |
version_node_churn |
compare() |
Version-node migration fan-out — every export moves LIB_1.0 → LIB_2.0 → n symbol_moved_version_node findings (the LLVM 17→18 36,991-finding shape) |
versioned_rename_churn |
compare() (collapse on) |
Versioned-symbol-scheme detection and collapse over 2×n churn (ICU/OpenSSL u_*_NN) |
nested_types |
compare() |
Transitive type-ancestor closure |
opaque_filter |
compare() |
Opaque-handle size filter (the known O(candidates × functions) residual) |
suppression_audit |
SuppressionList.audit() |
Rule-vs-finding matching (O(rules × findings)) |
severity |
categorize_changes() |
Severity categorization of findings |
serialize |
snapshot_to_json → from_dict |
Snapshot serialize/load round-trip (dump-pipeline proxy) |
report_html |
generate_html_report() |
HTML document assembly |
report_sarif |
to_sarif_str() |
SARIF JSON assembly |
report_junit |
to_junit_xml() |
JUnit XML assembly |
Peak memory¶
Every measurement also records the peak tracked heap (peak_mb, via
tracemalloc) of the timed call. The inputs are built outside the traced
window, so the figure attributes only the call's own allocations. The memory
pass also runs cold: process-wide caches warmed by the timing loop (e.g. the
functools.lru_cache demanglers) are cleared first, so input-scaled cache
growth is counted rather than hidden behind a warm cache. A flat per-item time
alongside a rising peak_mb flags an intermediate O(n²) space blow-up that a
wall-clock-only gate would miss. Disable with --no-memory (timing only); gate
with --max-memory-mb <budget>.
Alongside it, each point records the process peak RSS (rss_mb, via
resource.getrusage). Unlike peak_mb, which only sees Python-heap
allocations, RSS also counts native memory — pyelftools parse buffers and
c++filt subprocess pages — which is what dominates real libraries (the field
eval observed ~330 MiB RSS at LLVM scale, invisible to tracemalloc).
ru_maxrss is a process high-water mark, so it is monotonic across sizes and
the largest/last value is the true peak (it overstates a single call's own
footprint, since inputs are built in-process); gate the peak with
--max-rss-mb <budget>. RSS is unavailable on Windows (resource is
Unix-only), where the column is simply absent.
Coverage beyond compare()¶
The original sweep (PR #331) only covered compare() post-processing. A
follow-up gap analysis extended it to the two other stages that build the
largest data structures from the finding set:
- Suppression audit (
suppression.py,SuppressionList.audit) tests every rule against every change — O(rules × findings). Thesuppression_auditscenario holds the rule count fixed (a project's ruleset is roughly fixed while its library grows) and scales findings, so it stays linear in findings; a regression that makes per-finding matching itself super-linear (e.g. recompiling a pattern per change) shows up as a rising exponent. - Reporting —
to_markdown/to_jsonwere already guarded byslowtests;report_htmlandreport_sarifextend that to the HTML and SARIF renderers, which assemble the largest output documents. Both are linear.
Measured scaling (after fixes)¶
Most scenarios are linear at the sizes a real library reaches (per-change cost
roughly flat); type_churn and enum_churn are mildly super-linear (~1.7) but
bounded and tracked:
Figures are indicative local timings (absolute seconds vary with runner speed —
the tail exponent is the portable signal). The first group times
compare(); the second group, added in PR #336, times the suppression and
reporting stages (see Coverage beyond compare()).
| Scenario | time @ size | tail exponent |
|---|---|---|
add_remove |
0.32 s @ n=4000 | ~0.9 (linear) |
var_churn |
0.06 s @ n=4000 | ~1.0 (linear) |
elf_namespace |
0.33 s @ n=4000 | ~1.1 (linear) |
pe_churn / macho_churn |
<0.05 s @ n=500 | ~1.0 (linear) |
wide_struct |
0.1–0.2 s @ n=500 | ~1.0 (linear) |
typedef_churn / union_churn / vtable_churn |
0.7–1.1 s @ n=4000 | ~1.0 (linear, after fix #8) |
enum_churn |
1.0 s @ n=4000 | ~1.0 (linear, after fix #8 — was ≈1.8) |
type_churn |
0.38 s @ n=4000 | ~1.0 (linear, after fix #9 — was ≈1.6) |
opaque_filter |
0.45 s @ n=1000 | ~1.2 (linear at tracked sizes after fix #8) |
rename_churn |
2.1 s @ n=1000, capped above | bounded |
fuzzy_rename_churn |
0.39 s @ n=4000 | ~1.0 (linear) |
version_node_churn |
0.86 s @ n=10000 (10 k moves) | ~1.0 (linear) |
versioned_rename_churn |
0.87 s @ n=8000 (16 k changes) | ~1.1–1.2 (mild) |
nested_types |
0.70 s @ n=400 | inherent for deep chains |
suppression_audit |
0.09 s @ n=2000 (fixed 40-rule set) | ~1.0 (linear in findings) |
severity |
<0.01 s @ n=1000 | ~1.0 (linear) |
serialize |
0.12 s @ n=1000 | ~1.0 (linear) |
report_html / report_sarif / report_junit |
≤0.04 s @ n≤2000 | ~1.0 (linear) |
CI integration¶
.github/workflows/performance.yml
runs the scaling benchmark and the slow performance tests. Now that every
compare() scenario is linear, the lane is gating:
- Triggers: weekly schedule, manual
workflow_dispatch(with size / budget inputs), and every PR (opened/reopened/synchronize/labeled) — but the expensive jobs (scaling,regression,header-graph-perf— schedule/dispatch only, see below — andl2-cli-perf, which also runs the header-graph PR-vs-base gate) only actually run when aclassifyjob decides the PR touches performance-sensitive code, the classifier-job pattern (scripts/classify_perf_paths.py,tests/test_classify_perf_paths.py). This replaced an earlierpull_request.paths:trigger-level filter, for a real, not theoretical, reason: a trigger-level filter means the workflow's check run never registers at all on a non-matching PR, and the pattern list lived only in unversioned, untestable YAML glob strings — a real, perf-relevant change (the P0.2 Bazelaquery/cqueryroot-target scoping, the P0.3 auto-applied-L3-context-to-L2-headers change) merged with no Performance check run at all because the list hadn't been updated to cover it (see "Coverage gaps this workflow does not close" below). Theclassifyjob always runs, diffs the PR's changed files againstPERF_SENSITIVE_PATTERNS— the detector core (abicheck/diff_*.py,checker.py,post_processing.py,demangle.py,binary_fingerprint.py,surface.py, ...), all ofabicheck/buildsource/**, compare/dump orchestration (dry_run_estimate.py/service_input_resolution.pyand the otherservice_*/cache modules —service_scan.pywas renamed todry_run_estimate.pyandscan_engine.pywas deleted outright, along with the rest ofscan, in ADR-068 Phase 6), the benchmark scripts, and the perf tests — and reports arunoutput the four downstream jobs each gate on (if: needs.classify.outputs.run == 'true'). Adding theperformancelabel force-runs the lane regardless of changed paths (fixed alongside the classify job — the label previously had no effect unless the changed paths also happened to match, despite an existing comment claiming otherwise); for a PR that touches neither, run it on demand withworkflow_dispatch. - Armed budgets: the scaling step runs with
--max-exponent 1.4(the tail, largest-two-size slope) and--max-rss-mb 1024; theregressionjob blocks on a PR-vs-base slowdown exceedingmax(15%, 100ms)per (scenario, size) point (--regress-tolerance 0.15 --regress-min-delta-seconds 0.1— the "stable synthetic PR scenario" tier; a flat 50 % tolerance is an emergency stop, not real regression protection — several consecutive +15 % merges would otherwise nearly double the runtime before anything caught it).continue-on-erroris dropped on both, so a catastrophic regression fails the lane;scripts/benchmark_scaling.py's ownmain()additionally fails closed if--baselineloads to zero points, or if zero (scenario, size) points end up shared between base and head — so a scenario-set/--sizesmismatch can't silently turn the comparison into a no-op that reports a clean pass. The thresholds are CLI flags so the budget lives in the workflow, not the script — loosen a threshold rather than re-addingcontinue-on-errorif normal drift ever flakes a lane. - Per-scenario tuned sweeps, not one global
--sizes. The PR/schedule triggers no longer pass--sizesat all, so each scenario's own tuned default sweep (SCENARIOSinbenchmark_scaling.py) applies —onedal_large_surfacereaches its tuned 5k-20k range,nested_typesgets its full multi-point sweep (so a tail exponent is actually computable), andrename_churn/opaque_filterkeep all three of their tuned points. A previous version of this workflow always passed a single500 1000 2000 4000ladder to every scenario regardless, silently overriding all of the above.--sizes(and theworkflow_dispatchinput of the same name) still works for a manual, deliberately-scoped run. - Median, not fastest-of-N. Each point runs one untimed warmup, then 5
timed repeats on every trigger (7 on the weekly schedule, since it isn't
blocking a merge) — the reported/gated figure is the median of those
repeats, with min/max/p95/coefficient-of-variation recorded alongside it
(
scripts/perf_measurement.py). Keeping the fastest of a couple of runs (the previous behaviour,--repeat 2+min()) systematically favours the run least likely to have hit GC/scheduler noise — i.e. the one least representative of what a regression would actually look like — and gives no noise signal at all. - Base/head identity in the JSON report.
--meta git_sha=... --meta side=base|head --meta os_image=...stamps the report each side was actually measured on, so a report downloaded later can be matched back to the exact commit/runner it came from. - The
--max-exponentgate is per-scenario opt-out:nested_typesis an inherently super-linear embedding chain, so it carriesgate_exponent=Falseand is exempted (its tail slope is still printed for visibility, just not gated). Every other scenario is gated. - The exponent gate has a noise floor, checked on the slope's lower
endpoint. A tail slope is only as precise as the smaller of its two
points, so the gate applies only when both tail points take at least
EXPONENT_FLOOR_SECONDS(0.2 s,scripts/perf_measurement.py). Until 2026-10 the floor was checked against the scenario's peak. A cheap scenario whose 4000 point had just crossed 0.2 s was then gated on a slope whose 2000 point sat at ~0.1 s in runner jitter, andreport_sarifread 1.48 against the 1.4 budget on unchanged code (locally the same code read 1.18–1.35). The cheap gated scenarios (pe_churn,macho_churn,var_churn, the threereport_*,fuzzy_rename_churn,onedal_mass_removal) now sweep one size step higher (TAIL_ABOVE_FLOOR_SIZES), so both tail points clear the floor and they stay gated. A scenario whose lower tail point later drifts under the floor printsexponent gate inactivewith its tail value, never silently; raise that scenario's sizes to re-arm it. - Publishes the scaling table to the job summary and uploads the JSON.
slow regression guards also live in
tests/test_performance.py
— TestTypeChurnScaling (compare back to genuine O(n²)),
TestSuppressionAuditScaling (audit stays linear in findings), and the
HTML/SARIF cases in TestReporterScaling. They run in the existing slow lane
with generous thresholds, so a catastrophic regression fails fast without
flaking on normal drift.
The same workflow also carries a second, independent pair of jobs for the
G31 Phase D header-graph attach-cost gate
(scripts/check_header_graph_perf.py,
see the G31 Phase D follow-up plan):
header-graph-perf is report-only trend data on schedule/dispatch (on a
pull request the trend point is the header-graph gate's own head
measurement, uploaded under the same performance-header-graph artifact
name, so a PR does not spend a second runner re-measuring the same head;
no stable committed baseline
number would survive a runner/toolchain change, the same reasoning
check_mutation_score.py's SURVIVOR_BASELINE bootstrap avoids);
The header-graph PR-vs-base gate (steps of the l2-cli-perf job since 2026-10, formerly its own header-graph-regression job on a separate runner) follows this page's own --baseline/--regress-tolerance
same-runner base-vs-head pattern (see Baseline regression
below) and gates from day one, since that pattern never needs a stale
committed number to begin with.
Three measurement levels — which number means what¶
The single most common way to misread this page is to quote one level's number as another's. There are three harnesses, they measure genuinely different things, and none of them is a substitute for the others:
| Level | Harness | What it measures | Compiler? | Interpreter startup? |
|---|---|---|---|---|
| 1. Synthetic, in-process | scripts/benchmark_scaling.py |
compare(), suppression audit, severity, serialization and the HTML/SARIF/JUnit renderers over hand-built snapshots. Scaling exponents, peak heap, process RSS. |
no | no |
| 2. Real L2, in-process | scripts/check_header_graph_perf.py |
A real dumper.dump() of a real .so + synthetic header sweep, plus service._attach_header_graph's marginal cost. Gated as three separate metrics: dump_ms, attach_ms, total_ms. |
yes | no |
| 3. Full CLI | scripts/check_l2_cli_perf.py |
The whole abicheck CLI as a subprocess against a real compiled C++ fixture, across the six supported L2 forms. |
yes (except the stored-operand forms, which must run none) | yes |
Boundaries worth stating explicitly, because each has been a real source of confusion:
- Level 2's
total_msis not a CLI wall time. It isdump+attachin-process, and it excludes interpreter startup, config resolution, input resolution, serialization, comparison and rendering. It was previously namedbaseline_ms, which invited reading it as "the baseline cost of a run"; it is the cost of one phase. - Level 3's measured window is the subprocess's whole lifetime. Fixture compilation, snapshot pre-dumping for a stored-operand scenario, artifact download and dependency installation are all setup, measured separately and excluded. Every correctness check runs after the timed window closes.
- Old binary-only numbers are not L2 numbers. The ~5 s
libonedal_core.sodump quoted in this page's TL;DR is a historical binary-plus-DWARF figure. An L2 run of the same library additionally parses its public headers, and mixing the two is how a "dump got 10x slower" conclusion gets manufactured.
What the full-CLI harness asserts besides duration¶
A timing harness that checks only duration rewards the worst regression available to it: getting faster by doing less. Level 3 therefore gates correctness alongside time, and a scenario whose validation fails is a failure, never a fast data point. Per scenario, outside the timed window:
- the resolved evidence depth really reached
headerson every side (a binary-only fallback still accepts--depth headers, and it is faster); - public scoping resolved and did not fall back;
- the expected declarations survived into the snapshot (names, not counts — a count can be matched by a degraded parse that found a different set);
- the header call, include and type graph extractor passes all ran and none is recorded as degraded, and the include graph collected at least one edge;
- the fixture's deliberate break produced both a removal-family and a layout-family finding — a much stronger claim than "the verdict was BREAKING", which a single unrelated finding can satisfy;
- the unchanged control produced neither, i.e. no false positive;
--no-baselineread as an audit, with no manufactured compatibility verdict.
Native invocations are observed, not assumed¶
perf_receipt.NativeInvocationSpy prepends a directory of exec-ing shims to
PATH, one per spied tool, each logging its own argv before handing off to the
real binary. classify_invocation then bins each observed invocation as header
extraction, include pass, or version probe — a distinction that is load-bearing
rather than cosmetic: a cold L2 dump makes three castxml calls and only
one of them is a parse (the others are --version and -dumpmachine), so a
flat per-tool count cannot tell a cached parse from a skipped probe.
That turns three otherwise-unfalsifiable claims into measurements:
- a stored-snapshot/stored-snapshot comparison performs zero header extractions and zero include passes — measured, not assumed;
- a stored/live comparison performs exactly one side's worth, so the operand handed over as a snapshot is demonstrably not re-extracted;
- a run labelled "warm cache" really was served by a cache, proven by its extraction count dropping rather than by the fact that it ran second.
The last one matters because on a small fixture a second run served by nothing is indistinguishable from a warm one by wall time alone.
Three rules about which runs those measurements cover, each of which started as a defect where one favourable observation certified a batch:
- a
forbiddencontract is a zero over every invocation kind, not only the extraction bucket. "No compiler ran" is the claim, so an include pass or a version probe falsifies it exactly as a parse does. - a live contract additionally requires the include-graph pass to have run, at
least once per live side. Header-AST extraction is not the whole of the
measured L2 work, and a run that stops doing the
clang -Mpass is faster while still resolving depthheadersand still finding the deliberate break. - every cold/warm repetition is checked individually, paired by index, and the
reported cache service is the worst repetition. Reducing each batch with
min()let one warm repetition certify a scenario whose others re-extracted in full, so the gated median could describe an uncached run under a receipt claiming a served cache.
The same "every repetition, not the lucky one" rule governs correctness: the scenario's semantic validation runs at the end of each repetition, against the outputs that repetition just wrote, and every file an invocation is declared to produce is deleted beforehand and required afterwards. Validating once at the end inspected only the final report, so an earlier repetition that emitted degraded evidence while still writing a file and exiting with an allowed code kept its faster timing in the median whenever the last repetition happened to be correct.
Cache states are three things, not two¶
The harness separates, and never conflates:
- cold application cache — a fresh process with a fresh
XDG_CACHE_HOME. This is explicitly not a cold disk: the OS page cache still holds the fixture and the interpreter, and no attempt is made to drop it. Dropping a CI runner's page cache needs privilege and would measure the host. - warm AST cache — a fresh process against the same cache root and byte-identical inputs.
- invalidated — the same again after a transitive dependency header changes. The cache must not serve; if it does, the product is reusing stale evidence, which the harness reports as a failure rather than a speedup.
Which one actually served is read off the counters
(observed_cache_service: none / partial / full), never inferred from run
order.
Memory is reported as two differently-derived numbers¶
sampled_peak_tree_bytes is the largest simultaneous sum of RSS across the
measured process and every descendant alive at one sampling instant, taken from
a parent-side sampler at a recorded interval. It is deliberately not called
a peak:
- a spike shorter than the sampling interval is missed entirely;
- pages shared between processes are counted once per process, so it can also overstate real physical usage;
- a descendant that exits between two samples is never seen.
ru_maxrss_bytes is carried separately because it is a different thing: the
kernel's own high-water mark, which never misses a spike but is a maximum over
individual processes and so cannot see two live children's combined footprint.
Both are reported, labelled, with the interval and the observation limits —
picking one would hide the other's failure mode.
ru_maxrss_bytes comes from RUSAGE_CHILDREN, which is cumulative over every
child the harness has reaped, so it is reported only for the step that raised
it. A step that did not raise it gets null plus a ru_maxrss_scope naming the
earlier, heavier child that holds the mark. Without that rule, one 2 GB step
makes every following step report 2 GB as its own RSS — which is exactly what the
first published PVXS receipt did, showing 1.9 GB against a run whose sampled tree
peak was 438 MB. Read sampled_peak_tree_bytes for such a step.
Measured cost of the full-CLI lanes¶
All figures local (gcc 13.3.0 / castxml 0.7.0 / clang 18.1.3, 4 CPUs, Linux,
--repeat 3 unless stated). Reproduce with the commands in the harness's own
module docstring. These are lane costs, not per-scenario costs; per-scenario
numbers are in the receipt.
| Lane | Wall | User CPU | Native invocations | Fixture build (setup, excluded) |
|---|---|---|---|---|
--suite pr |
44–48 s | 40–43 s | 28 header extractions, 14 include passes, 100 probes | ~0.9 s |
--suite extended (--repeat 1) |
~65 s | — | — | ~4.7 s |
--suite extended (--repeat 3) |
~180 s | — | — | ~5.1 s |
The PR lane's cost is dominated by interpreter startup, not by analysis: the lane makes roughly 30 CLI invocations (8 gated steps plus setup and resolution steps, times 3 repeats) at ~0.6 s of startup each, so well over a third of the lane is spent before any evidence work happens.
Per-repetition validation (each repetition's own outputs checked, rather than only the final report) triples the validation work, all of it outside every timed window. It does not materially change the lane: re-measured at 43 s wall, inside the range above. That measurement was taken while the extended suite ran concurrently on the same 4-CPU host, so read it as an upper bound — which is what makes it usable here, since an upper bound inside the existing range is enough to say the range still holds. The range is deliberately left as it was rather than narrowed to a contended number.
Per-scenario gated full_cli medians on that fixture, with the coefficient of
variation that sets the gate's noise floor (--repeat 3, except the two-format
row, re-measured at --repeat 1 after the export grammar landed):
| Scenario | Step | Median | cv |
|---|---|---|---|
dump_l2 |
dump | 1.074 s | 6.4% |
compare_live_live |
compare | 1.219 s | 3.9% |
compare_stored_live |
compare | 1.389 s | 5.3% |
compare_stored_stored |
compare | 1.272 s | 15.9% |
compare_no_baseline |
audit | 1.125 s | 11.6% |
compare_two_formats |
compare_exporting_two_formats | 1.140 s | — |
compare_live_live (unchanged) |
compare | 1.327 s | 6.1% |
Those cv figures (up to ~16%) are why the lane's absolute floor is 0.5 s rather than 0: a purely relative 30% tolerance on a ~1.2 s measurement would be only ~0.36 s, inside what this fixture's own run-to-run variance already covers.
On the extended axes (--repeat 1): templates ~1.36 s, 8 headers ~2.48 s,
32 headers ~10.4 s, and five libraries ~1.15–1.50 s each — the same per-library
cost whether they share one dependency header or each have their own, because
the header-frontend invocation count follows the top-level headers rather than
their dependencies. The shared arm resolves through one physical file at a common
include root, and header_contexts is counted from the resolved paths actually
built rather than from the flag that asked for them; an earlier version gave each
library its own byte-identical copy, which (the AST cache keying on resolved
path) made the "shared" arm a second distinct-path workload.
The multi-library set is measured as one cache lifecycle — reset once before
the set, not before each member — since otherwise each library's comparison
starts from an empty cache and cross-library reuse is unobservable by
construction, which is the only thing separating the shared arm from the distinct
one. With that in place the measurement says something it previously could not:
at --repeat 2 every member of both arms performs 4 header extractions and 2
include passes, identical in the shared and distinct arms, so no cross-library
reuse happens today even when five libraries resolve one physical dependency
header through a common include root. That is an observation about the product,
not a harness gap, and it is recorded here rather than acted on: this work
deliberately changes no caching strategy (see "Scope" above). It is the
cross-library half of the ⚠️ row for bundle/multi-library orchestration in the
coverage table below.
One comparison, two artifacts. -o FORMAT=DESTINATION is repeatable and
every export renders the one completed analysis (ADR-068 slices 7m/7n), so a
JSON report and a human report come from a single compare. The harness
verifies that rather than assuming it: over a stored-old/live-new pair the
two-export invocation performs exactly one side's worth of header extraction, so
the second renderer provably does not re-run the analysis.
Instrumentation overhead, and a worked example of why ordering matters.
Measured by running the identical lane with and without --no-spy.
A single unordered pair (spy, then no-spy) gave +2.1 s wall (+4.6%) and +2.3 s CPU — a plausible-looking result, and one it would have been easy to publish. Repeating it in ABBA order (spy, no-spy, no-spy, spy), so drift across the sweep cannot be read as a configuration difference, gave medians of 45.2 s with the spy against 46.0 s without it: −0.7 s (−1.6%), i.e. the opposite sign, against a largest within-configuration spread of 2.2 s.
So the honest statement is that the spy's overhead is not resolvable above
run-to-run noise on this host, bounded by roughly ±5% of the lane, and the
first measurement's +4.6% was noise wearing a plausible number. Mechanically
that is what one expects: ~142 extra shim invocations per lane, each a /bin/sh
startup plus a printf plus an exec, against castxml parses that each cost
hundreds of milliseconds.
Note also that --no-spy disables every extraction-count assertion, so it is a
measurement aid, never a cheaper way to run the lane.
L2 scaling gate (headers x libraries)¶
scripts/check_l2_scaling_perf.py runs in the l2-cli-perf PR job after the
PR-vs-base comparison. That comparison measures one small fixture, so it
catches a constant-factor regression but not a change in shape. This gate
sweeps two axes through the real CLI and gates how cost grows:
| Axis | Operation | Sizes | Budget (marginal exponent) | Measured (2026-09, 4 CPUs) |
|---|---|---|---|---|
headers |
compare --depth headers, one library |
1 / 4 / 12 / 30 headers | 1.4 | 0.95 (1.3 s → 2.6 s) |
libraries |
directory compare (release fan-out), 2 headers each |
1 / 3 / 6 / 10 libraries | 1.7 | 1.39 (1.2 s → 8.6 s); 1.64 (→ 11.5 s) before the fixes below |
Peak process-tree RSS is gated at 1024 MB per point (observed ≤ 365 MB).
The whole gate takes about 75 s at --repeat 3.
The gated number is the marginal exponent: the slope of
log(wall(n) - wall(1)) against log(n - 1). Every CLI run pays a ~1 s fixed
floor (startup, imports, config, report writing), so a raw log-log slope at
these sizes reads close to 0 whatever the product does. Subtracting the n = 1
floor measures the work the axis adds: a linear step reads ~1.0, and a step
that redoes all previous units' work reads ~2.0. If the largest point is less
than 0.25 s above the floor, the sweep fails as unfittable rather than passing.
The library axis is still super-linear, and the budget does not bless that. Every member of a directory compare is dumped against the release's union header and include set, so any per-member step that walks that set grows with the member count. The total then grows faster than linearly. The two largest such steps are now memoized:
- contract fingerprinting's path resolution (
comparability_fields), about 35% of wall time at 16 libraries; - the C++20 dialect scan (
extract/header_scan_memo.py), which ran several times per dump over the identical set.
The first measurement blamed bundle_symbol_status → qualified_name_segments_walk.
That was a profiling artifact: cProfile attributed worker threads' time
wrongly, and py-spy corrected it. Before → after, 4 CPUs, cold cache:
6 libraries 5.0 s → 4.2 s, 10 libraries 11.5 s → 8.6 s,
and 16 libraries 32.2 s → 19.1 s.
The 1.7 budget catches regression; lower it toward ~1.1 once members stop
receiving the union set. Recorded in
Known gaps.
Real-integration profiles (oneDAL, SVS, PVXS)¶
scripts/l2_real_profiles.py pins the live integrations declaratively: each
profile's revisions (and how each side's operands are obtained), the libraries
and headers actually in L2 scope, the required tools and approximate build cost,
and reproducible prepare commands. They are periodic/manual only — an
ordinary PR must never download and build oneDAL (~120 build-minutes, ~25 GB).
SVS carries two profiles, deliberately not folded into one: svs compares
the released v0.4.0 runtime distribution against PR #387's head — the
comparison the integration actually gates on, and the one that exposed the real
ABI change — while svs_pr_base compares the PR's merge base against its head.
The latter is a useful additional smoke test and cannot substitute for the
former, since it cannot expose a change that entered the branch before the merge
base. The released side is consumed, not rebuilt: rebuilding a release from
its tag measures the rebuilding host's toolchain rather than the artifact
consumers received, and when the release and CI artifacts already exist there is
no reason to rebuild at all.
The rule the module exists to enforce: an unavailable profile is reported
PARTIAL/BLOCKED/NOT_RUN with a concrete reason, never silently replaced by a
synthetic substitute, and a synthetic number is never published under a real
project's name. Five further constraints it encodes:
- No declarative L2 bundle capability exists today. A multi-library profile is measured as the supported set of per-library L2 operations — the set's total, each library's own cost, the number of distinct header contexts, and how much work was genuinely repeated across them. That is not a bundle scan and the module never calls it one. Adding a bundle capability is product work.
- A library with no public API of its own is not an L2 case. oneDAL's
libonedal_threadis recorded as a non-case with a stated reason, rather than inflated into a sixth L2 library by pointing it at someone else's headers. - A historical baseline must be historical.
validate_side_headersrejects a plan that resolves both sides' headers to one root — the easy accidental substitution (check out the new revision, build both binaries, point both--headersets at the working tree) runs fine, is faster, and is not a temporal L2 comparison. PVXS must also not be measured by running the project's own script with--depth source: that is an L4/L5 measurement. It must also be the baseline the integration declares — which is why SVS's release-to-candidate and PR-base comparisons are separate profiles. - Readiness is not a measurement.
resolve_statusanswers only whether a host could measure a profile, so its positive answers areREADYandPARTIAL. It once returnedMEASUREDfor any request whoseprepared_rootmerely existed — an empty directory, neither side's library, neither side's headers, no comparison run. Every declared operand is now checked on both sides (missing_inputs), andMEASUREDis reachable only throughpromote_to_measured, which requires a completed timed window with validated output. - A scenario's findings mean what that scenario says they mean. Every
declared scenario carries an expectation (
SCENARIO_EXPECTATIONS). "Any finding is a false positive by construction" holds for a literal self-comparison and for nothing else: two independent builds under one controlled contract are to be investigated against the recorded compiler, flags, dependencies and artifact evidence, and two intentionally different build variants differ in contract on purpose. SVS's own PR artifacts make the point — the default and public-only builds have byte-identical runtime headers and materially different exported-symbol sets — so identical header text plainly does not guarantee identical binary evidence. Treating all three alike risks reading correct detection of a build-induced ABI change as a scanner defect, or suppressing it to satisfy the wrong expectation.
oneDAL solo L2 compare across the 2026-09 perf series¶
User-supplied receipts for one identical command (solo run, 125 header roots, fresh cache) at four revisions. The operands are not in this repository and no CI lane reproduces this run, so these are reference expectations for the manual real-integration profile, not gated numbers:
| Revision | Wall | Peak RSS | Exit | Verdict | lambda at spellings |
|---|---|---|---|---|---|
| 0.6.0 | 2:43:53 | 15,176,116 KB | 4 | BREAKING | 146 |
main e9d820797 |
11:47.46 | 16,239,532 KB | 4 | BREAKING | 290 |
main 963138528 |
11:04.71 | 16,525,332 KB | 2 | API_BREAK | 0 |
main 577d856a4 |
7:49.88 | 16,525,160 KB | 0 | NO_CHANGE | 0 |
Read it as three signals, not one. Wall time fell ~21x since 0.6.0 and ~33%
across the last step. Peak RSS did not fall (15.2 → 16.5 GB) — the series
bought time, not memory, and ~16 GB sits at the edge of a nominal 16 GB host
(see memory.md). And the verdict moved BREAKING → API_BREAK →
NO_CHANGE as checkout-path-dependent lambda at spellings were stripped
(#1343, #1355): at this scale, a correctness regression that re-introduces
path-dependent identity shows up first as a spurious verdict, which no
synthetic lane below would catch. A re-measurement should record all five
columns, not wall time alone.
Complexity and cost gates beyond wall-clock time¶
Timing exponents need several sizes, repeats and the slow lane. Most real
regressions also have a cheaper, exact symptom, so these gates measure that
instead and run where noted. All of them share the synthetic workloads in
tests/_compare_workloads.py (signature, rename, enum, variable, nested-type
and type churn, add/remove), whose every entity population grows with n.
| Gate | What it pins | Lane |
|---|---|---|
tests/test_compare_call_complexity.py |
No first-party function's call count grows faster than (size ratio)^1.5 in compare(), for every workload in three modes (default, --contract evaluation, pattern verdicts + surface metrics). Names the path:line(function). Oracle is the input's size ratio, not a recorded baseline. |
unit |
tests/test_pipeline_call_complexity.py |
The same, for snapshot serialization round trips, every report format (JSON, Markdown, SARIF, HTML, JUnit) over real findings, and release reconciliation as the member count grows. | unit |
tests/test_compare_cost_budgets.py + tests/perf_call_budgets.json |
Exact ratchet budgets per mode: same-argument repeat calls of a reviewed list of expensive functions (graph/surface/idiom builders, demangle_batch, ...) and child processes per compare(), which must also not grow with input size. A figure above or below its budget fails; re-record with python scripts/audit_repeated_calls.py --write-budgets. |
unit |
tests/test_extract_call_complexity.py + tests/_cpp_corpus.py |
Call-count complexity of dump() over a generated real C++ library (namespaces, virtuals, overloads, templates, typedef chains), and of compare() over two dumped versions at 10% and 100% churn. |
integration |
perf-antipatterns (scripts/perf_antipatterns.py) |
No new list-membership, re.compile, json.loads/deepcopy, self-copying accumulation, subprocess call, loop-invariant re-sort/copy or str += inside a loop under abicheck/; existing sites are per-function counts in scripts/perf_antipatterns_baseline.json. |
ai-readiness |
tests/test_history_scaling.py |
build_longitudinal_history stays linear in the number of releases (call counts, K=5 vs 20) and its time exponent over up to 50 releases stays sub-quadratic. |
unit + slow |
tests/test_compare_scaling_shapes.py |
Wall-clock exponent per workload shape -- catches a linear number of calls whose per-call cost grows. | slow |
The first of these found a real functions x types scan in
--pattern-verdicts (per-finding rebuild of every type name), and the
repeated-call audit found the namespace detectors demangling each snapshot
once per detector per side; both are fixed.
Investigating by hand¶
- Which functions are called repeatedly with the same arguments?
python scripts/audit_repeated_calls.py --workload type_churn --n 400 --top 30(plain values compare by value, other objects by identity; generator resumptions are not counted). - Which call counts grow with input? In a test or REPL, call
profile_call_counts(tests/_call_counts.py) at two sizes and pass both tables tosuperlinear_call_sites. Salt each run's names (the workloads'tagargument): demangling and canonical-spelling caches are process-wide. - Where does the time go?
python -m cProfile -o out.prof -m abicheck compare OLD NEW, thenpython -m pstats out.prof(sort cumulative,stats 40); for a flame graph of a live run,py-spy record -o flame.svg -- abicheck compare OLD NEW(py-spyis not a dependency; install it ad hoc). For memory, see memory.md andABICHECK_MEMORY_TRACE. - Every anti-pattern site, baseline or not:
python scripts/perf_antipatterns.py --all.
Coverage gaps this workflow does not close¶
An external performance audit (2026-08) found that compare()/dump/scan
scaling has real CI protection (this page's own subject), but two adjacent
lanes had false-green failure modes — a check reporting success while
actually verifying nothing. Both are fixed; recorded here so a future reader
doesn't rediscover them from scratch.
The eval-suite.yml source-tier (L3/L4/L5) lane reported success while
scanning zero libraries. skills-src/evaluation/field/runner.py's _dump_sources() called
abicheck dump --sources ... --depth full — a rung retired from the public
CLI (ADR-043 D2, collapsed into --depth source; see the "Scan-level
scalability sweep" note above for the same retirement). Every source-tier
scan therefore failed identically with the same click.BadParameter, and the
source-tier job's blanket continue-on-error: true (deliberately set,
since one library's own build/network flakiness is meant to be tolerated —
see the job's own comment) meant a 0-of-N-scanned run still reported as a
passing, green job, with REPORT.md's source-tier table still published
looking like real coverage. Fixed two ways, matching this page's own
"gate on the systemic signal, not the individual one" pattern: the --depth
full → --depth source argv fix itself, and a new
skills-src/evaluation/field/runner.py --fail-on-empty-source gate (source_tier_broken()) that
fails only when the whole tier is broken — zero libraries scanned
successfully, or every "successful" scan captured zero real L3 build
evidence — while still tolerating one library's own failure exactly as
before. continue-on-error is no longer set at the job level; the new gate
is what now distinguishes "tolerable per-library flake" (still passes) from
"the tool itself is broken" (now fails loudly).
Still open, deliberately not attempted in this pass (each needs its own
scoped design, per this repo's "known gaps over risky reactive patches"
convention — root AGENTS.md):
- True interleaved base/head measurement. The
regressionjob runs the base branch to completion, then the head branch to completion, on the same runner — not alternating scenario-by-scenario. Interleaving would average out slow drift over the run's wall-clock (thermal throttling, a noisy neighbour) that a strictly-sequential base-then-head comparison cannot distinguish from a real regression. Doing this soundly needs eachbenchmark_scaling.pyinvocation to run a single repeat and be invoked alternately from the workflow (or a driver script that shells out to both venvs in turn), then aggregates medians across rounds — a real, separate piece of orchestration, not a flag on the existing single-shot invocation. - No scheduled real-scale lane. Every gated lane is synthetic and small
(≤ 20k functions in-process, ≤ 30 headers / 10 libraries through the CLI
— see "L2 scaling gate" above). The
oneDAL receipts above — minutes of wall time, ~16 GB RSS, and a verdict that
depends on path-independent identity — are reproducible only by hand via
scripts/l2_real_profiles.py, andscripts/bench_release_memory.py,bench_graph_materialization.pyandbench_extraction_scope.pyare wired into no workflow. A peak-RSS regression on a real multi-GB AST, or a scale-only identity/verdict drift, therefore has no automated guard. - A maintained end-to-end depth/backend matrix. The
slowperf tests and the scaling/header-graph harnesses covercompare()and the L2 attach cost well; there is no equivalent maintained CI matrix over binary / headers (clang + castxml) / build (CMake + Bazel scoped/fallback) / source (seeded + unseeded) × cold/warm cache.skills-src/evaluation/field/scan_level_scaling.pysweeps the level axis but is manual-only (real clang time), andskills-src/evaluation/field/scaling.py'sABICHECK_L4_JOBSsweep is likewise manual. - Per-run performance receipts — now implemented for the L2 CLI path
(
scripts/perf_receipt.py, consumed byscripts/check_l2_cli_perf.py). A run now emits a versioned receipt carrying wall/user/sys time with its CPU accounting scope named, sampled concurrent process-tree RSS alongside a separately-labelledru_maxrss, nested phase windows, per-kind native invocation counts (header extraction vs. include pass vs. version probe), output sizes, correctness-validation status, and the effective thresholds that gated the run. What is still open is the wider depth/backend coverage below, and an L3/L4/L5 equivalent: the receipt layer is generic, but only the L2 CLI harness feeds it today. - Repeated L3 collection under
dump --sources/--build-info(P0.3) — open, accepted cost, not this PR's to close. Include seeding and compile-context derivation (derive_l2_include_dirs/derive_l2_compile_context,abicheck/buildsource/l2_seed.py) each independently trigger a full L3 collection (collect_inline_pack(layers=("L3",))) for the exact same(sources, build_info)inputs — up to three L3 collection passes per side for some input shapes, counting the embed step's own separate collection. A sentinel-basedevidence=sharing mechanism was prototyped in an earlier revision of this PR, then dropped:mainhad independently landed a parallel, far more heavily reviewed P0.3 pass in the interim (10+ review rounds onderive_l2_compile_contextalone — forced-language interaction with a matched compile unit's derived-std=, MSVC/std:-vs--std=precedence, ambiguity-signature narrowing order, ...) whose own_L2SeedPackArgsdocstring already considered and explicitly declined folding these two calls into one, calling the double collection "an accepted, documented cost" so eachderive_l2_*function's return shape stays independently additive. Re-litigating that call from inside a merge-conflict resolution, in code this delicate and this recently stabilized, was judged the wrong place to relitigate it — see this repository's own "known gaps over risky reactive patches" convention. This is now in scope for theclassifyjob'sPERF_SENSITIVE_PATTERNS(see "CI integration" above), so a change to it will run this workflow even though no benchmark scenario targets its specific cost yet. Closing it for real needs its own dedicated pass, reviewed on its own terms rather than folded into an unrelated CI-tooling PR. - Per-job shard classification. The
classifyjob (see "CI integration" above) reports one sharedrunoutput gating every downstream PR job uniformly — the exact same set of jobs a changed path used to trigger together under the oldpaths:filter, just moved into tested code. SplittingPERF_SENSITIVE_PATTERNSinto separate shards (e.g. one gating onlyscaling/regression, another gating only the twoheader-graph-*jobs) would cut CI cost for a PR that only touches one area, but was deliberately not attempted:check_header_graph_perf.pyimportsabicheck.buildsource.header_graph(reached throughservice._attach_header_graph), so even a directory as specifically-named asabicheck/buildsource/is relevant to the header-graph jobs, not just the compare()-scaling jobs its name suggests — a real per-job split needs a verified transitive-import trace from each script's own entry point, not a guess from a module's directory or filename. A wrong split would silently under-cover one shard, exactly the failure mode the classify job exists to close, so it's tracked here as future work rather than attempted speculatively.
Coverage gap analysis & remaining gaps¶
A second pass (continuation of PR #331) audited the whole pipeline for scaling risk and extended the harness to the highest-value uncovered paths plus per-call peak-memory tracking and PR-vs-base drift detection. Current status:
| Path | Status | Notes |
|---|---|---|
compare() post-processing |
✅ covered | Original PR #331 scenarios. |
| Suppression audit | ✅ covered | suppression_audit scenario + slow test. O(rules × findings); linear in findings for a fixed ruleset. |
| HTML / SARIF / JUnit reporting | ✅ covered | report_html / report_sarif / report_junit scenarios + slow tests; all linear. (to_markdown/to_json already guarded.) |
| Enum / typedef / union / wide-struct / vtable diffing | ✅ covered | enum_churn, typedef_churn, union_churn, wide_struct, vtable_churn. Sweeping typedef/union/vtable/enum across sizes (the original table only sampled n=500, so no exponent was ever computed) exposed a genuine ≈O(n²) in the affected-symbol enrichment — see fix #8; all four are linear after it, and opaque_filter dropped from ≈1.7 to ≈1.2 as a side effect (its cost was the enrichment, not _filter_opaque_size_changes). |
| PE/COFF & Mach-O diff arms | ✅ covered | pe_churn / macho_churn build pe=/macho= snapshots so diff_platform's PE/Mach-O detectors run. |
| Opaque-handle pointer-only check | ✅ covered | Was the O(candidates × functions) residual (_is_pointer_only_type, regex per pair, surfaced by type_churn ≈1.6); fix #9 linearized it via _opaque_usage_index (one indexed pass). Both type_churn and opaque_filter are now linear. |
| Versioned-symbol-scheme collapse (ICU/OpenSSL) | ✅ covered | versioned_rename_churn reproduces the field-eval P08 ICU 75→78 shape (16 k removed/added churn findings + the scheme-collapse pass). Profiling it surfaced a per-finding name re-tokenization in the namespace detectors (diff_namespaces._segments), now fast-pathed for plain names. ~1.1–1.2 tail exponent; the residual is the post-processing detector fan-out, not the scheme recogniser. |
| Severity categorization | ✅ covered | severity scenario over categorize_changes; linear. |
| Fuzzy rename matching (accept path) | ✅ covered | fuzzy_rename_churn — every symbol genuinely renamed → one func_likely_renamed per pair, the cost driver P11-refined identified (ICU 2134 renames = 94.5 s; rename detection, not symbol count, dominates). The pre-existing rename_churn only exercised the reject path (disjoint names, zero matches). Linear at ICU scale (≤8 k). |
| Version-node migration fan-out (LLVM bump) | ✅ covered | version_node_churn — every export moves LIB_1.0 → LIB_2.0, reproducing the LLVM 17→18 36,991-symbol_moved_version_node shape and the post-processing fan-out over it. Linear to 50 k. |
| Peak memory (all scenarios) | ✅ covered | tracemalloc peak_mb column + --max-memory-mb gate (cold-cache pass), plus process rss_mb (resource.getrusage) + --max-rss-mb gate — RSS catches native (pyelftools / c++filt) allocations tracemalloc cannot see (the ~330 MiB LLVM-scale figure). |
| Historical / PR-vs-base memory regression | ✅ covered (gating) | --regress-memory-tolerance/--regress-min-delta-mb + the memory half of the regression workflow job compare peak tracked heap against the base branch under max(20%, 4 MiB). Before it, memory was gated only against an absolute ceiling, which a doubling well under that ceiling passed silently. See Memory regression. |
| Historical / PR-vs-base regression | ✅ covered (now gating) | --baseline/--regress-tolerance + the regression workflow job measure the base branch and PR head on the same runner and flag scenarios that got slower by more than the tolerance — catching gradual drift the per-run exponent misses. continue-on-error is dropped, so it now blocks. See Baseline regression. |
| Dump / snapshot creation (DWARF/PE/PDB) | ⚠️ partial | The synthetic harness can't run the real parsers. The ELF symbol-table parse and the DWARF debug-info parse (-g build) are now guarded by tests/test_perf_dump_scaling.py (integration, gcc-only) — DWARF being the dominant real-library dump cost (ICU 18.6 MB snapshot, openblas 23 MB / 9.5 s). The serialize scenario proxies the rest of the pipeline. PE/COFF + PDB parsing remains unbenchmarked — those need a committed binary or a synthetic byte-stream generator (no Linux-only toolchain produces them). |
| Appcompat HTML / stack analysis / appcompat filtering | ⚠️ not benchmarked | stack_checker runs one compare() per dependency (inherent). Appcompat filtering uses set-membership lookups (appcompat.py — O(1) per change, likely already fine) and appcompat_html.py is linear by inspection; neither is timed. |
| Directory multi-library compare (release fan-out) | ✅ covered (scaling) | check_l2_scaling_perf.py's library axis: 1 / 3 / 6 / 10 libraries through the real CLI, marginal exponent gated at 1.7 (measured 1.39). |
| Bundle / environment-matrix compare | ⚠️ not benchmarked | Per-library cost is covered; bundle and environment-matrix orchestration is not. |
Recommended next steps (in priority order)¶
- ~~Wire a budget gate~~ — done: the lane now runs
--max-exponent 1.4(nested_typesexempt viagate_exponent=False) and--max-rss-mb 1024, andcontinue-on-erroris dropped on both the scaling andregressionjobs. The--regress-tolerance 0.5PR-vs-base check also blocks now; loosen a threshold rather than re-addingcontinue-on-errorif runner variance flakes a lane. - Extend the dump/parse guard to PE/PDB — the ELF symbol-table and DWARF
parses are now covered (
tests/test_perf_dump_scaling.py,integration, gcc +-g); the PE/COFF and PDB parsers still need a committed binary or a synthetic byte-stream generator behind theintegrationmarker (no Linux-only toolchain emits them). - Benchmark bundle / environment-matrix orchestration — directory multi-library scaling is covered by the L2 scaling gate. Bundle and environment-matrix orchestration, including appcompat/stack fan-out, is still untimed.
- ~~Optimise the super-linear residuals~~ — done: fix #8 linearized the
enrichment (typedef/union/vtable/enum/opaque), fix #9 the opaque pointer-only
check (
type_churn). No quadraticcompare()path remains at tracked sizes.
Baseline regression¶
The per-run scaling exponent catches catastrophic blow-ups but not a gradual 15–20 % slowdown. To catch drift, the harness can compare against a baseline:
# On the base branch / a prior commit, capture a baseline (5 repeats -> a
# meaningful median/cv; --meta records who/what this measurement is):
python scripts/benchmark_scaling.py --repeat 5 \
--meta git_sha=$(git rev-parse HEAD) --meta side=base \
--json-out base.json
# On the PR head, measure and compare (fails if any shared point regresses
# by more than max(15%, 100ms) vs. its baseline):
python scripts/benchmark_scaling.py --repeat 5 \
--meta git_sha=$(git rev-parse HEAD) --meta side=head \
--baseline base.json \
--regress-tolerance 0.15 --regress-min-delta-seconds 0.1
Memory regression¶
The rule above gates time. Peak memory has had the same treatment since the
memory gate was added (originally its own memory-regression workflow job,
now the second half of the regression job):
# Same two-step shape, with memory tracking left ON (no --no-memory), and
# --repeat 1 because a tracemalloc peak is an allocation count, not a
# wall-clock duration, so repeats buy far less than they do for timing.
python scripts/benchmark_scaling.py --repeat 1 --json-out base-memory.json
python scripts/benchmark_scaling.py --repeat 1 \
--baseline base-memory.json \
--regress-tolerance 100 \
--regress-memory-tolerance 0.20 --regress-min-delta-mb 4
A point regresses once its peak exceeds the baseline's by more than
max(--regress-memory-tolerance x baseline, --regress-min-delta-mb) — the
same combined relative/absolute rule as the timing gate, computed by the same
perf_measurement.combined_regression_threshold, so the two cannot drift
apart. The memory defaults are tighter (20 % / 4 MiB, against 50 % / 0 s):
a tracemalloc peak counts bytes the interpreter actually allocated and does
not move with GC timing, scheduler preemption or a cold cache, so the noise
that forces a loose timing tolerance is largely absent. Baseline peaks below
an 8 MiB floor are skipped, for the same reason the timing rule ignores
sub-50 ms points: at that size, fixture and import allocations dominate the
figure rather than the code under test.
The memory baseline is read from the same --baseline report — a
benchmark_scaling.py report already carries peak_mb on every point, so
there is no second file to keep in sync. A baseline produced with
--no-memory (or by a build predating this gate) carries no peak_mb at all;
that is reported as an inactive memory gate rather than treated as "allocated
nothing", which would otherwise flag every point in every run. The timing gate
still fails closed on an empty baseline, so a wholly missing or malformed
baseline is still caught.
Why this is a separate measurement run (the memory half of the
regression job, not another flag on its timing run): tracing the heap is not timing-neutral. measure()'s memory
pass clears every live lru_cache between sizes and runs an extra untimed
cold call, which is precisely the bias the timing job's own --no-memory
comment documents. Timing and memory therefore cannot be gated from one run.
In the memory run memory is on for both sides, so that bias applies equally
and cancels; the timing tolerance there is explicitly neutralised
(--regress-tolerance 100) so a distorted timing can never fail the memory
gate. The two runs need two runs, not two runners: they share one job,
one checkout pair and one pair of venvs, and the memory steps start only after
the timing steps have finished, so tracing never overlaps a timed
measurement. (Until 2026-10 they were two jobs, which checked out and
installed both sides twice on two runners for no measurement benefit.)
What it closes. peak_mb was recorded long before it was gated, and was
checked only against the absolute ceilings --max-memory-mb/--max-rss-mb.
An absolute ceiling catches a regression only once it crosses the ceiling, so
a change doubling a scenario's allocation from 200 MiB to 400 MiB passed the
then-2048 MiB ceiling in silence — exactly the gradual drift the timing side had
had a base-branch comparison for since PR #768.
Each point's reported/gated figure is the median of its timed repeats (plus
one untimed warmup) — not the fastest one; see scripts/perf_measurement.py
for why "keep the minimum across a couple of runs" systematically hides
regressions rather than catching them, and reports no run-to-run noise signal
at all (min/max/p95/coefficient-of-variation are recorded alongside the
median for exactly that reason).
Only scenarios present on both sides are compared (a scenario new in the PR
has no baseline and is skipped), and baseline times below a 50 ms noise floor
are ignored regardless of the threshold rule. A point regresses once its
absolute slowdown exceeds max(tolerance x baseline, --regress-min-delta-
seconds) — the combined relative/absolute rule protects a small baseline (a
few ms) from flagging on ordinary noise while still catching a genuinely large
percentage regression on a large baseline; --regress-min-delta-seconds 0
(the default) reduces to a pure percentage tolerance. A --baseline that
loads to zero points, or that shares zero (scenario, size) points with what
was actually measured, is a hard failure, not a silently-skipped comparison
that reports a clean pass — this closes a real incident where the regression
job's base-measurement step swallowed its own failure with || true, so a
PR could merge having "compared" against a baseline that was never actually
captured.
The regression
workflow job automates this on PRs: it installs the base branch and the PR head
into separate venvs on the same runner, runs both, and prints the regressions to
the job summary. It gates (a regression past the threshold above fails the
job) — loosen --regress-tolerance/--regress-min-delta-seconds rather than
re-adding continue-on-error if runner variance proves noisy.
The header-graph attach-cost gate
(scripts/check_header_graph_perf.py, see below) follows the identical
median/warmup/combined-threshold design over its own --repeat/
--regress-tolerance/--regress-min-delta-ms flags — the two scripts share
the underlying statistics via scripts/perf_measurement.py so their
regression math can't independently drift.
Scan level cost model: one cliff at L4¶
A real scan-level sweep on two UXL libraries (oneTBB v2021.12→.13, C++;
UMF v0.10→v0.11, C; raw data in skills-src/evaluation/validation/data/uxl_scan_results_2026-06.json)
shows the cost has one cliff, at the L4 AST-replay boundary, and the cheap
tier below it is dominated by the binary dump + always-on pattern scan, not by
the source layer:
| Level | Reaches | oneTBB (C++, 40 TUs) | UMF (C, 50 TUs) |
|---|---|---|---|
s0 diff classifier |
— (L0/L1 + pattern) | ~29 s | ~17 s |
s1 compile-DB |
+L3 | ~29 s | ~17 s |
s3 lexical |
(pattern only) | ~29 s | ~17 s |
s4 symbol/graph index |
+L3 +L5 | ~29 s | ~17 s |
s5 targeted AST |
+L4 (changed TUs) +L5 | ~222 s | ~22 s |
s6 full AST |
+L4 (all TUs) | ~215 s | ~21 s |
Rules of thumb:
- The cliff height is a C++ phenomenon. L4 cost = clang per-TU AST replay; it
scales with C++ template/STL instantiation depth, not
.soor TU count. Heavy C++ (oneTBB) jumps ~7× (29→222 s); plain C (UMF) barely moves (~1.3×, 17→21 s). Budget L4 by how templated the source is. - The cheap tier (s0–s4) is one price. All four cost the same — the floor is
the DWARF dump + lexical scan of the tree. Pick by coverage you need, not
cost:
s0≈s3(L0/L1 + pattern only),s1adds L3,s4adds the L5 reachability graph without paying for L4.s4is the structure sweet spot. s5is only cheaper thans6with a diff seed. Without--since/--changed-paththe changed-TU set is empty ands5replays every TU — same cost ass6. With a one-file seed, oneTBBs5dropped from 222 s to 11.5 s (~19×) for the identical verdict. This scoping applies only to thesource-changedcollect mode — i.e.s5and--mode pr. The other AST modes replay full scope regardless of any seed:--mode pr-deepresolves tograph-full, and--mode baseline/s6to full (source_replay.CI_MODE_TO_SCOPE:source-changed→changed,graph-full→full), so pinning those in CI will not produce the scoped speedup.auditcosts the same as the baseline modes — the wall-clock is L4/L5 collection of the new side, not the baseline diff.
The verdict was identical across all levels on both libraries: the authoritative L0/L1 binary diff sets the gate; L3–L5 add coverage/localization, not a different pass/fail. For a CI gate, the cheap tier suffices; spend on L4 only when you want source-body semantics or PR localization for humans.
Scan-level scalability sweep¶
The UXL run above fixed the corpus (two real libs) and varied the level. The
complementary question — how each level scales as a project's complexity
grows — is swept by skills-src/evaluation/field/scan_level_scaling.py,
a self-contained harness (no network/repo) that synthesises STL/template-heavy
C++ trees of increasing TU count, builds them with the host compiler, and runs
scan at each level against a slightly-changed baseline — recording wall time
and peak child RSS (os.wait4) per (size, level).
Two results (measured back when the harness still had a separate full (s6)
sweep entry alongside seedless source (s5) — see the note at the end of this
section for why that entry was later removed, without invalidating either
finding below):
- The cheap tier is flat in TU count.
binary/headers/build/graph(s4) cost the same at 4, 8, and 16 TUs (tail exponent ≈0) — they are priced on the binary dump + L2 header AST + L3 compile-DB parse, none of which grow with the number of.cppfiles. Seedless (full-tree-replay) source-level scanning is linear in TU count (every TU is replayed). Both as expected. - Seedless
--depth source(s5) used to hide a full-tree cost — now fixed. It cost ~2× the wall time and ~2.5× the RSS of the seeded run for the identical L4 coverage (both reportL4=1/1), and the gap widened with TU count. The seed scopes both the L4 replay and the L5 clang call-graph pass to the changed TU; without a seed the L4 replay fell back to headers-only (one TU) but the call-graph pass ran over the whole compile DB — a secondclang -ast-dump=jsonover every TU. The unseeded call-graph pass now scopes to the same compile units the L4 replay used (headers-only), so it is consistent with the L4 surface and no longer scales with the tree (~2.4× faster on a synthetic n=8 tree, identical verdict). Seeded runs are unchanged.
Why there is no separate full sweep entry anymore. ADR-043 D2 retired
--depth full from the public CLI ladder entirely — it collapsed into
--depth source, since replay scope (a change seed present vs. absent), not
evidence depth, was the only thing distinguishing them, and scan itself now
resolves that scope from whether --since/--changed-path was given (see
abicheck/model/evidence_depth_levels.py's EvidenceDepth.FULL/SourceScope
docstrings). The harness's own seedless "source" entry is the shape the
old "full" entry measured — keeping both would just re-run the identical
--depth source argv twice under two names. (Concretely, before this fix
--depth full was a hard click.BadParameter — a harness bug in the same
family as the skills-src/evaluation/field/runner.py one described in "Coverage gaps this workflow
does not close" below, just in a manual-only harness rather than a scheduled
CI lane, so it never produced a false-green.)
That whole-DB call-graph pass shells out to the same multi-GiB
clang -ast-dump=json as the L4 replay, but its worker count
(call_graph._call_graph_jobs) was CPU-bound only — it lacked the
RAM-aware, cgroup-aware clamp the L4 replay grew (_l4_jobs → _l4_mem_cap)
after the UXL oneTBB/oneDNN OOM. On a constrained host the L4 pass was protected
but the unseeded call-graph pass was not. _call_graph_jobs now shares the L4
memory cap (_call_graph_mem_cap → _l4_mem_cap, same ABICHECK_L4_JOB_MEM_GIB
budget); ABICHECK_CALL_GRAPH_JOBS still overrides the CPU count but memory wins
over an over-eager override, exactly like _l4_jobs.
L2 header-scan deadline enforcement (pathological headers)¶
A real-world field report (Intel SVS) found the cheap tier's flatness above has
an exception: a pathological header (deep #include/template complexity) can
make the L2 clang/castxml AST dump itself run far longer than its on-disk size
suggests — the report's own scan --dry-run estimate read 0.51 s for a header
set whose actual parse ran over 15,000 s and 3+ GiB RSS before an external
SIGKILL, because --budget was checked only once, after the whole scan had
already finished, and the clang/castxml subprocess.run(timeout=120) call had
no process-group isolation (a timeout only killed the direct child, orphaning
any compiler-driver grandchild).
The fix (abicheck/deadline.py) threads a shrinking --budget deadline down
to the L2 subprocess boundary (checked before each clang/castxml invocation,
not only at the end) and runs that subprocess in its own process group so a
timeout kills the whole tree. This is a bounding fix, not a speedup — a
genuinely pathological header still costs whatever clang/castxml need, up to
whatever --budget is given; it now fails cleanly at that boundary instead of
running unbounded.
Regression/perf-tracking coverage, deliberately without needing the SVS corpus itself (see "Extract minimal synthetic fixtures" guidance):
tests/test_deadline.py— fast, synthetic (sh/sleep), proves the process-group kill and mid-stage budget check mechanisms directly.tests/test_header_scan_deadline_integration.py— real clang, self-skips if absent.test_pathological_header_aborts_within_bounded_time_under_tiny_budgetreproduces the SVS shape with a genuinely expensive (not simulated) 4-line header: a recursive template chain whose clang-ast-dump=jsonoutput grows steeply super-linearly with recursion depth (calibrated locally: depth 100 → ~40 MB/0.2 s, depth 200 → ~280 MB/0.6 s, depth 300 → ~900 MB/1.5 s — kept at depth 150 in the test to stay CI-safe), and asserts a tiny budget bounds it. Theslow-marked companiontest_pathological_header_natural_cost_is_trackedrecords that header's unbudgeted natural cost so a future regression (lost disk cache, a clang upgrade changing dump behaviour) shows up in the existing per-test duration trend (tests/conftest.py'sABICHECK_DURATIONS_JSONhook →scripts/summarize_test_durations.py→ the CI run summary) — the same mechanism this page already relies on for thecompare()-scaling story, rather than a new bespoke benchmark harness.scripts/benchmark_scaling.pyis deliberately not the home for this: it is pure-Python by design ("no compiler/castxml" — see its module docstring) and this concern is inherently compiler-driven.
Fixed (follow-up): the L2 path (dumper._clang_header_dump, via the new
dumper_clang_errors.run_clang_to_ast_file) now spills clang's AST-dump
stdout straight to a temp file, mirroring the L4 per-TU replay
(source_extractors/clang.py's _run_ast_to_file) instead of capturing it
into a Python str. The calibration above showed a tiny header can
legitimately produce hundreds of MB to multiple GB of AST-dump output, which
capture_output=True would buffer on top of the parsed dict this code also
builds; measured ~27% lower Python-heap peak (tracemalloc) on the depth-150
fixture (364.5 MB → 267.1 MB) with the fix. The same deadline.run_bounded
treatment (shrinking --budget deadline, process-group kill on timeout) was
also extended to preprocessor_facts.py's live extractor and both L4 source
extractors (source_extractors/clang.py, source_extractors/castxml.py),
which previously used the same fixed-timeout/no-process-group pattern the P0
fix closed for L2 — and a --budget deadline expiring during a PE/Mach-O
header-scoped dump (service._try_header_scoped_dump) is no longer silently
swallowed by the broad except Exception that falls back to export-table
mode for a merely-unavailable header backend.
L2 acquisition: cache-key walk, AST cache write, include-map fan-out¶
Three costs around the L2 header parse that are not the parse itself. All
three were measured locally (this host: 4 CPUs, clang 20, Python 3.13) against
current-main function bodies and synthetic-but-realistically-shaped inputs --
not through a full oneDAL/SVS/PVXS dump/compare run, so read them as
component figures, not an end-to-end speedup claim.
Request-local context sharing¶
Scalar comparison and release-member fan-out now own one request-local L2
acquisition table. A content-addressed frontend context has one in-flight
clang/CastXML producer and one export-neutral SemanticIR normalization. Each
binary still runs the concrete parser's legacy declaration construction against
its own export set: those parsers own inclusion, fallback identity, and surface
facts, so reconstructing their binding from a neutral parse was experimentally
rejected after it lost real removal findings. Waiters have deadline-bounded queue
time, do not cancel a producer another member still needs, and failed or
input-mutated acquisitions are never retained as reusable results.
For Intel DPC++ the ordinary host context is also acquired directly with
-fsycl -fsycl-host-only. This prevents the compiler from serializing a full
unused device JSON document before the shared host evidence can be normalized.
An explicit device request continues through the multi-context framing route:
host-only and device ASTs are different parse contexts and must never share a
cache entry merely because their entry headers match.
The compact result retained after each release-member comparison also carries
that member's already-validated ElfMetadata and recorded native filename.
Final bundle-graph assembly consumes those fields directly; it falls back to a
stored-snapshot decode or live ELF parse only for a member that did not resolve
earlier. Filesystem alias probing remains enabled only for genuinely live paths,
so carrying metadata forward does not re-resolve a stored identity against the
caller's current working directory.
A six-library local control used one real shared public header declaring six C functions and six separately compiled DSOs, each exporting a different one. Cold-cache clang acquisition changed from 6 compiler + 6 normalization calls, 4.253 s to 1 compiler + 1 neutral normalization, 3.375 s; CastXML 0.6.11 changed from 6 + 6, 1.900 s to 1 compiler + 1 neutral normalization, 1.492 s. Legacy declaration binding remains one pass per member and is not included in the eliminated-call claim. Every member retained all six header declarations and only its own export was marked binary-exported. These are small-fixture local measurements, not oneDAL results; the periodic real-profile lane remains the owner of oneDAL/SVS/PVXS claims.
- Include-tree inventory for cache keys (
extract/cache_header_scan.py). Every header-parse cache key walks each include root to fold in (path, mtime) for every header-like descendant --dumper_ast_config. _cache_keyfor the AST cache,snapshot_cachefor the whole-snapshot one.Path.rglob("*")materializes aPathper entry visited and then sorts those objects; carrying path strings through the traversal and building aPathonly for the surviving, already-ordered entries measured 2.4x faster (/usr/include, 4188 matching entries: 47 ms -> 19 ms; a generated 15k-header tree: 151 ms -> 62 ms) with a byte-identical key. Two traversal behaviours are load-bearing and were matched deliberately rather than "cleaned up": the order is component-wise (PurePathcompares normcased parts, soa/b/c.hprecedesa/b.h) and symlinked directories are not descended into (which is also what makes a symlink loop terminate). Asorted(str(p) ...)"simplification" changes every cache key on disk while looking like a no-op --tests/test_cache_header_walk.pypins both against asorted(rglob("*"))oracle over generated trees. - The clang AST cache write (
dumper_cache._atomic_write_json).json.dumpnever reaches the C encoder:JSONEncoder.iterencodeselectsc_make_encoderonly under_one_shot, which isdumps' path. So the DPC++ host/device cache write was encoding a large AST a fragment at a time in pure Python. Descending the large containers in Python and handing each bounded subtree to the one-shot C encoder is byte-identical and measured 5x faster on a deep-template AST fixture (0.25 s -> 0.05 s for 2.1 MB of output; the ratio grows with document size). The obvious_atomic_write(path, json.dumps(obj).encode())is not the fix -- it holds a second full-size copy of a tree that can be multiple GB, which is why the streaming write exists at all. The plain (non-DPC++) path still streams a raw file copy of clang's own stdout and pays none of this. - Per-header
clang -Minclude-map fan-out (buildsource/include_graph_workers.py). The probes are independent subprocesses, so the pass was wall-clock-bound on nothing but serialization: 32 wrapper headers over a shared include took 1.71 s at one worker, 0.44 s at four (3.9x), with byte-identical depfiles; eight workers bought nothing measurable on 4 CPUs, which is why the auto default is CPU-derived rather than a fixed 8. This trades memory for wall time rather than removing compiler work -- child CPU time is unchanged -- so the pool is clamped by the sameprocess_resourcesRAM probe the L4 pool uses (at a much smaller per-worker budget:clang -Mis preprocess-only). One One sizing subtlety worth keeping: the shared gate is sized from the host budget alone and never replaced, while each pool's own size is that budget narrowed tomin(..., unit_count). Deriving the gate from a pool's size instead is what a review round caught here -- two sides with differing header counts resolve differing sizes, and rebuilding the gate for the second left the first pool holding an orphaned semaphore, so both admitted their full quota at once (tests/test_include_graph_parallel.py:: test_sides_with_different_unit_counts_still_share_one_gate, andperf.shared_resource_gate_keyed_on_a_per_caller_valueintests/regressions/manifest.py). Two further review rounds on the same gate are worth reading before touching it: the serial path (one unit, orjobs=1) must take a slot too -- one child is not no children, and an ungated one alongside the other side's full quota ishost_limit + 1-- and the wait for a slot must be bounded by the run's aggregate and--budgetdeadlines, or a short-budget request blocks past its own budget behind an unrelated request's long probe and reports the overrun only afterwards. The third round is why the gate is a counter plus a condition variable rather than aBoundedSemaphoreat all: a semaphore holds a size, and a size cached at construction cannot notice the host budget shrinking under a long-lived process -- every pool re-reads the budget and narrows itself while the stale gate keeps admitting the old number. Re-reading the limit on each admission removes the staleness instead of resizing around it, which would need to detect when every previous holder had drained. One thing that is not a valid shortcut here, and was ruled out with a real clang: replacing the per-header probes with a single umbrella TU. A header with an include guard that a previous header in the umbrella already pulled in disappears from the second header's dependency list entirely, so the umbrella silently loses edges the per-header probes see. The safe win is scheduling the same invocations better, never changing the compilation context.
L4 source-replay (dump-side) performance¶
The scaling harness above is pure-Python and times the compare pipeline. The
dump-side L4 source ABI replay (clang per-TU AST extraction) is a separate
cost, timed by skills-src/evaluation/field/scaling.py
on real source trees (it needs clang + a built tree, so it is manual, not in CI).
Knobs and the reasoning behind them (abicheck/buildsource/source_replay.py):
ABICHECK_L4_JOBS— worker count for the per-TU extract pool. Auto =min(TUs, cpu_count, 8). An explicit override is clamped tomax(8, 2×cpu_count)(logged when it fires) so a stray=64can't oversubscribe a host into thrash (skills-src/evaluation/field/SCALING.mdalready saw jobs=8 on 4 CPUs regress). Set=1to force serial (determinism).- Memory cap (auto + override). A single template-heavy C++ TU's
clang -ast-dump=jsonoutput — and its in-Python parse — can reach several GiB, so the worker count is also capped by available RAM (min(…, available / ABICHECK_L4_JOB_MEM_GIB), default3.0GiB/worker, Linux only). "Available" is the smaller of hostMemAvailable(/proc/meminfo) and the cgroup memory headroom (v2memory.max−memory.current, or v1memory.limit_in_bytes−memory.usage_in_bytes), so a container/pod confined to a small cgroup on a large host sizes its workers to what it is actually allowed to use rather than to host RAM. On a low-memory host this stops N concurrent giant ASTs from exhausting one process and getting the whole replay OOM-killed (the kernel SIGKILLs it →exit -9, all L4 work lost — observed on the UXL oneTBB/oneDNNs5/s6full-target replays on a 15 GiB host). The clamp is logged. For a template-heavy tree on a constrained host, prefer a seeded/scoped scan (--since/--changed-path→ a handful of TUs) over a full-targets5/s6; it sidesteps both the time and the memory cliff.ABICHECK_L4_JOB_MEM_GIBtunes the per-worker budget (lower = more workers). ABICHECK_L4_EXECUTOR(threaddefault /process) — after clang returns, the extractor parses clang's large JSON AST dump and builds structural fingerprints: pure-Python, GIL-bound work. A thread pool parallelizes only the clang subprocess wait, so that post-processing serializes on the GIL — part of the ~60–83 % "serial fraction" inskills-src/evaluation/field/SCALING.md.processruns the extract phase in aProcessPoolExecutor, parallelizing the AST work too (at the cost of pickling eachSourceAbiTuand per-process spawn). It is opt-in pending a measured win — compare the curves withpython skills-src/evaluation/field/scaling.py --jobs 1,2,4 --executor processvsthread. The driver falls back to serial if a process pool can't start (sandbox, spawn import error), so it never aborts L4.- Concurrent AST memory: clang's output is spilled to a temp file, not captured.
A template-heavy TU's
clang -ast-dump=jsonoutput can be multiple GiB. Capturing it (capture_output=True) holds the whole AST string in the heap from the moment clang finishes — and because the Cjsonparse holds the GIL, the default thread pool serializes parsing, so all N workers sit holding their giant AST strings (≈ N × text) while queued behind the GIL. Spilling clang's stdout to a temp file keeps those payloads on disk until each worker's turn to parse, so the heap holds roughly one payload at a time instead of N.json.loadstill reads the file back to parse, so a single TU's parse peak is unchanged (≈ serialized text + tree) — this is a concurrency win, not a per-TU one — and it also drops thetext=Truedecode copy (bytes parse) and frees the tree before the macro pass. The per-TU tree itself (~2–5× the AST text) is irreducible without a streaming JSON parser (a dependency the project avoids); for a template-heavy tree on a constrained host, a seeded/scoped scan is still the structural win. ABICHECK_L4_CACHE_DIR— persists the per-TU cache (SourceAbiCache, content-addressed + per-included-file dependency-hash invalidation) acrossdump --sourcesruns. Previously the inline path passed no cache, so every dump re-extracted every TU; wiring this dir makes a cold run (evalE4: zstd 48.6 s) reuse the warm cache (3.4 s). Point it at a CI cache directory restored viaactions/cacheto start every CI run warm. The cache validation phase is serial, so the dependency digest is memoized per replay pass — a public header included by N TUs is hashed once, not N times.
S2 preprocessor pre-scan performance (scan --depth build)¶
scan's S2 preprocessor pre-scan (ADR-035 D2, buildsource/preprocessor_facts.py)
is the conditional tier that runs once L3 build evidence is available: per-TU
ABI-macro-value capture (clang -E -dM) and public-header-leak detection
(clang -M). It is advisory-only (never a verdict on its own), but on a
real-world build it dominated scan --depth build's wall time: a reported
4-minute-to-20-minute jump on a library with ~2000 translation units, ~920s
spent in this tier alone — one serial clang -E -dM invocation per compile
unit, with no cap, even though L3 ingestion itself (bazel aquery) cost only
single-digit seconds.
Parallel probing, semantics-preserving for the ABI-macro-value/leak facts
this tier reports: each probe is one I/O-bound clang -E/-M subprocess
wait — unlike L4's AST parse, it is not GIL- or RAM-heavy — so probes run
concurrently via a thread pool, one probe per compile unit / public header
(unchanged from before this fix).
Deliberately NOT deduped by compile context. An earlier revision of this
fix probed once per distinct (language, cwd, flags) signature and fanned
the result out to every compile unit sharing it, on the premise that the
curated ABI-macro list (_GLIBCXX_USE_CXX11_ABI, NDEBUG,
_ITERATOR_DEBUG_LEVEL, …) is almost always driven by command-line/
predefined macros, not by a TU's own source text. Reverted (Codex review,
fresh evidence): clang -E -dM reflects the TU's own source and #include
chain too — two TUs sharing identical flags can legitimately resolve a
curated macro differently (a debug-only TU #undef-ing NDEBUG, a
conditionally-included config header) — and that is exactly the shape of
bug find_macro_divergence exists to catch. Flag-based dedup would silently
fan the first such TU's value out over the rest, hiding a real divergence
rather than merely under-covering a rare case — a detection regression in
the tier's own core purpose, worse than the perf win it bought. See this
repo's "known gaps over risky reactive patches" convention (root
AGENTS.md).
Knobs (abicheck/buildsource/preprocessor_facts.py), mirroring the L4
conventions above:
ABICHECK_PREPROCESSOR_SCAN_JOBS— worker count for the probe pool. Auto = one worker per compile unit / public header, capped by the samemax(8, 2×cpu_count)ceiling L4 uses. Set=1to force serial (determinism).ABICHECK_PREPROCESSOR_SCAN_MAX_PROBES(default 512) — caps the number of compile units / public headers probed, bounding worst-case cost on a build with an unusually large number of either. Truncation is reported in the scan's diagnostics and folds the coverage row down topartial(never silent — this file's "no silent caps" convention).ABICHECK_PREPROCESSOR_SCAN(default on) — set=0to skip the S2 tier entirely while keeping L3 for the other tiers. Reported as an honestnot_collectedcoverage row, the same as a missing compile DB or missingclang— never silently counted as clean.
The remaining, undeduped cost (~920s serial → roughly that divided by the
worker count, parallel) is a real, unavoidable floor for this tier's design:
every compile unit's macro-affecting #include chain can only be observed by
actually running the preprocessor over it. A caller that wants the D2
coverage row without this cost has the disable knob above; there is no
"fast and still divergence-complete" third option.
Why not precompiled headers (PCH) / modules?¶
A natural idea to cut the repeated per-TU header parse is a PCH over the public
headers. It does not apply here: clang -Xclang -ast-dump=json does not
re-emit declarations that came from a PCH, so loading one would silently drop the
very header surface L4 exists to capture — a correctness bug, not a speedup. The
right levers for repeated-parse cost are therefore the per-TU cache and the
replay scope (changed/target), both already in place.
--budget mid-step preemption gap (found live, pvxs full-version-matrix scan)¶
Real-world evidence: scan --depth
source on a real 62-TU library, --ast-frontend clang (no castxml on that
host), did not complete within a 3+ minute --budget on a 4-core host — RSS
climbed past 4.9 GiB before an external kill, with no indication --budget
itself was actually bounding the work.
Investigation traced this to a gap in the loop driver, not the deadline
mechanism itself: deadline.check() is already threaded per-TU inside each
extractor (source_extractors/clang.py/castxml.py call it before/after the
AST load, and bounded_timeout() shrinks the subprocess's own timeout to
whatever's left of the scan-wide deadline before the clang/castxml process is
even spawned) — so an already-dispatched TU always self-aborts quickly once
the budget is gone. The gap was that _extract_cache_misses's serial
fallback loop (used when jobs<=1, a single miss unit, or when the
opt-in ProcessPoolExecutor fails to start) and _replay_cache_lookup's
per-unit cache-key loop kept iterating to the next unit without checking
the deadline first — each subsequent unit still self-aborted fast once
dispatched, but only after paying the interpreter/extractor-startup overhead
to get there. Fixed: both loops now call deadline.check() before each
iteration (abicheck/buildsource/source_replay.py), so a hundred-unit miss
list under an already-exhausted budget costs roughly one unit's overhead, not
a hundred — regression-guarded by
tests/test_source_replay.py::test_extract_cache_misses_serial_path_stops_dispatching_once_budget_is_gone
(synthetic, fast — proven to fail without the fix by temporarily reverting
it). The default parallel pool.map() path is unchanged: it still relies
on each already-dispatched unit's own fast self-abort rather than a
stop-enqueuing check, since pool.map submits its whole batch up front — a
"don't submit further work once budget is gone" gate there would need a
bespoke, non-pool.map dispatch loop, a larger change than this fix.
Not resolved by this fix, still an open question: whether the underlying
per-TU cost on this real 62-TU pvxs binary was itself pathological (something
superlinear in this specific run) or is simply what clang-frontend L4 replay
genuinely costs per TU on a template-heavy real C++ codebase without castxml
— the first report's equivalent pass completed in 129s, but that ran with
castxml, not available on the pvxs scan's host. Distinguishing those two
needs a dedicated profiling pass on a real or skills-src/evaluation/field/scan_level_scaling.py-
synthesized multi-TU tree with a --budget sweep added to that harness
(mirroring how the L2 pathological-header investigation above was profiled),
not assumed from a single real-world data point. Until profiled, the safe
recommendation for a clang-only (no castxml) CI runner on a library this
size remains: skip scan --depth source in favor of compare for the L1/L2
release gate, or scope it with --since/--changed-path to just the
changed files rather than the whole library.
Weekly optimization report¶
python scripts/perf_report.py [--corpus N] [-o FILE] ranks hot functions, costed same-argument repeats (audit_repeated_calls.py --by-cost), calls per declaration (per_decl_x10:* budgets) and anti-pattern sites. performance.yml runs it weekly into the step summary. Record what you decide about a candidate in the findings registry.