Skip to content

Benchmark & Tool Comparison

This document explains how each ABI checking tool works, what it measured on the examples/ catalog, and why the numbers come out the way they do.

Note: abicheck's exact, up-to-date change-kind count is tracked in the Change Kind Reference. The examples/ catalog currently has 197 cases (catalog/ground_truth.json is the source of truth — see examples/README.md). Two benchmarks run against it:

  • A pinned 74-case cross-tool subset (case01-case73 + case26b), frozen so accuracy numbers stay reproducible release to release. See Pinned vendor benchmark summary (marked historical, superseded by the full-catalog benchmark below).
  • A full-catalog sweep scoring every case, with SKIP/ERROR/TIMEOUT counted as misses. See Full-catalog benchmark.

Which denominator is which. Of the 197 catalog cases, 159 are compilable v1/v2 shared-library (.so) pairs that abidiff/ABICC can also run against — abicheck's own competitor benchmark builds and scores these through the normal build → dump → compare pipeline. The remaining 38 don't fit that shape (10 single-artifact audit/cross-source checks, 15 build-source-pack (L3-L5) replays, 6 committed snapshot-pair fixtures, 5 multi-library bundle directories, 1 kernel-BTF blob, 1 Python stub-pair) and have no abidiff/ABICC equivalent, so they're scored by abicheck alone through dedicated test lanes instead of the tool-vs-tool tables. This split is derived directly from each case's mode/ bundle/fixtures/skip fields in ground_truth.json (the same fields scripts/benchmark_comparison.py's _try_special_case() routes on), so it stays accurate as the catalog grows — recompute it with:

python3 -c "
import json
v = json.load(open('catalog/ground_truth.json'))['verdicts']
special = sum(1 for e in v.values() if e.get('mode') == 'audit' or e.get('skip')
              or e.get('bundle') is True or e.get('category') == 'bundle'
              or e.get('mode') in ('snapshot-pair', 'reconcile')
              or e.get('fixtures') == ['old.json', 'new.json'] or e.get('stub_pair'))
print(f'{len(v)} total, {len(v) - special} .so-pair, {special} dedicated-lane')"

Why the tools disagree. The accuracy gaps below are mostly an evidence story: each tool sees a different subset of the binary/debug/header inputs. For the conceptual model — which evidence detects which change class — see Evidence & Detectability.


Current scan-quality snapshot

Examples Validation is the workflow for the runnable compare-mode catalog. It validates abicheck's current compare-mode coverage separately from the pinned vendor benchmark below: the catalog lanes answer "what does abicheck currently cover?", while the pinned 74-case subset answers "how does abicheck compare to ABICC/libabigail on a stable cross-tool corpus?" Volatile catalog lane counts are maintained only in the canonical Examples Validation status, generated from the latest CI artifacts; this table records methodology and interpretation rather than a second manual snapshot.

Scan Scope Execution Result Quality signal
Catalog metadata 197 ground-truth entries catalog/ground_truth.json + tests/test_evidence_tiers.py 159 binary competitor .so lanes + 38 dedicated non-.so lanes Single source of truth for examples, verdicts, expected kinds, and minimum evidence; split recomputed directly from ground_truth.json's mode/bundle/fixtures/skip fields (see the "Which denominator is which" note above)
Build/autodiscovery catalog integration suite python -m pytest tests/test_example_autodiscovery.py -v --tb=short -m integration Current CI result Green default single-library build lane; skipped items are covered by dedicated bundle/source/audit/BTF tests
Full example proof matrix catalog cases skills-src/evaluation/validation/scripts/collect_full_example_matrix.py over CI artifacts + bundle/G20/L3-L5/BTF proofs Current CI result Full-catalog source of truth; a SKIP in one lane is accepted only when a dedicated lane proves the case
Default/debug verdicts catalog cases PYTHONPATH=. python tests/validate_examples.py --toolchain {gcc,clang} --json Current CI result Single-library debug lane; dedicated non-.so cases skip here by design; XFAIL is not green full-matrix scope
Bundle release verdicts 5 bundle cases PYTHONPATH=. python skills-src/evaluation/validation/scripts/run_bundle_examples.py --json 5 PASS Runs the multi-library bundle examples through abicheck compare old/ new/
Runtime smoke catalog cases PYTHONPATH=. python skills-src/evaluation/validation/scripts/run_example_runtime_smoke.py --json Current CI result No BUILD_ERROR; runtime signal is evidence, not policy verdict
Release headers catalog cases validate_examples.py --artifact-variant release-headers --json in CI artifact Current CI result Reduced-evidence informational lane; false-positive guard passed
Stripped headers catalog cases validate_examples.py --artifact-variant stripped-headers --json in CI artifact Current CI result Reduced-evidence informational lane; known signal-loss rows remain visible there
Build/source proof fixed 10-case proof set validate_examples.py case01 case04 case98 case105 case122 case129 case130 case131 case132 case133 --artifact-variant build-source --json in CI artifact Current CI result Blocking fixed-set proof: every expected result must be present and PASS
Binary competitor scan 159 shared-library pairs × 2 external tools (4 tool/mode combinations) abicc (dumper + xml) and libabigail abidiff (+headers) over built .so pairs 636 tool invocations attempted; per-tool correct/accuracy in the full-catalog benchmark below Competitor .so lane only; the 38 dedicated non-.so cases are represented in their own lanes, not as missing .so results
Scan-depth matrix not independently re-run this pass abicheck compare OLD NEW --depth {binary,headers,build,source}, once per depth per target see prior methodology note below Compare-style status by depth; full-catalog audit/cross-source/bundle/BTF/snapshot cases are covered by dedicated lanes

The volatile catalog-lane status is intentionally maintained in the canonical Examples Validation block linked above; do not copy its counts here. The scan-depth matrix specifically needs a fresh run of abicheck compare --depth across the current comparable-target set (it was previously pinned to 141 targets against an older, smaller catalog) — that regeneration is a tracked follow-up, not fabricated here.

case97_api_depends_on_consumer_env and case105_concept_tightening are resolved: the former is proven by its own source_smoke oracle at the default compiler lanes, the latter by the build/source (L4) lane. The one case not proven by a direct detector/CLI match is case111_enumerable_thread_specific_lambda_ambiguity: every evidence tier (L0-L5) currently reaches COMPATIBLE, a real tracked detector gap (see its README), so it is credited in the full example matrix via known-gap-oracle provenance — its own source_smoke proves the canonical API_BREAK — rather than direct coverage. See the validation runbook for the direct-vs-known-gap-oracle accounting.

Current stripped-header signal-loss cases: case103_toolchain_flag_drift, case117_no_unique_address, case129_struct_return_convention, case60_base_class_position_changed, and case69_trivial_to_nontrivial.

Release and stripped full-catalog lanes remain reported-only. The fixed ten-case build/source proof is blocking. A complete build/source run over every applicable L3-L5 case remains an extended/manual validation path because it is much heavier than the default/debug full-catalog gate.


How each tool analyses ABI

abicheck (compare mode)

.so (v1) ──► ELF reader: exported symbols, SONAME, visibility
             castxml (Clang AST): types, methods, vtable, noexcept
             DWARF reader: size cross-check
          ──► snapshot (JSON)
                              ├──► checker engine ──► verdict
.so (v2) ──► (same) ──► snapshot (JSON) ┘

Analysis basis: ELF symbol table + Clang AST via castxml + DWARF. Header requirement: Yes — headers are passed to castxml for full type analysis. Compiler requirement: None — castxml runs separately as a standalone tool.

This gives abicheck three independent data sources per symbol: ELF (what is exported), AST (what the C++ type contract says), and DWARF (actual compiled layout for cross-check).


Verdict vocabulary comparison

Verdict abicheck compare abidiff ABICC
NO_CHANGE ✅ ✅ (exit 0) ⚠️ reports 100% compat
COMPATIBLE ✅ ✅ (exit 4) ⚠️ reports 100% compat
API_BREAK ✅ ❌ ❌
BREAKING ✅ ✅ (exit 8+) ✅

API_BREAK = source-level break, binary-compatible. Example: parameter renamed, access level changed, pure API contract violation with no ABI binary change. Only abicheck compare can emit this verdict.


Why abicheck leads the matrix

abicheck uses three independent analysis passes per comparison:

  1. ELF pass — symbol table diff: detects visibility changes, SONAME, symbol binding, symbol version policy, added/removed/renamed exported symbols
  2. castxml pass — Clang AST diff: detects noexcept, static qualifier, const qualifier, method-became-static, pure virtual additions, access level, parameter/return type changes that are invisible in ELF/DWARF
  3. DWARF cross-check — validates actual compiled type sizes, struct/class member offsets, vtable slot offsets, base class offsets, and #pragma pack / -march-sensitive alignment that header analysis alone may compute incorrectly

Neither abidiff nor ABICC runs all three passes. abidiff has no AST (misses noexcept, static, const). ABICC has no ELF pass (misses SONAME, visibility). ABICC(dump) has no AST (same gaps as abidiff plus instability on complex C++).


Benchmarking by evidence tier

The cross-tool matrix above answers "how does abicheck compare to other tools when each is given its best input?" A second, orthogonal benchmark answers "how much of the catalog can be discovered from each source of information?" — i.e. how detection grows as you feed abicheck more of the five sources.

This is tracked in two layers: catalog/ground_truth.json records the minimum evidence layer for each case, while a dedicated benchmark mode empirically scans the runnable cases at progressively richer artifact layers:

python3 scripts/benchmark_comparison.py --evidence-tiers
# restrict to specific cases/suite as usual:
python3 scripts/benchmark_comparison.py --evidence-tiers --cases case01 case07 case34

This is the slow path: it builds each case once and then runs the full dump+compare pipeline up to four times per case (L0-L3), so scope it with --cases/--suite for quick iteration.

For each case it builds the libraries once, then runs the full dump+compare pipeline four times:

Tier abicheck input --dry-run mode Active detectors
L0 binary only stripped .so, no -H Symbols-only ≈ 6 / 30
L1 + debug info -g .so, no -H DWARF-only ≈ 24 / 30
L2 + public headers -g .so, -H include/ Full (AST + DWARF) 30 / 30
L3 + build context L2 plus -p build/ (when a compile DB exists) Full + build evidence 30 / 30 + L3

The /30 denominator above is a point-in-time snapshot from an earlier run and has not been refreshed since (the registered-detector count is now 56, per detector_registry.registry — see abicheck/detector_registry.py). --dry-run also no longer reports a detector-enabled fraction at all (it now lists which Lx layers are present, with basic per-layer stats). Re-run python3 scripts/benchmark_comparison.py --evidence-tiers (needs castxml + gcc/g++) for current per-tier numbers rather than trusting this table.

L4 (source ABI replay) uses the build/source pack produced by collect. The tiered benchmark runner does not exercise that mode yet, so the empirical L0-L3 run still reports L4-only cases as not reached until source-pack support is added. The table below includes the L4 minimum from ground_truth.json.

Which source discovers what

Each case in catalog/ground_truth.json carries a min_evidence field — the weakest source at which abicheck reaches every one of the case's cataloged expected_kinds, not just its verdict — derived by scripts/evidence_tiers.py (compute_min_evidence() takes the strongest tier across all expected_kinds, by design: "the whole break is only fully visible once every contributing kind is") and validated by tests/test_evidence_tiers.py. Aggregating min_evidence over the catalog's 186 compare-style cases (everything except the 11 single-artifact audit/cross-source/BTF checks, which have no old-vs-new concept to place on an evidence staircase) yields the cumulative minimum-evidence coverage below. One of those 186, case111, has no min_evidence at all — it is the one documented detector gap where no tier currently reaches the canonical verdict — so it's excluded from the 185-case denominator rather than miscounted against a tier. Recompute this table directly from ground_truth.json any time with:

python3 -c "
import json
from collections import Counter
v = json.load(open('catalog/ground_truth.json'))['verdicts']
cs = {k: e for k, e in v.items() if not (e.get('mode') == 'audit' or e.get('skip'))}
counts = Counter(e.get('min_evidence') for e in cs.values() if e.get('min_evidence') not in (None, 'none'))
total = sum(counts.values())
cum = 0
for tier in ['L0', 'L1', 'L2', 'L3', 'L4', 'L5']:
    cum += counts[tier]
    print(f'{tier}: +{counts[tier]:<3} cumulative {cum}/{total} ({cum/total:.0%})')
"
Source provided Layer Cases first detectable here Cumulative Representative cases
Just the binary L0 64 64 / 185 (35%) symbol removal (01), SONAME (05), visibility (06), symbol-version removed (65), all 5 bundle cases
+ Debug symbols L1 69 133 / 185 (72%) struct layout (07), enum value (08), vtable (09), calling convention (64), bitfield (63), toolchain flag drift (103), templated-base detail:: leak (77)
+ Public headers L2 24 157 / 185 (85%) access level (34), default arg removed (123), class final (125), detail:: leaks (74–76), scoped-internal no-change (118–120)
+ Build data L3 10 167 / 185 (90%) build-mode flips: exceptions (130), RTTI (131), thread-safe statics (132), TLS model (133), enum size (152), struct packing (153), LTO (154), char signedness (155), C++ standard floor (98)
+ Sources L4 5 172 / 185 (93%) uninstantiated template (122), public macro removed (156), inline function removed (157), concept tightening (105), public typedef removed (158)
+ Source graph L5 13 185 / 185 (100%) public API internal dependency (160), target dependency added (161), exported symbol source owner changed (162), private-field/base/parameter-type leaks (187–189, 191), call-graph reachability through suppression (192), reconciled internal-declaration rename/ambiguous-rename/move/identity-reconciliation (194–197)

Why L3 now matters. Earlier snapshots had no standalone L3-only catalog cases. The current compare-mode catalog includes build-mode flips whose relevant facts come from build context when artifact metadata is insufficient: exceptions, RTTI, thread-safe statics, TLS model, enum size, struct packing, LTO, and char signedness policy.

Why L5 is listed. L5 is a derived source graph, not a sixth input. It is included here because ground_truth.json uses it as the minimum evidence for source-to-symbol reachability cases.

Crediting rule. A tier only counts as discovering a case when it emits the cataloged change kind with the right verdict, not merely a matching verdict — otherwise a weak tier that returns a bare COMPATIBLE/NO_CHANGE (the "found nothing" defaults) would be miscredited. Active BREAKING/API_BREAK verdicts are genuine findings, so a verdict match suffices there (and avoids penalising tier-appropriate variant kinds such as L0's func_removed_elf_only).

L5's "first detectable" column is a kind-set floor, not a verdict floor, for 4 of its 13 cases. case187/188/191 land in L5 here even though their BREAKING verdict is empirically reachable at L1, and case189 at L0 (verified with --evidence-tiers --cases case187 case188 case189 case191) — each already fires from a real, artifact-level structural break (a field/base/parameter type change). The one L5 kind in their expected_kinds, public_api_internal_dependency_added, is correlated context on that already-detected break — naming which internal type the new dependency reaches — not what makes the verdict fire. They're credited to L5 here purely because the crediting rule above requires every cataloged kind, not because the source graph is required to catch the break.

Not the same number as the full-catalog benchmark below. This staircase is a discoverability floor (the weakest source that reaches the correct verdict per case, credited from ground_truth.json labels); it does not penalize a tier for over-calling elsewhere in the catalog. The full-catalog benchmark below is the stricter, empirically-measured number — it scores all 193 cases including false positives, which is why L3-L5 reads 99.5% there rather than the 100% this table's L5 row shows (the full-catalog run also treats SKIP on the 34 dedicated-lane cases as no-signal until their own dedicated lane proves them, whereas this staircase credits them by their cataloged min_evidence label directly).

Two directions matter, not just one:

  • Discovery. Most layout and source-only breaks are simply invisible without the right source — a struct-field insertion is NO_CHANGE at L0 and BREAKING only once L1 debug info is present.
  • False-positive suppression. More evidence also removes spurious breaks: the scoped-internal cases (118–120) change an internal struct that looks like a layout break at L1, and only L2 header scoping lets abicheck correctly return NO_CHANGE.

Caveat. The L2/L3 columns require castxml (and, for L3, a compile_commands.json) to be present in the benchmark environment; where a source is unavailable the runner records the tier as n/a/ERROR rather than a miss, so read the tiered numbers together with the evidence-coverage report for the run.


Full-catalog benchmark (2026-07-18, all 193 cases)

Every catalog case scored, with SKIP/ERROR/TIMEOUT/incapacity all counted as misses — a tool that hung, crashed, or simply has no mode for a case shape scores exactly like a wrong verdict. This is a stricter (and more honest) denominator than "accuracy over cases the tool managed to complete," so read it as the answer to "if I pointed this tool at the whole catalog blind, how often would it tell me the truth?"

Reproducibility envelope. abicheck 0.5.0, code commit ffa860c — the benchmark numbers below were measured against this commit, which is on main and stable across a squash-merge, unlike a branch-local docs commit. main has since moved on past ffa860c (this doc's branch was rebased onto it): 56055ac split abicheck/service.py's output-rendering helpers into a service_render.py leaf module (behavior-preserving, fixes an AI-readiness file-size gate) and also fixed real detector bugs — three macOS-only Itanium-mangled-name normalization fixes (no effect on Linux, where this benchmark ran) and one platform-agnostic fix to how public_api_internal_dependency_added findings get surface-filtered. The numbers below are accurate for ffa860c but have not been re-verified against 56055ac; the platform-agnostic fix could in principle change results for cases involving that finding kind — treat a re-run against current main as a tracked follow-up, not yet done. ground_truth.json sha256 7836d8b79f96. All six lanes below (abicheck, abicheck_full, abidiff, abidiff_headers, abicc_dumper, abicc_xml) were regenerated live against the current 193-case catalog on 2026-07-18 — no frozen/carried-over data (ABICC's two modes are each frozen right after their own live run since they can't run concurrently with themselves, then merged into the same live pass that runs the other four tools; see commands below). Tool versions: castxml 0.6.3, libabigail abidiff 2.4.0, abi-compliance-checker 2.3. Wall time 1396s (~23 min) for the live abicheck/abidiff pass; peak RSS 708.5 MiB.

# ABICC lanes are frozen ahead of time (each mode run alone, ABICC hangs
# on some cases when run concurrently with itself):
python3 scripts/benchmark_comparison.py --tools abicc_dumper --freeze abicc_dumper
python3 scripts/benchmark_comparison.py --tools abicc_xml --freeze abicc_xml

# abicheck/abicheck_full/abidiff/abidiff_headers run live; the frozen
# abicc_dumper/abicc_xml columns above merge in automatically:
python3 scripts/generate_benchmark_report.py \
  --tools abicheck abicheck_full abidiff abidiff_headers --check
Tool Correct / 193 Accuracy False positives False negatives Total time
abicheck (L2, headers) 185 95.9% 0 8 199s (~3 min)
abicheck (L3-L5, +sources) 192 99.5% 0 1 916s (~15 min)
libabigail (abidiff) 55 28.5% 5 133 1.2s
libabigail + headers 55 28.5% 5 133 5.5s
ABICC (abi-dumper) 86 44.6% 8 99 872s (~15 min)
ABICC (xml/legacy) 78 40.4% 7 108 1871s (~31 min)

ABICC is roughly 340-727× slower than libabigail for the identical 193-case catalog — abi-dumper/abidiff is ~727× (872s vs 1.2s), xml/abidiff_headers is ~340× (1871s vs 5.5s) — while scoring lower on accuracy than abicheck's L2 lane. This is why ABICC/libabigail results are frozen (--freeze) into scripts/frozen_competitor_results.json — a committed reference file merged into every subsequent run automatically — rather than re-run on every abicheck iteration; nothing in a competitor's own verdict changes when abicheck itself is patched.

Reading the false-positive/false-negative split: a false positive is a tool over-calling severity (reporting a worse verdict than the true one — crying wolf); a false negative is under-calling it (silence on a real break, including every SKIP/ERROR/TIMEOUT, since a tool that cannot tell you about a break failed to warn just as surely as one that said COMPATIBLE).

  • libabigail's misses are overwhelmingly false negatives (133/193, DWARF has no view into noexcept/static/const/layout-invisible changes) — it rarely cries wolf (FP=5), it mostly stays silent. 34 of those misses are a flat SKIP on the dedicated-lane cases (audit/cross-source, bundle, BTF, snapshot-pair, build-source-pack, stub-pair — see the "Which denominator is which" note up top) that have no ELF pair for abidw/abidiff to read at all.
  • ABICC's misses skew false-negative too (99-108/193) for the same reason plus its own timeout/error behavior: the same 34 non-.so cases SKIP outright, and a further 5 (abi-dumper, plus 1 ERROR on case16_inline_to_non_inline) / 16 (xml) hit the 90s per-case timeout in this environment — case09, case81, case104, case105, case109, case114, case129-case133 are among the routine offenders on the xml mode here.
  • abicheck's false positives are 0 on both lanes. The L3-L5 lane's raw string mismatches include 6 cases (case16, case47, case54, case62, case99, case185) where the harness correctly credits a COMPATIBLE → COMPATIBLE_WITH_RISK promotion as evidence enrichment rather than a miss (by the authority rule: the source-replay lane sees a real risk signal — a reserved-field reuse, a stale-inlined-body risk, symbol-binding/ownership drift — that a binary/header-only lane structurally cannot see). The one genuine remaining miss on both lanes is case111_enumerable_thread_specific_lambda_ambiguity (API_BREAK expected, every evidence tier from L0 through L5 currently reaches COMPATIBLE — a documented detector gap, see its README, not a harness artifact).
  • abicheck L2's other 7 misses (case98, case105, case122, case130-case133) are structurally below the L2 lane's evidence floor per ground_truth.json's min_evidence — build-mode flips and concept/source-replay facts an L2 (headers, no -p build/) lane cannot see by design, not by gap. The L3-L5 lane resolves all seven.

Methodology history. Earlier passes of this benchmark scored substantially lower for the L3-L5 lane (as low as 61%) due to a mix of benchmark-harness bugs — a forced -include crashing legal type redefinitions, a build-source-pack helper silently bypassed after a CLI refactor, inconsistent per-case source-file naming defeating rename detection — and a couple of real product fixes (a C-function inline-removal false positive, a field_renamed classification gap). Each was root-caused and fixed individually; see git/PR history around commit 1d2487c82ec5 for the full account rather than a narrated blow-by-blow here. The only genuine, currently unresolved gap across both abicheck lanes is case111_enumerable_thread_specific_lambda_ambiguity.


Rule-family accuracy (160 demonstrated rule families)

Every table above scores accuracy per raw case — three demonstrations of the same rule (a canonical case plus its confirmed duplicate/variant siblings, see scripts/catalog_rule_registry.py) count as three ABI concepts, so a rule with more sibling fixtures than another contributes more to (or costs more from) a tool's apparent accuracy for a reason that has nothing to do with how many rules it actually understands. docs/contribute/plans/examples-catalog-split.md's "What is left" item 3 tracked this as a real, separate change: adding a rule-family dimension changes what is measured, not just how the existing number is labelled.

This table answers a different question: for how many of the catalog's 160 demonstrated compatibility rules does a tool get every member case right — the canonical demonstration and every confirmed duplicate/variant sibling? One miss anywhere in a family makes the whole family a miss, mirroring how a maintainer actually judges "does this tool understand this rule" rather than "how many near-identical fixtures did it happen to get right." benchmark_comparison._rule_family_accuracy() computes it by joining each run's per-case results against catalog_rule_registry. build_families(); scripts/generate_benchmark_report.py renders it as the table below and drift-checks it against this section the same way it drift-checks the flat table above (parse_rule_family_table/ diff_rule_family_against_doc).

Only catalog_rule_registry.STATUS_DEMONSTRATED families are scored — a "referenced-only" family (named only by a scenario's related_rules, a mechanism no single-library case demonstrates alone yet) has no rule-entity case of its own to attribute a verdict to, so it is out of scope for this table by construction, not by gap; see docs/contribute/catalog-coverage.md for that count.

Reproducibility envelope. Measured 2026-09-06 against commit 75aa966f9a08, catalog/ground_truth.json sha256 56ece287a3d8. Only abicheck (L2, headers) was live-measured this pass, over the full 197-case catalog: gcc/g++ 13.3.0 + castxml 0.7.0. libabigail/ABICC are omitted rather than shown at their last frozen values — this catalog has grown since scripts/frozen_competitor_results.json was last refreshed against those tools (its ground_truth_sha256 no longer matches, the same staleness the "Cache state & status detail" appendix these reports emit already reports as n/a for exactly this reason), and re-deriving a family-level number from a stale per-case cache would misrepresent it as current. abicheck (L3-L5, +sources) is omitted for a different reason: this pass's environment has no working contrib/abicheck-clang-plugin build, so every compiled case in that lane reported ERROR rather than a real verdict — publishing that pass's number would record an environment gap as a product regression against the 99.5% per-case accuracy the full-catalog benchmark above already measured for that lane, rather than a genuine family-level result. Re-running all four lanes together (the same live-plus-freshly-frozen pass the full-catalog benchmark above describes) to populate this table's remaining rows is a tracked follow-up, not fabricated here.

Tool Correct / 160 families Family accuracy
abicheck (L2, headers) 152 95.0%

Reading the family-vs-case gap. abicheck (L2, headers) misses 9/197 cases individually (95.4% per-case, from this same run) but only 8/160 rule families (95.0% per-family) — the two counts differ by exactly one, and for a revealing reason rather than because two misses share a family: case111_enumerable_thread_specific_lambda_ambiguity (the one documented detector gap named above) is a scenario-entity case, not a rule-entity one — it names constructor-overload-ambiguity only in its own related_rules, a referenced-only family with no rule-entity case of its own — so its miss is outside this dimension's scope entirely, by the same "only demonstrated families are scored" rule stated above, not because it was scored and happened to overlap another miss. The remaining eight misses are each in their own single-case family, so none of this run's family misses overlap either: seven are the L2-evidence-floor misses named above (case98, case105, case122, case130-case133), plus case115_bit_int_width_changed, which reports ERROR rather than BREAKING in this run's environment for the reason its own module docstring documents — no GCC 14+ available to build its C23 _BitInt fixture, a toolchain gap rather than a detector gap. The two totals track each other closely for this particular tool and run purely because none of its misses happen to share a family or fall outside the rule-entity scope more than once. The two numbers diverge more sharply for a tool with uneven family coverage — e.g. one that gets a rule's canonical case right but a documented variant (language, public-surface, symbol-versioning) wrong scores that whole family as a miss even though its per-case tally looks almost identical.


Pinned vendor benchmark summary (2026-07-18, 74-case subset)

Historical. Superseded by the full-catalog benchmark above, which covers all 193 cases with a stricter denominator (SKIP/ERROR/TIMEOUT count as misses) plus an FP/FN breakdown. Kept here because the original 74-case release-pinned methodology stays useful as a small, fast, stable corpus for spot-checking a tool change without paying the full-catalog runtime (this refresh: 86s for abicheck, vs 199s for the same lane over the full 193-case catalog above). The harness has since dropped the standalone abicheck_compat/abicheck_strict tool lanes — --tools only accepts abicheck, abicheck_full, abidiff, abidiff_headers, abicc_dumper, abicc_xml now, so the original 2026-05-19 run's compat (71/74, 95%) and strict (62/74, 83%) numbers can no longer be reproduced verbatim, and the compat command itself was removed in 0.6.

Release-pinned scan status from python3 scripts/benchmark_comparison.py --suite pinned74 --abicc-mode both on the original 74-case benchmark subset (same code commit ffa860c and ground_truth.json as the full-catalog run above — see that section's reproducibility envelope for why the code commit, not a branch-local docs commit, is the stable reference).

Tool Correct / 74 Accuracy False positives False negatives Total time
abicheck (L2, headers) 74 100% 0 0 86.0s
abicheck (L3-L5, +sources) 74 100% 0 0 454.1s
libabigail (abidiff) 21 28.4% 2 51 0.5s
libabigail + headers 21 28.4% 2 51 3.1s
ABICC (abi-dumper) 48 64.9% 2 24 700s (~12 min)
ABICC (xml/legacy) 47 63.5% 1 26 240s (~4 min)

Scan-status matrix

abicheck, abicheck_full, abidiff, and abidiff_headers complete all 74/74 cases cleanly. Only the two ABICC lanes leave cases unscored — the Correct/Accuracy columns above already fold this in, but not which specific cases: abicc_dumper completes 71/74 (case09_cpp_vtable, case59_func_became_inline timeout; case16_inline_to_non_inline error); abicc_xml completes 72/74 (case16_inline_to_non_inline, case60_base_class_position_changed timeout).

Commands used

python3 scripts/benchmark_comparison.py \
  --suite pinned74 \
  --abicc-mode both

A single run now covers all six tools; the previous multi-invocation sequence (separate --skip-abicc and per-mode --abicc-timeout 20 calls) was a workaround for an older, flakier ABICC integration and is no longer necessary — --abicc-timeout still exists if you need to bound a hang more aggressively than the 90s default.


Run the benchmark yourself

# Fresh benchmark for the current checkout
python3 scripts/benchmark_comparison.py --abicc-mode both
# Skip ABICC (CI-friendly, ~15s total)
python3 scripts/benchmark_comparison.py --skip-abicc
# Select specific cases or tools
python3 scripts/benchmark_comparison.py --cases case01 case09 case21
python3 scripts/benchmark_comparison.py --tools abicheck abidiff

Choosing the right tool

Scenario Recommended
New CI pipeline, full accuracy abicheck compare
Strict gate (any addition = fail) abicheck compare --severity-preset strict
Debug build available, DWARF check abicheck compare (castxml already better)
Quick ELF-only sanity check abidiff (fast, 28% (21/74) but catches symbol removals)