Benchmark & Tool Comparison¶
This document explains how each ABI checking tool works, what it measured on the
examples/ catalog, and why the numbers come out the way they do.
Note: abicheck's exact, up-to-date change-kind count is tracked in the Change Kind Reference. The
examples/catalog currently has 197 cases (catalog/ground_truth.jsonis the source of truth — seeexamples/README.md). Two benchmarks run against it:
- A pinned 74-case cross-tool subset (
case01-case73+case26b), frozen so accuracy numbers stay reproducible release to release. See Pinned vendor benchmark summary (marked historical, superseded by the full-catalog benchmark below).- A full-catalog sweep scoring every case, with SKIP/ERROR/TIMEOUT counted as misses. See Full-catalog benchmark.
Which denominator is which. Of the 197 catalog cases, 159 are compilable
v1/v2shared-library (.so) pairs that abidiff/ABICC can also run against — abicheck's own competitor benchmark builds and scores these through the normal build → dump → compare pipeline. The remaining 38 don't fit that shape (10 single-artifact audit/cross-source checks, 15 build-source-pack (L3-L5) replays, 6 committed snapshot-pair fixtures, 5 multi-library bundle directories, 1 kernel-BTF blob, 1 Python stub-pair) and have no abidiff/ABICC equivalent, so they're scored by abicheck alone through dedicated test lanes instead of the tool-vs-tool tables. This split is derived directly from each case'smode/bundle/fixtures/skipfields inground_truth.json(the same fieldsscripts/benchmark_comparison.py's_try_special_case()routes on), so it stays accurate as the catalog grows — recompute it with:python3 -c " import json v = json.load(open('catalog/ground_truth.json'))['verdicts'] special = sum(1 for e in v.values() if e.get('mode') == 'audit' or e.get('skip') or e.get('bundle') is True or e.get('category') == 'bundle' or e.get('mode') in ('snapshot-pair', 'reconcile') or e.get('fixtures') == ['old.json', 'new.json'] or e.get('stub_pair')) print(f'{len(v)} total, {len(v) - special} .so-pair, {special} dedicated-lane')"Why the tools disagree. The accuracy gaps below are mostly an evidence story: each tool sees a different subset of the binary/debug/header inputs. For the conceptual model — which evidence detects which change class — see Evidence & Detectability.
Current scan-quality snapshot¶
Examples Validation is the workflow for the runnable compare-mode catalog. It
validates abicheck's current compare-mode coverage separately from the pinned vendor
benchmark below: the catalog lanes answer "what does abicheck currently cover?",
while the pinned 74-case subset answers "how does abicheck compare to
ABICC/libabigail on a stable cross-tool corpus?" Volatile catalog lane counts
are maintained only in the canonical Examples Validation status,
generated from the latest CI artifacts; this table records methodology and
interpretation rather than a second manual snapshot.
| Scan | Scope | Execution | Result | Quality signal |
|---|---|---|---|---|
| Catalog metadata | 197 ground-truth entries | catalog/ground_truth.json + tests/test_evidence_tiers.py |
159 binary competitor .so lanes + 38 dedicated non-.so lanes |
Single source of truth for examples, verdicts, expected kinds, and minimum evidence; split recomputed directly from ground_truth.json's mode/bundle/fixtures/skip fields (see the "Which denominator is which" note above) |
| Build/autodiscovery | catalog integration suite | python -m pytest tests/test_example_autodiscovery.py -v --tb=short -m integration |
Current CI result | Green default single-library build lane; skipped items are covered by dedicated bundle/source/audit/BTF tests |
| Full example proof matrix | catalog cases | skills-src/evaluation/validation/scripts/collect_full_example_matrix.py over CI artifacts + bundle/G20/L3-L5/BTF proofs |
Current CI result | Full-catalog source of truth; a SKIP in one lane is accepted only when a dedicated lane proves the case |
| Default/debug verdicts | catalog cases | PYTHONPATH=. python tests/validate_examples.py --toolchain {gcc,clang} --json |
Current CI result | Single-library debug lane; dedicated non-.so cases skip here by design; XFAIL is not green full-matrix scope |
| Bundle release verdicts | 5 bundle cases | PYTHONPATH=. python skills-src/evaluation/validation/scripts/run_bundle_examples.py --json |
5 PASS | Runs the multi-library bundle examples through abicheck compare old/ new/ |
| Runtime smoke | catalog cases | PYTHONPATH=. python skills-src/evaluation/validation/scripts/run_example_runtime_smoke.py --json |
Current CI result | No BUILD_ERROR; runtime signal is evidence, not policy verdict |
| Release headers | catalog cases | validate_examples.py --artifact-variant release-headers --json in CI artifact |
Current CI result | Reduced-evidence informational lane; false-positive guard passed |
| Stripped headers | catalog cases | validate_examples.py --artifact-variant stripped-headers --json in CI artifact |
Current CI result | Reduced-evidence informational lane; known signal-loss rows remain visible there |
| Build/source proof | fixed 10-case proof set | validate_examples.py case01 case04 case98 case105 case122 case129 case130 case131 case132 case133 --artifact-variant build-source --json in CI artifact |
Current CI result | Blocking fixed-set proof: every expected result must be present and PASS |
| Binary competitor scan | 159 shared-library pairs × 2 external tools (4 tool/mode combinations) | abicc (dumper + xml) and libabigail abidiff (+headers) over built .so pairs |
636 tool invocations attempted; per-tool correct/accuracy in the full-catalog benchmark below | Competitor .so lane only; the 38 dedicated non-.so cases are represented in their own lanes, not as missing .so results |
| Scan-depth matrix | not independently re-run this pass | abicheck compare OLD NEW --depth {binary,headers,build,source}, once per depth per target |
see prior methodology note below | Compare-style status by depth; full-catalog audit/cross-source/bundle/BTF/snapshot cases are covered by dedicated lanes |
The volatile catalog-lane status is intentionally maintained in the canonical
Examples Validation block linked above; do not copy its counts here. The
scan-depth matrix specifically needs a fresh run of abicheck compare --depth
across the current comparable-target set (it was previously pinned to 141
targets against an older, smaller catalog) — that regeneration is a tracked
follow-up, not fabricated here.
case97_api_depends_on_consumer_env and case105_concept_tightening are
resolved: the former is proven by its own source_smoke oracle at the default
compiler lanes, the latter by the build/source (L4) lane. The one case not
proven by a direct detector/CLI match is
case111_enumerable_thread_specific_lambda_ambiguity: every evidence tier
(L0-L5) currently reaches COMPATIBLE, a real tracked detector gap (see its
README), so it is credited in the full example matrix via known-gap-oracle
provenance — its own source_smoke proves the canonical API_BREAK — rather
than direct coverage. See
the validation runbook for
the direct-vs-known-gap-oracle accounting.
Current stripped-header signal-loss cases: case103_toolchain_flag_drift,
case117_no_unique_address, case129_struct_return_convention,
case60_base_class_position_changed, and case69_trivial_to_nontrivial.
Release and stripped full-catalog lanes remain reported-only. The fixed ten-case build/source proof is blocking. A complete build/source run over every applicable L3-L5 case remains an extended/manual validation path because it is much heavier than the default/debug full-catalog gate.
How each tool analyses ABI¶
abicheck (compare mode)¶
.so (v1) ──► ELF reader: exported symbols, SONAME, visibility
castxml (Clang AST): types, methods, vtable, noexcept
DWARF reader: size cross-check
──► snapshot (JSON)
├──► checker engine ──► verdict
.so (v2) ──► (same) ──► snapshot (JSON) ┘
Analysis basis: ELF symbol table + Clang AST via castxml + DWARF. Header requirement: Yes — headers are passed to castxml for full type analysis. Compiler requirement: None — castxml runs separately as a standalone tool.
This gives abicheck three independent data sources per symbol: ELF (what is exported), AST (what the C++ type contract says), and DWARF (actual compiled layout for cross-check).
Verdict vocabulary comparison¶
| Verdict | abicheck compare | abidiff | ABICC |
|---|---|---|---|
NO_CHANGE |
✅ | ✅ (exit 0) | ⚠️ reports 100% compat |
COMPATIBLE |
✅ | ✅ (exit 4) | ⚠️ reports 100% compat |
API_BREAK |
✅ | ❌ | ❌ |
BREAKING |
✅ | ✅ (exit 8+) | ✅ |
API_BREAK = source-level break, binary-compatible. Example: parameter renamed,
access level changed, pure API contract violation with no ABI binary change.
Only abicheck compare can emit this verdict.
Why abicheck leads the matrix¶
abicheck uses three independent analysis passes per comparison:
- ELF pass — symbol table diff: detects visibility changes, SONAME, symbol binding, symbol version policy, added/removed/renamed exported symbols
- castxml pass — Clang AST diff: detects noexcept, static qualifier, const qualifier, method-became-static, pure virtual additions, access level, parameter/return type changes that are invisible in ELF/DWARF
- DWARF cross-check — validates actual compiled type sizes, struct/class member offsets,
vtable slot offsets, base class offsets, and
#pragma pack/-march-sensitive alignment that header analysis alone may compute incorrectly
Neither abidiff nor ABICC runs all three passes. abidiff has no AST (misses noexcept, static, const). ABICC has no ELF pass (misses SONAME, visibility). ABICC(dump) has no AST (same gaps as abidiff plus instability on complex C++).
Benchmarking by evidence tier¶
The cross-tool matrix above answers "how does abicheck compare to other tools when each is given its best input?" A second, orthogonal benchmark answers "how much of the catalog can be discovered from each source of information?" — i.e. how detection grows as you feed abicheck more of the five sources.
This is tracked in two layers: catalog/ground_truth.json records the minimum
evidence layer for each case, while a dedicated benchmark mode empirically scans
the runnable cases at progressively richer artifact layers:
python3 scripts/benchmark_comparison.py --evidence-tiers
# restrict to specific cases/suite as usual:
python3 scripts/benchmark_comparison.py --evidence-tiers --cases case01 case07 case34
This is the slow path: it builds each case once and then runs the full
dump+comparepipeline up to four times per case (L0-L3), so scope it with--cases/--suitefor quick iteration.
For each case it builds the libraries once, then runs the full dump+compare
pipeline four times:
| Tier | abicheck input | --dry-run mode |
Active detectors |
|---|---|---|---|
| L0 binary only | stripped .so, no -H |
Symbols-only | ≈ 6 / 30 |
| L1 + debug info | -g .so, no -H |
DWARF-only | ≈ 24 / 30 |
| L2 + public headers | -g .so, -H include/ |
Full (AST + DWARF) | 30 / 30 |
| L3 + build context | L2 plus -p build/ (when a compile DB exists) |
Full + build evidence | 30 / 30 + L3 |
The
/30denominator above is a point-in-time snapshot from an earlier run and has not been refreshed since (the registered-detector count is now 56, perdetector_registry.registry— seeabicheck/detector_registry.py).--dry-runalso no longer reports a detector-enabled fraction at all (it now lists whichLxlayers are present, with basic per-layer stats). Re-runpython3 scripts/benchmark_comparison.py --evidence-tiers(needscastxml+gcc/g++) for current per-tier numbers rather than trusting this table.L4 (source ABI replay) uses the build/source pack produced by
collect. The tiered benchmark runner does not exercise that mode yet, so the empirical L0-L3 run still reports L4-only cases as not reached until source-pack support is added. The table below includes the L4 minimum fromground_truth.json.
Which source discovers what¶
Each case in catalog/ground_truth.json
carries a min_evidence field — the weakest source at which abicheck reaches
every one of the case's cataloged expected_kinds, not just its verdict —
derived by
scripts/evidence_tiers.py
(compute_min_evidence() takes the strongest tier across all expected_kinds,
by design: "the whole break is only fully visible once every contributing
kind is") and validated by tests/test_evidence_tiers.py. Aggregating
min_evidence over the catalog's 186 compare-style cases (everything
except the 11 single-artifact audit/cross-source/BTF checks, which have no
old-vs-new concept to place on an evidence staircase) yields the cumulative
minimum-evidence coverage below. One of those 186,
case111,
has no min_evidence at all — it is the one documented detector gap where no
tier currently reaches the canonical verdict — so it's excluded from the
185-case denominator rather than miscounted against a tier. Recompute this
table directly from ground_truth.json any time with:
python3 -c "
import json
from collections import Counter
v = json.load(open('catalog/ground_truth.json'))['verdicts']
cs = {k: e for k, e in v.items() if not (e.get('mode') == 'audit' or e.get('skip'))}
counts = Counter(e.get('min_evidence') for e in cs.values() if e.get('min_evidence') not in (None, 'none'))
total = sum(counts.values())
cum = 0
for tier in ['L0', 'L1', 'L2', 'L3', 'L4', 'L5']:
cum += counts[tier]
print(f'{tier}: +{counts[tier]:<3} cumulative {cum}/{total} ({cum/total:.0%})')
"
| Source provided | Layer | Cases first detectable here | Cumulative | Representative cases |
|---|---|---|---|---|
| Just the binary | L0 | 64 | 64 / 185 (35%) | symbol removal (01), SONAME (05), visibility (06), symbol-version removed (65), all 5 bundle cases |
| + Debug symbols | L1 | 69 | 133 / 185 (72%) | struct layout (07), enum value (08), vtable (09), calling convention (64), bitfield (63), toolchain flag drift (103), templated-base detail:: leak (77) |
| + Public headers | L2 | 24 | 157 / 185 (85%) | access level (34), default arg removed (123), class final (125), detail:: leaks (74–76), scoped-internal no-change (118–120) |
| + Build data | L3 | 10 | 167 / 185 (90%) | build-mode flips: exceptions (130), RTTI (131), thread-safe statics (132), TLS model (133), enum size (152), struct packing (153), LTO (154), char signedness (155), C++ standard floor (98) |
| + Sources | L4 | 5 | 172 / 185 (93%) | uninstantiated template (122), public macro removed (156), inline function removed (157), concept tightening (105), public typedef removed (158) |
| + Source graph | L5 | 13 | 185 / 185 (100%) | public API internal dependency (160), target dependency added (161), exported symbol source owner changed (162), private-field/base/parameter-type leaks (187–189, 191), call-graph reachability through suppression (192), reconciled internal-declaration rename/ambiguous-rename/move/identity-reconciliation (194–197) |
Why L3 now matters. Earlier snapshots had no standalone L3-only catalog cases. The current compare-mode catalog includes build-mode flips whose relevant facts come from build context when artifact metadata is insufficient: exceptions, RTTI, thread-safe statics, TLS model, enum size, struct packing, LTO, and char signedness policy.
Why L5 is listed. L5 is a derived source graph, not a sixth input. It is included here because
ground_truth.jsonuses it as the minimum evidence for source-to-symbol reachability cases.Crediting rule. A tier only counts as discovering a case when it emits the cataloged change kind with the right verdict, not merely a matching verdict — otherwise a weak tier that returns a bare
COMPATIBLE/NO_CHANGE(the "found nothing" defaults) would be miscredited. ActiveBREAKING/API_BREAKverdicts are genuine findings, so a verdict match suffices there (and avoids penalising tier-appropriate variant kinds such as L0'sfunc_removed_elf_only).
L5's "first detectable" column is a kind-set floor, not a verdict floor, for 4 of its 13 cases.case187/188/191land inL5here even though theirBREAKINGverdict is empirically reachable atL1, andcase189atL0(verified with--evidence-tiers --cases case187 case188 case189 case191) — each already fires from a real, artifact-level structural break (a field/base/parameter type change). The oneL5kind in theirexpected_kinds,public_api_internal_dependency_added, is correlated context on that already-detected break — naming which internal type the new dependency reaches — not what makes the verdict fire. They're credited toL5here purely because the crediting rule above requires every cataloged kind, not because the source graph is required to catch the break.Not the same number as the full-catalog benchmark below. This staircase is a discoverability floor (the weakest source that reaches the correct verdict per case, credited from
ground_truth.jsonlabels); it does not penalize a tier for over-calling elsewhere in the catalog. The full-catalog benchmark below is the stricter, empirically-measured number — it scores all 193 cases including false positives, which is whyL3-L5reads 99.5% there rather than the 100% this table'sL5row shows (the full-catalog run also treatsSKIPon the 34 dedicated-lane cases as no-signal until their own dedicated lane proves them, whereas this staircase credits them by their catalogedmin_evidencelabel directly).
Two directions matter, not just one:
- Discovery. Most layout and source-only breaks are simply invisible
without the right source — a struct-field insertion is
NO_CHANGEat L0 andBREAKINGonly once L1 debug info is present. - False-positive suppression. More evidence also removes spurious breaks:
the scoped-internal cases (118–120)
change an internal struct that looks like a layout break at L1, and only L2
header scoping lets abicheck correctly return
NO_CHANGE.
Caveat. The L2/L3 columns require
castxml(and, for L3, acompile_commands.json) to be present in the benchmark environment; where a source is unavailable the runner records the tier asn/a/ERRORrather than a miss, so read the tiered numbers together with the evidence-coverage report for the run.
Full-catalog benchmark (2026-07-18, all 193 cases)¶
Every catalog case scored, with SKIP/ERROR/TIMEOUT/incapacity all counted as misses — a tool that hung, crashed, or simply has no mode for a case shape scores exactly like a wrong verdict. This is a stricter (and more honest) denominator than "accuracy over cases the tool managed to complete," so read it as the answer to "if I pointed this tool at the whole catalog blind, how often would it tell me the truth?"
Reproducibility envelope. abicheck
0.5.0, code commitffa860c— the benchmark numbers below were measured against this commit, which is onmainand stable across a squash-merge, unlike a branch-local docs commit.mainhas since moved on pastffa860c(this doc's branch was rebased onto it):56055acsplitabicheck/service.py's output-rendering helpers into aservice_render.pyleaf module (behavior-preserving, fixes an AI-readiness file-size gate) and also fixed real detector bugs — three macOS-only Itanium-mangled-name normalization fixes (no effect on Linux, where this benchmark ran) and one platform-agnostic fix to howpublic_api_internal_dependency_addedfindings get surface-filtered. The numbers below are accurate forffa860cbut have not been re-verified against56055ac; the platform-agnostic fix could in principle change results for cases involving that finding kind — treat a re-run against currentmainas a tracked follow-up, not yet done.ground_truth.jsonsha2567836d8b79f96. All six lanes below (abicheck,abicheck_full,abidiff,abidiff_headers,abicc_dumper,abicc_xml) were regenerated live against the current 193-case catalog on 2026-07-18 — no frozen/carried-over data (ABICC's two modes are each frozen right after their own live run since they can't run concurrently with themselves, then merged into the same live pass that runs the other four tools; see commands below). Tool versions: castxml0.6.3, libabigailabidiff2.4.0,abi-compliance-checker2.3. Wall time 1396s (~23 min) for the live abicheck/abidiff pass; peak RSS 708.5 MiB.
# ABICC lanes are frozen ahead of time (each mode run alone, ABICC hangs
# on some cases when run concurrently with itself):
python3 scripts/benchmark_comparison.py --tools abicc_dumper --freeze abicc_dumper
python3 scripts/benchmark_comparison.py --tools abicc_xml --freeze abicc_xml
# abicheck/abicheck_full/abidiff/abidiff_headers run live; the frozen
# abicc_dumper/abicc_xml columns above merge in automatically:
python3 scripts/generate_benchmark_report.py \
--tools abicheck abicheck_full abidiff abidiff_headers --check
| Tool | Correct / 193 | Accuracy | False positives | False negatives | Total time |
|---|---|---|---|---|---|
| abicheck (L2, headers) | 185 | 95.9% | 0 | 8 | 199s (~3 min) |
| abicheck (L3-L5, +sources) | 192 | 99.5% | 0 | 1 | 916s (~15 min) |
libabigail (abidiff) |
55 | 28.5% | 5 | 133 | 1.2s |
| libabigail + headers | 55 | 28.5% | 5 | 133 | 5.5s |
| ABICC (abi-dumper) | 86 | 44.6% | 8 | 99 | 872s (~15 min) |
| ABICC (xml/legacy) | 78 | 40.4% | 7 | 108 | 1871s (~31 min) |
ABICC is roughly 340-727× slower than libabigail for the identical
193-case catalog — abi-dumper/abidiff is ~727× (872s vs 1.2s), xml/abidiff_headers
is ~340× (1871s vs 5.5s) — while scoring lower on accuracy than abicheck's
L2 lane. This is why ABICC/libabigail results are frozen
(--freeze) into scripts/frozen_competitor_results.json — a committed
reference file merged into every subsequent run automatically — rather than
re-run on every abicheck iteration; nothing in a competitor's own verdict
changes when abicheck itself is patched.
Reading the false-positive/false-negative split: a false positive is a tool over-calling severity (reporting a worse verdict than the true one — crying wolf); a false negative is under-calling it (silence on a real break, including every SKIP/ERROR/TIMEOUT, since a tool that cannot tell you about a break failed to warn just as surely as one that said COMPATIBLE).
- libabigail's misses are overwhelmingly false negatives (133/193, DWARF
has no view into noexcept/static/const/layout-invisible changes) — it
rarely cries wolf (FP=5), it mostly stays silent. 34 of those misses are a
flat
SKIPon the dedicated-lane cases (audit/cross-source, bundle, BTF, snapshot-pair, build-source-pack, stub-pair — see the "Which denominator is which" note up top) that have no ELF pair forabidw/abidiffto read at all. - ABICC's misses skew false-negative too (99-108/193) for the same
reason plus its own timeout/error behavior: the same 34 non-
.socasesSKIPoutright, and a further 5 (abi-dumper, plus 1ERRORoncase16_inline_to_non_inline) / 16 (xml) hit the 90s per-case timeout in this environment —case09,case81,case104,case105,case109,case114,case129-case133are among the routine offenders on the xml mode here. - abicheck's false positives are 0 on both lanes. The L3-L5 lane's raw
string mismatches include 6 cases (
case16,case47,case54,case62,case99,case185) where the harness correctly credits aCOMPATIBLE→COMPATIBLE_WITH_RISKpromotion as evidence enrichment rather than a miss (by the authority rule: the source-replay lane sees a real risk signal — a reserved-field reuse, a stale-inlined-body risk, symbol-binding/ownership drift — that a binary/header-only lane structurally cannot see). The one genuine remaining miss on both lanes iscase111_enumerable_thread_specific_lambda_ambiguity(API_BREAKexpected, every evidence tier from L0 through L5 currently reachesCOMPATIBLE— a documented detector gap, see its README, not a harness artifact). - abicheck L2's other 7 misses (
case98,case105,case122,case130-case133) are structurally below the L2 lane's evidence floor perground_truth.json'smin_evidence— build-mode flips and concept/source-replay facts an L2 (headers, no-p build/) lane cannot see by design, not by gap. The L3-L5 lane resolves all seven.
Methodology history. Earlier passes of this benchmark scored substantially lower for the L3-L5 lane (as low as 61%) due to a mix of benchmark-harness bugs — a forced
-includecrashing legal type redefinitions, a build-source-pack helper silently bypassed after a CLI refactor, inconsistent per-case source-file naming defeating rename detection — and a couple of real product fixes (a C-function inline-removal false positive, afield_renamedclassification gap). Each was root-caused and fixed individually; see git/PR history around commit1d2487c82ec5for the full account rather than a narrated blow-by-blow here. The only genuine, currently unresolved gap across both abicheck lanes iscase111_enumerable_thread_specific_lambda_ambiguity.
Rule-family accuracy (160 demonstrated rule families)¶
Every table above scores accuracy per raw case — three demonstrations of
the same rule (a canonical case plus its confirmed duplicate/variant
siblings, see scripts/catalog_rule_registry.py)
count as three ABI concepts, so a rule with more sibling fixtures than
another contributes more to (or costs more from) a tool's apparent accuracy
for a reason that has nothing to do with how many rules it actually
understands. docs/contribute/plans/examples-catalog-split.md's "What is
left" item 3 tracked this as a real, separate change: adding a rule-family
dimension changes what is measured, not just how the existing number is
labelled.
This table answers a different question: for how many of the catalog's
160 demonstrated compatibility rules does a tool get every member case
right — the canonical demonstration and every confirmed duplicate/variant
sibling? One miss anywhere in a family makes the whole family a miss,
mirroring how a maintainer actually judges "does this tool understand this
rule" rather than "how many near-identical fixtures did it happen to get
right." benchmark_comparison._rule_family_accuracy() computes it by
joining each run's per-case results against catalog_rule_registry.
build_families(); scripts/generate_benchmark_report.py renders it as the
table below and drift-checks it against this section the same way it
drift-checks the flat table above (parse_rule_family_table/
diff_rule_family_against_doc).
Only catalog_rule_registry.STATUS_DEMONSTRATED families are scored — a
"referenced-only" family (named only by a scenario's related_rules, a
mechanism no single-library case demonstrates alone yet) has no rule-entity
case of its own to attribute a verdict to, so it is out of scope for this
table by construction, not by gap; see
docs/contribute/catalog-coverage.md
for that count.
Reproducibility envelope. Measured 2026-09-06 against commit
75aa966f9a08,catalog/ground_truth.jsonsha25656ece287a3d8. Onlyabicheck (L2, headers)was live-measured this pass, over the full 197-case catalog: gcc/g++ 13.3.0 + castxml 0.7.0.libabigail/ABICCare omitted rather than shown at their last frozen values — this catalog has grown sincescripts/frozen_competitor_results.jsonwas last refreshed against those tools (itsground_truth_sha256no longer matches, the same staleness the "Cache state & status detail" appendix these reports emit already reports asn/afor exactly this reason), and re-deriving a family-level number from a stale per-case cache would misrepresent it as current.abicheck (L3-L5, +sources)is omitted for a different reason: this pass's environment has no workingcontrib/abicheck-clang-pluginbuild, so every compiled case in that lane reportedERRORrather than a real verdict — publishing that pass's number would record an environment gap as a product regression against the 99.5% per-case accuracy the full-catalog benchmark above already measured for that lane, rather than a genuine family-level result. Re-running all four lanes together (the same live-plus-freshly-frozen pass the full-catalog benchmark above describes) to populate this table's remaining rows is a tracked follow-up, not fabricated here.
| Tool | Correct / 160 families | Family accuracy |
|---|---|---|
| abicheck (L2, headers) | 152 | 95.0% |
Reading the family-vs-case gap. abicheck (L2, headers) misses 9/197
cases individually (95.4% per-case, from this same run) but only 8/160
rule families (95.0% per-family) — the two counts differ by exactly one,
and for a revealing reason rather than because two misses share a family:
case111_enumerable_thread_specific_lambda_ambiguity (the one documented
detector gap named above) is a scenario-entity case, not a rule-entity
one — it names constructor-overload-ambiguity only in its own
related_rules, a referenced-only family with no rule-entity case of its
own — so its miss is outside this dimension's scope entirely, by the same
"only demonstrated families are scored" rule stated above, not because it
was scored and happened to overlap another miss. The remaining eight
misses are each in their own single-case family, so none of this run's
family misses overlap either: seven are the L2-evidence-floor misses named
above (case98, case105, case122, case130-case133), plus
case115_bit_int_width_changed, which reports ERROR rather than
BREAKING in this run's environment for the reason its own module docstring
documents — no GCC 14+ available to build its C23 _BitInt fixture, a
toolchain gap rather than a detector gap. The two totals track each other
closely for this particular tool and run purely because none of its misses
happen to share a family or fall outside the rule-entity scope more than
once. The two numbers diverge more sharply for a tool with uneven family
coverage — e.g. one that gets a rule's canonical case right but a
documented variant (language, public-surface, symbol-versioning) wrong
scores that whole family as a miss even though its per-case tally looks
almost identical.
Pinned vendor benchmark summary (2026-07-18, 74-case subset)¶
Historical. Superseded by the full-catalog benchmark above, which covers all 193 cases with a stricter denominator (SKIP/ERROR/TIMEOUT count as misses) plus an FP/FN breakdown. Kept here because the original 74-case release-pinned methodology stays useful as a small, fast, stable corpus for spot-checking a tool change without paying the full-catalog runtime (this refresh: 86s for
abicheck, vs 199s for the same lane over the full 193-case catalog above). The harness has since dropped the standaloneabicheck_compat/abicheck_stricttool lanes —--toolsonly acceptsabicheck,abicheck_full,abidiff,abidiff_headers,abicc_dumper,abicc_xmlnow, so the original 2026-05-19 run's compat (71/74, 95%) and strict (62/74, 83%) numbers can no longer be reproduced verbatim, and thecompatcommand itself was removed in 0.6.
Release-pinned scan status from python3 scripts/benchmark_comparison.py --suite pinned74 --abicc-mode both
on the original 74-case benchmark subset (same code commit ffa860c and ground_truth.json as the
full-catalog run above — see that section's reproducibility envelope for why the code commit, not a
branch-local docs commit, is the stable reference).
| Tool | Correct / 74 | Accuracy | False positives | False negatives | Total time |
|---|---|---|---|---|---|
| abicheck (L2, headers) | 74 | 100% | 0 | 0 | 86.0s |
| abicheck (L3-L5, +sources) | 74 | 100% | 0 | 0 | 454.1s |
libabigail (abidiff) |
21 | 28.4% | 2 | 51 | 0.5s |
| libabigail + headers | 21 | 28.4% | 2 | 51 | 3.1s |
| ABICC (abi-dumper) | 48 | 64.9% | 2 | 24 | 700s (~12 min) |
| ABICC (xml/legacy) | 47 | 63.5% | 1 | 26 | 240s (~4 min) |
Scan-status matrix¶
abicheck, abicheck_full, abidiff, and abidiff_headers complete all
74/74 cases cleanly. Only the two ABICC lanes leave cases unscored — the
Correct/Accuracy columns above already fold this in, but not which
specific cases: abicc_dumper completes 71/74 (case09_cpp_vtable,
case59_func_became_inline timeout; case16_inline_to_non_inline error);
abicc_xml completes 72/74 (case16_inline_to_non_inline,
case60_base_class_position_changed timeout).
Commands used¶
A single run now covers all six tools; the previous multi-invocation
sequence (separate --skip-abicc and per-mode --abicc-timeout 20 calls)
was a workaround for an older, flakier ABICC integration and is no longer
necessary — --abicc-timeout still exists if you need to bound a hang more
aggressively than the 90s default.
Run the benchmark yourself¶
# Fresh benchmark for the current checkout
python3 scripts/benchmark_comparison.py --abicc-mode both
# Select specific cases or tools
python3 scripts/benchmark_comparison.py --cases case01 case09 case21
python3 scripts/benchmark_comparison.py --tools abicheck abidiff
Choosing the right tool¶
| Scenario | Recommended |
|---|---|
| New CI pipeline, full accuracy | abicheck compare |
| Strict gate (any addition = fail) | abicheck compare --severity-preset strict |
| Debug build available, DWARF check | abicheck compare (castxml already better) |
| Quick ELF-only sanity check | abidiff (fast, 28% (21/74) but catches symbol removals) |