Evidence & Detectability: What Each Method Can and Cannot See¶
One idea drives this whole page: different methods observe different evidence, and no single method detects every compatibility issue.* A tool can only report what its inputs let it see. Feed it symbols only and it sees symbol changes; feed it debug info and it sees layout; feed it headers and it sees source-level API. Some changes (
#definemacros, inline/template bodies, uninstantiated templates) are invisible to any* artifact comparison.
This topic in three pages — you are on Model
Model — this page: the L0–L5 evidence layers, what each can and
cannot see, and the --depth dial that collects them.
Worked example — What Each Level Sees: one
tiny library walked up every level, with the actual data.
Flags — Evidence Depth: the --depth
flag and recipe reference.
This page is the conceptual companion to the practical Limitations and Tool Comparison pages; for the teaching-track version — which break families need which evidence, with worked example cases — see Detecting Breaks (Step 6 of the series). It answers the question users ask most often:
"Why did tool A catch this and tool B didn't?"
Almost always, the answer is evidence: the two tools were looking at different inputs.
0. The five sources of information¶
Canonical evidence-model reference. The
L0–L5table below is the source of truth for the evidence layers. Other pages (getting-started, choose-your-workflow, scan levels, architecture, build & source data) show the same model tailored to their context and link back here — when the model changes, update this table first.
A model independent of any one tool¶
Before the abicheck-specific vocabulary: any compatibility checker — this one or another — draws its evidence from some subset of seven generic categories, each answering a genuinely different question about the library. This split matters because it is what lets you reason about a tool you've never used: ask which of these seven categories it actually consumes, and you already know its blind spots.
| Generic evidence category | Question it answers |
|---|---|
| Artifact evidence | What did the compiler actually emit — symbols, loader metadata? |
| Debug/type evidence | What layout facts did the compiler record about emitted types? |
| Declared-interface evidence | What does the source say the interface is (headers, an AST)? |
| Build evidence | Under what compiler, flags, and environment was this built? |
| Source evidence | What do macros, inline bodies, and templates actually contain? |
| Consumer evidence | What does one real, specific consumer actually use? |
| Runtime evidence | What actually happens when the code runs? |
abicheck's own L0–L5 layer codes are one tool's concrete realization
of five of these — runtime evidence is structurally out of reach for any
static checker, which is exactly the gap
Assurance Beyond Static Checking covers instead:
| Generic category | abicheck layer |
|---|---|
| Artifact evidence | L0 |
| Debug/type evidence | L1 |
| Declared-interface evidence | L2 |
| Build evidence | L3 |
| Source evidence | L4 |
| (derived from L3/L4, not itself provided) | L5 |
| Consumer evidence | (named contract only — see --used-by/--required-symbol) |
| Runtime evidence | (not modeled — see Assurance Beyond Static Checking) |
The rest of this page, and every other page that cites "L0–L5," is
using abicheck's own vocabulary for this generic model — worth keeping
distinct in your head from the model itself, since a different tool
realizes the same seven categories with different names, different
granularity, or gaps in different places.
abicheck's five inputs¶
A release engineer can hand abicheck up to five different
sources of information about a library, ordered from the least to the most.
Each one adds facts the previous cannot see; none of them is complete on its
own. abicheck names them with the layer codes L0–L4. A sixth layer,
L5, is not something you hand over — it is a source/build graph abicheck
derives from L3 (and any L4 surface) to localize and explain findings. So
the full model is six evidence layers, L0–L5 (matching
Build Info & Sources), of which the five L0–L4
are inputs you provide and L5 is derived. This section covers the five you
provide; the derived L5 layer is detailed below and in
Build Info & Sources. You can see which artifact
layers (L0–L2) a given input exposes
with abicheck dump --dry-run (its "Available data layers" section reports
L0–L5 presence/absence without writing a snapshot); the build/source layers
(L3/L4) are not reported there — they surface in the pack-aware compare
layer_coverage table once you supply a build/source pack:
| # | Source you provide | Layer | abicheck input | What it newly reveals | Authority |
|---|---|---|---|---|---|
| 1 | Just the binary | L0 | a stripped .so/.dll/.dylib |
Exported symbols, SONAME/install-name, symbol versions, visibility, binding, DT_NEEDED/LC_LOAD_DYLIB dependencies |
Authoritative |
| 2 | + Debug symbols | L1 | a -g build (DWARF/PDB) or sidecar debug file |
Type layout: struct/class sizes, field offsets, enum values, vtable slots, calling convention, packing/alignment | Authoritative (matched to binary) |
| 3 | + Public headers | L2 | -H include/ (parsed by castxml or clang — --ast-frontend) |
Source-level API: signatures, overloads, access (public/private), final/explicit/noexcept, templates, declared default args, public/internal scoping |
Authoritative for header-visible API |
| 4 | + Build system data & options | L3 | -p build/ (compile DB, CMake/Ninja/Bazel/Make) |
The flags the library was actually built with: -std, _GLIBCXX_USE_CXX11_ABI, -fvisibility, -fabi-version, toolchain/sysroot, target graph, export maps |
Corroborating |
| 5 | + Sources | L4 | a build/source pack (per-TU source ABI replay) | Facts that never reach the binary: macro constants, constexpr values, default-argument values, inline/template bodies, uninstantiated templates |
Corroborating (→ API_BREAK/risk) |
Read this staircase-shaped for the common case: each step up the table usually both finds breaks the step below is blind to and prevents false positives the step below would raise. A struct-field insertion is invisible at L0 but obvious at L1 (case07); an internal-struct change that looks like a break at L1 is correctly dismissed once L2 headers reveal the struct is non-public (case118).
But it is not a strict ranking where every higher layer is more authoritative than every lower one for every question — the table's own rightmost column already says why: L0-L2 are marked Authoritative, L3/L4 only Corroborating. A higher layer adds scope (what's public, what flags applied) that can correctly overrule a lower layer's naive reading — but for the one fact a lower artifact layer directly observed (a symbol is present, a struct is this many bytes), a higher layer's absence of evidence is never grounds to override it. That asymmetry is the authority rule below in full, and it is why the next section's own false-positive/false-negative table is not monotonic on both axes at every step (L1 alone adds false positives L0 was too blind to raise) even though the false-negative axis is. Read "staircase" as more evidence, generally fewer misses — not as every layer strictly dominates the one below it on every question.
What each layer buys: fewer false negatives and fewer false positives¶
Fact owner for current numbers. A per-depth accuracy count (binary /
headers / build / source, each with its own eval-target count, correct-
verdict coverage, false positives, and false negatives) is exactly the kind
of volatile, machine-checkable fact this repo's docs contract says must have
one owner — Tool Comparison's "Current scan-quality
snapshot"
(the "Scan-depth matrix" row) is that owner; see that page for the run
status and target count. Do not re-add a specific eval-target count,
run-freshness claim, or per-depth percentage table to this page — a second
copy is exactly how this page and that one disagreed before.
Qualitatively — and this part doesn't drift, because it follows from what
each layer can and cannot observe, not from a specific run's numbers — the
shape holds regardless of the exact counts: binary (L0 alone) is the cheap
floor with many invisible API/header/source-only breaks; headers (L0+L2)
is the best low-cost gate for two distinct reasons, not one — the declared
header AST itself recovers misses L0/L1 are structurally blind to
(source-only signature/access/noexcept changes), and, separately, the
public/private boundary that same AST supplies removes false positives an
L1-only run would over-call on internal churn; build (+L3) adds
build-context corroboration on top; source (+L4) has the highest recall,
since source-smoke proofs additionally cover consumer-only API hazards no
artifact tier can see. (Whether a given layer introduces zero new false
positives is itself a measured, catalog- and run-specific result — see the
fact owner above for the current number, not a guarantee this qualitative
description makes on its own.) An earlier full rung —
whole-library replay, as opposed to source's changed-TU replay — scored
identically to source on the comparable-target set the matrix was last
measured against, which is why the two were collapsed into one public
source rung; see
Removed scan axes.
Across the full staircase, adding evidence drives both error axes down — it
is not a trade-off where you must choose between missing breaks and crying wolf.
The gain is not perfectly monotonic at every single step (a middle layer can
see a change before it has the context to scope it — L1 below is exactly that),
but each higher layer either recovers a false negative or removes a false
positive the layer beneath could not. abicheck tracks this as a CI
gate (scripts/check_tier_accuracy.py): it runs one labelled change per case at
each evidence level and records, per level, whether the tool under-calls it
(a false negative — the layer is structurally blind to a real break) or
over-calls it (a false positive — the layer sees the change but lacks the
context to tell public from internal):
| Change (ground truth) | L0 | L1 | L2 | L3 |
|---|---|---|---|---|
| public struct grew — breaking | ❌ FN | ✅ | ✅ | ✅ |
| C function parameter widened — breaking | ❌ FN | ✅ | ✅ | ✅ |
| public enum value changed — breaking | ❌ FN | ✅ | ✅ | ✅ |
| internal struct grew — non-breaking | ✅ | ❌ FP | ✅ | ✅ |
| internal enum value changed — non-breaking | ✅ | ❌ FP | ✅ | ✅ |
| cross-stdlib embed, same size — risk | ❌ FN | ❌ FN | ❌ FN | ✅ |
Read the columns as a story:
- L0 (symbols only) is blind: it misses every layout / signature / enum break (false negatives) — yet raises no layout false positives, precisely because it sees no layout at all. Low false positives here are an artefact of blindness, not of accuracy.
- L1 (+ debug info) catches the real breaks L0 missed — and, seeing layout for the first time, now over-calls internal-type churn it cannot tell apart from public churn (false positives appear).
- L2 (+ public headers) knows the public/private boundary, so it removes those false positives by scoping internal churn out — while keeping every real break. This is the single biggest false-positive reduction.
- L3 (+ build context) catches a last class of break no artifact tier can
see: a public type embedding
std::by value across two different stdlib implementations at the same size — invisible until the build flags reveal the mismatch.
So the honest shape is not "false positives fall monotonically" — L1 actually introduces false positives that L0 was too blind to raise, and L2 clears them. What holds monotonically, and what the gate enforces, is the false-negative side: more evidence never hides a break a weaker tier already caught (the authority rule — corroborating evidence may scope away a false positive, but never delete an artifact-proven break). With full evidence every case is correct (0 FP, 0 FN); CI publishes this matrix on every run, so each layer's contribution is a tracked number, not a claim.
What about L4/L5? The tracked matrix stops at L3 because L4 (source replay) and the derived L5 graph cannot be projected from a synthetic binary snapshot — they need real source. But they move both axes just as strongly. L4 is the only layer that can catch a macro /
constexpr/ default-argument / inline- or template-body change — a false negative invisible to every artifact tier L0–L3 (a stripped binary, its debug info, and its headers all compile the same emitted ABI). And L4/L5 cut false positives by proving which declarations are genuinely reachable and exported (the cross-source checks —exported_not_public,private_header_leak). Their accuracy is tracked separately: by the cross-check FP/FN corpus (also incheck_fp_rate.py) and by each example'smin_evidencetier incatalog/ground_truth.json.The derived sixth layer,
L5. Beyond the five sources above, abicheck derives anL5source/build graph (include/type/call reachability) from L3 (and any L4 surface) to localize and explain findings and prioritize cross-symbol impact. It is covered with the other build/source layers in Build Info & Sources.Layers (
L) vs. the depth dial. TheL0–L5codes name evidence layers — what abicheck sees and how much that evidence is trusted. Theabicheck comparecommand has one knob,--depth(binary|headers|build|source— exactly four public rungs), that selects how far down these layers to collect. The--depthdial section below explains the mapping (and the removeds0–s6/--mode/--source-methodaxes it replaced — see the appendix).
How they combine¶
The layers are independent and additive, not a fallback chain — abicheck overlays every source you give it and lets the strongest evidence win, under one rule, the authority rule — this is its definition, and every other page links here rather than restating it:
Artifact-backed evidence (L0/L1/L2) is authoritative for the shipped-ABI verdict. Build/source evidence (L3/L4) explains, localizes, scopes, or adds confidence to a finding, and can raise source-/API-level findings of its own — but it never silently deletes an artifact-proven break.
Concretely: L0 says a symbol changed; L1 says its layout changed by N
bytes; L2 says and the public declaration that names it changed too; L3 says
and it was built with a different -std, so expect churn; L4 says and the
macro it expands actually changed value. The verdict is computed worst-wins
across all of them. The design of how the layers are collected and
reconciled is in Architecture;
the per-case evidence each example needs is benchmarked in
Tool Comparison §Benchmarking by evidence tier.
Best input you can give abicheck: old + new library, matching public headers, debug info, and the build's compile database — L0+L1+L2+L3 together. With less, abicheck degrades down the staircase and tells you exactly which layers it had via the
dump --dry-run/layer_coveragereport.
Why call it "evidence"?¶
First, concretely: "evidence" is just the umbrella term for the sources of information in the table above. The artifact sources are the binary (L0), its debug info (L1), and its public headers (L2); the additional sources are the project's build-system data (L3 — compile flags, toolchain, target graph), its source tree (L4 — per-TU source ABI replay), and a source/build graph (L5 — include/type/call reachability). When the docs say "build/source evidence (L3/L4/L5)", that is exactly what they mean.
The umbrella word is a deliberate forensic metaphor, not decoration: abicheck treats "is this compatible?" as something it must prove from facts, the way a case is built from evidence, rather than as a single computation over one data source. Three properties of evidence are exactly the properties abicheck needs, and "tier" or "level" would imply the wrong ones:
- Independent and partial. Each source contributes some facts and none is complete on its own — a binary shows symbols but not layout, headers show API but not what was actually built. Evidence is additive and overlaid, not a ranked ladder you fall back down. (Call them "tiers" and readers assume a fallback chain; they aren't one.)
- Different authority. Just like physical vs. circumstantial evidence in a
courtroom, not all of it carries equal weight. Artifact evidence (L0–L2) is
what was actually built and shipped, so it is authoritative — only it can
declare a binary
BREAKING. Build/source evidence (L3/L4/L5) is corroborating — it explains, localizes, scopes, adds confidence, removes false positives, and can raise its own source-/API-level findings, but it can never overturn or silently delete an artifact-proven break. This is the authority rule. - Honest about what it had. Because the verdict is only as strong as the
evidence behind it, every run reports the evidence it actually collected (the
layer_coveragetable and the "checks enabled… and why others are not" capability report). The output literally says "here is the evidence I had, so here is what I could and couldn't check."
So "evidence" + the authority rule is the mental model that lets abicheck keep
adding sources for more accuracy without ever letting a weaker source override
a proven break. This four-way authority split (artifact-proven / corroborating
source-level / corroborating risk / consumer-demonstrated) is exactly what each
finding's evidence_status field spells out in machine-readable form — see
Output Formats § Per-finding epistemic status.
The --depth dial: how much evidence to collect¶
The layers above describe what abicheck can see. dump and compare
share one knob that decides how much of it to gather — --depth,
each rung named by the evidence you get and additive over the
one below it. As of the pre-1.0 CLI reset, the ladder has exactly four
public rungs — no more, no fewer:
--depth |
Reaches | Needs |
|---|---|---|
binary |
L0 exported symbols + binary metadata + L1 debug info (DWARF/PDB/BTF/CTF types and layouts) when the binary carries it (no L2 AST) + the always-on pattern scan | just the artifact(s) |
headers |
+ L2 header AST (the public/internal boundary) | a public-header directory + a C/C++ frontend |
build |
+ L3 build context (flag/toolchain drift) | a compile DB / build dir |
source |
+ L4 source-ABI replay + the L5 graph | sources and clang |
There is no fifth full rung. The old full depth (whole-library L4 replay,
as opposed to source's changed-TU replay) has been collapsed into
source — the two rungs only ever differed in replay scope, never in
which evidence layer they reached, so keeping both as separate public options
was pure surface area. See the appendix
for the full removal list and migration mapping.
Scope rule — which translation units --depth source actually replays:
- On
dump,--depth sourcealways uses TARGET scope — it replays the whole current library target.dumptakes no seed, so there is no narrowing to apply. - On
compare,--depth sourceuses CHANGED scope (just the TUs touched by a--since/--changed-pathseed) when a valid seed is present, and TARGET scope (the whole current library) otherwise — never an empty replay. That fallback is a deliberate bug fix: pinning--depth sourcewith no seed could once silently collect zero translation units and report clean by omission.
So a seeded compare --depth source analysed less than the whole
library. If you need the whole target replayed, omit the seed.
Omit --depth for auto — the default. auto names the
state "you didn't pin a rung"; it resolves to the fixed headers rung,
the same default compare has always used. It is not risk-driven: through
2026-09-09 an omitted depth was scored from the --since/--changed-path
seed and could escalate to build/source on a high-risk diff, but 0.6 retired that along with --risk-rules.
Nothing escalates on your behalf any more — a run that needs L3-L5
evidence must pin --depth build or --depth source explicitly, or it will
not collect it. A seed (--since/--changed-path) now only scopes a rung
you pinned; it no longer selects one.
compare --no-baseline is the one-build audit/hygiene/source-consistency
run — not a separate --audit flag (there isn't one), simply what declaring
no baseline means; supply an OLD operand instead and the run compares the
two sides.
A pinned deep depth is a contract (fail-loud)
Pinning --depth build|source with no source input
(--sources/--build-info) is an error, not a silent shallow scan: there
is nothing to collect L3/L4/L5 from. Pass the evidence, or pin
--depth binary/--depth headers for a shallower run that is honest
about its rung. Omitting --depth is not the way to ask for a
best-effort binary run any more: since 0.6 it resolves to a fixed headers, not to whatever the inputs
happen to support.
The resolved depth selects an internal collection mode, which decides which L-layers get collected and at what replay scope:
flowchart LR
subgraph D["--depth · the dial (how deep)"]
d0["binary"]:::cheap
d1["headers"]:::cheap
d2["build"]:::cheap
d3["source"]:::exp
end
subgraph L["L-axis · evidence (what)"]
L01["L0/L1 artifact (authoritative)"]
L2e["L2 header AST"]
L3e["L3 build context"]
L45["L4 replay + L5 graph"]
end
d0 --> L01
d1 --> L2e
d2 --> L3e
d3 --> L45
classDef cheap fill:#e6f4ea,stroke:#34a853;
classDef exp fill:#fce8e6,stroke:#ea4335;
Three properties of the dial worth internalizing:
- There is no
graphrung. The L5 reachability graph is an internal consequence of--depth source, never its own user-facing rung — you do not select the graph directly. - Cost has exactly one cliff, at L4.
binary/headers/buildare one cheap price;sourcepays for clang per-TU AST replay, and the cliff height tracks C++ template/STL instantiation depth, not TU count. Oncompare, a--since/--changed-pathseed keeps that replay to the changed TUs (CHANGED scope); without one — and always ondump, which takes no seed — it pays the cliff for the whole target (TARGET scope). Flag-level detail: Evidence Depth; measured numbers: Performance § scan-level cost model. - Coverage is honest. A run can request a deep level and only reach a shallow one (clang missing, no sources); abicheck never reports that as "scan failed" — every scan states the L-depth it actually reached and, for each disabled check, the precise input or tool to add (the capability report in Build Info & Sources § Evidence coverage; worked illustration: case147).
Combining two layers can also resolve a finding that is invisible or ambiguous to either alone: case148 crosschecks L2 header macros against L3 build flags; case149 crosschecks two L4 per-TU layouts; case150 crosschecks the L0 export table against L2 declarations in both directions.
Migrating an old command line? The removed s0…s6/--mode/--source-method/
--max axes map onto --depth in the
Removed scan axes.
1. The detectability matrix¶
The most important table on this page. Read it as: given only this evidence, what can a checker conclude — and what is it structurally blind to?
| Evidence available | Detects well | Cannot detect well |
|---|---|---|
| Exported symbol table only (stripped binary, no headers) | Removed/added exported symbols, symbol versions, visibility, SONAME/install-name, dependency (DT_NEEDED) changes |
Struct layout, enum values, calling convention, source-only API changes, macro changes, inline/template body changes |
| Debug info (DWARF / PDB / BTF) | Type layout, field offsets, enum values, class sizes, vtables, calling convention, packing/alignment | Source-only API intent, macros, default arguments, some template/header-only changes |
| Headers / AST (CastXML / Clang) | Source signatures, overloads, default args, access/final/explicit/noexcept, templates visible in headers |
Inline body semantics, macro expansion policy (unless modeled), runtime behavior |
| Source diff / compiler-based API extraction | Macros, inline function bodies, constexpr bodies, uninstantiated templates, source-level API |
The binary layout actually emitted into a shipped library (unless paired with the binary/debug info) |
| Runtime app swap / integration test | Real loader/linker behavior and tested execution paths | Untested public API, future consumers, silent layout corruption (unless a test happens to expose it) |
| Bundle scan (multi-library) | Cross-DSO dependency / provider / entry-point problems | Pure source compatibility and semantic behavior not represented in artifacts or manifests |
The first four rows are the artifact + source sources of §0 (L0/L1/L2 and the L4 source row); L3 build-context is a separate corroborating layer and is intentionally not a row here. The last two — runtime app swap and bundle scan — are orthogonal evidence axes, not extra rungs on the staircase.
Why abicheck combines layers¶
abicheck is strongest because it does not rely on a single row. It overlays
the five independent, additive sources of
§0 above — plus the derived L5 graph —
for six evidence layers in all (L0–L5; see the
§0 table for what each layer reveals, and
Architecture and Source & Build Data for how they are reconciled).
The best input you can give it is therefore:
old library + new library + matching public headers + debug info + build context — L0+L1+L2+L3 together.
With less, abicheck degrades gracefully down the staircase — a stripped binary
with no headers collapses toward symbol-only checking, where layout and
source-only breaks are invisible. See
Recommendation: feed .so + debug info + headers.
2. Methods compared, by the evidence they use¶
Each method is good at what its evidence exposes and blind to the rest. None is a complete contract check on its own.
a. Build an app and swap the library¶
The most realistic consumer-level test — but not a complete contract check. It only exercises what one app imports and runs.
| Strength | Example |
|---|---|
| Loader/linker failures | App fails because a required symbol is missing |
| Real runtime behavior | App crashes when it calls into changed ABI |
| Consumer-specific risk | App doesn't use the removed function, so this app still works |
| End-to-end deployment validation | RPATH/RUNPATH, search path, symbol versions all exercised |
| It misses | Why |
|---|---|
| Unused public APIs | The app only tests what it imports/executes |
| Silent data corruption | Tests may pass while layout is subtly wrong |
| Source compatibility | Binary may run, but recompiling may fail |
| Future consumers | One app is not the whole public contract |
| Header-only / source-only breaks | Existing binary doesn't exercise changed source |
This maps to abicheck's compare --used-by
scoping (an application-scoped view folded into compare, not a separate
command). See §4
for its exact scope.
b. libabigail, ABICC, and abicheck, by evidence¶
abidiff is DWARF-first (falls back toward symbol-only on a stripped
release, and a header directory is a public-symbol filter there, not an
AST), ABICC's two workflows are DWARF-based or GCC-header-based (each
missing the other's facts), and abicheck overlays every source it is given.
The per-tool capability and per-case results are owned by
Tool Comparison.
e. Methods beyond ABI diff tools¶
ABI diffing is one tool in a release-engineering kit. Complementary methods:
| Method | What it adds |
|---|---|
| Downstream rebuilds | Detect source API breaks by recompiling real consumers |
| Runtime smoke / probe tests | Detect loader errors and common runtime failures |
| ABI/API snapshot baselines | Treat release snapshots as immutable contract records |
| Symbol-version script / export-map linting | Enforce the intended public/private boundary |
| Header/source API extraction | Catch macros, inline definitions, template surface |
| Fuzz / integration tests | Catch behavioral changes behind a stable ABI |
| Reverse-dependency CI | Ecosystem/distribution-wide validation |
| Security-hardening scanners | Catch non-ABI deployment regressions (RELRO/PIE/canary/FORTIFY) |
The security-hardening check is the clean example of "not ABI, but still a release-compatibility risk": an ABI-compatible upgrade can weaken hardening while a normal ABI gate stays green. abicheck reports that as deployment risk, not an ABI break.
3. Traditional shared libraries vs header-only libraries¶
This distinction trips people up constantly, so it gets its own section.
Traditional .so / .dll / .dylib¶
There is a real binary contract to compare — exported symbols, symbol versions, dependency metadata, layout in debug info, public declarations in headers. abicheck's model is strongest here:
For compiled shared libraries, ABI compatibility is mainly about whether existing, already-built consumers can keep linking, loading, and calling into the new binary using the old contract.
Header-only libraries¶
A header-only library often has no exported library ABI — the code is compiled into each consumer. Compatibility is therefore mostly:
| Compatibility type | Meaning |
|---|---|
| Source API compatibility | Will existing users recompile? |
| Generated ABI compatibility | Will rebuilt objects stay compatible with other objects? |
| Semantic compatibility | Does inline/constexpr/template behavior still mean the same thing? |
| Configuration compatibility | Do macros/features/flags produce the same public surface? |
abicheck can still help in some cases:
| Case | How abicheck helps |
|---|---|
| Header-only API also gates a shared-library boundary | Header-AST comparison catches some API changes |
Explicit template instantiations shipped in a .so |
The emitted instantiations can be checked |
| Header constants / default args / source signatures in the AST | Some source-level API breaks are found |
| App links a runtime helper library | App mode (compare --used-by) checks the app's imported symbols |
But it cannot fully validate a pure header-only library: implicit header-only template instantiations are not in any shipped artifact (the documented mitigation is explicit instantiation of public templates that form part of the ABI — see Template Instantiation).
Header-only compatibility strategy
Use source API extraction, compile tests across supported compilers/standards, downstream rebuilds, and behavioral tests. Use abicheck for emitted artifacts, explicit template instantiations, or companion runtime libraries — not as the sole gate for header-only code.
4. App mode: consumer-scoped vs library-compare: contract-scoped¶
compare --used-by (repeatable; folds the old
appcompat command) answers a deliberately narrow question: will this
application still work with the new library? It parses the app's required
symbols, runs the full library comparison once, checks new-symbol
availability, and reports the app's own result beside the primary
(full-library) verdict — informational context, never a substitute for it.
That scope cuts both ways:
| App mode can say | App mode cannot say |
|---|---|
| "This app doesn't import the removed symbol." | "The library is generally ABI-compatible." |
| "This app needs symbol version X and the new lib lacks it." | "All future consumers are safe." |
| "This app is unaffected by this library-wide break." | "Header-only source users can recompile." |
| "This deployment path is OK for this app." | "No semantic behavior changed." |
App mode is consumer-scoped compatibility. Library
compareis product-contract compatibility. Use both: a plaincompareprotects the library contract;compare --used-byprotects a specific consumer deployment.
For header-only libraries, app mode is less central unless there's a companion runtime library — an existing app binary already contains the header-only code it compiled earlier, so swapping a library may not exercise the changed header-only implementation at all.
5. What ABI tools cannot prove¶
Even with perfect evidence, artifact comparison has hard boundaries. These are not abicheck's job — they need tests, specs, or source-AST tooling. Treat this as a guard against over-trusting any ABI tool (see Limitations for the authoritative list):
| Case | Why it's invisible / out of scope |
|---|---|
| Macro-only changes | Macros are preprocessor behavior; not in the artifact |
| Inline function body changed, same signature | No exported ABI change; body is compiled into the consumer |
constexpr behavior changed |
Source/semantic compatibility, no symbol change |
| Template body changed but not instantiated | No emitted artifact to compare |
| Uninstantiated template signature change | Not in the shipped .so unless instantiated (case122) |
| Header-only change not affecting exports | There may be no shared-library ABI surface |
| Stripped binary, no headers/debug | Mostly symbol-level comparison only |
| Header/binary mismatch | The tool may analyze a contract the binary wasn't built with — false results |
Static archives (.a / .lib) as archive containers |
abicheck analyzes linkable images/shared libraries/objects, not archive containers (details) |
| Pure behavioral / semantic changes | Same ABI/API, different meaning — needs tests/spec review |
| Ownership / lifetime / thread-safety guarantee changes | A signature can be byte-identical while the contract it implements flips |
The takeaway is the same one Part 0 opens with: a stable ABI is necessary but not sufficient for a compatible release. ABI tools prove the binary contract held; behavioral compatibility still needs your tests and your specification. See Assurance Beyond Static Checking for what to run alongside it — consumer rebuild tests, binary-swap tests, golden/differential tests, ASan/TSan lifecycle and concurrency tests — and what each one actually proves.
Four compatibility dimensions live entirely on the far side of this boundary, and each has its own page that starts from here rather than re-arguing it: behavioral and semantic compatibility (same signature, different meaning), data, wire and storage compatibility (a layout that outlives both binaries), ownership and lifetime contracts (who frees what, for how long), and concurrency and initialization contracts (thread-safety and init order).
6. Stored snapshots answer from stored evidence¶
A snapshot records where each declaration came from — a source_header per
function, type and variable, and (at L3) the source file of each compile unit.
Those recorded paths are provenance, not a licence to re-read the current
filesystem.
That distinction matters because the usual CI shape is stored baseline versus
live build: the OLD side is a .json snapshot published weeks ago, from a
checkout that no longer exists on the runner. If a source-derived check
re-opened the paths that snapshot names, it would characterise the historical
side from whatever happens to sit at the same path today — a different branch,
an edited header, or an unrelated file. Every answer it produced would be a
statement about the present dressed as history.
So abicheck reads a side's recorded source paths only when it is entitled to. A side can carry two source-evidence sources — its declared headers, and the compile units of an embedded L3 build pack — and each is judged on its own provenance, because one can be today's and the other historical at the same time:
| Evidence source on a side | What happens |
|---|---|
Declared headers, extracted in this run from headers (you passed -H, so the AST frontend opened those files) |
Read normally — those paths are today's paths |
| Declared headers, extracted in this run from the binary alone (DWARF or the symbol table) | Not read. DW_AT_decl_file names a path on the build machine, which this run never opened |
Build pack collected in this run (--sources, or a --build-info build directory) |
Read normally — this run resolved those compile units |
Build pack that came off disk (a pre-captured --build-info pack, or one embedded in a stored snapshot) |
Not read. Its recorded compile-unit paths are as historical as a stored snapshot's |
| Anything loaded from a stored snapshot | Not read at all. The check reports that the historical evaluation was not possible |
| Loaded, with an explicitly supplied and verified source context | Read, on the caller's stated provenance — the caller has asserted the recorded tree is the one on disk, which covers both sources |
A side that is licensed for one source and not the other reports what it read
and still establishes no absence: the unlicensed paths stay visible in the
coverage account as not_licensed, so a construct that only the unread half
could have contained never reads as introduced.
The second row is the one that surprises people. A snapshot built from debug
info records where each declaration was compiled from, not a file this run
has seen — and for a downloaded or previously-built binary that path either
does not exist locally or belongs to something else entirely. So a headerless
compare old.so new.so reports these advisory facts as not evaluated. Pass the
headers (-H) if you want them.
The consequence you will see in a report: the lexical pattern and
preprocessor pre-scan blocks of a stored-versus-stored (or
stored-versus-live) comparison state their coverage as not established for
the stored side, and every construct they track reads not_evaluated rather
than introduced or resolved. That is the honest answer. introduced is a
claim that the construct was absent before, and an absence claim needs
evidence about the OLD side — not an inference from a file the runner happens
to be holding.
Sufficiency is also answered per check rather than per run, because the three
checks rest on different evidence: the lexical scan on a set of files, macro
divergence on one clang -E -dM probe per compile unit, private-header leaks
on one clang -M probe per public header. A build whose compile units exceeded
the probe cap has not established the absence of a macro divergence, but its
public headers may still have been probed completely — so the leak check can be
established while the macro check is not, and neither answer is allowed to
stand in for the other.
The same rule governs coverage generally. Sufficiency for an absence claim is
computed from the set of inputs a check expected, with every one of them
accounted for — scanned, missing, unreadable, unsupported, deliberately
excluded, or not licensed — and any gap leaves the absence unestablished. A
declared input that no longer exists is a gap, never silent full coverage; so
is a directory that could not be read, which is easy to miss because a failed
directory walk reports nothing at all unless you ask it to.
Presence is the asymmetric case: a construct the scan actually saw is there,
whatever else the scan failed to read, so persistent survives partial
coverage where introduced and resolved do not.
If you need the historical side genuinely re-characterised, re-dump it from a checkout of its own commit; comparing two snapshots will not silently do it for you.
Removed scan axes¶
Earlier releases selected evidence with --source-method s0…s6, --mode
and --max; all three are gone. The migration table lives with
the other retired-surface maps, in
Migrating to the Current CLI § Removed scan axes.
See also: Part 0 — Compatibility as a Product Contract · Limitations · Tool Comparison · Application Compatibility · Multi-Binary Releases.
Ladder: ← Contract-Aware Compatibility · Concepts c2 · The evidence model · What Each Level Sees →