Snapshot Format (.abi.json)¶
abicheck dump writes a snapshot — a serializable, JSON representation of a
library's ABI surface — and abicheck compare reads two snapshots (or a live
binary and a saved snapshot) to produce a verdict. Checking a snapshot into your
repository as a baseline is the recommended way to detect ABI drift over time
(see Baseline Management).
This page documents the snapshot contract: its schema version, its compatibility rules, and its top-level structure.
Snapshots are not reports. A snapshot describes one library's ABI surface. The JSON that
compareemits is a separate comparison report with its own version field (report_schema_version). The two are versioned independently — see Two contracts below.
What dump writes today¶
A snapshot is a single JSON document with a sectioned envelope: three top-level keys, and every field this page describes lives inside one of the named sections.
{
"schema_version": 57,
"sections": {
"binary": {"section_kind": "binary", "section_schema_version": 1, "payload": {"...": "..."}},
"declarations": {"section_kind": "declarations", "section_schema_version": 1, "payload": {"...": "..."}},
"types": {"section_kind": "types", "section_schema_version": 1, "payload": {"...": "..."}},
"layout": {"...": "..."},
"debug": {"...": "..."},
"build": {"...": "..."},
"graph": {"...": "..."},
"provenance": {"...": "..."},
"semantic_ir": {"...": "..."}
},
"section_schema_versions": {
"binary": 1, "build": 1, "debug": 1, "declarations": 1, "graph": 1,
"layout": 1, "provenance": 1, "semantic_ir": 2, "types": 1
}
}
(Excerpt reduced from a real abicheck dump libfoo.so.1 -H foo.h -o snap.json
on this repository's own test fixture — the section list and both version
maps are verbatim; each payload is elided.)
| Key | Meaning |
|---|---|
schema_version |
The document's version — one integer, currently 57. Top-level so a loader can read it without parsing the rest. |
sections |
The nine named sections. Each carries its own section_kind, section_schema_version and payload. |
section_schema_versions |
A flat map of the same per-section versions, so a reader can check them without walking sections. |
Every dump invocation produces this, with no flag. Compression
(--compression, see Storage encoding) wraps it
but never changes it.
This is not the same number as the report's.
schema_versionversions one library's snapshot. The JSONcompareemits is a separate document with its ownreport_schema_version, and the product version (abicheck --version) is a third, independent number. See Two contracts.
Read a snapshot through the public loader¶
The physical layout is an implementation detail and has already changed once (see below). Consumers should go through the supported API, which unwraps the envelope and every legacy shape transparently:
from abicheck.serialization import load_snapshot
snap = load_snapshot("baseline.abi.json")
snap.library, snap.version, snap.functions, snap.types
load_snapshot / save_snapshot / write_snapshot are the public
compatibility surface. A consumer doing json.load(f)["functions"] is
reading a physical layout it does not own.
Reading an older snapshot¶
The envelope changed at schema v42. Two separate compatibility questions follow from that, and they have different answers:
| Question | Answer |
|---|---|
Can the current abicheck read a pre-v42 flat .abi.json? |
Yes. snapshot_from_dict/load_snapshot detect the shape and read it exactly as before. Stored baselines keep working. |
| Can a pre-v42 abicheck read a current snapshot? | No. v42 was bumped specifically so an older reader hard-rejects instead of silently reading every field as absent. |
| Does an external JSON consumer written against the flat shape still work? | No. It sees sections where it expected functions. Use the loader above. |
Loading an older snapshot warns, and the warning is the point:
UserWarning: Snapshot schema_version 8 predates this abicheck's schema_version 57:
header_cv_facts_reliable, param_kind_facts_reliable are marked unreliable on this
snapshot, so the affected detectors will decline to trust these stale facts rather
than risk a false positive purely from this tool upgrade.
Loading and re-saving does not upgrade the evidence. A re-saved snapshot
carries schema_version: 57 and the current envelope, but the warning
persists — it then says so explicitly — and the affected facts stay
unestablished. Serialization cannot invent evidence an older extractor never
collected. If you need those facts, re-run dump against the artifact.
This is why a release baseline should be regenerated when you upgrade
abicheck, rather than round-tripped.
Schema version history¶
schema_version is a single integer, not MAJOR.MINOR.
The current value is 57. See
abicheck/storage/snapshot_schema_versions.py's SCHEMA_VERSION for the
authoritative, up-to-date value and the full per-version comment.
This section is a lookup aid, not reading material — skip it unless you are tracing one specific field back to the version that introduced it.
Per-version history (v12 onward)¶
Two bumps were not additive. v42 replaced the flat document with the
sectioned envelope described in What dump writes
today, which breaks an external consumer reading
the physical layout. v49 replaced the shape of one field's payload:
surface_graph is now the compact graph-table encoding (see "Fields"), so a
reader of the old per-entity nodes/edges objects cannot read a v49 graph
by ignoring unknown keys. Every other bump added fields without changing
the meaning of existing ones — provenance metadata, PE/Mach-O
support, build-mode capture, declaration provenance (source_header/origin),
embedded build/source evidence, CastXML CV-qualifier reliability, the hybrid
AST frontend's per-fact producer map, the resolved AST toolchain identity,
(v12) the owner class of a hidden friend (Function.hidden_friend_owner),
(v13) the CastXML version-gate outcome (ast_toolchain_supported /
ast_toolchain_unsupported_reasons), (v14) extraction-contract fingerprints
proving two snapshots were compared under a comparable profile/scope
(AbiSnapshot.contract — verdict-blocking: see the
compatibility table below), (v15) structured compile-context provenance
for the header-AST parse (ast_resolved_standard, ast_cplusplus_macro,
ast_compile_args, ast_sysroot), (v16) DWARF-vs-header-AST layout
coherence (dwarf_layout_coherence, dwarf_layout_coherence_mismatches —
see "Compile-context provenance" below), (v17) which SYCL/DPC++ AST pass
a header-AST snapshot was built from (frontend_context_kind),
(v18) whether dump's default toolchain/system-header exclusion was
applied (dependency_scope, see dumper_scoping.py), (v19) whether the
direct-clang backend's deprecated/is_scoped facts are reliable
(clang_deprecation_facts_reliable, G31 Phase C — those facts became
genuinely populated by the clang backend at this version; see
dumper_clang.py), and (v20) whether the direct-clang backend's
TypeField.default (default member initializer) facts are reliable
(clang_field_initializer_facts_reliable, G31 Phase C — same shape as v19,
one fact and one version later; see dumper_clang_expr._field_initializer_value),
and (v21) whether the direct-clang backend's RecordType.vtable/
vptr_offset_bits facts are reliable (clang_vtable_facts_reliable, G31
Phase C — the direct-clang backend's vtable/vptr reconstruction; see
dumper_clang_vtable.py), and (v22) whether the direct-clang backend's
Param.is_restrict facts are reliable (clang_restrict_facts_reliable, G31
Phase C — castxml was that fact's only producer until this version; see
dumper_clang._clang_param_is_restrict), and (v23) whether the direct-clang
backend's Param.is_va_list facts are reliable
(clang_va_list_facts_reliable, G31 Phase C continued — no backend had
populated this fact at all before this version, and only for the x86-64
System V spelling; see dumper_clang_qualifiers._clang_param_is_va_list),
and (v24) whether the castxml backend's Variable.access facts are
reliable (castxml_var_access_facts_reliable, G31 Phase C continued — no
backend had populated this fact at all before this version; see
dumper_castxml._CastxmlParser._access_level), and (v25) a fully-qualified-
name-keyed twin of typedefs (AbiSnapshot.typedefs_qualified, G31 Phase C
continued) that closes a bare-name collision between two member typedefs
sharing a spelling in different classes/namespaces — needs no reliability
flag, since an empty dict degrades identically to "no typedefs at all" for
a pre-v25 snapshot, unlike the real-but-wrong scalar defaults v19-v23 above
guard against, (v26) Fact[T] siblings for RecordType.bases_fact/
virtual_bases_fact/vtable_fact/vptr_offset_bits_fact and
Param.is_va_list_fact (see storage/fact_codec.py),
(v27) Function.is_compiler_generated — closes a castxml L4 extractor bug
where a compiler-synthesized implicit special member leaked into the source
graph as if it were genuine public API; needs no reliability flag, since
None (a pre-v27 snapshot's default) degrades cleanly to today's inclusive
behavior rather than being misread as "confirmed user-written", (v28) each
declaration's entity_id carrier persisted through its own codec
(storage/entity_id_codec.py), (v29) AbiSnapshot.surface_graph — the
unconditional public-surface/L5 evidence graph (see
"Fields" below) persisted through its own to_dict() encoding, not
asdict()'s naive recursion (from v49 in the compact graph-table encoding
described in "Fields" below), and (v30) RecordType.is_final_fact — the
fact/capability registry's (abicheck/model/fact_registry.py) first registered Fact[T] conversion;
needs no reliability flag, since is_final's own None already
unambiguously means "not captured", (v31) EntityId sidecars for
typedefs and constants (AbiSnapshot.typedef_entity_ids/
constant_entity_ids) — additive twins of
typedefs_qualified/constants carrying the structural identity both
header-AST backends already resolve while walking their own intermediate
representation, keyed identically to their partner dict; empty on a
DWARF-only snapshot or one predating this field, and (v32)
RecordType.is_abstract_fact/
data_size_bits_fact/is_standard_layout_fact/is_trivially_copyable_fact/
qualified_name_fact/source_header_fact — the same registry's next batch
of case-(b) conversions (fields already X | None-typed, so the existing
"None already unambiguously means not captured" bridge applies directly);
qualified_name_fact is the one field in this batch both header backends
construct as an explicit Fact.present(...) rather than relying on the
generic bridge, since a None qualified name at global scope is itself a
confirmed determination, not missing evidence, and (v33)
EnumType.qualified_name_fact/source_header_fact — the identical
case-(b) pattern applied to EnumType's own twin fields, and (v34)
Variable.source_header_fact/alignment_bits_fact/elf_binding_fact —
the same case-(b) pattern applied to Variable's own three fields;
elf_binding_fact's decoded value is reconstructed as a real
SymbolBinding enum member rather than left as the bare JSON string
(storage/fact_codec.py's decode_variable_facts), since existing
readers unconditionally access .value on it, and (v35)
Function.contract_attributes_fact/is_explicit_fact/
is_hidden_friend_fact/source_header_fact/is_variadic_fact/
exception_spec_fact/is_override_fact/hidden_friend_owner_fact/
elf_binding_fact/is_compiler_generated_fact — the same case-(b)
pattern applied to Function's own ten remaining fields, closing out this dataclass's Fact[T] conversion; elf_binding_fact gets the
identical SymbolBinding reconstruction as Variable.elf_binding_fact,
and (v36) AbiSnapshot.ast_resolved_standard_fact — the same case-(b)
pattern applied to the last remaining case-(b) field outside the four
declaration dataclasses, and (v37) ElfMetadata.dynamic_flags_fact/
has_init_fact/has_fini_fact, PeMetadata.delay_imports_fact,
MachoMetadata.rpaths_fact — the same case-(b) pattern applied to the
three binary-format metadata blocks, schema-version-driven rather than
backend-driven since each block is parsed by exactly one backend;
dynamic_flags_fact's decoded value is reconstructed as a real
frozenset rather than left as the bare JSON list
(snapshot_platform_blocks.elf_from_dict), the identical reconstruction
Variable.elf_binding_fact/Function.elf_binding_fact already need for
their own non-JSON-native value type, and (v38) AbiSnapshot.semantic_ir
— the canonical, backend-independent IR, plus its
semantic_ir_conflicts sibling. Encoded by storage/semantic_ir_codec.py
as a list of {"occurrence": …, "entity": …} entries, not a JSON
object: the IR is keyed by an OccurrenceId dataclass, and rendering that
key as a string would lose the typed ScopePath inside it exactly as v28's
own entity_id encoding would have. Absent (the key is dropped, not
written as null) for a snapshot whose extraction produced no IR, which is
every snapshot written by a backend not yet narrowed onto the shared
normalizer. semantic_ir_conflicts is sparse the same way, so for such a
snapshot a v38 document differs from the v37 one only in the version
stamp. Then (v39) TypeField.is_const_fact/
is_volatile_fact/is_mutable_fact — the first case-(a)
conversion: unlike every _fact sibling above, these three fields' own
values (a plain False) carry no availability signal at all, so the
snapshot-level header_cv_facts_reliable flag is what a pre-v39 document's
load consults (storage/fact_codec.py's apply_case_a_fact_backfill)
before a blanket False may be read as Fact.present(False) rather than
Fact.not_collected(), and (v40) Function.deprecated_fact/
Variable.deprecated_fact/RecordType.deprecated_fact/
EnumType.deprecated_fact plus EnumType.is_scoped_fact — the rest of the
case-(a) family clang_deprecation_facts_reliable guards, converted the
same way (TypeField.deprecated's own sibling landed one version earlier,
with the rest of that dataclass's fields), and (v41) Param.is_restrict_fact
and Variable.access_fact — the last two fields the fact registry tracked as eligible-but-unconverted, closing that phase's field-by-field
conversion. access_fact's decoded value is rebuilt into a real
AccessLevel member, the same reconstruction elf_binding_fact needs.
Then (v42) the on-disk wire format itself changed, not just a field:
The storage redesign made snapshot_to_json() write
storage.sectioned_document's single-file sectioned envelope instead of a flat
document. Bumped specifically so a pre-Phase-8 reader (whose own
SCHEMA_VERSION was already 41) hits the hard-rejection path below instead of
silently reading every top-level field as absent/empty — a same-numbered
envelope change would have given that reader no signal at all. This build
itself reads the envelope transparently regardless of version, per
snapshot_from_dict's own is_sectioned_document check. Then (v43)
Variable.is_static persisted — closes the plain-C/extern "C" same-named
static-vs-external variable identity collision tu_merge._variable_key's own
docstring long documented as a known, accepted limitation; missing on a pre-v43
snapshot loads as False, matching every prior reader's implicit assumption
since the field did not exist. Then (v44) AbiSnapshot.header_only
persisted — the explicit marker for a snapshot built by the binary-less
header-AST dump path (workstream F S1, "Header-only comparison"; no
SO_PATH, no --sources/--build-info); missing on a pre-v44 snapshot
loads as False, matching every prior snapshot's implicit "this has a
binary, or is a pre-existing source-only dump" status. (v46) Function/Variable each gained three Fact[bool] siblings --
declared_in_headers_fact, in_public_contract_fact,
binary_exported_fact -- splitting the three independent facts visibility
used to conflate (a declaration exists in the parsed headers; it belongs to
the promised public contract; a binary symbol is exported). They are absent
on a pre-v46 snapshot, where abicheck/model/surface_facts.py derives all
three from the stored visibility value as a PARTIAL, diagnostic-stamped
reading -- never as a confirmed negative, so a headerless snapshot still
reads "declaration not established" rather than "no declaration". Then
(v48) AbiSnapshot.excluded_header_matching persisted — which rule the
recorded exclusion patterns were matched by: "glob" for the native
--exclude-header (fnmatch, plus a */<pattern> try), "abicc" for a
descriptor's <skip_headers> under ABICC's own three rule classes
(basename, component-boundary path/directory, compiled pattern; written only
by the since-removed ABICC compat front end, still read), and the legacy "exact" for a
descriptor snapshot written before those rule classes existed, when a
descriptor skip was a plain basename-or-path membership test. The same
pattern text is not the same scope under any two of them — fftw/fftw.h
excludes nothing under "exact" and takes the header under "abicc" — so
recording the text alone let the comparability gate accept two snapshots
covering different surfaces. Absent on a pre-v48 snapshot, which loads as
"unknown" — not as "glob": v47 already recorded a descriptor's
exact-matched <skip_headers> alongside the native fnmatch ones, so
assuming fnmatch for a mode-less snapshot would let a baseline holding
include/foo.h compare clean against a native glob snapshot that excluded a
different set of headers. An unrecorded rule is refused rather than guessed.
A mode-less snapshot carrying no patterns is unaffected — there is nothing
for a rule to have matched. Then (v49) AbiSnapshot.surface_graph is
written in the compact graph-table encoding (storage/graph_table_codec.py,
see "Fields" below) and holds observed evidence only; a pre-v49 graph still
loads.
(v50) The persisted header/L5 source graph (surface_graph,
build_source.source_graph) uses the evidence-entity-model invariant-I1 node
ids (SourceGraphSummary.schema_version 3): one id per entity, from
abicheck/model/graph_entity_identity.py. A C-linkage declaration is keyed on
its linker name, an identity-less entity (a castxml constructor/destructor
placeholder, an unmangled overload, a type spelling two declarations share)
is an explicit unresolved:// node, a flat-path type is keyed on its
qualified name, and a new identity_aliases map records second spellings of
one entity. A pre-v50 snapshot loads unchanged, graph included; comparing its
graph against a v50 one reports the L5 layer as not compared (on the
coverage row and as a warning) rather than diffing ids that name entities
differently. Re-dump the older side to restore the graph comparison.
(v51) dwarf may carry struct_odr_conflicts/enum_odr_conflicts (name →
list of further, layout-distinct definitions another compile unit gave a
struct/enum already in structs/enums) and odr_conflicts_observed (the
DWARF walk looked for them). The keys are written only when that walk ran,
so a snapshot without one encodes exactly as v50. A pre-v51 snapshot loads
with odr_conflicts_observed false: its lack of conflicts means "not looked
for", never "none". The debug-type join (compare/debug_type_join.py)
reports a conflicted name as ambiguous rather than trusting the first
definition.
(v56) elf.symbols[].code_hash — a 128-bit BLAKE2b hex digest of an
exported function's code bytes, written only when computed. The stripped-binary
rename check reads it: equal hashes confirm a same-size rename and break a tie
between two same-size candidates. Unequal hashes are not treated as evidence,
because a relative call or data reference changes when the same function moves.
An absent key means "not hashed".
(v57) semantic_ir function occurrences carry the five signature facts
return_type_spelling, parameter_type_spellings, parameter_kinds,
ref_qualifier and is_variadic, inside an IR document stamped
"version": 3. The spellings are the producer's own, not canonicalized. A
pre-v57 document decodes them as not collected, and loading fills them from
the snapshot's own functions, so an older baseline compares exactly as before.
(v55) public_header_identifiers_fact — every identifier token the public
header set's raw text spells, every preprocessor branch included (comments and
literals stripped). --contract public reads it to tell "no public header
names this export" from "the parsed branches did not declare it". An
evidence-free not_collected fact is omitted on write; an absent key loads as
not_collected, on a pre-v55 snapshot as on a current one.
(v54) Each type slot's resolved identity: Function.return_type_identities_fact
and Param/Variable/TypeField.type_identities_fact, each a Fact holding
the qualified names of the records/enums the slot resolves to. A slot's own
spelling is the bare source text (Cache *), which cannot say which of two
same-leaf records (ns1::Cache/ns2::Cache) it names. These facts record
the compiler's answer, and the public contract domain confirms a break on
the reached record with it. castxml writes them; every other producer leaves
the key out, so its documents encode exactly as v53. On load, an absent key
reads as NOT_COLLECTED in a v54 document and as absent in a pre-v54 one;
either way the evaluator keeps its pre-v54 answer.
(v52) extraction_scope — the ownership rules a header-derived snapshot's
declarations were classified under (see Extraction scope and ownership
below). Declaration lists encode exactly as v51. A pre-v52 snapshot loads with
the field absent — unrecorded, never read as "no rules" — and every
declaration's owner unknown.
(v47) AbiSnapshot.excluded_header_patterns persisted — the
--exclude-header PATTERN values a snapshot was dumped under. The parsed
surface is narrower than the operand names and nothing else recorded that,
so a stored baseline dumped with an exclusion, compared later against a full
dump of the same library, reported every declaration the excluded header
carried as appearing out of nowhere. Absent on a pre-v47 snapshot, which
loads as the empty tuple — correct for every such snapshot, since the flag
did not exist. Before
that (v45)
Param.kind_fact persisted — closes the last case-(a) field the "field-by-field conversion complete" note missed: neither header-AST
backend had ever determined a parameter's indirection kind
(value/pointer/reference/rvalue-reference) before this version, so a
pre-v45 header-derived snapshot's blanket "value" is a placeholder, not
a confirmed reading — AbiSnapshot.param_kind_facts_reliable marks it.
Forward / backward compatibility¶
abicheck loads a snapshot best-effort and never migrates it in place. The rule
is determined entirely by comparing the file's schema_version against the
SCHEMA_VERSION the running abicheck supports:
File schema_version |
Behavior on load |
|---|---|
| Missing | Treated as 1 (the pre-versioning format) and loaded normally. |
Older or equal to this build (<= 54) |
Loaded cleanly. Fields introduced by newer versions are absent and fall back to their defaults (None, empty, or a tri-state None that suppresses the detectors depending on that evidence). No warning. |
Newer than this build, and < 14 |
Loaded best-effort with a UserWarning ("Data may be incomplete or misinterpreted. Upgrade abicheck…"). The load is not aborted — unrecognised keys are ignored and recognised keys are read. |
Newer than this build, and >= 14 |
Hard-rejected — IncompatibleSnapshotSchemaError — instead of warn-and-continue. |
Two consequences worth internalising:
- Reading is version-tolerant in both directions. An older baseline produced by an earlier abicheck loads without error against a newer abicheck; missing fields simply take defaults. This is what makes checked-in baselines durable across tool upgrades.
- A newer snapshot usually warns rather than fails — but not once a
verdict-blocking field exists. Prior to v14 every bump was purely
additive, so an older reader can safely ignore a field it doesn't
recognise. Starting at v14,
AbiSnapshot.contractmakes a bump verdict-blocking: a reader that silently dropped it could compare two possibly-incomparable snapshots and produce an ordinary, wrong verdict.snapshot_from_dicttherefore hard-rejects (rather than warns-and-loads) any fileschema_versionthat is both newer than the running build'sSCHEMA_VERSIONand>= 14(_MIN_SCHEMA_VERSION_REQUIRING_HARD_REJECTIONinabicheck/serialization.py) — this only protects readers built from that guard's introduction onward; see theSCHEMA_VERSIONhistory comment for the full explanation of what it can and cannot retroactively protect. Upgrade abicheck to read a newer snapshot faithfully.
Storage encoding¶
Everything above describes the logical snapshot — the decoded JSON
payload. On disk, that payload may be stored plain, gzip-compressed, or
zstd-compressed; compression is a storage/transport envelope around the
identical JSON, never a new schema, and never changes schema_version,
AbiSnapshot.contract, evidence depth, build_source, or the verdict two
snapshots produce when compared.
| Encoding | Canonical suffix | Notes |
|---|---|---|
| plain | .abicheck.json / .abi.json |
debugging, small Git-reviewable snapshots |
| gzip | .abicheck.json.gz / .abi.json.gz |
universal interoperability |
| zstd | .abicheck.json.zst / .abi.json.zst |
preferred for baseline/release/cache storage |
compare and the Python API
(abicheck.serialization.load_snapshot) all read every encoding
transparently — detected from magic bytes, not just the filename suffix.
abicheck dump produces one: it infers the encoding from -o/--output's
suffix by default (--compression auto), or accepts an explicit
--compression {none,gzip,zstd}; write_snapshot is the Python API
equivalent for writing. The full storage-envelope model (determinism, atomic writes,
decompression limits, and what's still deferred).
Sectioned packaging¶
Orthogonal to compression: the envelope shown in
What dump writes today is this page's field set
packaged into named, independently versioned sections
(binary/declarations/types/layout/debug/build/graph/
provenance/semantic_ir). Every dump/write_snapshot invocation writes
it; no flag selects it.
snapshot_from_dict/load_snapshot unwrap it before any field below is
read, and an older flat .abi.json is still read exactly as it always was —
the field-level contract below is unaffected either way; only the outermost
envelope differs. The directory-backed content-addressed package is a
separate, advanced storage shape with its own page: see
Project Snapshot Format.
Field reference¶
The keys below are the AbiSnapshot model's fields
(abicheck/model/snapshot.py), as encoded by the serializer. In the current
envelope they live inside the matching sections[...] payload; in a
pre-v42 flat document they are top-level. load_snapshot presents them
identically either way, which is why this reference is written against the
model rather than against either physical layout. Optional keys are omitted or null when there is no data
(for example, a pure-ELF dump has no dwarf or build_source).
Identity and provenance¶
| Key | Type | Meaning |
|---|---|---|
schema_version |
int | Snapshot format version (currently 57). |
library |
string | Library identity, e.g. libfoo.so.1. |
version |
string | Library version string, e.g. 1.2.3. |
source_path |
string | null | Original path the snapshot was taken from. |
platform |
string | null | elf, pe, macho, or null. |
language_profile |
string | null | c, cpp, sycl, or null. |
git_commit |
string | null | Git SHA captured at dump time. |
git_tag |
string | null | Git tag (e.g. v2.0.0), supplied or auto-detected. |
created_at |
string | null | ISO 8601 timestamp set at dump time. |
build_id |
string | null | Opaque CI identifier (run ID, build number). |
contract |
object | null | Extraction-contract fingerprints (schema v14, verdict-blocking — see "Forward / backward compatibility" above): profile_fingerprint/scope_fingerprint plus their named resolved sub-inputs, proving two snapshots were extracted under a comparable profile/scope. null when no producer populated it yet. |
dependency_scope |
string | null | (schema v18) "filtered" when the toolchain/system-header exclusion (dumper_scoping.py) was applied, "full" when opted out via --include-system-declarations. Every front end filters by default (include_dependencies=False): the dump/compare CLI, service.run_dump and InputSpec all share that default since the 0.6 defaults-alignment pass — before it, a typed-API caller that omitted the field got the unfiltered surface while the identical CLI invocation got the filtered one, and the two were not comparable. null on any pre-v18 snapshot or any snapshot with no header-derived declarations. comparability.check_contracts_comparable raises ScopeMismatchError only when BOTH sides carry an explicit, non-null value and they differ — null is deliberately NOT treated as "full" (an ordinary pre-v18 baseline is usually already-filtered content that simply predates this tag; assuming "full" for it would spuriously flag the routine "compare a cached baseline against a fresh dump" workflow), so a genuinely ambiguous untagged snapshot is left unchecked on this axis rather than guessed at. |
header_only |
boolean | (schema v44) true only for a snapshot built by the binary-less header-AST dump path — either dump -H api.h (no SO_PATH/--sources/--build-info) or a pathless dump --dump-manifest m.yaml naming real header-AST roots (workstream F S1, "Header-only comparison"). Explicit, not inferred from platform/from_headers: a pre-existing --sources/--build-info source-only dump also has platform: null, but carries no header-AST declarations at all. public_header_dirs alone does NOT select this path — it is a declaration-provenance (public-vs-internal) classifier only, never a source of headers to parse; a binary-less request naming only public_header_dirs (no header file, no manifest) is rejected before extraction. false (the default) for every snapshot predating this field and every ordinary binary dump. |
Compile-context provenance (schema v15, header-AST parses only)¶
Populated only when the snapshot came from a header-AST parse (from_headers
true); null/empty on a DWARF/symbols-only or binary-only snapshot, and on
any pre-v15 snapshot. ast_compile_args and ast_sysroot are redacted via
the same RedactionPolicy every L3 build-evidence adapter applies (secret-
looking -D values and absolute home-prefixed paths are stripped/normalized
before persistence — see abicheck/buildsource/redaction.py).
| Key | Type | Default | Meaning |
|---|---|---|---|
ast_resolved_standard |
string | null | null |
The C/C++ standard actually used for the header parse: an explicit -std=/--std=//std: value verbatim, or "gnu++20" when the requires/concept heuristic forced it. null means the frontend's own unpinned default was used (never guessed at). |
ast_cplusplus_macro |
string | null | null |
The standard-mandated __cplusplus literal for ast_resolved_standard (e.g. "201703L" for "gnu++17"), looked up from a static ISO-standard table. null when ast_resolved_standard is unset or not a recognized C++ edition. |
ast_compile_args |
array of strings | [] |
The ordered extra compiler arguments passed to the header frontend (--compiler-option tokens, then a shlex-split composed-flags string derived from the compilation database --build-info resolves to), redacted. |
ast_sysroot |
string | null | null |
The --sysroot passed to the header frontend, if any, redacted. |
ast_toolchain (dict[str, str], populated since schema v9) carries the
exact tool identity behind the header-AST parse. It is untyped/free-form —
new keys are additive and never require a schema bump — but these keys are
stable and machine-checked by tests/test_tool_identity.py/
tests/test_castxml_policy.py:
| Key | Meaning |
|---|---|
selected / compiler_selected |
The exact frontend/host-compiler executable path selected from PATH (or an explicit --compiler). |
realpath / compiler_realpath |
The same path with symlinks resolved. |
sha256 / compiler_sha256 |
SHA-256 of the executable's file contents, so a same-version binary rebuild/repackage still changes provenance. |
version / compiler_version |
The raw, bounded --version transcript for that exact executable revision. |
target_triple / compiler_target_triple |
The <tool> -dumpmachine output for that executable (GCC/G++/Clang/Clang++ only — omitted, not empty, for a tool that doesn't support the flag, e.g. castxml itself or MSVC cl.exe). |
castxml_version |
CastXML's own release version (e.g. "0.7.0"), parsed from version — castxml-producer snapshots only. |
castxml_bundled_clang_version |
The bundled/linked Clang's major.minor (e.g. "18.1"), parsed from version — castxml-producer snapshots only. Kept separate from castxml_version since the two floors (MIN_CASTXML, MIN_CASTXML_CLANG_MAJOR in castxml_policy.py) are independently enforced by the version gate. |
A hybrid snapshot (ast_producer == "hybrid") namespaces every key from
both runs instead of picking one — castxml_selected, castxml_version,
castxml_castxml_version, clang_selected, clang_target_triple, and so
on (dumper_hybrid.py's merge is a generic castxml_/clang_-prefixed
dict union, so a key that already started with castxml_ on the castxml
side is not special-cased).
DWARF-vs-header-AST layout coherence (schema v16)¶
The clang L2 header backend is layout-blind (no size_bits/alignment_bits/
field offset_bits) — when the binary being dumped also carries DWARF debug
info, dumper_layout_backfill.backfill_dwarf_layout() backfills that layout
from the same binary's DWARF, but only for a record it can corroborate as
the same declaration (matching name, kind, and field/base overlap — see
that function's docstring for the exact rules). These two fields make that
corroboration outcome visible instead of silent; they never change what
gets backfilled, only report on it.
| Key | Type | Default | Meaning |
|---|---|---|---|
dwarf_layout_coherence |
string | null | null |
One of "matched" (every record eligible for backfill was corroborated, or none needed it), "partial" (some corroborated, some had no DWARF candidate at all — benign, e.g. declared-but-never-instantiated), "mismatch" (at least one record found a uniquely-named DWARF candidate but the two disagreed — backfill already refused to merge that record's layout), or "unavailable" (the clang backend ran but the binary carried no usable DWARF at all). null on any snapshot not built via the clang L2 backend (a castxml snapshot computes layout directly — not a coherence question) and on any pre-v16 snapshot. |
dwarf_layout_coherence_mismatches |
array of strings | [] |
Header record names backfill found a uniquely-named DWARF candidate for but rejected as uncorroborated — populated only when dwarf_layout_coherence == "mismatch". |
SYCL/DPC++ frontend context (schema v17, header-AST parses only)¶
| Key | Type | Default | Meaning |
|---|---|---|---|
frontend_context_kind |
string | null | null |
Which AST pass ("host" or "device") this header-AST snapshot's clang backend selected via --frontend-context (sycl_context.py). null on any non-SYCL/DPC++ invocation and on any pre-v17 snapshot. |
| public_header_identifiers_fact | object | absent | absent | v55. A Fact: status, value (sorted identifier list when present), diagnostics, producer ("public_header_text"). failed when a named header is missing or unreadable, unsupported when a header uses ## token pasting. |
Extraction scope and ownership (schema v52)¶
Written for every snapshot a run extracts from headers; absent on a binary- or debug-only snapshot and on any snapshot a run loaded (a stored baseline keeps the scope it was dumped under).
| Key | Type | Meaning |
|---|---|---|
extraction_scope.ownership_rules |
object | target_roots, dependencies (name, header_roots), private_headers, private_namespaces, dependency_evidence. Roots are POSIX paths relative to the project root when they lie under it, absolute otherwise. |
extraction_scope.dependency_evidence |
string | What was kept of dependency declarations. Always "full" today. |
extraction_scope.prefilter |
object | null | Reserved for a frontend prefilter; always null today. |
extraction_scope.fingerprint |
string | sha256: over the three keys above, canonicalized (sorted, de-duplicated). What comparability and the configuration digest (surface.ownership) compare. |
extraction_scope.diagnostics |
array | The classifier's namespace-mismatch diagnostics (a target file declaring into a dependency's namespace). Omitted when empty. |
extraction_scope.entity_ownership.decisions |
array | Interned [owner, contract, rule_id] triples. |
extraction_scope.entity_ownership.<functions\|variables\|types\|enums> |
array of int | One index into decisions per declaration of that list, in list order; -1 for a declaration that was not classified. A list whose length disagrees with the snapshot's own list is ignored on load (its declarations stay unclassified) rather than misattributed. |
owner is target, dependency:<name>, toolchain or unresolved;
contract is public, private, external or unresolved. An
unresolved owner is a real answer (no root claims the file); an
unclassified declaration has no decision at all, and readers treat it as
unknown.
ABI surface¶
| Key | Type | Meaning |
|---|---|---|
functions |
array | Exported functions (name, mangled name, return type, params, virtuality, access, provenance). Since v54 a castxml-dumped function's return type, each parameter (and each variable's type, each record field) may carry a *type_identities_fact: the qualified record/enum names that slot resolves to, which its bare spelling cannot express. |
variables |
array | Exported global/static variables. |
types |
array | Records (struct/class/union) with fields, bases, vtable, and layout descriptors. |
enums |
array | Enumerations with members and underlying type. |
typedefs |
object | Typedef name → underlying type. Bare-name-keyed; two distinct member typedefs sharing a spelling in different classes/namespaces collide onto one key (see typedefs_qualified). |
typedefs_qualified |
object | Fully-qualified-name-keyed twin of typedefs (schema v25) — collision-free. Empty for a pre-v25 snapshot or one produced without per-class qualified typedef scoping (e.g. DWARF-only). |
constants |
object | Preprocessor/compile-time constants (qualified name → value). |
typedef_entity_ids |
object | EntityId sidecar for typedefs_qualified (schema v31), keyed identically — the typed ScopePath/kind/leaf-name identity a dict[str, str] cannot carry on a declaration object. Empty for a pre-v31 or DWARF-only snapshot. |
constant_entity_ids |
object | EntityId sidecar for constants (schema v31), keyed identically. Empty for a pre-v31 or DWARF-only snapshot. |
Evidence-tier and mode flags¶
| Key | Type | Meaning |
|---|---|---|
elf_only_mode |
bool | True when dumped without headers (all functions carry ELF-only provenance). |
from_headers |
bool | True when the surface was parsed from public headers (drives the header-aware evidence tier). Omitted from the file when it was only inferred on load, so a reload re-runs the same inference. |
scope_fallback |
string | null | Public-scope fallback marker. |
parsed_with_build_context |
bool | True when parsed with build-context evidence. |
Platform and debug metadata (optional)¶
| Key | Type | Meaning |
|---|---|---|
elf |
object | null | ELF metadata: SONAME, DT_NEEDED, version defs/reqs, symbols, imports, hardening flags. |
pe |
object | null | PE/COFF metadata (Windows DLL exports, machine, characteristics). |
macho |
object | null | Mach-O metadata (dylib exports, CPU slices, install name). |
dwarf |
object | null | DWARF struct/enum layout (v51: plus ODR conflicts, see above). |
dwarf_advanced |
object | null | Toolchain, calling conventions, value-ABI traits. |
sycl |
object | null | SYCL plugin-interface metadata. |
dependency_info |
object | null | Resolved dependency graph (nodes, edges, unresolved). |
build_mode |
object | null | Normalized compiler/stdlib/standard capture (ADR build-mode work). No dump path writes it today; it is read back only from a document that carries one (see known-gaps.md). |
Embedded build/source evidence (optional)¶
| Key | Type | Meaning |
|---|---|---|
build_source_pack |
object | null | Reference to an out-of-band build/source pack. Older snapshots may store this under the legacy key evidence_pack, which the loader still reads. |
build_source |
object | null | Inline-embedded build/source facts for single-artifact workflows. Omitted when nothing was embedded. |
surface_graph |
object | omitted | (v29) The unconditional public-surface/L5 evidence graph — never gated on build_source, unlike the row above. The key is omitted entirely (not written as null) for a snapshot predating this field, a binary-only snapshot, or one whose headers were never parsed — encode_surface_graph() pops the key rather than writing a null placeholder. When build_source.source_graph is the identical object, it is omitted from build_source's own encoding rather than written twice; the loader restores that alias on read. From v49 the value is storage/graph_table_codec.py's compact encoding ("encoding": "graph-table/1"): one strings table, nodes/edges objects of equal-length index columns (id/src/dst, kind, label for nodes, and facts), and deduplicated attrs and facts tables. Only observed evidence is written: nothing the loader rederives (indexes, graph_id, finalize-owned coverage counts, and each entity's resolved/conflicts/occurrences/attrs/provenance/confidence, all rebuilt from its facts), and no fact from a producer that projects the snapshot's own records (the public-surface builder's declaration/type/symbol/binary_symbol/debug_type nodes and declares/references/declares_linker_name/exports/debug_type_of edges), which a reader rebuilds on demand. A value without encoding is the pre-v49 per-entity form and still loads. The graph is decoded on first access, not at load (a corrupt value raises then). |
build_context_defines |
array of strings | The build's active -D macro set, harvested from a compile database. Empty when no compile database was supplied. |
conditional_fields |
object | {type: {field: {guard, type, is_bitfield, ...}}} registry of record fields guarded by a single positive #ifdef/#if defined(...), including fields a context-free header parse pruned from types[].fields. Feeds the build-context reconciliation diff pass, which runs unconditionally whenever both snapshots carry this field (one-comparison-product.md Phase 7i); empty when no compile database was supplied at dump time. |
Internal cache fields on the model (
_func_by_mangled,_var_by_mangled,_type_by_name) and the runtime-onlyfrom_headers_inferredqualifier are never serialized.
Two contracts: snapshot vs report¶
schema_version and report_schema_version are different fields on different
files:
Snapshot (dump) |
Comparison report (compare -o json=-) |
|
|---|---|---|
| Version field | schema_version |
report_schema_version |
| Type | integer (currently 57) |
string MAJOR.MINOR (e.g. 1.0) |
| Describes | one library's ABI surface | the diff between two snapshots |
A snapshot has no report_schema_version, and a report has no
schema_version; the two version numbers evolve independently. For the report
contract and its stability policy, see
Output Formats.
Stability guidance¶
- Check baselines into version control. A saved
.abi.jsonis the intended input tocompare; storing one per release lets CI diff each build against the last shipped ABI. See Baseline Management. - Older baselines stay readable. Because loading fills missing newer fields with defaults, a baseline written by an earlier abicheck compares correctly against a live binary dumped by a newer one — no regeneration required for a routine tool upgrade.
- Regenerate when you want new evidence. Fields added in a newer
schema_version(e.g. build-mode or embedded source evidence) are only present in freshly-dumped snapshots. Re-dump the baseline to benefit from detectors that rely on that evidence. - Pin the abicheck version in CI if a
UserWarningabout a newerschema_versionwould be treated as an error in your pipeline.
See also¶
- Baseline Management — producing, storing, and comparing snapshots as ABI baselines.
- Output Formats — the comparison-report JSON and
report_schema_version.
The build/source pack envelope¶
- The pack is content-addressed and versioned independently
(
evidence_pack_version) from the ABI snapshot schema, so it never bloats an ordinary dump. The snapshot stores only a lightweightevidence_packreference (content hash + coverage summary); old readers ignore it. - Every extractor writes both a raw artifact (under
raw/, for provenance/debugging) and an abicheck-owned normalized fact model (e.g.build/build_evidence.json). Only normalized facts feed comparison and the content hash. - Command lines and paths are redacted (home prefixes, secret-looking
-Dmacros) before they are persisted.
See Source & Build Data for the full model.