Skip to content

Snapshot Format (.abi.json)

abicheck dump writes a snapshot — a serializable, JSON representation of a library's ABI surface — and abicheck compare reads two snapshots (or a live binary and a saved snapshot) to produce a verdict. Checking a snapshot into your repository as a baseline is the recommended way to detect ABI drift over time (see Baseline Management).

This page documents the snapshot contract: its schema version, its compatibility rules, and its top-level structure.

Snapshots are not reports. A snapshot describes one library's ABI surface. The JSON that compare emits is a separate comparison report with its own version field (report_schema_version). The two are versioned independently — see Two contracts below.


What dump writes today

A snapshot is a single JSON document with a sectioned envelope: three top-level keys, and every field this page describes lives inside one of the named sections.

{
  "schema_version": 57,
  "sections": {
    "binary":       {"section_kind": "binary",       "section_schema_version": 1, "payload": {"...": "..."}},
    "declarations": {"section_kind": "declarations", "section_schema_version": 1, "payload": {"...": "..."}},
    "types":        {"section_kind": "types",        "section_schema_version": 1, "payload": {"...": "..."}},
    "layout":       {"...": "..."},
    "debug":        {"...": "..."},
    "build":        {"...": "..."},
    "graph":        {"...": "..."},
    "provenance":   {"...": "..."},
    "semantic_ir":  {"...": "..."}
  },
  "section_schema_versions": {
    "binary": 1, "build": 1, "debug": 1, "declarations": 1, "graph": 1,
    "layout": 1, "provenance": 1, "semantic_ir": 2, "types": 1
  }
}

(Excerpt reduced from a real abicheck dump libfoo.so.1 -H foo.h -o snap.json on this repository's own test fixture — the section list and both version maps are verbatim; each payload is elided.)

Key Meaning
schema_version The document's version — one integer, currently 57. Top-level so a loader can read it without parsing the rest.
sections The nine named sections. Each carries its own section_kind, section_schema_version and payload.
section_schema_versions A flat map of the same per-section versions, so a reader can check them without walking sections.

Every dump invocation produces this, with no flag. Compression (--compression, see Storage encoding) wraps it but never changes it.

This is not the same number as the report's. schema_version versions one library's snapshot. The JSON compare emits is a separate document with its own report_schema_version, and the product version (abicheck --version) is a third, independent number. See Two contracts.

Read a snapshot through the public loader

The physical layout is an implementation detail and has already changed once (see below). Consumers should go through the supported API, which unwraps the envelope and every legacy shape transparently:

from abicheck.serialization import load_snapshot

snap = load_snapshot("baseline.abi.json")
snap.library, snap.version, snap.functions, snap.types

load_snapshot / save_snapshot / write_snapshot are the public compatibility surface. A consumer doing json.load(f)["functions"] is reading a physical layout it does not own.

Reading an older snapshot

The envelope changed at schema v42. Two separate compatibility questions follow from that, and they have different answers:

Question Answer
Can the current abicheck read a pre-v42 flat .abi.json? Yes. snapshot_from_dict/load_snapshot detect the shape and read it exactly as before. Stored baselines keep working.
Can a pre-v42 abicheck read a current snapshot? No. v42 was bumped specifically so an older reader hard-rejects instead of silently reading every field as absent.
Does an external JSON consumer written against the flat shape still work? No. It sees sections where it expected functions. Use the loader above.

Loading an older snapshot warns, and the warning is the point:

UserWarning: Snapshot schema_version 8 predates this abicheck's schema_version 57:
header_cv_facts_reliable, param_kind_facts_reliable are marked unreliable on this
snapshot, so the affected detectors will decline to trust these stale facts rather
than risk a false positive purely from this tool upgrade.

Loading and re-saving does not upgrade the evidence. A re-saved snapshot carries schema_version: 57 and the current envelope, but the warning persists — it then says so explicitly — and the affected facts stay unestablished. Serialization cannot invent evidence an older extractor never collected. If you need those facts, re-run dump against the artifact. This is why a release baseline should be regenerated when you upgrade abicheck, rather than round-tripped.

Schema version history

schema_version is a single integer, not MAJOR.MINOR. The current value is 57. See abicheck/storage/snapshot_schema_versions.py's SCHEMA_VERSION for the authoritative, up-to-date value and the full per-version comment.

This section is a lookup aid, not reading material — skip it unless you are tracing one specific field back to the version that introduced it.

Per-version history (v12 onward)

Two bumps were not additive. v42 replaced the flat document with the sectioned envelope described in What dump writes today, which breaks an external consumer reading the physical layout. v49 replaced the shape of one field's payload: surface_graph is now the compact graph-table encoding (see "Fields"), so a reader of the old per-entity nodes/edges objects cannot read a v49 graph by ignoring unknown keys. Every other bump added fields without changing the meaning of existing ones — provenance metadata, PE/Mach-O support, build-mode capture, declaration provenance (source_header/origin), embedded build/source evidence, CastXML CV-qualifier reliability, the hybrid AST frontend's per-fact producer map, the resolved AST toolchain identity, (v12) the owner class of a hidden friend (Function.hidden_friend_owner), (v13) the CastXML version-gate outcome (ast_toolchain_supported / ast_toolchain_unsupported_reasons), (v14) extraction-contract fingerprints proving two snapshots were compared under a comparable profile/scope (AbiSnapshot.contract — verdict-blocking: see the compatibility table below), (v15) structured compile-context provenance for the header-AST parse (ast_resolved_standard, ast_cplusplus_macro, ast_compile_args, ast_sysroot), (v16) DWARF-vs-header-AST layout coherence (dwarf_layout_coherence, dwarf_layout_coherence_mismatches — see "Compile-context provenance" below), (v17) which SYCL/DPC++ AST pass a header-AST snapshot was built from (frontend_context_kind), (v18) whether dump's default toolchain/system-header exclusion was applied (dependency_scope, see dumper_scoping.py), (v19) whether the direct-clang backend's deprecated/is_scoped facts are reliable (clang_deprecation_facts_reliable, G31 Phase C — those facts became genuinely populated by the clang backend at this version; see dumper_clang.py), and (v20) whether the direct-clang backend's TypeField.default (default member initializer) facts are reliable (clang_field_initializer_facts_reliable, G31 Phase C — same shape as v19, one fact and one version later; see dumper_clang_expr._field_initializer_value), and (v21) whether the direct-clang backend's RecordType.vtable/ vptr_offset_bits facts are reliable (clang_vtable_facts_reliable, G31 Phase C — the direct-clang backend's vtable/vptr reconstruction; see dumper_clang_vtable.py), and (v22) whether the direct-clang backend's Param.is_restrict facts are reliable (clang_restrict_facts_reliable, G31 Phase C — castxml was that fact's only producer until this version; see dumper_clang._clang_param_is_restrict), and (v23) whether the direct-clang backend's Param.is_va_list facts are reliable (clang_va_list_facts_reliable, G31 Phase C continued — no backend had populated this fact at all before this version, and only for the x86-64 System V spelling; see dumper_clang_qualifiers._clang_param_is_va_list), and (v24) whether the castxml backend's Variable.access facts are reliable (castxml_var_access_facts_reliable, G31 Phase C continued — no backend had populated this fact at all before this version; see dumper_castxml._CastxmlParser._access_level), and (v25) a fully-qualified- name-keyed twin of typedefs (AbiSnapshot.typedefs_qualified, G31 Phase C continued) that closes a bare-name collision between two member typedefs sharing a spelling in different classes/namespaces — needs no reliability flag, since an empty dict degrades identically to "no typedefs at all" for a pre-v25 snapshot, unlike the real-but-wrong scalar defaults v19-v23 above guard against, (v26) Fact[T] siblings for RecordType.bases_fact/ virtual_bases_fact/vtable_fact/vptr_offset_bits_fact and Param.is_va_list_fact (see storage/fact_codec.py), (v27) Function.is_compiler_generated — closes a castxml L4 extractor bug where a compiler-synthesized implicit special member leaked into the source graph as if it were genuine public API; needs no reliability flag, since None (a pre-v27 snapshot's default) degrades cleanly to today's inclusive behavior rather than being misread as "confirmed user-written", (v28) each declaration's entity_id carrier persisted through its own codec (storage/entity_id_codec.py), (v29) AbiSnapshot.surface_graph — the unconditional public-surface/L5 evidence graph (see "Fields" below) persisted through its own to_dict() encoding, not asdict()'s naive recursion (from v49 in the compact graph-table encoding described in "Fields" below), and (v30) RecordType.is_final_fact — the fact/capability registry's (abicheck/model/fact_registry.py) first registered Fact[T] conversion; needs no reliability flag, since is_final's own None already unambiguously means "not captured", (v31) EntityId sidecars for typedefs and constants (AbiSnapshot.typedef_entity_ids/ constant_entity_ids) — additive twins of typedefs_qualified/constants carrying the structural identity both header-AST backends already resolve while walking their own intermediate representation, keyed identically to their partner dict; empty on a DWARF-only snapshot or one predating this field, and (v32) RecordType.is_abstract_fact/ data_size_bits_fact/is_standard_layout_fact/is_trivially_copyable_fact/ qualified_name_fact/source_header_fact — the same registry's next batch of case-(b) conversions (fields already X | None-typed, so the existing "None already unambiguously means not captured" bridge applies directly); qualified_name_fact is the one field in this batch both header backends construct as an explicit Fact.present(...) rather than relying on the generic bridge, since a None qualified name at global scope is itself a confirmed determination, not missing evidence, and (v33) EnumType.qualified_name_fact/source_header_fact — the identical case-(b) pattern applied to EnumType's own twin fields, and (v34) Variable.source_header_fact/alignment_bits_fact/elf_binding_fact — the same case-(b) pattern applied to Variable's own three fields; elf_binding_fact's decoded value is reconstructed as a real SymbolBinding enum member rather than left as the bare JSON string (storage/fact_codec.py's decode_variable_facts), since existing readers unconditionally access .value on it, and (v35) Function.contract_attributes_fact/is_explicit_fact/ is_hidden_friend_fact/source_header_fact/is_variadic_fact/ exception_spec_fact/is_override_fact/hidden_friend_owner_fact/ elf_binding_fact/is_compiler_generated_fact — the same case-(b) pattern applied to Function's own ten remaining fields, closing out this dataclass's Fact[T] conversion; elf_binding_fact gets the identical SymbolBinding reconstruction as Variable.elf_binding_fact, and (v36) AbiSnapshot.ast_resolved_standard_fact — the same case-(b) pattern applied to the last remaining case-(b) field outside the four declaration dataclasses, and (v37) ElfMetadata.dynamic_flags_fact/ has_init_fact/has_fini_fact, PeMetadata.delay_imports_fact, MachoMetadata.rpaths_fact — the same case-(b) pattern applied to the three binary-format metadata blocks, schema-version-driven rather than backend-driven since each block is parsed by exactly one backend; dynamic_flags_fact's decoded value is reconstructed as a real frozenset rather than left as the bare JSON list (snapshot_platform_blocks.elf_from_dict), the identical reconstruction Variable.elf_binding_fact/Function.elf_binding_fact already need for their own non-JSON-native value type, and (v38) AbiSnapshot.semantic_ir — the canonical, backend-independent IR, plus its semantic_ir_conflicts sibling. Encoded by storage/semantic_ir_codec.py as a list of {"occurrence": …, "entity": …} entries, not a JSON object: the IR is keyed by an OccurrenceId dataclass, and rendering that key as a string would lose the typed ScopePath inside it exactly as v28's own entity_id encoding would have. Absent (the key is dropped, not written as null) for a snapshot whose extraction produced no IR, which is every snapshot written by a backend not yet narrowed onto the shared normalizer. semantic_ir_conflicts is sparse the same way, so for such a snapshot a v38 document differs from the v37 one only in the version stamp. Then (v39) TypeField.is_const_fact/ is_volatile_fact/is_mutable_fact — the first case-(a) conversion: unlike every _fact sibling above, these three fields' own values (a plain False) carry no availability signal at all, so the snapshot-level header_cv_facts_reliable flag is what a pre-v39 document's load consults (storage/fact_codec.py's apply_case_a_fact_backfill) before a blanket False may be read as Fact.present(False) rather than Fact.not_collected(), and (v40) Function.deprecated_fact/ Variable.deprecated_fact/RecordType.deprecated_fact/ EnumType.deprecated_fact plus EnumType.is_scoped_fact — the rest of the case-(a) family clang_deprecation_facts_reliable guards, converted the same way (TypeField.deprecated's own sibling landed one version earlier, with the rest of that dataclass's fields), and (v41) Param.is_restrict_fact and Variable.access_fact — the last two fields the fact registry tracked as eligible-but-unconverted, closing that phase's field-by-field conversion. access_fact's decoded value is rebuilt into a real AccessLevel member, the same reconstruction elf_binding_fact needs.

Then (v42) the on-disk wire format itself changed, not just a field: The storage redesign made snapshot_to_json() write storage.sectioned_document's single-file sectioned envelope instead of a flat document. Bumped specifically so a pre-Phase-8 reader (whose own SCHEMA_VERSION was already 41) hits the hard-rejection path below instead of silently reading every top-level field as absent/empty — a same-numbered envelope change would have given that reader no signal at all. This build itself reads the envelope transparently regardless of version, per snapshot_from_dict's own is_sectioned_document check. Then (v43) Variable.is_static persisted — closes the plain-C/extern "C" same-named static-vs-external variable identity collision tu_merge._variable_key's own docstring long documented as a known, accepted limitation; missing on a pre-v43 snapshot loads as False, matching every prior reader's implicit assumption since the field did not exist. Then (v44) AbiSnapshot.header_only persisted — the explicit marker for a snapshot built by the binary-less header-AST dump path (workstream F S1, "Header-only comparison"; no SO_PATH, no --sources/--build-info); missing on a pre-v44 snapshot loads as False, matching every prior snapshot's implicit "this has a binary, or is a pre-existing source-only dump" status. (v46) Function/Variable each gained three Fact[bool] siblings -- declared_in_headers_fact, in_public_contract_fact, binary_exported_fact -- splitting the three independent facts visibility used to conflate (a declaration exists in the parsed headers; it belongs to the promised public contract; a binary symbol is exported). They are absent on a pre-v46 snapshot, where abicheck/model/surface_facts.py derives all three from the stored visibility value as a PARTIAL, diagnostic-stamped reading -- never as a confirmed negative, so a headerless snapshot still reads "declaration not established" rather than "no declaration". Then (v48) AbiSnapshot.excluded_header_matching persisted — which rule the recorded exclusion patterns were matched by: "glob" for the native --exclude-header (fnmatch, plus a */<pattern> try), "abicc" for a descriptor's <skip_headers> under ABICC's own three rule classes (basename, component-boundary path/directory, compiled pattern; written only by the since-removed ABICC compat front end, still read), and the legacy "exact" for a descriptor snapshot written before those rule classes existed, when a descriptor skip was a plain basename-or-path membership test. The same pattern text is not the same scope under any two of them — fftw/fftw.h excludes nothing under "exact" and takes the header under "abicc" — so recording the text alone let the comparability gate accept two snapshots covering different surfaces. Absent on a pre-v48 snapshot, which loads as "unknown" — not as "glob": v47 already recorded a descriptor's exact-matched <skip_headers> alongside the native fnmatch ones, so assuming fnmatch for a mode-less snapshot would let a baseline holding include/foo.h compare clean against a native glob snapshot that excluded a different set of headers. An unrecorded rule is refused rather than guessed. A mode-less snapshot carrying no patterns is unaffected — there is nothing for a rule to have matched. Then (v49) AbiSnapshot.surface_graph is written in the compact graph-table encoding (storage/graph_table_codec.py, see "Fields" below) and holds observed evidence only; a pre-v49 graph still loads.

(v50) The persisted header/L5 source graph (surface_graph, build_source.source_graph) uses the evidence-entity-model invariant-I1 node ids (SourceGraphSummary.schema_version 3): one id per entity, from abicheck/model/graph_entity_identity.py. A C-linkage declaration is keyed on its linker name, an identity-less entity (a castxml constructor/destructor placeholder, an unmangled overload, a type spelling two declarations share) is an explicit unresolved:// node, a flat-path type is keyed on its qualified name, and a new identity_aliases map records second spellings of one entity. A pre-v50 snapshot loads unchanged, graph included; comparing its graph against a v50 one reports the L5 layer as not compared (on the coverage row and as a warning) rather than diffing ids that name entities differently. Re-dump the older side to restore the graph comparison.

(v51) dwarf may carry struct_odr_conflicts/enum_odr_conflicts (name → list of further, layout-distinct definitions another compile unit gave a struct/enum already in structs/enums) and odr_conflicts_observed (the DWARF walk looked for them). The keys are written only when that walk ran, so a snapshot without one encodes exactly as v50. A pre-v51 snapshot loads with odr_conflicts_observed false: its lack of conflicts means "not looked for", never "none". The debug-type join (compare/debug_type_join.py) reports a conflicted name as ambiguous rather than trusting the first definition.

(v56) elf.symbols[].code_hash — a 128-bit BLAKE2b hex digest of an exported function's code bytes, written only when computed. The stripped-binary rename check reads it: equal hashes confirm a same-size rename and break a tie between two same-size candidates. Unequal hashes are not treated as evidence, because a relative call or data reference changes when the same function moves. An absent key means "not hashed".

(v57) semantic_ir function occurrences carry the five signature facts return_type_spelling, parameter_type_spellings, parameter_kinds, ref_qualifier and is_variadic, inside an IR document stamped "version": 3. The spellings are the producer's own, not canonicalized. A pre-v57 document decodes them as not collected, and loading fills them from the snapshot's own functions, so an older baseline compares exactly as before.

(v55) public_header_identifiers_fact — every identifier token the public header set's raw text spells, every preprocessor branch included (comments and literals stripped). --contract public reads it to tell "no public header names this export" from "the parsed branches did not declare it". An evidence-free not_collected fact is omitted on write; an absent key loads as not_collected, on a pre-v55 snapshot as on a current one.

(v54) Each type slot's resolved identity: Function.return_type_identities_fact and Param/Variable/TypeField.type_identities_fact, each a Fact holding the qualified names of the records/enums the slot resolves to. A slot's own spelling is the bare source text (Cache *), which cannot say which of two same-leaf records (ns1::Cache/ns2::Cache) it names. These facts record the compiler's answer, and the public contract domain confirms a break on the reached record with it. castxml writes them; every other producer leaves the key out, so its documents encode exactly as v53. On load, an absent key reads as NOT_COLLECTED in a v54 document and as absent in a pre-v54 one; either way the evaluator keeps its pre-v54 answer.

(v52) extraction_scope — the ownership rules a header-derived snapshot's declarations were classified under (see Extraction scope and ownership below). Declaration lists encode exactly as v51. A pre-v52 snapshot loads with the field absent — unrecorded, never read as "no rules" — and every declaration's owner unknown.

(v47) AbiSnapshot.excluded_header_patterns persisted — the --exclude-header PATTERN values a snapshot was dumped under. The parsed surface is narrower than the operand names and nothing else recorded that, so a stored baseline dumped with an exclusion, compared later against a full dump of the same library, reported every declaration the excluded header carried as appearing out of nowhere. Absent on a pre-v47 snapshot, which loads as the empty tuple — correct for every such snapshot, since the flag did not exist. Before that (v45) Param.kind_fact persisted — closes the last case-(a) field the "field-by-field conversion complete" note missed: neither header-AST backend had ever determined a parameter's indirection kind (value/pointer/reference/rvalue-reference) before this version, so a pre-v45 header-derived snapshot's blanket "value" is a placeholder, not a confirmed reading — AbiSnapshot.param_kind_facts_reliable marks it.

Forward / backward compatibility

abicheck loads a snapshot best-effort and never migrates it in place. The rule is determined entirely by comparing the file's schema_version against the SCHEMA_VERSION the running abicheck supports:

File schema_version Behavior on load
Missing Treated as 1 (the pre-versioning format) and loaded normally.
Older or equal to this build (<= 54) Loaded cleanly. Fields introduced by newer versions are absent and fall back to their defaults (None, empty, or a tri-state None that suppresses the detectors depending on that evidence). No warning.
Newer than this build, and < 14 Loaded best-effort with a UserWarning ("Data may be incomplete or misinterpreted. Upgrade abicheck…"). The load is not aborted — unrecognised keys are ignored and recognised keys are read.
Newer than this build, and >= 14 Hard-rejected — IncompatibleSnapshotSchemaError — instead of warn-and-continue.

Two consequences worth internalising:

  • Reading is version-tolerant in both directions. An older baseline produced by an earlier abicheck loads without error against a newer abicheck; missing fields simply take defaults. This is what makes checked-in baselines durable across tool upgrades.
  • A newer snapshot usually warns rather than fails — but not once a verdict-blocking field exists. Prior to v14 every bump was purely additive, so an older reader can safely ignore a field it doesn't recognise. Starting at v14, AbiSnapshot.contract makes a bump verdict-blocking: a reader that silently dropped it could compare two possibly-incomparable snapshots and produce an ordinary, wrong verdict. snapshot_from_dict therefore hard-rejects (rather than warns-and-loads) any file schema_version that is both newer than the running build's SCHEMA_VERSION and >= 14 (_MIN_SCHEMA_VERSION_REQUIRING_HARD_REJECTION in abicheck/serialization.py) — this only protects readers built from that guard's introduction onward; see the SCHEMA_VERSION history comment for the full explanation of what it can and cannot retroactively protect. Upgrade abicheck to read a newer snapshot faithfully.

Storage encoding

Everything above describes the logical snapshot — the decoded JSON payload. On disk, that payload may be stored plain, gzip-compressed, or zstd-compressed; compression is a storage/transport envelope around the identical JSON, never a new schema, and never changes schema_version, AbiSnapshot.contract, evidence depth, build_source, or the verdict two snapshots produce when compared.

Encoding Canonical suffix Notes
plain .abicheck.json / .abi.json debugging, small Git-reviewable snapshots
gzip .abicheck.json.gz / .abi.json.gz universal interoperability
zstd .abicheck.json.zst / .abi.json.zst preferred for baseline/release/cache storage

compare and the Python API (abicheck.serialization.load_snapshot) all read every encoding transparently — detected from magic bytes, not just the filename suffix. abicheck dump produces one: it infers the encoding from -o/--output's suffix by default (--compression auto), or accepts an explicit --compression {none,gzip,zstd}; write_snapshot is the Python API equivalent for writing. The full storage-envelope model (determinism, atomic writes, decompression limits, and what's still deferred).

Sectioned packaging

Orthogonal to compression: the envelope shown in What dump writes today is this page's field set packaged into named, independently versioned sections (binary/declarations/types/layout/debug/build/graph/ provenance/semantic_ir). Every dump/write_snapshot invocation writes it; no flag selects it.

snapshot_from_dict/load_snapshot unwrap it before any field below is read, and an older flat .abi.json is still read exactly as it always was — the field-level contract below is unaffected either way; only the outermost envelope differs. The directory-backed content-addressed package is a separate, advanced storage shape with its own page: see Project Snapshot Format.


Field reference

The keys below are the AbiSnapshot model's fields (abicheck/model/snapshot.py), as encoded by the serializer. In the current envelope they live inside the matching sections[...] payload; in a pre-v42 flat document they are top-level. load_snapshot presents them identically either way, which is why this reference is written against the model rather than against either physical layout. Optional keys are omitted or null when there is no data (for example, a pure-ELF dump has no dwarf or build_source).

Identity and provenance

Key Type Meaning
schema_version int Snapshot format version (currently 57).
library string Library identity, e.g. libfoo.so.1.
version string Library version string, e.g. 1.2.3.
source_path string | null Original path the snapshot was taken from.
platform string | null elf, pe, macho, or null.
language_profile string | null c, cpp, sycl, or null.
git_commit string | null Git SHA captured at dump time.
git_tag string | null Git tag (e.g. v2.0.0), supplied or auto-detected.
created_at string | null ISO 8601 timestamp set at dump time.
build_id string | null Opaque CI identifier (run ID, build number).
contract object | null Extraction-contract fingerprints (schema v14, verdict-blocking — see "Forward / backward compatibility" above): profile_fingerprint/scope_fingerprint plus their named resolved sub-inputs, proving two snapshots were extracted under a comparable profile/scope. null when no producer populated it yet.
dependency_scope string | null (schema v18) "filtered" when the toolchain/system-header exclusion (dumper_scoping.py) was applied, "full" when opted out via --include-system-declarations. Every front end filters by default (include_dependencies=False): the dump/compare CLI, service.run_dump and InputSpec all share that default since the 0.6 defaults-alignment pass — before it, a typed-API caller that omitted the field got the unfiltered surface while the identical CLI invocation got the filtered one, and the two were not comparable. null on any pre-v18 snapshot or any snapshot with no header-derived declarations. comparability.check_contracts_comparable raises ScopeMismatchError only when BOTH sides carry an explicit, non-null value and they differ — null is deliberately NOT treated as "full" (an ordinary pre-v18 baseline is usually already-filtered content that simply predates this tag; assuming "full" for it would spuriously flag the routine "compare a cached baseline against a fresh dump" workflow), so a genuinely ambiguous untagged snapshot is left unchecked on this axis rather than guessed at.
header_only boolean (schema v44) true only for a snapshot built by the binary-less header-AST dump path — either dump -H api.h (no SO_PATH/--sources/--build-info) or a pathless dump --dump-manifest m.yaml naming real header-AST roots (workstream F S1, "Header-only comparison"). Explicit, not inferred from platform/from_headers: a pre-existing --sources/--build-info source-only dump also has platform: null, but carries no header-AST declarations at all. public_header_dirs alone does NOT select this path — it is a declaration-provenance (public-vs-internal) classifier only, never a source of headers to parse; a binary-less request naming only public_header_dirs (no header file, no manifest) is rejected before extraction. false (the default) for every snapshot predating this field and every ordinary binary dump.

Compile-context provenance (schema v15, header-AST parses only)

Populated only when the snapshot came from a header-AST parse (from_headers true); null/empty on a DWARF/symbols-only or binary-only snapshot, and on any pre-v15 snapshot. ast_compile_args and ast_sysroot are redacted via the same RedactionPolicy every L3 build-evidence adapter applies (secret- looking -D values and absolute home-prefixed paths are stripped/normalized before persistence — see abicheck/buildsource/redaction.py).

Key Type Default Meaning
ast_resolved_standard string | null null The C/C++ standard actually used for the header parse: an explicit -std=/--std=//std: value verbatim, or "gnu++20" when the requires/concept heuristic forced it. null means the frontend's own unpinned default was used (never guessed at).
ast_cplusplus_macro string | null null The standard-mandated __cplusplus literal for ast_resolved_standard (e.g. "201703L" for "gnu++17"), looked up from a static ISO-standard table. null when ast_resolved_standard is unset or not a recognized C++ edition.
ast_compile_args array of strings [] The ordered extra compiler arguments passed to the header frontend (--compiler-option tokens, then a shlex-split composed-flags string derived from the compilation database --build-info resolves to), redacted.
ast_sysroot string | null null The --sysroot passed to the header frontend, if any, redacted.

ast_toolchain (dict[str, str], populated since schema v9) carries the exact tool identity behind the header-AST parse. It is untyped/free-form — new keys are additive and never require a schema bump — but these keys are stable and machine-checked by tests/test_tool_identity.py/ tests/test_castxml_policy.py:

Key Meaning
selected / compiler_selected The exact frontend/host-compiler executable path selected from PATH (or an explicit --compiler).
realpath / compiler_realpath The same path with symlinks resolved.
sha256 / compiler_sha256 SHA-256 of the executable's file contents, so a same-version binary rebuild/repackage still changes provenance.
version / compiler_version The raw, bounded --version transcript for that exact executable revision.
target_triple / compiler_target_triple The <tool> -dumpmachine output for that executable (GCC/G++/Clang/Clang++ only — omitted, not empty, for a tool that doesn't support the flag, e.g. castxml itself or MSVC cl.exe).
castxml_version CastXML's own release version (e.g. "0.7.0"), parsed from version — castxml-producer snapshots only.
castxml_bundled_clang_version The bundled/linked Clang's major.minor (e.g. "18.1"), parsed from version — castxml-producer snapshots only. Kept separate from castxml_version since the two floors (MIN_CASTXML, MIN_CASTXML_CLANG_MAJOR in castxml_policy.py) are independently enforced by the version gate.

A hybrid snapshot (ast_producer == "hybrid") namespaces every key from both runs instead of picking one — castxml_selected, castxml_version, castxml_castxml_version, clang_selected, clang_target_triple, and so on (dumper_hybrid.py's merge is a generic castxml_/clang_-prefixed dict union, so a key that already started with castxml_ on the castxml side is not special-cased).

DWARF-vs-header-AST layout coherence (schema v16)

The clang L2 header backend is layout-blind (no size_bits/alignment_bits/ field offset_bits) — when the binary being dumped also carries DWARF debug info, dumper_layout_backfill.backfill_dwarf_layout() backfills that layout from the same binary's DWARF, but only for a record it can corroborate as the same declaration (matching name, kind, and field/base overlap — see that function's docstring for the exact rules). These two fields make that corroboration outcome visible instead of silent; they never change what gets backfilled, only report on it.

Key Type Default Meaning
dwarf_layout_coherence string | null null One of "matched" (every record eligible for backfill was corroborated, or none needed it), "partial" (some corroborated, some had no DWARF candidate at all — benign, e.g. declared-but-never-instantiated), "mismatch" (at least one record found a uniquely-named DWARF candidate but the two disagreed — backfill already refused to merge that record's layout), or "unavailable" (the clang backend ran but the binary carried no usable DWARF at all). null on any snapshot not built via the clang L2 backend (a castxml snapshot computes layout directly — not a coherence question) and on any pre-v16 snapshot.
dwarf_layout_coherence_mismatches array of strings [] Header record names backfill found a uniquely-named DWARF candidate for but rejected as uncorroborated — populated only when dwarf_layout_coherence == "mismatch".

SYCL/DPC++ frontend context (schema v17, header-AST parses only)

Key Type Default Meaning
frontend_context_kind string | null null Which AST pass ("host" or "device") this header-AST snapshot's clang backend selected via --frontend-context (sycl_context.py). null on any non-SYCL/DPC++ invocation and on any pre-v17 snapshot.

| public_header_identifiers_fact | object | absent | absent | v55. A Fact: status, value (sorted identifier list when present), diagnostics, producer ("public_header_text"). failed when a named header is missing or unreadable, unsupported when a header uses ## token pasting. |

Extraction scope and ownership (schema v52)

Written for every snapshot a run extracts from headers; absent on a binary- or debug-only snapshot and on any snapshot a run loaded (a stored baseline keeps the scope it was dumped under).

Key Type Meaning
extraction_scope.ownership_rules object target_roots, dependencies (name, header_roots), private_headers, private_namespaces, dependency_evidence. Roots are POSIX paths relative to the project root when they lie under it, absolute otherwise.
extraction_scope.dependency_evidence string What was kept of dependency declarations. Always "full" today.
extraction_scope.prefilter object | null Reserved for a frontend prefilter; always null today.
extraction_scope.fingerprint string sha256: over the three keys above, canonicalized (sorted, de-duplicated). What comparability and the configuration digest (surface.ownership) compare.
extraction_scope.diagnostics array The classifier's namespace-mismatch diagnostics (a target file declaring into a dependency's namespace). Omitted when empty.
extraction_scope.entity_ownership.decisions array Interned [owner, contract, rule_id] triples.
extraction_scope.entity_ownership.<functions\|variables\|types\|enums> array of int One index into decisions per declaration of that list, in list order; -1 for a declaration that was not classified. A list whose length disagrees with the snapshot's own list is ignored on load (its declarations stay unclassified) rather than misattributed.

owner is target, dependency:<name>, toolchain or unresolved; contract is public, private, external or unresolved. An unresolved owner is a real answer (no root claims the file); an unclassified declaration has no decision at all, and readers treat it as unknown.

ABI surface

Key Type Meaning
functions array Exported functions (name, mangled name, return type, params, virtuality, access, provenance). Since v54 a castxml-dumped function's return type, each parameter (and each variable's type, each record field) may carry a *type_identities_fact: the qualified record/enum names that slot resolves to, which its bare spelling cannot express.
variables array Exported global/static variables.
types array Records (struct/class/union) with fields, bases, vtable, and layout descriptors.
enums array Enumerations with members and underlying type.
typedefs object Typedef name → underlying type. Bare-name-keyed; two distinct member typedefs sharing a spelling in different classes/namespaces collide onto one key (see typedefs_qualified).
typedefs_qualified object Fully-qualified-name-keyed twin of typedefs (schema v25) — collision-free. Empty for a pre-v25 snapshot or one produced without per-class qualified typedef scoping (e.g. DWARF-only).
constants object Preprocessor/compile-time constants (qualified name → value).
typedef_entity_ids object EntityId sidecar for typedefs_qualified (schema v31), keyed identically — the typed ScopePath/kind/leaf-name identity a dict[str, str] cannot carry on a declaration object. Empty for a pre-v31 or DWARF-only snapshot.
constant_entity_ids object EntityId sidecar for constants (schema v31), keyed identically. Empty for a pre-v31 or DWARF-only snapshot.

Evidence-tier and mode flags

Key Type Meaning
elf_only_mode bool True when dumped without headers (all functions carry ELF-only provenance).
from_headers bool True when the surface was parsed from public headers (drives the header-aware evidence tier). Omitted from the file when it was only inferred on load, so a reload re-runs the same inference.
scope_fallback string | null Public-scope fallback marker.
parsed_with_build_context bool True when parsed with build-context evidence.

Platform and debug metadata (optional)

Key Type Meaning
elf object | null ELF metadata: SONAME, DT_NEEDED, version defs/reqs, symbols, imports, hardening flags.
pe object | null PE/COFF metadata (Windows DLL exports, machine, characteristics).
macho object | null Mach-O metadata (dylib exports, CPU slices, install name).
dwarf object | null DWARF struct/enum layout (v51: plus ODR conflicts, see above).
dwarf_advanced object | null Toolchain, calling conventions, value-ABI traits.
sycl object | null SYCL plugin-interface metadata.
dependency_info object | null Resolved dependency graph (nodes, edges, unresolved).
build_mode object | null Normalized compiler/stdlib/standard capture (ADR build-mode work). No dump path writes it today; it is read back only from a document that carries one (see known-gaps.md).

Embedded build/source evidence (optional)

Key Type Meaning
build_source_pack object | null Reference to an out-of-band build/source pack. Older snapshots may store this under the legacy key evidence_pack, which the loader still reads.
build_source object | null Inline-embedded build/source facts for single-artifact workflows. Omitted when nothing was embedded.
surface_graph object | omitted (v29) The unconditional public-surface/L5 evidence graph — never gated on build_source, unlike the row above. The key is omitted entirely (not written as null) for a snapshot predating this field, a binary-only snapshot, or one whose headers were never parsed — encode_surface_graph() pops the key rather than writing a null placeholder. When build_source.source_graph is the identical object, it is omitted from build_source's own encoding rather than written twice; the loader restores that alias on read. From v49 the value is storage/graph_table_codec.py's compact encoding ("encoding": "graph-table/1"): one strings table, nodes/edges objects of equal-length index columns (id/src/dst, kind, label for nodes, and facts), and deduplicated attrs and facts tables. Only observed evidence is written: nothing the loader rederives (indexes, graph_id, finalize-owned coverage counts, and each entity's resolved/conflicts/occurrences/attrs/provenance/confidence, all rebuilt from its facts), and no fact from a producer that projects the snapshot's own records (the public-surface builder's declaration/type/symbol/binary_symbol/debug_type nodes and declares/references/declares_linker_name/exports/debug_type_of edges), which a reader rebuilds on demand. A value without encoding is the pre-v49 per-entity form and still loads. The graph is decoded on first access, not at load (a corrupt value raises then).
build_context_defines array of strings The build's active -D macro set, harvested from a compile database. Empty when no compile database was supplied.
conditional_fields object {type: {field: {guard, type, is_bitfield, ...}}} registry of record fields guarded by a single positive #ifdef/#if defined(...), including fields a context-free header parse pruned from types[].fields. Feeds the build-context reconciliation diff pass, which runs unconditionally whenever both snapshots carry this field (one-comparison-product.md Phase 7i); empty when no compile database was supplied at dump time.

Internal cache fields on the model (_func_by_mangled, _var_by_mangled, _type_by_name) and the runtime-only from_headers_inferred qualifier are never serialized.


Two contracts: snapshot vs report

schema_version and report_schema_version are different fields on different files:

Snapshot (dump) Comparison report (compare -o json=-)
Version field schema_version report_schema_version
Type integer (currently 57) string MAJOR.MINOR (e.g. 1.0)
Describes one library's ABI surface the diff between two snapshots

A snapshot has no report_schema_version, and a report has no schema_version; the two version numbers evolve independently. For the report contract and its stability policy, see Output Formats.


Stability guidance

  • Check baselines into version control. A saved .abi.json is the intended input to compare; storing one per release lets CI diff each build against the last shipped ABI. See Baseline Management.
  • Older baselines stay readable. Because loading fills missing newer fields with defaults, a baseline written by an earlier abicheck compares correctly against a live binary dumped by a newer one — no regeneration required for a routine tool upgrade.
  • Regenerate when you want new evidence. Fields added in a newer schema_version (e.g. build-mode or embedded source evidence) are only present in freshly-dumped snapshots. Re-dump the baseline to benefit from detectors that rely on that evidence.
  • Pin the abicheck version in CI if a UserWarning about a newer schema_version would be treated as an error in your pipeline.

See also


The build/source pack envelope

  • The pack is content-addressed and versioned independently (evidence_pack_version) from the ABI snapshot schema, so it never bloats an ordinary dump. The snapshot stores only a lightweight evidence_pack reference (content hash + coverage summary); old readers ignore it.
  • Every extractor writes both a raw artifact (under raw/, for provenance/debugging) and an abicheck-owned normalized fact model (e.g. build/build_evidence.json). Only normalized facts feed comparison and the content hash.
  • Command lines and paths are redacted (home prefixes, secret-looking -D macros) before they are persisted.

See Source & Build Data for the full model.