Skip to content

Graph Coverage & Negative Evidence

The optional embedded L5 source graph can prove a positive: "public entry X reaches internal declaration Y" (a DECL_CALLS_DECL/DECL_REFERENCES_DECL/... edge exists). It is much harder to trust the graph for a negative: "no public entry reaches Y" is only true if the collection pass that built the graph actually looked everywhere it needed to. This page explains why absence of an edge is not always proof of absence of a dependency, and how abicheck's suppression gate reflects that honestly instead of guessing.

Why an absent edge isn't automatically proof

SourceGraphSummary (the in-memory L5 graph abicheck's snapshot carries) records, per extractor pass, whether its own coverage was complete: the pass either ran over the full project scope with no errors (an edge family it covers is trustworthy for presence and absence), ran over a narrowed scope (a found edge is real; a missing one proves nothing about what it never looked at), or degraded on collection errors (the edges it found are real; the ones it did not are an untracked gap). The three pass states and the fields that carry them are documented in Source Graph Schema § Coverage pass states.

Two collection strategies commonly produce exactly this shape:

  • Header-only collection (the L2 header-only graph, attached automatically whenever a supported dump/compare run has header evidence at --depth headers or deeper — including runs that also provide real build/source evidence, not just header-only ones) sees declarations and signatures but never a function body, so it cannot see a DECL_CALLS_DECL edge a public inline function's body creates into an internal specialization — the graph is real, just structurally unable to answer that question.

This graph is always built from a clang whole-translation-unit AST, whatever --ast-frontend selected. A castxml-frontend run therefore still invokes clang -ast-dump=json once per side for the graph (service_header_graph_attach.py) and pays its cost even though the snapshot's declarations came from castxml. Without a usable clang the graph is simply not attached — the dump still succeeds, with the L5 header graph absent rather than empty. A warm run skips that cost through the header-graph projection sidecar stored next to the clang AST cache entry; a cold castxml run pays the full clang parse. Budget memory for it the way you would for a clang-frontend run: the document is compacted and ASCII-escaped when cached, but the first parse still reads clang's full output. - A collector-upgrade (old snapshot dumped header-only, new snapshot with a real --build-info compile database) is not a "new dependency appeared" signal — it is the same project seen through two different lenses. abicheck's source-graph diff findings account for this asymmetry rather than reporting phantom PUBLIC_API_INTERNAL_DEPENDENCY_ADDED churn every time collection tooling improves.

Tri-state reachability

Because of this, Change.reachability_state is not the boolean Change.public_reachable alone — it is one of three states:

State Meaning
reachable (PROVEN_REACHABLE) A walk positively found a path from the public surface to this change.
unreachable (PROVEN_UNREACHABLE) A walk examined this change and found no path — and the walk's own coverage was trustworthy for that verdict (the type-layout walk always is; the call-graph walk is, as long as it wasn't the only signal available while flagged narrowed/degraded).
unknown No walk reached a verdict at all, or the only walk that could have was itself narrowed/degraded coverage — the honest "we don't know" answer.

MarkReachability (the pipeline step that computes this, before suppression runs) sets this alongside the existing public_reachable boolean.

What this means for suppression

The suppression reachability: unreachable-only default (the common case for a broad namespace/source_location rule) keeps its original, boolean-only semantics for backward compatibility: it treats unreachable and unknown identically, exactly as it always has. That is deliberately unchanged — most projects have no embedded L5 graph at all, and the type-layout walk (which has no coverage caveat) already dominates that common case.

For a project that does rely on L5 graph evidence and wants a suppression rule to require actual proof, opt into the stricter gate with reachability: proven-unreachable-only — see Suppressions § Proven vs. unknown reachability for the rule syntax and the suppression_reachability_unknown diagnostic it produces when coverage isn't good enough to prove a match.

Typed absence across every relationship

The same rule applies beyond the L5 graph. Every relationship abicheck relates evidence through answers one of three things for a given subject: present, proven absent, or unknown (compare/edge_query.py, evidence-entity-model invariant I4). "Proven absent" needs the producer of that relationship to have covered the scope it was asked about:

Relationship Absence is proven when Otherwise unknown, for example
exports (a declaration and the export table) every export table the snapshot owes was read no table captured; a platform block that was never parsed
exports (an export and the declarations) the library's own headers were parsed a binary-only or DWARF-only dump; a toolchain export whose system headers were filtered out
debug_type_of (a header type and the debug info) the binary carries debug info a stripped binary
declares / references the header AST ran no header AST; a type spelling that names several types
L5 source-graph edges the pass covered the scope (see above) a narrowed, degraded, or header-only body-blind pass

The compare report states this per side in its Relationship coverage section (JSON: edge_coverage). It lists only the relationships whose absence is unknown, with the producer status behind each, and says so explicitly when every producer covered its scope.

Migration: header-graph is now default-on

Before G29 Phase A, the L2 header-only graph (and its COMPILE_UNIT_INCLUDES_FILE include-file extension) only got built if you explicitly passed --header-graph/--header-graph-includes to dump or compare. As of G29 Phase A, --depth headers (the default depth) always builds it automatically — there is no flag to remember and nothing to opt into. The two flags still exist but are hidden, deprecated no-ops kept only for a transition window before removal.

This doesn't change how you should reason about completeness: whether the graph saw everything it needed to is still reported through the coverage fields described above (the full, narrowed and degraded pass states and the tri-state reachability status), never through whether a flag was passed. A header-only collection degrades the same way it always did (declarations and signatures only, no function bodies) — it is just no longer possible to accidentally run without it when depth headers or deeper evidence is available.

Canonical entity identity and rename/move reconciliation

The header-only graph and a build-integrated graph can identify the same declaration differently depending on which pass saw it first. Without any reconciliation, an old/new comparison sees a renamed internal declaration as an unrelated node removal plus an unrelated node addition — a reader has to notice the two facts independently and infer by hand that they describe the same entity.

abicheck.buildsource.entity_identity computes a canonical identity for every graph declaration/type node, in preference order:

  1. canonical — a compiler-provided stable identity (a clang USR, when a producer supplies one) or a real Itanium/MSVC mangled name.
  2. normalized — a fully-qualified semantic signature (qualified name + kind + arity/parameter types) when no mangling is available.
  3. reduced — a source-relative identity (file + enclosing scope + name, always an alias, never the primary key) or, when nothing else is available at all, a clearly-marked synthetic:sha256:... fallback.

abicheck.buildsource.graph_reconcile then reconciles an old/new graph diff's added/removed nodes using that identity: an exact canonical-id match, an exact (bidirectionally-unambiguous) alias match, or — as a last resort — a match on unique structural position when even the qualified name changed. Ambiguous evidence never resolves to a guess: if two candidates share the same alias or structural position, neither is reconciled — both stay a plain add/remove, exactly as before reconciliation existed. A match produces a declaration_renamed, declaration_moved, or declaration_identity_reconciled finding — pure enrichment, RISK-tier, never overriding or suppressing an artifact-proven finding elsewhere in the comparison (the same authority rule as everywhere else on this page).

See examples/case194_header_graph_rename_reconciled/ examples/case195_header_graph_ambiguous_rename_not_reconciled for a reconciled rename and its deliberately-unreconciled ambiguous counterpart.


Ladder: ← Source & Build Data · Concepts c3 · Internals · Unified Impact Assessment →