Architecture Decision Records¶
Status field convention¶
A documentation-lifecycle review (2026-07) found that a single "Status" word
per ADR routinely conflated three independent facts — whether the decision
was accepted, whether it's implemented, and whether that implementation
claim has been verified against the current code — which is exactly how
several ADRs went stale silently (e.g. ADR-022 said "implemented" when only
one of its four backends had shipped). Introducing a separate structured
frontmatter schema (decision_status/implementation_status/
verification_status fields, as some ADR tooling does) was considered and
deferred: every ADR here uses a single plain-prose **Status:** line, and
retrofitting 40+ files with fabricated metadata (owners, PR numbers, "last
verified" dates for documents nobody actually re-audited line-by-line) would
trade one inaccuracy for another. Instead, the convention going forward
is to keep encoding the same three facts in that one line, explicitly:
**Status:** <decision: Proposed | Accepted | Superseded by ADR-NNN | Deprecated>
— <implementation: implemented | partially implemented (name what's missing)
| not implemented>. <optional amendment note: what's stale, what superseded
it, where the current behavior is documented instead>
The **Verified:** receipt (opt-in)¶
A 2026-08 status review found the residual half of the same problem, which no amount of status prose solves: ADR-049's Status said its shadow evaluator "is not called from any pipeline stage" and that nothing was wired into the CLI or reports, and five merged PRs later that was simply false. The status line was internally consistent, agreed with the index row, and was wrong — because the code moved and nobody re-read it. Comparing two documents cannot catch that; only re-reading the code can.
So a second, optional metadata line records when someone last did:
placed directly after the Status paragraph. It means "a maintainer checked this ADR's Status claims against the tree at that commit" — nothing more.
Name a commit on the default branch, not the branch you're writing the
receipt on. A branch commit stops existing once the PR squash-merges, and
because the ai-readiness job checks out full history, an unresolvable sha
is an error rather than a tolerated skip — so a receipt anchored to a PR
commit passes on that PR and then fails on main permanently. Since a PR
that adds a receipt normally doesn't change the code being attested, the
right sha is the main commit the branch is based on. adr-status-sync
enforces this. It
is opt-in precisely so it stays truthful: an ADR without one has simply never
been checked, which is the honest state of most of the files below and is not
an error. Adding one you didn't actually perform is worse than having none.
Name a module in full if you want it watched — abicheck/foo.py, not
_foo.py after a sibling. Family shorthand is deliberately not guessed at:
the gate would have to infer whether the family is replaced or appended, and
a wrong inference either watches an unrelated file or silently watches
nothing. A full path is exact, auditable from the Status line alone, and
keeps working after the file is renamed or deleted (which is when a claim is
most likely to have gone stale).
What it buys is a tripwire rather than a promise. adr-status-sync
(implemented in scripts/adr_status_sync.py, registered as a check by
scripts/check_ai_readiness.py) reads the first-party module paths the Status
paragraph names and warns when commits after the recorded sha have touched
any of them — i.e. when the code a claim describes has moved out from under
it. That's a WARN, not an ERROR: a changed module doesn't prove the claim went
stale, it proves nobody has re-read it since. Move the line forward when you
re-check, or correct the claim.
The same check ERRORs on a flat contradiction between an ADR's own Status and its row in the table below — one side claiming nothing is implemented while the other claims something is (which is exactly how ADR-056's row went stale). It deliberately does not require the two to paraphrase each other: the index cell is an abridgement of a paragraph that often runs a dozen lines, and an earlier prototype that compared them more strictly flagged 15 of 56 ADRs, nearly all false positives.
Good examples already in this table: ADR-022 ("partially implemented" +
naming exactly which backend shipped), ADR-037 (distinguishes "the contract
is implemented" from "enforcement is advisory until 1.0"), ADR-025 ("Proposed,
but substantially implemented/generalized elsewhere" + pointers to the ADRs
that absorbed it). When you touch an ADR and confirm a claim against current
code, update its **Status:** line rather than leaving the reader to infer
freshness from the file's git history. scripts/check_usecase_docs_sync.py
and the adr-index-nav-sync AI-readiness check keep the registry and nav
mechanically honest; the status line itself is still maintainer-verified
prose, not generated — treat a status claim you haven't personally checked
against the code as unverified, regardless of how confident it reads.
| # | Title | Status |
|---|---|---|
| 001 | Technology Stack — Python + pyelftools + castxml | Accepted — implemented, substantially amended |
| 002 | Multi-binary / release compare UX and architecture | Accepted — implemented |
| 003 | Data Source Architecture — checks, instruments, and binary types (+ exploratory binary fingerprint extension) | Accepted — implemented; conceptually extended by the L0–L5 model (ADR-028–031, 041) |
| 004 | Report Filtering, Deduplication, and Leaf-Change Mode | Accepted — implemented |
| 005 | Application Compatibility Checking | Accepted — implemented |
| 006 | Package-Level Comparison | Accepted — implemented |
| 007 | BTF and CTF Debug Format Support | Accepted — implemented |
| 008 | Full-Stack Dependency Validation | Accepted — implemented |
| 009 | Verdict System and Exit Code Contract | Accepted — implemented |
| 010 | Policy Profile System | Accepted — implemented |
| 011 | ABI Change Classification Taxonomy | Accepted — implemented |
| 012 | ABICC Drop-In Compatibility Layer | Deprecated — Retired: compat command removed |
| 013 | Suppression System Design | Accepted — implemented |
| 014 | Output Format Strategy | Accepted — implemented |
| 015 | Snapshot Serialization and Schema Versioning | Accepted — implemented |
| 016 | Three-Tier Visibility Model | Accepted — implemented; extended by ADR-024's two-axis surface model |
| 017 | GitHub Action Design | Accepted — implemented |
| 018 | Cross-Platform Binary Format Support | Accepted — implemented |
| 019 | Testing Strategy and Parity Validation | Accepted — implemented |
| 020a | Build-Context Aware Header Extraction | Accepted — implemented |
| 020b | SYCL and Heterogeneous Computing Stack Support | Accepted — implemented |
| 021a | Debug Artifact Resolution Subsystem | Accepted — implemented |
| 021b | MCP Security Model | Deprecated — Retired: MCP interface removed |
| 022 | Baseline Registry and Snapshot Distribution | Superseded by ADR-043 D4 — not implemented; the filesystem backend and baseline command group that once shipped were deleted, and D4 records recreating a registry as an explicit non-goal, not a deferred rebuild |
| 023 | Bundle-Aware Multi-Binary ABI Analysis | Accepted — implemented |
| 024 | Public ABI Surface Resolution and False-Positive Traceability | Accepted — implemented |
| 025 | PR-Diff-Aware ABI Evaluation (Source Diff as Trigger and Localizer) | Proposed; D1–D3 absorbed by ADR-033/035, D4 still future work |
| 026 | Source-Only Changes and the Evidence-Tier Boundary | Accepted — substantially superseded by ADR-028/030/035/038 (its "no embedded Clang" conclusion was reversed) |
| 027 | API Surface Intelligence — Structure Metrics, Idiom Detection, Cross-Library Reasoning, Pattern-Aware Verdicts | Accepted — Phases 0-5 implemented; --pattern-verdicts default-on flip deferred pending release-cycle FP-rate/parity validation |
| 028 | Optional Source and Build Evidence Pack Architecture | Accepted — implemented |
| 029 | Build Graph and Toolchain Context Capture | Accepted — implemented |
| 030 | Source ABI Replay and Linked Source Surface | Accepted — implemented |
| 031 | Source and Implementation Graph Augmentation | Accepted — implemented |
| 032 | Evidence Extractor Plugin Interface and Security Model | Accepted — implemented |
| 033 | CI Rollout, Performance, Caching, and Validation Strategy | Accepted — implemented |
| 034 | Managed-Runtime and Non-C ABI Frontends | Proposed |
| 035 | PR-Tier Source Intelligence and Cross-Source Validation | Accepted — implemented (G19, D1–D10) |
| 036 | Report view-model and canonical report severity | Accepted — core implemented (Increments 1-2); Increment 3 (routing html_report.py/pr_comment.py through ReportModel) remains optional cleanup |
| 037 | CLI Interface Contract, Configuration Balance, and Extension Policy | Accepted — implemented (G22) |
| 038 | Working With Sources — Full-Scan and Two Build-Injection Flows | Accepted — implemented |
| 039 | Build-Context Reconciliation of Context-Free Header-Parse Artifacts | Accepted — implemented |
| 040 | compare Surface Reduction — Side-Aware Flags, Config Demotion, Run Profiles |
Accepted — phased implementation substantially complete (Phase A landed then reversed and removed by ADR-068 D5, --profile is gone; Phase B landed, Phase C Lever-1 landed except the ast-frontend carve-out, Phase D landed as a constraint-aware subset) |
| 041 | Compiler-Facts Semantic Impact Graph — Roadmap and P0 Slice | Accepted — P0 slices 1-4, the header-only-graph addendum, and P1 items 1-5 implemented; remainder is roadmap, not a shipping commitment |
| 042 | Formal separation of CompatibilityDecision and GateDecision | Accepted — implemented for JSON/SARIF/compare-release gate summaries and html_report.py's CI Gate card; mcp_server.py/junit_report.py still compute an exit code inline in places |
| 043 | Pre-1.0 CLI Surface Reset — Root Command Collapse, Depth Ladder Narrowing, and Dry-Run Unification | Accepted — implemented |
| 044 | Reachability-Aware Suppression and the Effective Public ABI | Accepted — P0, P1, and P2 all implemented (see ADR for exact scope) |
| 045 | Identity-Based Old/New Entity Matching | Accepted — implemented for RecordType and EnumType |
| 046 | Source Graph Identity v2 — USR-Based Entity Resolution and Evidence-Preserving Merge | Accepted — D1-D6 all implemented, each to the documented scope (D4 is a deliberately scoped subset of the originally sketched full rewrite; see ADR for exactly what's covered per decision) |
| 047 | GitHub Actions Integration Model — Project Lifecycle Over Aggregate-Centric Design | Accepted — substantially implemented (P0 and the main P1 lifecycle implemented; P2 partially implemented, see ADR); see G30 |
| 048 | Canonical Entity Identity and Graph Reconciliation (G31 Phase B) | Accepted — implemented |
| 049 | Contract Relevance and Compatibility Configuration | Accepted (2026-07-26) — Phases 0-6 implemented (vocabulary, typed config/resolver/packs, finding identity, evaluator in all three contract domains, persisted evidence context + replay, one resolved config per front end, --contract domain selection); Phase 7 mostly implemented — coverage exit landed for compare, scan --against, the release fan-out, and aggregate (workflows/aggregate/fold.py's own contract_coverage_exit), and (2026-09-02) the contract decision is authoritative (not shadow) whenever --contract is given, via contract_pipeline.py's ContractEvaluationStage running before the verdict; the default flip (applying --contract without it being asked for) is not yet made — public-contract-default.md's Phase 6 relevance defects that could lose a known public break are both closed (the template-instantiated-parameter seed mismatch by threading directly-referenced stdlib spellings; the ambiguous_namespaced_leaf identity gap on 2026-10-01 by schema v54's castxml-captured slot identities), leaving one explained, evidence-limited public loss (the spelling-only shape the clang/DWARF backends still produce) and the package/real_binaries lanes covered by integration tests rather than the always-on measurement — a coverage bound awaiting acceptance before the flip |
| 050 | Comparability Contract — Profile/Scope Fingerprints and the Multi-TU Manifest | Accepted — implemented (Phase 0 and Phases A-E; D1-D6); see G32 |
| 051 | Documentation Operational Model (Ownership Registry + Docs-Contract Gate) | Accepted — Stages 1-4 implemented; Stage 5 explicitly deferred |
| 052 | Unified Impact Assessment Model (G29 Phase 3, slices 1-11) | Accepted — slices 1-11 implemented |
| 053 | TU → Link-Unit → DSO Source-Evidence Attribution | Accepted — implemented (core algorithm + validator; CLI/Action pipeline wiring deferred, see D5) |
| 054 | CLI Project-Integration Surface Consolidation | Accepted — implemented |
| 055 | Typed Request/Result Completeness and a Schema-Version Registry | Accepted — implemented (D1-D4), including D1's structural half: the CLI and typed API share one input resolution. D4 (MCP dedup) and other MCP-specific claims are historical — the MCP server was later removed |
| 056 | Multi-Artifact / Library-Set scan |
Superseded by 068 (2026-09-06) — the scan command it extends is being retired. Never formally signed off; the slice that shipped ahead of sign-off (Phases 1-4's engine/detector/CLI/Action work, see G35) is deleted with the command, and its deferred items (example catalog, --dry-run estimator, the never-started member-identity manifest) are cancelled rather than carried. The capability survives: an old-side-less N-library audit is compare --no-baseline DIR over 065's acquisition/selection model, and the declared provider manifest becomes project configuration |
| 057 | Consumer Graph and the Consumer/Source Impact Join (G29 Phase 4, slice 1) | Accepted — slice 1 implemented (consumer graph, the join, ADR-046 D6's tier-1 selector, the --used-by overlay wiring); slice 2's use-case manifest/graph is implemented, and compare --use-cases has a real report-level use_case_impact block attributing findings to declared use cases; runtime-trace ingestion and a per-finding Change.affected_use_cases schema field remain not implemented; see G29 |
| 058 | Native Compatibility Agent Skills — User-Task-First Domain Layer | Accepted — partially implemented (G36 P0.1–P0.3/P0.6–P0.8: the generator and gates shipped; P0.9 partial, dogfooding deferred; P0.4/P0.5 product-surface items remain; of P1, P1.6 is done and the rest remains). Portfolio reset to one skill (2026-08-20 amendment, superseding the 2026-08-11 four-skill freeze): review-native-library-change (renamed from native-binary-compatibility-review) is the sole published skill, an unvalidated internal candidate; the other three are removed from skills-src/ and every generated tree, recoverable from git history. A second, same-date amendment ("PR 2") then rewrote that skill's workflow content — customer-outcome framing, a ten-step decision procedure, an integrated named-consumer branch, a narrowed v0.1 validated scope, a structured decision-report contract — still unvalidated. A third, same-date amendment ("PR 3.5") renamed it again to check-abi-compatibility (a user-outcome name rather than a mechanism name) and recorded, as design intent only, that PR 4's external-distribution deliverable should be an npm/npx-installable package published from this repository rather than a separate distribution repo. A fourth, same-date amendment ("PR 3") landed the full G37 evaluation corpus (12 scenarios) and a real 48-run pilot under the new name; its dominant finding is a harness confound (a 12-turn budget cutting off 31% of runs, asymmetrically by arm) rather than a skill-quality result — see skills-src/evaluation/agents/skills/pilot-results/README.md. A fifth amendment ("Harbor task battery") added a generated Harbor task directory per scenario (skills-src/evaluation/agents/skills/harbor/tasks/) — schema-validated against the real harbor package and end-to-end verified for Category A reference solutions, but never run through an actual Harbor trial (no working container/sandbox runtime in this environment). A sixth, same-day amendment ("Harbor made canonical") then decided Harbor is the surface for all new scenario/trial work going forward; runners/claude_code.py is kept only as the historical record of the existing pilot — see skills-src/evaluation/agents/skills/harbor/CLAUDE.md. Still unvalidated; see G36 / G37 |
| 059 | Compressed Snapshot Storage Envelope | Accepted — implemented (core snapshot I/O, dump, compare/scan --against/Python API, the snapshot cache, actions/baseline, the root Action's dump-mode snapshot-compression input, compressed-release-asset auto-fetch, and both publish workflows' snapshot-compression input; resolve-baseline needed no changes at all — already transparent); baseline-set manifest v2, a deterministic .tar.zst packager, and the wider docs sweep deferred, see the ADR's own "What this ADR does not (yet) close" |
| 060 | Synthetic-Consumer Compile-Probe Layer — Deferred | Accepted — not implemented (decision to defer), and no follow-up phase is scheduled; see G31 Phase D |
| 061 | Responsibility-Package Architecture and Flat-Namespace Migration | Accepted — partially implemented. The named slices in Phases 0, 1, 2, 3 and 5 have landed: architecture enforcement plus the aggregation migration; a ReportDocument boundary and a fact-vs-formatting split for every output format (JSON, SARIF, JUnit, --stat, HTML, all four Markdown views), with the gate decision, the per-finding verdict and every post-render fold moved before rendering; typed Request -> ResolvedPlan -> Result artifact workflows with all three service pipelines workflows-owned and both binary formats executing through one execute_dump_request; the frontends package with cli.py down to a 140-line registration facade, 47 direction violations closed and policy_file.py classified policy via the PolicyFileProtocol/ReclassifyRuleProtocol pair (taking service.py from 1,763 to 283 lines); and Phase 5's parser split, source-graph separation, change-catalog repartition and cycle-exception cleanup. Phase 4 and repository-wide convergence remain open: a landed phase label covers the slices that phase named, not the repository-wide guarantee. Six acceptance gaps are tracked in the ADR's own "Remaining acceptance gaps" section and worked in its bounded six-package closure sequence — dynamic first-party imports hiding forbidden edges from the checks (including the workflows -> frontends rendering back-edge), public_root_surfaces conflating a supported path with an owning layer, six separately-built report documents rather than one per completed evaluation, request/plan types carrying no selection/inventory/acquisition state, unresolved storage-vs-model ownership (serialization.py's own remaining codec logic and bundle_facts.py still open; its backfill and probe_harness.py's compare -> storage edge closed), and legacy modules with no recorded disposition. Capability placement follows ADR-068, result semantics the vision workstreams, storage work ADR-062/063 |
| 062 | Project Snapshot Storage v2 — Content-Addressed Sections, Explicit Fact Availability, and Occurrence-Preserving Identity | Proposed — partially implemented; Phase 0 primitives implemented (abicheck/storage/: fact availability, occurrence-preserving identity, canonical encoding, separated version axes). (a) Default single-artifact format: implemented and wired — the D8 sectioned document (storage.sectioned_document) is now dump/compare/scan's default write/read shape (Phase 8 redesign), with no CLI flag required; an older flat .abi.json stays readable. (b) Multi-artifact input/output reachability: partially implemented — a real directory-backed PackageManifest/ObjectStore (abicheck/project_snapshot_store.py) exists; A1.7 landed (compare's directory/package release fan-out accepts a multi-artifact ProjectSnapshot package as either operand, with --old-variant/--new-variant selection), but the .tar.zst transport form (part of A1.1) is not produced by anything, so nothing exercises A1.7 against one. (c) Shared-evidence storage: A1.4 implemented (BundleFacts/baseline sets fold onto the sectioned representation, both the persisted-document and live-object paths reconciled onto one physical layout per Track 1); digest-deduplicated shared evidence beyond that and BuildSourcePack/project source-graph dedup (A1.5) remain not implemented; bundle_variants: variant capture (A1.6: project capture-variants, one VariantRef per declared variant with separate declared/captured maps, and comparison_scope.variant_pairing on a stored/stored release comparison) and non-ELF artifact membership (A1.8) have landed. (d) Performance/transport: the .tar.zst transport form (A1.1's remainder) and all of Phase 2 (lazy loading, streaming encode, cache migration, indexes) are not implemented; see storage format v2 plan |
| 063 | One Semantic Pipeline — Unifying Application, Fact, Identity, and Outcome Models | Accepted — roadmap ADR, partially implemented (sub-phases 2B-8B complete; Phase 10 accounting open). Phase 0 (Fact[T]/FactStatus infrastructure, plus every known reader migrated off the legacy fields onto it) is complete; Phase 1's dump/scan typed-API convergence has landed its real dump execution routing onto execute_dump_request for both binary formats — ELF first, then PE/Mach-O the identical way (that half verified via mock-based CLI/unit tests only, no real PE/Mach-O toolchain), with cli_buildsource.dump_source_only() a named permanent exception; Phase 2's EntityId/ScopePath primitive and typed scope tracking have landed and populate entity_id, with wire-schema-v2 persistence, on both header-AST backends, DWARF, PE/Mach-O, and the ELF-symbol-only fallback, plus PDB's own RecordType/EnumType (types only, an unverified-against-real-MSVC heuristic) and BTF/CTF's own struct/enum types (types only) — every L1 producer now populates entity_id for its types, though PDB's and BTF/CTF's own function/variable identity remains unimplemented, so "every producer, every kind" is not yet accurate; Change.entity_id is wired at most function/type/variable diff sites and read as an additive finding_identity alias, though most DWARF/PE/Mach-O/ELF-only detectors and the diff_filtering.py/type_reachability.py string-identity call sites remain unmigrated (Phase 2B); Phase 3's infrastructure and its planned surface.py/export_surface.py migration have both landed: compare/surface_graph.py, policy/public_surface.py/public_surface_closure.py/public_surface_query.py, AbiSnapshot.surface_graph, and threaded EntityId resolution through compare()/compare_snapshots() have landed, and surface.py's own closure-walk traversal was migrated onto the new modules and deleted (not kept alongside) — but D5's own literal premise, that compute_public_surface() becomes a traversal through the graph, is deliberately not what shipped: three review rounds concluded the graph-reading design is unsafe, so the shipped closure walk never touches AbiSnapshot.surface_graph/GraphNode.attrs, only a pure function of the snapshot's own declarations — a considered departure from D5 as written, not an oversight, and not treated here as "phase complete"; export_surface.py's root-seeding now reads the evidence-entity-model Phase 2 observed exports join (compare/export_join.py) rather than a graph traversal, and type_reachability.py's stdlib-reference logic stays unmigrated by deliberate, documented design (to avoid a policy -> extract architecture violation), and the new graph builder's node ids still don't unify with the pre-existing L5 builder's (real, open follow-up work); Phase 4's AnalysisPlan pre-flight resolution has landed, closing the --build-target + pre-captured Bazel jsonproto gap (dump/compare/scan alike) and the matching .abicheck.yml-only dry-run parity gap for every CLI-reachable path (service_scan.run_scan_set's own direct typed-API call remains the one unwidened boundary), with the -H-on-wrong-mode scenario and the cli-contract/engine-cli-boundary gate widening named as deliberately out of scope; Phase 5's fact/capability registry is complete — infrastructure (abicheck/model/fact_registry.py, the fact-registry-completeness AI-readiness check, the generated docs/reference/fact-registry.md) plus the full field-by-field population, landed as nine conversion batches (schema v31-v40) that empty KNOWN_UNCONVERTED_ELIGIBLE_FACTS; no detector branches on FactStatus yet, which that phase scopes out by design; Phase 6 has landed nine slices — SemanticIR/CanonicalEntity and persistence, the header-AST normalizer for records/enums/typedefs/functions/variables/constants wired through every header-AST-backed platform (castxml/clang, ELF/PE/Mach-O, --ast-frontend hybrid) plus PE/Mach-O assembly, DWARF's own SemanticIR assembly (the first non-header-AST producer to reach it, on top of Phase 2's already-landed entity_id), CanonicalEntity.template_arguments for records, a --dump-manifest multi-TU dump's own occurrence-detail loss (each contributing TU's raw, pre-merge fragment is now normalized independently, disambiguated by cross-fragment source-location variance and TU-local linkage), PDB's own SemanticIR gap for RecordType/EnumType (types only, on top of Phase 2's already-landed entity_id), and BTF/CTF's own SemanticIR gap for struct/enum types (types only, via the new extract/debug_layout_semantic_ir.py; deliberately does not widen AbiSnapshot.types/.enums) — but is not complete: clang produces no occurrence at all for a template specialization, a function template's own arguments are unattempted, and — except for two landed cohorts — no detector or the checker itself reads SemanticIR yet (Phase 6B: compare/typedefs.py and compare/constants.py now decide per-side authority off a real SemanticIR when one is present, each side independently, no legacy index built for a side that carries one — landed as Track T3, 2026-09-05; since Phase 10 moved every declaration into the snapshot's SemanticIR, the other detector families read that IR-owned declaration store, and only checker reads of AbiSnapshot.dwarf/dwarf_advanced remain, on a shrink-only baseline); Phase 7 (RunOutcome and the last inline exit-code computation) is implemented; Phase 8 has landed its full D8 section split (every legacy document field partitioned across typed sections, not only SemanticIR) plus CLI wiring — a sectioned document is now dump/compare/scan's default write/read shape, with a directory-backed ProjectSnapshot package (still single-artifact) reachable as a typed-API primitive and a compare/scan --against input shape; Phase 8B (multi-artifact-canonical-storage) landed typed DTOs for every one of the eight D8 legacy sections (not only semantic_ir) across three prior PRs; a real multi-artifact PackageManifest writer/reader exists, reconciled onto one physical layout (Track 1, 2026-09-04) after briefly landing in two independent, non-interoperable forms — abicheck/bundle_facts_store.py (a live BundleFacts object) is now a thin wrapper over abicheck/storage/import_bundle_facts.py/import_baseline_set.py (an already-persisted BundleFacts/baseline-set document, via VariantRef.sections, landed 2026-09-04), so both entry points share the identical layout — see ADR-062's own text for the closure record; what remains per docs/_meta/one-semantic-pipeline-status.yaml's sectioned_storage entry is CLI/consumer wiring for the multi-artifact package half specifically — it is reachable only as a typed-API primitive and an explicit compare/scan --against input shape, never dump/compare/scan's own default path (which already uses the single-artifact sectioned document Phase 8 landed); Phase 9 (selector/suppression/reclassification consolidation) is complete — policy/selectors.py's SelectorSet is the one selector grammar Suppression/ReclassifyRule both build on, removing reclassify.py's importlib.import_module workaround for an import cycle that no longer exists; Phase 10 not yet. Phases 2B/4B/5B/6B/7B/8B adopted 2026-09-02 as the consumer-cutover/legacy-removal extension an external review found missing (all six closed 2026-09-30 — status owned by the implementation plan's own sub-phase table; the ledger, docs/_meta/one-semantic-pipeline-status.yaml, tracks per-concept authority one level up, not sub-phase status directly), alongside an accepted D5 amendment stating the shipped public-surface design (a deterministic SemanticReferenceIndex over snapshot declarations, not a traversal of the mergeable evidence graph) as the target. Generalizes and finishes ADR-042/046/048/049/050/055/061/062 rather than replacing any of them; see implementation plan |
| 064 | One Canonical Gate Algorithm and Exit-Decision Precedence | Accepted — substantially implemented. Formalizes CLI cleanup phase two's PR G2: removing --exit-code-scheme, keeping auto's inference as the only behaviour, and the full six-axis ExitDecision precedence (evidence-contract error, budget overflow, not-comparable, mode-dependent removed-required-library rank, gate, coverage/assurance floors). PR G1's three-axis core (compatibility gate, contract coverage, analysis assurance) shipped additively before this ADR (#789). Of this ADR's own two-stage plan, stage 1a landed complete: resolve_scan_exit_decision/resolve_release_exit_decision (abicheck/policy/exit_decision_precedence.py), pure functions reproducing the remaining axes' precedence (including a release's independent operational-error axis), unit-tested against the real code they model. Stage 1b landed partially: ExitDecision.to_dict serializes all five ADR-064 fields (report schema 2.47/1.22), scan's NOT_COMPARABLE outcome and the release fan-out's JSON summary both persist a real exit block now, verified to always agree numerically with the real, untouched exit-code functions they parallel. scan's _BudgetOverflow/_EvidenceContractError abort points now persist a decision for both the typed ScanResult API (abicheck.workflows.scan_abort_result.scan_abort_result_fields, SCAN_SCHEMA_VERSION 1.23) and the native scan --format json CLI path (cli_scan._emit_scan_abort_report, a ScanOutcome-envelope-compatible payload); a late _BudgetOverflow (after a real gate/coverage/assurance/audit decision already exists) also preserves those prior contributions instead of discarding them (attach_prior_on_budget_overflow, covering both the baseline-compare and audit-only branches). The release fan-out's GateOptions unification landed 2026-09-02 (ADR-064's own dedicated slice, not PR G2 — see that ADR's own "Landed" note). 2026-09-04: stage 2 landed — --exit-code-scheme, .abicheck.yml's exit_code_scheme: key, the kind: gate pack field, and the typed API's exit_code_scheme fields are all deleted; the gate algorithm is now fully automatic everywhere. The --artifact-set member-level evidence-contract signal question also closed the same day, reusing the single-binary path's dedicated exit 7. Still open: a typed request's own gate.* pack field (--pack stays CLI-only, ADR-049 D8); see cli-cleanup-phase-two.md for the authoritative status |
| 065 | Comparison Scope, Member Selection, and Input Completeness | Proposed — S1/S2/S3/S4 implemented (explicit member selection and the --dry-run plan view; the acquisition record, completeness axis on RunOutcome/ExitDecision, release no comparison completed outcome, degraded stranded-library marker and comparison_scope report section; package component inventories plus --support-promise findings; and deletion of the release fan-out's set-difference removal path), S0 open; design record for vision.md's partial-matrix/scope decisions (unmatched ≠ removed, ambiguity is a diagnostic, completeness is a run outcome, zero comparisons is never success); see vision workstream plan |
| 066 | Longitudinal Compatibility History and Project-Defined Versioning Policy | Proposed — S1 and S2 implemented; offline history over existing snapshots, a separable versioning-policy model (scheme/promise/windows/enforcement) that changes acceptance and never facts; does not revive ADR-022's registry; see vision workstream plan |
| 067 | Change-Intent Acknowledgment and the Policy-Disposition Audit | Accepted — S1, S2 and S3 implemented (one conserved disposition ledger behind every suppression application point, scalar compare through the release/bundle fan-out, aggregate, and the consumer-scoped --used-by/--required-symbol(s) paths; raw-versus-effective totals with rule provenance in every supported projection; reclassification and scope-exclusion reason breakdowns; not_evaluated detectors; the release recommendation reads the conserved delta; D5's bounded, ambiguity-checked acknowledgment records sharing the same selector grammar and ledger overlay mechanism as reclassification; D6's allow/warn/block additions review gate, an orthogonal exit axis that never reclassifies a finding — engine-level (checker.compare(acknowledgments=...)) and CLI exit-fold wiring only, no --acknowledgments CLI flag yet). scan reports still carry no disposition_audit block for aggregate to fold (only distinguished as "missing", never supplied) — outside S2's own scope. S4's base/head policy-delta and suppression-growth warnings remain unimplemented; see vision workstream plan |
| 068 | One Comparison Product — Retiring scan and Consolidating the CLI Conceptual Model |
Accepted — Phase 6 (scan retirement) has landed: the scan root command is deleted outright (abicheck scan exits 64, naming compare/compare --no-baseline), its scan-only modules/tests are deleted, and tests/parity/ is a compare-only regression corpus. All three flag demotions Phase 6 unblocked have since landed (each its own 2026-09-11 amendment): compare --env-matrix (.abicheck.yml's deployment: config key), dump --build-target (removed outright once scan's own removal cleared its only real blocker), and compare --require-complete-analysis (.abicheck.yml's assurance.require_complete config key) — none remain open. Retires the scan root command outright rather than reshaping it again, on the finding that compare could not reach scan's eleven cross-source checks, its pattern/preprocessor scans, changed-path localization or the abi3 audit at all (import-graph verified on main@309c8a82) — so migrating them was a capability gain for every compare user. Six root verbs in three tiers (compare/dump/deps; aggregate/project; frozen compat); one comparison product where baseline availability and cardinality are scope, with an audit spelled compare --no-baseline NEW (declared, never inferred from arity); one-sided checks become per-side comparison stages reported as introduced/resolved/persistent/not_evaluated, so a pre-existing problem is never manufactured into a new one; presentation may never change facts, gate or exit code (--explain-patterns stops implying --pattern-verdicts); per-run operands on the CLI and stable properties in .abicheck.yml, with no one-for-one YAML translation and no --set escape hatch; deps keeps its distinct runtime/deployment question and converges internally instead of merging; compat frozen and excluded from every metric; hard removal, no deprecation aliases. Supersedes 056; amends 037 D5/D7, 040 (reverses Lever 3's run profiles), 043 D1/D5, 047 §8's S5 routing, 055 (ScanRequest/ScanResult leave the registry) and 008 (states deps' user question); depends on 065 and protects 049's open false-negative prerequisites, which gate the contract-mechanism consolidation. See implementation plan for the capability-by-capability retirement map, the full flag inventory, the nine-phase migration sequence and the 28-scenario acceptance matrix |
| 069 | Name Shape Is Not Contract Membership | Accepted — implemented. One product rule: how a symbol or namespace is spelled (a vtable/typeinfo artifact, a detail/impl segment, a version segment) states its representation or a project's convention, never its membership in the compatibility contract — that is 049's question. Forces three changes, all landed: a name-derived report bucket may be counted but never used to discount a finding (the surface-breakdown note stops concluding RTTI/internal findings are "not public-API breaks" and stops labelling the remainder the "genuine public-surface" count — a vtable layout change on a public, user-derivable class is a genuine public ABI break); scope-convention classification consults only the scope that owns an entity, resolved structurally through the Itanium nested-name parser rather than by scanning the whole mangled string, which had been attributing a parameter type's namespace to the function; and a version namespace is not a stability promise, so v0 leaves DEFAULT_EXPERIMENTAL_NAMESPACES. That last one is a change to an existing default and is why this ADR exists (AGENTS.md's Authority rule) — it is not a loosening: EXPERIMENTAL_* is an overlay DetectNamespacePatterns appends, never a relabelling, so the underlying FUNC_REMOVED/BREAKING is emitted either way and the verdict and exit code are identical under both configurations — what the default drops is only the annotation asserting the removal was expected, which is the assertion a version segment cannot support. Affects only projects with a segment spelled exactly v0, and is restored exactly by one line in the new experimental_namespaces: policy key. Amends 044 (that key sits beside its internal_namespaces:); constrains what 049's consumers may claim without a resolved --contract. Deliberately does not address the fourth defect from the same report — Visibility.PUBLIC as a proxy for declaration presence, an evidence-selection problem needing an AbiSnapshot schema change inside 050's contract, recorded in known-gaps.md |
| 070 | The Action Layer Does Not Encode CLI Semantics | Accepted — D3 implemented for both extra-args tokenizers (action/run.sh and actions/check-target/action.yml no longer carry a hand-maintained value-option list; each queries the installed abicheck, scoped to the command it will invoke, and fails closed when it cannot — see below); D1, D2 and D4 not yet implemented. One boundary rule: the composite Action validates its own inputs and does not re-implement, restrict, or enumerate CLI semantics. Extends 037's D10.1 front-end/engine boundary one layer outward (from cli*.py-vs-Tier-1 to Action-vs-CLI) and constrains what 047's shell layer may assert alone. D1 input grammar only (a guard about the Action's own machinery survives — check-target's single project-wide header: input genuinely cannot stage a bundle baseline at headers depth); D2 the CLI owns flag acceptance, configuration resolution and precedence, with _is_cli_error()/exit-64 as the reporting path; D3 derive a needed CLI fact from the installed CLI rather than transcribing it — both extra-args tokenizers run after action.yml's own pip install, so a live query is available and is more correct than any snapshot, which cannot match the version the workflow installed; a committed generated artifact is permitted only for validate-inputs.sh, the one shell that runs pre-install; D4 a justification cites a symbol that exists or is deleted. Motivated by a 2026-09-12 audit that found all three copy classes rotted while review passed repeatedly — two guards rejecting what the CLI now accepts, twelve retired option names plus four missing live ones in both tokenizer tables, four exit-code comments citing the deleted cli_scan.py/scan_engine, and the tokenizer's own false "no live abicheck is reachable" premise, which was load-bearing enough that the audit's first revision adopted it unchecked. Accepted cost, stated rather than hidden: deleting a restriction mirror moves some failures after the toolchain install and replaces a tailored Action message with the CLI's, and an input the Action refuses today may succeed — a user-visible change, which is why this is an ADR and not a routine edit. D3 fails closed rather than falling back, which is a correction of this ADR's own first revision: it prescribed treating an undeterminable table as "nothing is value-taking" on the reasoning that under-recognition is safe, and a reviewer's counterexample disproved that — extra-args: --version --dry-run is argv the CLI accepts as --version's own value, so an opaque tokenizer invents a --dry-run, the Action skips its own output/sidecar injection as it must for a real dry run, and a full comparison then runs with no report written (silent, and an over-detection, so the direction argument was wrong too). An undetermined table is therefore fatal, scoped to a non-empty extra-args so a runner whose interpreter cannot import abicheck keeps working for every invocation that does not use the escape hatch. Evidence and phasing: action-cli-surface-drift.md |
| 071 | Release Analysis-Assurance Fold — One Axis, Any Cardinality | Accepted — implemented. Defines what assurance.require_complete means for a directory/package (release) compare, which until now raised a usage error (exit 64) on the premise that "the per-library fan-out has no single analysis_assurance result to gate on". It has one per compared member: the release-level answer is max over those members, folded into the existing 064 ExitReason.ANALYSIS_ASSURANCE axis rather than a new orthogonal one — over a single member the fold is the identity, so a one-member package gates and reports exactly as the scalar path does (AGENTS.md's "One model, any cardinality"), and a second field would hand a consumer reading analysis_assurance_contribution a 0 for a run that axis actually floored. max, never min: a member's incomplete analysis is not maskable by a complete sibling, the same monotonic direction 049's coverage floor and 065 D6's scope floor take. Distinct from that scope axis and orthogonal to it — scope asks whether every selected member was compared at all (inventory), this asks whether the comparisons that ran had complete evidence — so both apply and both are named independently in reasons. No new policy setting and no new exit number: assurance.require_complete defaults off, so every pre-existing release invocation's exit code, report bytes, and stderr are unchanged. Report-side the release JSON gains an analysis_assurance fold block naming the members that fell short and why, plus per-library analysis_assurance_status. The stored-BundleFacts operand takes the identical fold off its own per_library results. Retires four guards that existed only to stay ahead of the missing semantics (cli_compare_options._reject_set_input_flags, compare_bundle_facts_rejections, project_targets's kind: bundle run-plan rejection, actions/check-target's validate-inputs.sh _fail) |
| 072 | The PR Comment Is a Projection, Not a Second Report | Accepted — implemented. The sticky PR comment renders the completed report's own facts and never recomputes, re-derives or re-runs anything (D1). Two report/-owned projections back it: an entity-by-operation rollup (report/change_summary.py, canonical ChangeKindMeta dimensions, never kind-name matching) and an evidence projection (report/evidence_summary.py) (D2). Counting units and exactness are carried on the data, so a release report's capped per-library sample is disclosed as a floor rather than presented as a total (D3), and a confidence the report did not state is never invented (D4). Detector applicability keeps three distinct states over 067 D3's existing machinery (D5), and routine detector-disablement notes are split from material coverage_warnings structurally — by reconstructing each note through the one function that formats it — never by parsing prose (D6). pr-comment-on: changes now also posts for a material analysis limitation, posting eligibility only, with never still authoritative and no verdict, gate or exit code moved (D7). Size handling tightens a per-section row budget before downgrading detail, preserving headline, identity, exact counts, evidence, dispositions, exact omitted-row counts and navigation, and truncating only on line boundaries with <details> closed (D8); a grouped row always carries a complete member route and a single-symbol family is not aggregated (D9). 'View workflow run' and 'Download full report' are distinct, the latter rendered only from a caller-supplied successful upload URL (D10). report_summary's compatibility percentage is deliberately not propagated to the comment; its finding-count-over-symbol-count semantics are documented on CompatibilityMetrics and in known-gaps.md (D11) |
| 073 | Report-Only Publication and the Trusted-Reporter Boundary | Accepted — implemented. Adds actions/report (publish an already-produced canonical JSON report to a pull request as a sticky comment and/or job summary) and actions/verify-source-run (select and unpack the producer run a trusted workflow_run publisher reports on), so a fork PR can be reported on without any privileged job ever running contributor-controlled build logic. Publication is a capability separate from analysis and the Action installs abicheck and nothing else, enforced statically over its executable surface and behaviourally against a trapped PATH (D1); every decision — render, bound, create/update/clear/skip — lives in importable, credential-free Python and the shell keeps only argument marshalling and the API call (D2). A failed post fails the step with posted=false and is never downgraded to, or reported as, a clean compatibility result, and symmetrically a non-clean verdict never fails this reporting Action (D3). The sticky marker carries a monotonic run-id/attempt ordering guard, so a late-finishing older producer run skips with skipped-reason=stale rather than overwriting a newer result, while an unorderable marker publishes rather than freezing the comment (D4); a result that no longer holds is cleared in place — under on: changes too — rather than left standing, keeping the ordering marker a delete would lose (D5). Comment and job summary are bounded independently in bytes against the real platform limits (65,536 characters / 1 MiB per step) with cuts on line boundaries, <details> closed, and truncation disclosed in the body (D6). Run selection is one shared, tested boundary rather than a per-project recipe: repository/workflow/event/id/attempt/conclusion verified, the pull request resolved through the API with an artifact-supplied number only ever cross-checked, the PR head SHA kept distinct from the analysed merge commit and their association verified, artifacts taken only from that exact run (D7); the artifact is hostile input, with size/count/ratio caps, traversal/symlink/non-regular refusal, no executable bit preserved, and json.loads as the only consumer (D8). A trusted reporter does not make contributor-produced report contents trusted evidence — the boundary establishes recipient and execution safety only, and assurance/provenance keep coming from the report's own recorded evidence facts (D9). Finally, pr_comment.build_model gains the missing aggregate_schema_version branch, so a fan-in document renders as itself instead of falling through to the compare adapter and reporting "No ABI changes" for a run failing on every target in it; unavailable targets, not_comparable/operational_error legs, missing required targets, refused member reports and every axis shortfall render as explicit limitations and post under on: changes (D10). Extends 047; constrained by 070; publishes 072's projection unchanged |
| 074 | Logical Macro Definitions on the CLI (-D/--define) |
Accepted — implemented. Adds a narrow, repeatable -D/--define NAME[=VALUE] to dump and compare: a logical preprocessor definition for the L2 header parse, not a raw compiler argument. Amends rather than reverses 068's demotion of the L2 compiler/frontend family to .abicheck.yml's compile: block — that reduction's reason was toolchain identity is stable per project, which holds for --compiler/--sysroot/-std=/arbitrary --compiler-option but not for a feature macro that gates an opt-in public surface, which selects which surface is analysed exactly as -H/-I do. --gcc-options/--compiler-option are not restored under any spelling. Definitions apply to both sides of a compare identically, with no old=/new= form, since two sides parsed under different macro contexts are two different public surfaces (D1). Grammar is NAME/NAME=VALUE split on the first = only, ASCII-identifier names, NAME= distinct from NAME; whitespace-bearing and function-like definitions are rejected with a precise message and documented as limitations rather than spelled four different ways across CastXML/Clang/MSVC (D2). Precedence is a merge by macro name against compile.defines in the single existing fold, cli_options.merge_compile_config: a CLI-named macro is re-emitted last so it also beats a raw -DNAME in compile.options, every other config define keeps its exact position, and a run without -D produces a byte-identical token tail (D3). Injection is structurally impossible rather than blocklisted — one definition renders to exactly one define-prefixed argv token and the name must be a bare C identifier, so --define=-Xclang/@resp.txt/-DFOO are usage errors with targeted hints (D4). compile.defines stays the recommendation for stable CI; the macro set participates in extraction identity through the pre-existing ast_compile_args/macro_ops path, so a macro-off and a macro-on snapshot are refused as profile_mismatch rather than silently diffed (D5). The CastXML+MSVC row of its compatibility matrix is reasoned from this repository's own unconditional -D emission, not a fresh Windows probe, and is recorded as an unverified boundary |
| 075 | Target Ownership, Extraction Scope, and Ownership in the Graph | Accepted — not yet implemented. Records the ownership rules a snapshot was classified under (AbiSnapshot.extraction_scope, snapshot schema v52; absent loads as unrecorded, never as full) and stamps each declaration with a Fact owner/contract/rule id once in extract/, stored as an interned table (D1, D2). Refuses a pair whose dependency_evidence/prefilter differ, or whose rules differ under referenced; compares with a warning and a moved-findings report line when only the rules differ under full, and with a report line against an unrecorded baseline (D3). The rules' fingerprint joins the configuration digest as surface.ownership (D4). Adds three graph relations keyed on invariant-I1 node ids — owned_by/in_contract (derived, from the persisted per-entity fact) and provided_by (resolved_join, over BundleExportIndex) — none persisted (D5). dependency_evidence accepts only full; the scope keys drive classification, never retention (D6). Contract inputs become explicit CompatibilityEvaluationConfig.surface.ownership values with D7 provenance, and one obligation predicate excludes private/external declarations on the scalar, one-member and release paths alike (D7) |