ADR-061: Responsibility-Package Architecture and Flat-Namespace Migration¶
Date: 2026-08-24
Status: Accepted — partially implemented. The named foundation and
migration slices in Phases 0, 1, 2, 3, and 5 have landed; Phase 4 and
repository-wide convergence remain incomplete. A landed phase label means
the slices that phase named closed — not that the repository-wide
guarantee behind them holds. Each phase record under "Implementation plan"
states its own scope qualification, and the outstanding acceptance gaps
(real dependency enforcement, canonical result/report convergence,
supported-facade cleanup, storage/model ownership, and retirement or
explicit acceptance of legacy ownership debt) are stated once in
Remaining acceptance gaps, worked in the
bounded closure sequence that replaces the open-ended
Phase 4 narrative. Closing the last facade-size item would not by itself
satisfy the definition of done.
Authority: product capability placement follows
ADR-068; result
semantics follow the repository-root vision.md and its workstreams
(vision-api-abi-evolution.md);
storage-format work is coordinated with
ADR-062 and
ADR-063. This ADR owns implementation
ownership and dependency boundaries only.
Decision maker: abicheck maintainers
Context¶
abicheck has outgrown its predominantly flat package layout. Large modules
such as aggregate.py, analysis_assurance.py, and appcompat.py coexist
with expanding prefix families such as aggregate_*, cli_*, service_*,
diff_*, and reporter_*. Those prefixes provide visual proximity, but they
do not establish ownership, a supported entry point, or a permitted dependency
direction. They are packages in naming convention only.
The visible symptom is a collection of files near the repository's historical 2,000-line ceiling. The architectural problem is broader:
- Physical ownership is ambiguous. A contributor adding compare behavior
can plausibly choose
cli.py,cli_options.py, one of severalcli_compare_*files,service.py, or aservice_*pipeline. Nothing in the filesystem answers which module owns the behavior. - Mechanical splitting preserves coupling. Moving functions into a new sibling while re-exporting them from the old module reduces line count but leaves callers, monkeypatch targets, and reverse imports attached to the original owner.
- Architectural roles are mixed. Frontend translation, orchestration, extraction, comparison, policy evaluation, exit-code selection, and rendering can occur in one call chain without typed stage boundaries.
- Incidental paths acquire compatibility weight. Internal imports and private monkeypatch locations are often treated as if they were documented APIs. This prevents ownership transfer and makes every split permanent.
- Repository guidance compensates for the layout. Agent instructions have accumulated a detailed module inventory, implementation history, dynamic counts, and case-specific investigations because the package tree itself does not communicate where new work belongs.
The newer typed compare and dump paths demonstrate the better shape already: a typed request is resolved explicitly, executed, classified, and returned as a typed result. This ADR standardizes that shape across the repository and gives it a physical package model.
This is a repository architecture decision, not a mass-rename proposal. The target tree describes where code belongs after incremental migrations. Empty packages are not created merely to resemble the diagram.
Decision drivers¶
- A contributor or agent must be able to route a change without first reading a multi-thousand-line module or a historical instruction manual.
- Facts, compatibility findings, policy decisions, workflow results, and rendered output must each have one owner.
- Dependency direction must be machine-checkable and must not rely on a growing cycle allowlist.
- Existing documented imports and command behavior must remain compatible while implementation ownership moves.
- Migration must proceed as behavior-preserving vertical slices rather than a repository-wide flag day.
- Line-count enforcement must prevent new debt without confusing file size with architectural quality.
- Dry-run and normal execution must resolve configuration through the same path.
Decision¶
Adopt eight responsibility packages, arranged in three conceptual rings, and
freeze the flat abicheck/ namespace against new implementation families.
Imports point inward/downward. A reverse import is an architecture defect, not a reason to extend an exception list.
The end-to-end data flow is:
CLI / Python API
-> typed Request
-> resolved Plan
-> Extraction Result / Snapshot
-> Raw Findings
-> Policy Decision
-> Workflow Result
-> ReportDocument
-> JSON / Markdown / HTML / SARIF / JUnit
Each fact and decision is computed once. Later stages project or render it; they do not reconstruct it.
D1. Responsibility packages and dependency contracts¶
| Package | Owns | May depend on | Must not own |
|---|---|---|---|
model |
Immutable shared domain values and persisted/public identities | Standard library and lightweight typing dependencies | Filesystem access, subprocesses, Click, rendering, policy execution |
storage |
Snapshot/baseline serialization, cache behavior, snapshot/baseline schemas, migrations | model |
Extraction, compatibility decisions, report schemas, presentation |
extract |
Reading binary, debug, header, build, and source evidence into facts | model, storage |
Severity, suppression, gate decisions, user-facing output |
compare |
Comparability, old/new matching, identity, detectors, raw findings | model |
User policy, suppression, exit codes, rendering |
policy |
Effective configuration, contract relevance, suppression, classification, assurance, severity, gate decisions | model, compare |
Parsing artifacts, running compilers, rendering reports |
workflows |
Operation orchestration, sequencing, resource lifetime, and request/plan/result composition | model, storage, extract, compare, policy |
Click concepts and format rendering |
report |
The canonical immutable ReportDocument, report schemas, and pure format projections |
model, compare, policy, workflows |
Re-running comparison or changing findings, severity, verdicts, or gate state |
frontends |
CLI, typed-Python, and compatibility input translation and output selection | model, workflows, report |
Extraction algorithms, precedence rules, and business decisions |
The dependency list is exact for first-party responsibility packages. A package may use third-party libraries appropriate to its role, but a third-party import must not be used to bypass an architectural boundary.
errors.py, documented compatibility type modules, and root entry points may
remain at package root as explicitly classified public surfaces. They are not
an unbounded ninth layer.
D2. End-state physical layout¶
The following is a destination map. A directory is created only when at least one implementation and its tests move into it.
abicheck/
__init__.py documented public exports only
__main__.py frontend entry point
cli.py temporary/public facade
service.py typed-Python facade
api_types.py compatibility exports during migration
errors.py supported public exceptions
model/
entities.py
snapshot.py
findings.py
coverage.py
decision.py
change_catalog/
registry.py
symbols.py
types.py
platform.py
build.py
source.py
storage/
snapshot.py
baseline.py
cache.py
schema.py snapshot/baseline schema ownership
migrations.py
extract/
protocols.py
binary/{elf,pe,macho}.py
debug/{dwarf,pdb}/
debug/{btf,ctf}.py
headers/{castxml,clang}/
build/{compile_commands,cmake,bazel,make}.py
source/{graph,replay,provenance}.py
compare/
engine.py
comparability.py
identity.py
filtering.py
matching/{symbols,types,source_entities}.py
detectors/{symbols,types,cpp,platform,build,source}.py
bundle/{graph,matching,detectors}.py
policy/
effective_config.py
contract.py
suppression.py
classification.py
assurance.py
severity.py
gate.py
packs/
workflows/
artifact/{contracts,resolve,execute}.py
dump/{contracts,resolve,execute}.py
compare/{contracts,resolve,execute}.py
aggregate/{contracts,resolve,load,fold,reconcile,execute}.py
release/{discovery,matching,execute}.py
project.py
dependencies.py
appcompat.py
report/
document.py
build.py
grouping.py
schema.py report schema ownership
render/{json,markdown,html,sarif,junit}.py
frontends/
python_api.py
cli/root.py
cli/commands/{dump,compare,aggregate,release}.py
cli/options/{evidence,compiler,policy,output}.py
compat/abicc.py
compat/ retained public namespace; delegation only
There is deliberately no workflows/scan/ or cli/commands/scan.py in this
layout: ADR-068 retires
the command, and its surviving capabilities belong to the compare workflow
(D7). Until that retirement lands, scan's existing modules stay where they
are and are not migrated into a permanent home.
Every top-level responsibility package has a scoped AGENTS.md when it is
created. That file states purpose, allowed first-party imports, canonical
entry points, test locations, and prohibited responsibilities.
D3. Task routing is authoritative¶
The root contributor/agent contract must include this routing table near its beginning:
| Change | Owner |
|---|---|
| Read a new binary, debug, header, build, or source fact | extract/ |
| Add or change an ABI entity/value shared across stages | model/ |
| Match old/new entities or identify a change | compare/ |
| Decide relevance, suppression, classification, severity, or gate effect | policy/ |
| Coordinate dump, compare, scan, release, aggregate, project, or dependency behavior | workflows/ |
| Serialize snapshots/baselines, maintain their schemas or migrations, or manage caches | storage/ |
| Add a report field, report schema, or output format | report/ |
| Add a CLI flag, API adapter, or ABICC translation | frontends/ |
A new production file is not created until its owner can be selected from this table. If no row applies, the contributor must amend the architecture decision or explain why the behavior is a public root surface; inventing a new prefix family is not the fallback.
D4. Production module formation¶
Each module must have one completion sentence, such as "parses ELF dynamic
symbols," "matches old and new function identities," or "renders a
ReportDocument as Markdown." Descriptions such as "common utilities,"
"additional CLI logic," and "helpers used by compare" indicate that the
module has no stable owner.
New generic names are prohibited unless the architecture check carries a narrow, documented exception:
Names describe the responsibility instead: name_normalization.py,
compiler_flags.py, public_surface.py, report_grouping.py,
symbol_matching.py, or resource_lifetime.py.
No new root implementation sibling may extend a pseudo-package family, for
example cli_new_helpers.py, reporter_extra.py, service_scan_more.py,
diff_types_additional.py, or bundle_analysis_v2.py.
Module docstrings normally occupy 5–20 lines and state:
- what the module owns;
- what adjacent concern it does not own; and
- its canonical entry point, when one exists.
PR chronology, incident narratives, dynamic counts, temporary migration status, and individual known gaps do not belong in production docstrings. Durable rationale belongs in an ADR; active defects belong in issues or a small known-gap registry; user-visible changes belong in changelog material.
D5. Size is a pressure signal with a hard new-code ceiling¶
| Module type | Normal target | Review warning | Hard maximum for a new file |
|---|---|---|---|
| Normal production module | 100–400 | 500 | 800 |
| Compatibility facade | 20–100 | 120 | 150 |
Package AGENTS.md |
40–100 | 120 | 150 |
Root AGENTS.md |
150–250 | 300 | 350 |
| Test module | 100–600 | 800 | 1,200 |
| Parser or declarative catalog exception | 300–800 | 900 | 1,200 |
An exception above 800 lines is limited to generated code, a data-only declarative catalog, or a parser whose state machine remains one responsibility. It requires a debt record containing an owner, rationale, recorded line baseline, target, and review date. It may not silently grow.
Existing oversized files are governed by no-growth baselines rather than an immediate rewrite. The existing 2,000-line gate remains temporarily as a backstop until the debt ledger covers all legacy exceptions and the focused architecture check has demonstrated equivalent or stronger protection.
D6. Imports expose ownership¶
Across responsibility packages, production code uses explicit absolute imports from the canonical implementation module:
Within a package, relative imports are acceptable. New internal code must not
import behavior through abicheck.service, abicheck.cli, another legacy
facade, or a broad package re-export. A migrated implementation must never
import back through its old facade.
Import cycles are architecture defects. Permanent cycle allowlists and
TYPE_CHECKING imports used solely to hide a layering cycle are prohibited.
A temporary migration edge, if unavoidable, is recorded in debt.yaml with
an owner and expiry rather than added to the stable contract.
A dynamic import is still a dependency. An importlib.import_module
call naming a known first-party module is an import edge, whatever an AST
scan of ast.Import/ast.ImportFrom nodes can see. Deferring resolution
changes when the dependency is satisfied, never which layer depends on
which, so a dynamic bridge introduced specifically to keep a forbidden
direction out of the checks is unresolved debt with an owner and an expiry —
never evidence that a migration closed. Genuinely dynamic or plugin-style
loading is the narrow, documented exception; the architecture check is
expected to resolve literal import_module("...") calls and simple aliases
of them as edges. Today's bridges are catalogued in
gap A.
Package __init__.py files are small and inert. They may document and export
a narrow package surface through explicit __all__; they do not register
plugins, inspect the environment, touch files, load every submodule, or hold
product logic. Internal callers prefer the implementation module over a
package-wide re-export.
D7. Major workflows use Request -> ResolvedPlan -> Result¶
Every major operation is divided into contracts.py, resolve.py, and
execute.py (with responsibility-specific modules alongside them where
needed).
contracts.py contains typed inputs and outputs only:
@dataclass(frozen=True)
class CompareRequest:
old: ArtifactInput
new: ArtifactInput
scope: ComparisonScopeSelection
configuration: EvaluationConfiguration
@dataclass(frozen=True)
class ResolvedComparePlan:
old_plan: ResolvedArtifactPlan
new_plan: ResolvedArtifactPlan
acquisition: ScopeAcquisitionRecord
effective_configuration: EffectiveEvaluationConfig
@dataclass(frozen=True)
class CompareResult:
comparison: ComparisonResult
decision: ExitDecision
coverage: CoverageSummary
timings: StageTimings
The plan carries selection and acquisition state, not just two resolved operands: "which members were selected, which were acquired, and which were expected but never reached a completed comparison" is a resolved fact (ADR-065), not frontend orchestration.
- Requests express user intent without Click concepts.
- Plans contain normalized, fully resolved effective values and provenance.
- Results contain achieved facts and decisions, not formatted output.
- Values are immutable unless resource ownership requires a deliberately controlled context manager.
- Stages do not exchange large untyped dictionaries.
resolve.py turns a request into a plan. It owns precedence, validation,
normalized paths, backend/evidence selection, compiler/build configuration,
and resource preparation. It does not compare artifacts or render output.
execute.py consumes the plan. It owns stage ordering, resource lifetime,
extraction, comparison, policy evaluation, timings, and degradation
collection. A composition entry point has the conceptual form:
def run_compare(request: CompareRequest) -> CompareResult:
with resolve_compare_request(request) as plan:
return execute_compare_plan(plan)
scan is deliberately not the example here, and is not a future owner of
this shape: ADR-068
retires the command and moves its capabilities onto the one comparison
product. Do not migrate a command scheduled for retirement into a permanent
contract — route its surviving capabilities to the compare workflow instead.
Dry-run renders that same resolved plan. A separate estimator may summarize cost, but it may not independently predict effective backend, depth, policy, or configuration.
D8. Compatibility facades preserve public paths, not private coupling¶
Documented public modules such as abicheck.service, abicheck.cli, and
documented type modules may remain while implementation moves. A facade:
- stays below 150 lines;
- has explicit
__all__; - delegates or re-exports only;
- contains no domain logic;
- is used by external callers, not new internal code;
- documents whether its path is permanently supported or scheduled for removal.
A private re-export is not retained solely because a test monkeypatches it. The test moves with the implementation and patches the actual owner. Facade tests verify delegation and supported import compatibility; they do not retest the underlying algorithm.
A supported public path and an owning package are two separate answers.
architecture/modules.yaml's public_root_surfaces exempts a caller from
needing the imported module to be classified; it does not answer who owns
the implementation behind that path, and a module on the list still needs a
real owner. Canonical internal callers reach the owner, not the
compatibility route. Conflating the two is
gap B.
D9. Catalogs, parsers, and renderers have specific shapes¶
Catalogs. The change registry is partitioned by taxonomy, not into one
file per change kind. Declarative modules such as symbols.py, types.py,
platform.py, build.py, and source.py feed one registry.py, which
validates globally unique identifiers, complete metadata, valid references,
and non-contradictory defaults. Detection remains in compare; policy
algorithms remain in policy.
Parsers. Large backend parsers are divided by parsed entity or parser
state responsibility, never arbitrary line ranges. A CastXML package, for
example, may contain context.py, location.py, type_resolution.py,
functions.py, records.py, enums.py, templates.py, and backend.py.
backend.py coordinates traversal; entity modules parse one class of node
using shared context. They do not independently open input, resolve global
configuration, or create policy findings.
Renderers. Every renderer is a pure projection:
A renderer cannot remove findings, change severity, reconstruct a verdict,
calculate an exit code, repair workflow omissions, or mutate its input. All
formats consume the same immutable ReportDocument, built once from the
workflow result.
D10. Tests mirror responsibility ownership¶
The intended test topology is:
tests/
unit/{model,storage,extract,compare,policy,workflows,report,frontends}/
contract/{public_api,cli,schemas,compatibility_imports}/
integration/{extract,workflows,platforms}/
golden/reports/
factories/{snapshots,findings,artifacts}.py
fixtures/
Existing tests migrate with their production implementation rather than in a
separate cosmetic reorganization. Unit tests patch the owner module. Test
names describe behavior rather than private function names. Golden tests pin
stable output contracts but do not replace semantic assertions. Shared test
construction uses responsibility names (snapshot_factory.py), not generic
test_helpers.py. Large tests split by scenario or contract axis, never by
line number.
D11. Agent guidance is a routing contract, not a history database¶
The root AGENTS.md targets 200–300 lines (350 hard maximum) and contains:
- project purpose and supported/development Python versions;
- the task-to-package routing table;
- dependency direction and public compatibility rules;
- canonical verification commands;
- a change checklist; and
- links to architecture decisions, contributor documentation, and the current issue tracker.
It does not reproduce every module, current detector/test counts, bug investigations, implementation chronology, temporary migration state, or statements normalizing a large module as "intentionally" large.
Each responsibility package's scoped AGENTS.md is 60–120 lines and answers
only: purpose, permitted imports, canonical entry points, test locations, and
prohibited responsibilities. Tool-specific adapters such as CLAUDE.md and
Copilot instructions remain 5–20-line pointers to the canonical root and
nearest scoped instructions; they do not fork architecture policy.
Dynamic facts stay with generated repository facts. Durable design history stays in ADRs. Active defects stay in issues or a small machine-readable known-gap registry.
D12. Stable architecture and temporary debt are separate data¶
Create these files during Phase 0:
modules.yaml is the stable, desired dependency contract:
layers:
model:
path: abicheck/model
may_import: []
storage:
path: abicheck/storage
may_import: [model]
extract:
path: abicheck/extract
may_import: [model, storage]
compare:
path: abicheck/compare
may_import: [model]
policy:
path: abicheck/policy
may_import: [model, compare]
workflows:
path: abicheck/workflows
may_import: [model, storage, extract, compare, policy]
report:
path: abicheck/report
may_import: [model, compare, policy, workflows]
frontends:
path: abicheck/frontends
may_import: [model, workflows, report]
debt.yaml is the temporary migration ledger. Each entry records at least:
files:
- path: abicheck/aggregate.py
baseline_lines: <measured-at-adoption>
target: workflows/aggregate
rule: no_growth
category: workflow_monolith
owner: <team-or-maintainer>
rationale: <why-this-cannot-move-in-phase-0>
review_by: <date>
The implementation must measure baselines from the adoption commit; this ADR
does not hard-code guessed line counts. modules.yaml should remain stable as
the desired architecture. debt.yaml should shrink toward empty and must not
become a permanent import allowlist.
D13. A focused architecture check enforces the contract¶
Add scripts/check_architecture.py and route it through the existing
scripts/verify.py step catalog. It enforces:
- hard line limits for new files;
- no growth of recorded large files;
- no new forbidden root prefix files or root implementation packages;
- declared cross-package dependency direction;
- no new responsibility-package cycles;
- facade size, explicit-export, and delegation-only constraints;
- no unclassified first-party imports from migrated packages;
- no flat module occupying a target package name;
- scoped instruction presence for created responsibility packages; and
- schema and path validity for
modules.yamlanddebt.yaml.
The checker reports a precise import edge, file, and violated rule. It does not bury architecture enforcement inside a growing generic readiness script; that script may invoke the focused check but does not reimplement it.
The initial version operates only on migrated responsibility packages plus new files and debt baselines. It must not claim the legacy flat tree already conforms. Tightening coverage is a migration deliverable and is visible in the debt ledger.
Current-to-target ownership map¶
| Current area | Target owner |
|---|---|
aggregate.py, aggregate_findings.py, aggregate_manifest.py |
workflows/aggregate, with report projection in report |
bundle.py, bundle_analysis.py, release comparison code |
analysis in compare/bundle; discovery and fan-out in workflows/release |
buildsource/inline.py |
shared values in model; extraction in extract; orchestration in workflows/artifact |
buildsource/source_graph.py |
graph values in model; construction in extract/source; comparison in compare |
dumper_castxml.py, dumper_clang.py |
extract/headers/castxml and extract/headers/clang |
elf_metadata.py, pe_metadata.py, macho_metadata.py |
extract/binary |
diff_*, checker.py, comparability.py, finding_identity.py |
compare |
change_registry.py |
declarative model/change_catalog; classification algorithms in policy |
| assurance, suppression, severity, and contract configuration | policy |
reporter.py, reporter_markdown.py, HTML/SARIF/JUnit modules |
report/document, report/build, and report/render |
| compare/dump service pipelines | workflows |
scan's surviving capabilities |
the compare workflow, per ADR-068 — not a workflows/scan package |
bundle_facts.py and its serialization/store siblings |
values in model; persistence in storage; capture/comparison orchestration in workflows |
probe_harness.py's snapshot (de)serialization |
storage, reached from a workflow — comparison does not own persistence |
cli.py, cli_* |
frontends/cli; root cli.py becomes a facade |
service.py |
public facade over workflows |
compat/cli.py |
retained namespace delegating to frontends/compat |
| serialization, snapshot I/O, caches, and baselines | storage |
scripts/check_ai_readiness.py |
thin orchestration plus focused checks under scripts/quality/ where appropriate |
| historical/known-gap material in root instructions | ADRs, issues, architecture docs, or generated facts according to content type |
This table routes responsibilities; it does not require one commit per row or authorize a mechanical move without behavior tests. Rows still carrying legacy implementations need a recorded disposition — migrate, retain as a supported public module, or accept an explicit exception — rather than remaining unclassified indefinitely; see gap F.
Implementation plan¶
Each phase below is a closure record, not a chronology: what landed, what
the label does and does not mean, the durable lessons worth not
rediscovering, and the phase's own acceptance criteria. Active completion
tracking lives once, in
Remaining acceptance gaps and the
closure sequence, rather than inside each phase's
narrative — D11's rule applied to this document itself. The per-PR
investigation history these sections used to carry (review rounds,
superseded measurements, blockers that later re-measured away) is in git;
the last full-length revision is cfda6774. The technical findings that
outlive their PR are either summarized under each phase's "durable lessons"
here or already mirrored into
known gaps, their canonical home.
Phase 0 — stop new debt¶
- Accept this ADR and add the compact task-routing/dependency contract to root guidance.
- Add
architecture/modules.yaml,architecture/debt.yaml, and their schema/documentation. - Inventory existing oversized and prefix-family modules, record measured no-growth baselines, owners, targets, rationales, and review dates.
- Implement
scripts/check_architecture.pyfor new-file limits, frozen root families, debt no-growth, and dependency checks over migrated packages. - Register the focused check in
scripts/verify.pyand add focused unit tests for valid and invalid miniature trees. - Keep the old 2,000-line check as a temporary final backstop.
- Reduce tool-specific instruction adapters to pointers. Shorten the root instructions only after durable historical material has an explicit new home; do not delete unique operational knowledge during cleanup.
Acceptance: a new forbidden prefix sibling, an oversized new module, an undeclared responsibility import, or growth of a debt-tracked file fails the focused check with an actionable message. Existing debt remains runnable.
Phase 1 — prove the pattern with aggregation¶
Create real implementation modules, not empty scaffolding:
workflows/aggregate/
contracts.py
resolve.py
load.py
fold.py
reconcile.py
execute.py
report/
aggregate.py
Move typed contracts and behavior with their tests. Switch internal callers
immediately to the new owner. Retain in aggregate.py only documented public
exports that require compatibility. The new package cannot import the old
facade.
Acceptance: semantic results and JSON are exactly compatible; public imports covered by contract tests continue to work; internal imports use the new owner; no reverse facade import or duplicated aggregation decision exists; the relevant debt entries shrink or disappear.
Phase 2 — establish the canonical report document¶
Landed. report/document.py's immutable, JSON-shaped ReportDocument
exists, and every output format — JSON, SARIF, JUnit (via
report/render_xml.py), --stat, HTML (html_report.build_html_document +
report/render_html_document.py), and all four Markdown views
(report/render_markdown_document.py, report/render_markdown_alternate.py)
— builds one and projects it. Every format also has the fact-vs-formatting
split applied: a compute_* half that reads the DiffResult and decides,
and a render_* half that formats and decides nothing. Item 4 (decisions
before construction) closed for both halves it names — the gate decision
through policy/gate_decision.gate_decision_for_result, the per-finding
verdict through report/finding.py's ReportFinding — and item 5
(post-render mutation) closed for every fold, including the scoped-gate JSON
fold, now report/scoped_gate.apply_scoped_gate operating on the still-
mutable dict before the single render.
What the label does not mean. Each format builds its own document from
the DiffResult; service_render.render_output still dispatches the result,
snapshots, severity configuration, and presentation options down six
independent paths. That is a per-format boundary, and it is real — but
wrapping six separately-built projections in the same immutable type is not
proof that they cannot disagree. The stronger invariant this ADR's
definition of done actually states — one completed evaluation → one
completed semantic report document → every format — is gap C below, and is
duplication-and-convergence-assessment.md
Phase 4's ReportEnvelope target rather than a second design.
Gap C status (2026-09-07, updated through the JUnit slice — all five
named formats now converged): JSON, Markdown-full/review, HTML's default
view, SARIF's default view, and JUnit's default view all converge onto the
shared choke point (HTML, SARIF, and JUnit only partially — see their own
progress updates below, and the JUnit slice's overall closure summary
further down this section); Markdown-leaf/root-cause stays its own
separate, legitimate document by design (see the scope decision below).
report/build.py's
build_report_document(result, ...) is now the
single function that performs the full report_mode="full" build
(_build_json_base, _add_abi_surface_breakdown, _add_changes_block, the
gate decision, the side-facts fold, etc.) — moved out of reporter.to_json's
own inline body (now a thin build -> render_json wrapper) and out of
service_render._render_json_output (which calls it directly for
report_mode="full", bypassing to_json entirely). Verified against the
pre-refactor path via tests/unit/report/test_build_report_document.py
(byte-identical JSON for both the default and show_only-filtered cases,
plus a mock-based assertion that render_output("json", ...) calls the
shared build exactly once and never falls back to the legacy to_json
pipeline for full-mode JSON). to_json's --stat/leaf/root-cause
report modes are not routed through this choke point yet — they remain
each their own independent build, same as before this change; only
report_mode="full" (the default) was covered by that first slice.
Progress update (2026-09-07, later same day): full-mode Markdown and
review (unconditional-recommendation markdown) now also route through the
one shared build. service_render.render_output()'s markdown branch
(only for report_mode == "full" — --stat and the Markdown
leaf/root-cause alternate views deliberately do not, see the scope
decision below) and its review branch each now call
build_report_document(result, ...) exactly once and thread the resulting
ReportDocument down through to_markdown/to_review_digest into
report/render_markdown_document.py's build_markdown_document/
build_review_digest_document (both gained an optional report_document
parameter; a direct caller passing none keeps the prior, independent-build
behaviour, so this is additive, not a signature break). Those two functions
now reuse the shared document's disposition_audit field instead of a
second, independently-resolved call to compute_disposition_audit.
Everything else full-mode Markdown/review render (the headline table,
policy section, severity groupings, confidence section, and so on) stays
computed the way it already was, reading DiffResult directly through
reporter_markdown.py's existing compute_* functions — deliberately, not
an oversight: on inspection, none of it was actually a second, independently
decided value at risk of drifting from JSON's own decision. Every
classification these Markdown sections rely on (categorize_changes,
gate_eligible_changes, apply_show_only, _suppress_dangling_correlation_
notes, impact_for) was already the identical shared pure function JSON's
own build calls, not a parallel reimplementation — two calls to the same
deterministic function of (result, ...) cannot disagree, so the
byte-for-byte-safe, real convergence available here was structural (route
through one call, reuse what the shared document already carries in a
matching shape) rather than a rewrite of every Markdown section to read
JSON-shaped fields it doesn't have a matching presentation for
(disposition_audit is the one field whose shape matches exactly;
headline/policy/severity_groups/etc. have no JSON-document counterpart
at all, since JSON never renders them in that shape). Verified via
tests/ markdown/review suites plus the full golden suite, all byte-
identical to pre-change output (see this ADR's own PR history / the
duplication-and-convergence-assessment.md plan for the exact commit).
Progress update (2026-09-07, later still the same day): HTML's default
view now also routes through the one shared build, to the same depth as the
Markdown/review slice. service_render.render_output()'s html branch
now calls build_report_document(result, show_only=show_only,
show_impact=show_impact, severity_config=severity_config) once and forwards
the resulting ReportDocument into html_report.generate_html_report /
build_html_document (both gained an optional report_document parameter,
additive — a direct caller passing none keeps the prior, independent-build
behaviour). build_html_document reuses the shared document's
disposition_audit field (reconstructed via DispositionAudit.from_dict,
the same round-trip Markdown's own report_document handling already uses)
at both of its two call sites — the compat_html ABICC-clone layout's own
disposition-audit block, and compute_summary_table's audit argument —
instead of two independent calls to compute_disposition_audit over the
same ledger. HTML's remaining facts (bucketing changes into removed/changed/
added, the per-section ChangeRow tables, compat_html's ABICC severity-band
bucketing, the gate/scoped-verdict cards) were read in full while doing this
work and confirmed to be exactly the gap the prior assessment already
recorded: JSON's flat changes[] array plus summary severity block has no
matching shape for any of them today, so converging them would mean adding
new fields to the shared document first (the "genuinely new shared-document
design" the assessment below already named for HTML) — deliberately not
attempted in this slice, same reasoning as severity_groups staying
Markdown-side. Also closed in this slice: the ABICC-clone compat_html=True
layout previously had no golden test at all (a gap C acceptance-criteria
item this slice was asked to close alongside the wiring above); it now
has one (tests/golden/html_template/main_report_compat.html,
tests/test_html_template_golden.py), verified byte-identical on every
pre-existing case and passing on the new one. Verified via the HTML test
suite, the full golden suite (including the new compat_html case), and the
usual ruff/mypy/ai-readiness/architecture gates, all clean.
Progress update (2026-09-07, SARIF slice): SARIF now also routes through
the one shared build, same structural depth as the Markdown/review and HTML
slices. service_render.render_output()'s sarif branch now calls
build_report_document(result, show_only=show_only,
severity_config=severity_config) once (unconditionally, independent of
report_mode — SARIF's own report_mode="root-cause" only adds extra
per-result properties on top of the same shape, unlike Markdown's genuinely
separate leaf/root-cause documents, so there is no reason to gate the shared
build on it) and forwards the resulting ReportDocument into sarif.
to_sarif/to_sarif_str (both gained an optional report_document
parameter, additive only — a direct caller with none keeps the prior,
independent-build behaviour). to_sarif reuses the shared document's
disposition_audit field (read directly off report_document.to_mapping(),
already JSON-shaped so no reconstruction step is needed — unlike HTML's/
Markdown's own DispositionAudit.from_dict round-trip) instead of an
independent compute_disposition_audit call for its properties.
dispositionAudit block. SARIF's own shape — the rule catalog (rules_seen),
per-result level/location derivation, root-cause grouping, and the
scopedGate/severityGate/coverage-notification blocks — was re-read in
full during this slice and confirmed to still be exactly the gap the prior
assessment already recorded: none of it has a matching field in
build_report_document's JSON-shaped structure today, so genuinely
converging it means adding new shared-document fields first, deliberately
not attempted here — same reasoning as HTML's and Markdown's own remaining
facts. sarif.py already had SARIF's own compute/render split in substance
(to_sarif computes the SARIF dict; to_sarif_str composes it with
report.render_json.render_mapping_as_json, the same generic JSON-freeze-
and-render step SARIF's own report/AGENTS.md entry already names) — this
slice's job was wiring the shared build into the existing split's compute
half, not inventing a new one. Verified via the SARIF test suite (220
passed), a new TestSarifReusesSharedDocument class in
tests/unit/report/test_build_report_document.py (byte-identical
to_sarif_str output with and without a supplied report_document, equal
dispositionAudit values, and a render_output("sarif", ...)-calls-the-
shared-build-exactly-once assertion for both report_mode="full" and
"root-cause"), the full golden suite, the HTML template golden (shared-
code regression tripwire), and the usual ruff/mypy/ai-readiness/architecture
gates, all clean.
Progress update (2026-09-07, JUnit slice — the last of the five named
formats): JUnit now also routes through the one shared build, same
structural depth as HTML's/SARIF's own slices.
service_render.render_output()'s junit branch now calls
build_report_document(result, show_only=show_only,
severity_config=severity_config) once, unconditionally (independent of
report_mode, matching SARIF's own reasoning: JUnit's own "root-cause"
mode only adds rootCauseId/rootCause attributes to each <failure> on
top of the same shape, it does not restructure the per-symbol <testcase>
tree), and forwards the resulting ReportDocument into junit_report.
to_junit_xml/_build_testsuite (both gained an optional report_document
parameter, additive only — a direct caller with none keeps the prior,
independent-build behaviour). _add_disposition_audit_properties reuses
the shared document's disposition_audit field via the same
disposition_audit_dict_reusing_document helper SARIF's slice introduced,
instead of an independent compute_disposition_audit call for its
abicheck.detected_total/abicheck.effective_total/
abicheck.disposition.* testsuite properties.
JUnit's own remaining facts were re-read in full during this slice and
confirmed to be exactly the gap the prior assessment recorded, with one
addition the prior assessment did not have available to check yet: JUnit's
per-finding verdict/category resolution (_is_failure/_failure_type) was
already routed through report.finding's ReportFinding/
build_report_findings primitive by an earlier slice (ADR-061 Phase 2 item
4b) — the same canonical primitive a full convergence would want — but as a
separate call from what build_report_document computes, not a value
read off the shared document, because build_report_document's own
_add_changes_block does not itself build a ReportFinding set at all
(JSON's changes[] entries resolve each change's verdict inline via
effective_verdict_for_change, with no per-finding IssueCategory in the
JSON shape at all today). Threading that through would mean either (a)
rewriting JSON's own _change_to_dict to compute and carry IssueCategory
too — a change to a format whose output this slice must leave
byte-identical — or (b) inventing a second, JSON-object-external field on
ReportDocument keyed by finding_id for a fact only JUnit needs; neither
is the safely-mechanical, already-shaped substitution this slice's own
scope is (see "What remains open" below). JUnit's symbol/testcase tree and
its root-cause grouping are, as before, its own SARIF/JUnit-shaped
computation with no shared-document counterpart.
Verified via the JUnit test suite (156 passed), a new
TestJunitReusesSharedDocument class in tests/unit/report/
test_build_report_document.py (byte-identical to_junit_xml output with
and without a supplied report_document, equal disposition-audit testsuite
properties, and build-called-exactly-once assertions for both
report_mode="full" and "root-cause"), the full golden suite, the HTML
template golden (shared-code regression tripwire), and the usual
ruff/mypy/ai-readiness/architecture gates, all clean.
Two acceptance tests from the original gap-C task were added in this
slice, now that all five formats cross the shared-document boundary: a
TestRendererOrderIndependence class rendering the same completed
DiffResult through json/html/sarif/junit/markdown in two different
orders and asserting each format's own output is byte-identical regardless
of order, with build_report_document (mock-spied, wraps= the real
function) called exactly once per format render in either order; and a
TestSarifAndJunitDecisionBoundary class, the honest, narrower sibling of
test_render_html.test_render_html_imports_no_decision_making_module for
these two formats — see that class's own docstring for why an import-based
"no decision module reached at all" guard would be false for sarif.py/
junit_report.py as currently structured (neither has HTML's real
compute/render module split yet), and what the real, current boundary it
asserts instead is (the disposition-audit reuse path specifically, checked
both by call-count and by an AST scan of the actual call site).
Closure package 3 (ReportEnvelope): gap C's remaining half — one
document per evaluation, not one per format — has now landed. The five
progress updates above each routed one format through the shared build
function; what stayed open was that service_render.render_output called
it once per format branch, so N formats of one evaluation still meant N
documents that merely agreed. abicheck/report/envelope.py's
ReportEnvelope (the plan's Phase 4 design, not a second one) is now built
once, above format selection, by report/build.py's
build_report_envelope: it resolves the severity GateDecision first and
hands it to build_report_document (so the document's severity block and
the object SARIF/HTML read are the same object, not two agreeing calls),
resolves one ReportFinding per Change — result.changes and
scoped_only_changes — and carries the presentation-only RenderOptions.
service_render.render_envelope(fmt, envelope) then selects a pure
projection. Every previously-open decision in the paragraphs below is
closed by reading the envelope: SARIF's and HTML's own
gate_decision_for_result calls, HTML's and the review digest's own
report_findings_for calls, JUnit's own build_report_findings call, and
Markdown's surface_changes re-resolution. SARIF's invocation
exitCode/exitCodeDescription fold moved to report/sarif_invocation.py
(a renderer does not own exit behaviour) and JUnit's disposition-audit
properties to report/junit_disposition.py. Everything still listed as
open below was re-examined item by item and is presentation — an
arrangement of already-decided findings in a format-specific shape — or a
separate document by this ADR's own earlier scope decision (--stat/
oneline, Markdown's/JSON's leaf/root-cause); there is no third
category left unaccounted for. The process exit fold
(cli._exit_with_severity_or_verdict) deliberately stays in frontends:
the envelope carries the exit decision the report publishes, not the code
the CLI exits with. Verified by capturing 512 renders (4 change sets × 16
option sets × 8 formats, render_output end to end) against the
pre-refactor tree first and diffing byte-for-byte after — the same
"golden first, then refactor" discipline this section's own durable lessons
record — plus TestRendererOrderIndependence's new cases: one envelope
rendered into every format in several orders is byte-identical each time,
render_envelope agrees with render_output byte-for-byte, and a
call-count spy shows every decision function runs exactly once during
envelope construction and never again across ten subsequent projections.
What remains open. Markdown's leaf/root-cause alternate views (see
the scope decision immediately below — these are separate, legitimate
documents, same reasoning as JSON's own leaf/root-cause/--stat, not an
oversight left out of any slice), HTML's own bucketing/section/compat-mode
computation, SARIF's own rule-catalog/level-derivation/root-cause/
scoped-gate computation, and JUnit's own per-finding verdict/category
resolution, symbol/testcase tree, and root-cause grouping (see the
respective progress updates above — every format's shared build call itself
has now landed; only disposition_audit reuse was safely available beyond
that for any of the three non-JSON/Markdown formats). None of this is an
oversight: each is genuinely format-specific presentation, or would require
a genuinely new shared-document field this initiative deliberately declined
to invent mid-slice — see each progress update's own reasoning.
Gap C overall closure state, across all five named formats (JSON,
Markdown-full/review, HTML, SARIF, JUnit) — 2026-09-07, JUnit slice, the
last of the five. What is genuinely converged: the decision layer —
per-finding verdict/category (report/finding.py's ReportFinding, used
directly by JSON's severity JSON and by JUnit, wherever it is used, computed
via the one canonical primitive), the gate decision
(policy/gate_decision.gate_decision_for_result), and the disposition audit
(report/disposition_audit.py's compute_disposition_audit) — is resolved
exactly once per render via build_report_document and reused, never
re-derived, by JSON, Markdown/review, HTML, SARIF, and JUnit alike, for
every fact each format actually reuses today. What remains legitimately
format-specific presentation, precisely stated per format (not overclaimed):
Markdown's severity_groups headed-section grouping and its leaf/
root-cause alternate views; HTML's removed/added/changed bucketing,
per-section ChangeRow tables, compat_html's ABICC severity-band
bucketing, and its nav_bar/summary_table/gate_card/scoped_verdict
dataclasses; SARIF's rule catalog, per-result level/location derivation,
root-cause grouping, and scopedGate/severityGate/coverage-notification
blocks; and JUnit's per-finding verdict/category resolution (itself already
routed through the canonical ReportFinding primitive, just not read off
the shared document — see the JUnit progress update above for exactly why),
symbol/testcase tree, and root-cause grouping. As the HTML and Markdown
slices' own reports already found and this slice reconfirms for SARIF and
JUnit: most of each format's own section/layout logic was never actually a
second, independently-decided value at risk of drifting from another
format's — it is presentation-only computation over already-agreed facts,
which is what "converged" means in this ADR's sense (cannot disagree on a
decision), not "byte-identical internal implementation" across formats.
Closing any of the items in the paragraph above for real would mean adding
a new field to ReportDocument for that format's own shape first (as each
format's own progress update says), which is deliberately out of scope for
this initiative's five slices — a further, separately-scoped piece of work
if a future session judges it worth doing.
Durable lessons.
- A renderer that performs a registry lookup (
report_classifications,checker_policy.impact_for) is deciding, not formatting. The rendered bytes are identical either way, so no golden test can catch it; the guard is an AST scan of the renderer's own imports (test_render_html_imports_no_decision_making_module), verified to fail against the pre-fix module. - A section's
None(this section does not exist) is not its empty value. Collapsing the two renders an empty table for a result that carried no data at all. - Do not cache a policy-resolved value on the mutable
DiffResult.report_findings_forrecomputes per call because a caller may mutate the result between two renders; removing the hazard beat invalidating a cache against every mutation surface. - Filtering that looks like formatting stays compute-side when it is really a report decision — which summary rows are non-empty, which reclassify rules are still active (an expired waiver must not be disclosed as in effect).
- Every closure here was verified by capturing a byte-exact golden against
the pre-refactor path first, plus a property test stating the contract
a golden cannot (round-trip safety, render purity, escaping per field).
Two paths (
to_review_digest,--report-mode root-cause) had no golden at all before their slice; the fixture came first, not after. -
A whole-document projection that pushes its module past the 800-line new-file ceiling splits out a sibling (
render_html_document.py,render_markdown_alternate.py). It never trims to fit. -
Define immutable
ReportDocumentcontracts from existing report-model behavior rather than inventing a second schema. - Build the document once from a workflow result.
- Route JSON and Markdown first, then HTML, SARIF, and JUnit, through pure projections.
- Move all filtering, severity, verdict, and gate decisions before document construction.
- Delete output-specific verdict repair and post-render mutation after parity tests cover every format.
Acceptance: all renderers consume one document; format parity and golden
tests pass; mutability tests show renderers cannot alter the workflow result;
no renderer computes an exit code or compatibility decision. Items 1, 4, and
5 are met per format; item 2's "once" — one document shared by every format
of one evaluation — was gap C, and is met by closure package 3's
ReportEnvelope (see the gap C status section above).
Phase 3 — converge artifact workflows¶
Landed. The ArtifactRequest -> ResolvedArtifactPlan -> ArtifactResult
shape is real: workflows/artifact/contracts.py holds the plan type,
workflows/artifact/resolve.py decides a plan without running it, and
workflows/artifact/execute.py runs one and reports what it achieved. All
three service pipelines (service_dump_pipeline.py,
service_input_resolution.py, service_compare_pipeline.py) have
workflows owners and are free of CLI imports; dump resolves its request
once above the --dry-run branch, so dry-run renders the plan execution
consumes rather than one that merely agrees with it. Both binary-format
paths (ELF under CLI cleanup phase two PR C, PE/Mach-O under ADR-063 Phase
1) now execute through the one shared execute_dump_request.
What the label does not mean. This is the typed-resolution foundation,
not universal frontend convergence. CompareRequest/DumpRequest
themselves still carry no selection, inventory, or acquisition state
(ADR-065's scope model does not apply to a single-pair/single-artifact
request at all); the release fan-out's own multi-member selection,
inventory, and acquisition state now live on a sibling typed pair,
ReleaseCompareRequest/ReleaseComparePlan, resolvable with one direct
Python call — see gap D below for exactly what that closes and what is
still command-level orchestration. The same gap
vision-api-abi-evolution.md
records against scope convergence. The PE/Mach-O migration is verified by
mock-based CLI/unit tests only: the layering claim is proven, an end-to-end
run against a real PE/Mach-O binary is not.
Durable lessons.
- A blocker recorded once goes stale as the tree moves. Re-measure before scoping work against it — three separate blockers in this ADR turned out to describe a tree that had since changed.
- Engine code that cannot import upward writes private copies instead. The
depth ladder existed four times and the
abicheck_inputs/guard three, each copy's own comment explaining that it was a copy. The coupling was already being paid for in duplication before anything moved; the fix is a leaf both sides may import (evidence_depth.py,buildsource/pack_shape.py), not a facade. - Error contracts are part of the move and are preserved exactly, not
tidied:
ValidationError(usage, exit 64) andSnapshotError(operational, exit 1) mean different things to a CI consumer. Every code was measured against the real CLI, and the characterization tests were written and committed before the move. - An engine module owns no output stream.
on_outputreplaced aquietflag that was only meaningful to a caller holding a stream.
Use the pattern already emerging in the typed compare, dump, input-resolution, and artifact-plan code:
Route dump, both compare operands, the release fan-out, application compatibility, and dependency comparison through shared per-artifact resolution and execution contracts. Pair-wide decisions remain in the pair workflow; single-input resolution does not acquire artificial knowledge of both sides.
Acceptance: equivalent CLI and typed-API requests resolve equivalent plans; extraction occurs once per artifact; resource lifetimes cover execution; dry-run renders the same resolved plan normal execution consumes; achieved depth and degradation are result facts rather than frontend guesses. Met for the single-artifact path; not met for selection, inventory, and acquisition state (gap D).
Phase 4 — thin CLI and Python API¶
Landed. abicheck/frontends/ exists and holds the CLI: commands
(frontends/cli/commands/{dump,compare}.py), runtime (verbosity, output,
provenance, the exit decision), the option cluster
(frontends/cli/options/*). Root cli.py went from 1,959 lines to a
~120-line registration facade. It briefly also carried
frontends/cli/moved.py, a lazy alias table keeping ~80 private helpers
importable from abicheck.cli after they moved; that was retired once every
caller was migrated to the owning module, taking the table, the
__getattr__ resolving it and the module-class assignment guard with it. Classifying the whole cli_* family frontends
surfaced 47 real direction violations — the CLI reaching past the engine
into policy, compare, and extract — and all 47 were closed rather
than suppressed, each routed through a workflows re-export surface
(workflows/gate.py, extraction.py, findings.py, scan_config.py).
ENGINE_CLI_BOUNDARY_ALLOWLIST went 15 → 4 across Phases 3 and 4.
policy_file.py is now classified policy, unblocked by the structural
PolicyFileProtocol/ReclassifyRuleProtocol pair in
model/policy_file_protocol.py, which moved compare_snapshots,
load_suppression_and_policy, _validate_contract_mode, and
dedup_policy_override_warnings into workflows/compare_policy.py and took
service.py from 1,763 lines to 283.
What the label does not mean — this phase is open. service.py is 283
lines, not below 150, and the shortfall is no longer a blocker: it is nine
re-export blocks whose supported surface has not been audited, private
compatibility names retained for test patch locations (which D8 forbids),
and scan-shaped bindings whose future is ADR-068's to decide. Root
cli.py still registers scan, imports the legacy option hub, and applies
variant options at the root because the command module hit its size cap —
a placement decided by file-size pressure rather than ownership, which is
exactly what D5 says must not happen. That work is gap D and gap F below,
sequenced as closure packages 4 and 6.
Durable lessons.
checker_policy.py's model-vs-policy split is the hinge several other moves waited on:ChangeKindis defined there alongside real policy algorithms, so anythingmodel-owned that must name aChangeKindis blocked until it splits.- The
PolicyFileinvestigation's answer is a structuralProtocol, not a subclass or a data-only base. NarrowingDiffResult.policy_fileto a data-only type breaks every consumer that calls a method on it under the mypy gate, even though the runtime object is unchanged; a protocol satisfied structurally does not. Two mechanical requirements were each reproduced againstmypy --strictbefore being trusted: collection members must be read-only@propertydeclarations (a plain attribute is invariant and rejectsdictagainstMapping), and a protocol must declare the whole surface real callers use — including data attributes (to_verdict) and the exactlistvsSequenceshape a downstream parameter demands. Reclassifyingpolicy_file.pyascomparewas rejected outright:compute_verdictis policy logic by any reading, and mislabeling it only relocates the ambiguity this ADR exists to remove. check_architecture.pychecks a file's own imports only when that file has a classification;unclassified-importadditionally requiresmigrated_source. So a flat, classified module reporting zero findings says nothing about whether moving it is safe, and a deliberately unclassified leaf (reclassify.py) is exempt from every check by construction. Measure against the state after the move, not before it.- Every hand-taken count in this phase went stale or was wrong at least once. Numbers here are re-measured or dropped, never carried forward.
- A
monkeypatch.setattragainst a name resolved through a lazy__getattr__rebinds nothing the real caller reads, and a re-export surface binds its names at import time. Both are ordinary Python semantics, and both make a facade's tests silently inert. -
Trimming a facade's explanatory comments to hit a line count transfers no ownership and does not satisfy the criterion.
-
Move command input translation into
frontends/cli/commandsand reusable Click-only option declaration intofrontends/cli/options. - Make workflows the sole operation owners and reports the sole rendering owners.
- Reduce root
cli.pyto command registration/delegation and rootservice.pyto documented typed functions. - Derive every frontend's process response from the same
GateDecision. - Update tests to patch implementation owners, retaining facade tests only for supported public imports and delegation.
Acceptance: both root facades are below 150 lines, declare __all__, and
contain no product logic; frontend modules contain no extraction or policy
algorithm; CLI/API parity tests exercise shared workflows. cli.py meets the
size criterion; service.py does not, and the size criterion is the last
check of this phase rather than its definition — reduce the coupling first
(closure package 6), then measure.
Phase 5 — parsers and catalogs¶
Landed — all four named items.
- CastXML and Clang parsing split by entity and shared parser context —
closed on both backends (
extract/headers/castxml/*,extract/headers/clang/*), with parity held by the sharedparse_*surface behinddumper._header_ast_parser. - Source-graph values, construction, and comparison separated — closed
for every internal caller off the
buildsource/source_graph.pyfacade; the shared node/edge-classification predicates relocated intomodel/source_graph_query.py, andtemplate_graph.pyclosed last via a split intotemplate_graph_fold.py. - Change catalog repartitioned — all 397 entries moved into D9's
model/change_catalog/{symbols,types,platform,build,source}.pytaxonomy by which detector actually produces each kind, with all four registry-validation properties (global uniqueness, valid references, non-contradictory defaults, complete metadata) enforced. The eight empty flat siblings were deleted andchange_registry.pyis now a pure assembly point. - Superseded private re-exports, migration edges, and cycle exceptions removed — both slices landed with no stale allowlist entries remaining.
The model package and the *_metadata.py dataclass/parser split (each
format's facts in model/*_facts.py, re-exported by its parser) landed here
too — the split Phase 4 was blocked on.
What the label does not mean. Phase 5 closed the migrations it named; it
did not retire repository-wide legacy debt. The surviving flat parser
modules (pe_metadata.py, macho_metadata.py, dwarf_metadata.py,
symvers_metadata.py, and siblings) are still unclassified — their
extract classification is outstanding for every one of them — and the
storage/model ownership questions in gap E are untouched by it.
Durable lessons.
- A module that conflates a value with the code that produces it has no
valid single classification. The split (
model/*_facts.py+extract/*_metadata.py) is the general fix, and the same shape recurs for bundle facts and for snapshot persistence (gap E). - A shared decoder two layers both need moves to an inward leaf both may
import (
model/mangled_name.py), for the whole codebase-wide call-site set — not reactively, from whichever caller tripped the gate first. - Catalog completeness is enforced at construction, not reviewed: an entry
with no
impactfails at import time.
Acceptance: parser fixtures demonstrate byte/fact parity where applicable; catalog validation proves all four of D9's properties; no parser imports policy/report/workflows/frontends; the corresponding debt entries are removed. Met for the four items named above.
Remaining acceptance gaps¶
These are the differences between the landed phase slices above and this ADR's own definition of done. They are stated once, here, rather than tracked inside each phase's narrative. Every one is a dependency-and-ownership question; none of them is closed by a file getting shorter.
A. Real dependency violations, not their visibility to the checker¶
Closure package 2 re-measured this gap rather than trusting the list
above (this ADR's own repeatedly-learned lesson): a real, repo-wide AST
scan for every first-party importlib.import_module("...") call found 21
call sites, not the 6 this section used to name — ADR-063 track T10 had
already closed report/render_markdown_document.py's and
report/scoped_gate.py's bridges before this package started, and the
remaining 21 sort into three groups, not one:
- A real, forbidden-direction evasion (
workflows -> frontends, workflows may not import frontends):workflows/render.pyresolvingservice_render.pythroughimportlib.import_module("..service_render", __package__)inside each function body — the exact shape this section used to describe. Closed:workflows/render.pyretired;service.py(a flat,workflows-legacy-classified module, the one real caller) now importsservice_renderdirectly and statically. The edge itself is real and still crossesworkflows -> frontends— retiring the bridge module made it visible, it did not make the direction legal — so it is recorded as a revieweddependency-directionexception inarchitecture/debt.yaml's newdependency_direction_exceptions(see below), not silently passing. - A second real, forbidden-direction evasion this re-measurement
found that the original gap A text never named:
cli_dump_helpers.py(frontends) resolvingheader_conditionals.py(extract) the same way — frontends may only reach extract through workflows. Closed the same way: now a plain static re-export, recorded as a seconddependency_direction_exceptionsentry. - Legitimate same-layer or already-legal-direction bridges — the
other 19 call sites, all of the shape D6 already carves out
("genuinely dynamic or plugin-style loading is the narrow, documented
exception") or the same-layer back-compat re-export shim
AGENTS.md's own "Moving helpers out of a module that re-exports them?" guidance recommends:service.py'sservice_header_scopedbinding,workflows/input_resolution.py'sservice_dump_nativebinding,comparability.py<->comparability_profile.py,type_reachability.py<->type_reachability_stdlib_spellings.py,serialization.py<->bundle_facts_serialization.py,model/snapshot.py'ssemantic_ir_legacy_adapterassertion,policy/public_surface.py's two split-module re-exports,buildsource/source_graph.py/inline.py/template_graph.py's own split-module shims,reporter_markdown.py->report/ dispatch_markdown.py,annotations.py->annotations_step_summary.py,cli_buildsource.py->cli_graph.py/cli_buildsource_helpers.py,cli.py'sMOVED-table facade__getattr__, andfrontends/cli/commands/compare_bundle_facts.py's two bindings intoworkflows. Each one either stays inside one layer (so no direction is even at stake — the bridge exists purely to avoid growing the pre-existing, already-baselinedcli_buildsource/scan_engineimport cycle, or a same-layer back-compat split-module cycle) or crosses an already-legal direction (frontends -> workflows). None evadescheck_architecture.py's direction check in the sense this gap is about; each evades onlycheck_ai_readiness.py'simport-cycle-growthscan, which is deliberately a different, broader question (see the completion test below for why that scan is not widened here).detector_registry.py's plugin-discovery loop andpolicy/public_surface.py's dict-keyed target (aName, not a string literal) are D6's own named "genuinely dynamic" exception outright — no literal target to resolve at all.
The fix is composition at the outer boundary, not a better bridge —
for the service.py -> service_render edge specifically, this remains the
correct target, not yet fully reachable in one pass. workflows/render.py
has retired, and the import is real, static, and visible; what has not
yet happened is service.py itself ceasing to be workflows-classified
for this one responsibility, since service.py is also imported, directly
off the flat facade, by three other workflows-classified modules
(abicheck/l0_export_delta.py, abicheck/appcompat.py,
abicheck/service_scan.py) that would need their own D6 migration onto
the real workflow owner first — reclassifying service.py today would
just move today's invisible-bridge problem into three new, real
workflows -> frontends edges at those call sites instead of closing it.
That migration is recorded as the accepted exception's own stated
follow-up, not attempted in this pass. The cli_dump_helpers.py ->
header_conditionals.py edge has the same shape: closing it for real needs
a workflows-owned wrapper cli.py's and frontends/cli/commands/
dump.py's call sites route through instead of naming the extract-owned
functions directly.
Completion test — met for the direction check, deliberately not
extended to import-cycle-growth: scripts/check_architecture.py's
_imports() now resolves a literal importlib.import_module("...") call
(including a module-level alias such as _importlib = importlib, and the
__import__("importlib").import_module(...) chained form) as a real
import edge for the dependency-direction check, with focused unit tests
over miniature trees proving both directions: an evasion of a forbidden
direction fails, a same-layer or already-legal-direction bridge does not,
and a genuinely dynamic (non-literal) target is left alone rather than
guessed at. Every one of the 21 real call sites was re-checked against the
strengthened tool; the only two that turned into dependency-direction
findings are the two named above, both now resolved via the reviewed
dependency_direction_exceptions mechanism below rather than left dynamic
and unlisted.
check_ai_readiness.py's import-cycle-growth scan is deliberately not
widened the same way, on reconsideration of the task as originally
framed: that scan's own docstring already documents the identical
cli_buildsource -> cli_graph shim as its intended, narrow escape hatch
("If you switch a shim like that to a static import, expect this gate to
flag the cycle... Fix the direction... instead"), and AGENTS.md's own
"Moving helpers out of a module that re-exports them?" guidance
prescribes exactly this importlib.import_module pattern as the
correct way to preserve a back-compat re-export path without
recreating a real two-way file-level import cycle. Making that scan see
these edges would not surface a new architectural problem — every one of
the 19 legitimate bridges above is a deliberate, reviewed answer to a
real two-file cycle a split-for-file-size already created — it would
instead flag ~19 already-accepted, already-documented patterns across the
codebase as new cycle growth, which AGENTS.md's own "Don't extend
IMPORT_CYCLE_ALLOWLIST... as a routine step" rule and this closure
package's "no new IMPORT_CYCLE_ALLOWLIST entries" constraint together
rule out fixing by allowlisting. Widening that scan is not this gap's
target (gap A is about layer-direction violations hidden from
check_architecture.py, D6's own framing); doing so anyway would trade a
closed gap for a large, unrelated wave of allowlist churn against a
policy this repository has already, deliberately, decided the other way.
architecture/debt.yaml's new shape: dependency_direction_exceptions
is the ledger shape this section previously said closure package 2 owed —
distinct from the files/no_growth schema (rule: "dependency-direction",
keyed by (path, target), not a line-count baseline), holding exactly the
two edges above, each with an owner, a dated review, and a rationale
naming the follow-up migration that would close it for real.
scripts/check_architecture.py validates the new list's own schema (a
malformed entry suppresses nothing) and consults it only for the exact
(path, target) pair it names — every otherdependency-direction finding
still fails the gate.
B. Public compatibility surfaces separated from ownership exemptions¶
architecture/modules.yaml's public_root_surfaces currently lets a
migrated package import checker_policy, reclassify, and serialization
without those modules having an owning layer. Two different questions are
being answered by one list: is this import path supported for external
users? and which layer owns the implementation behind it? A "yes" to the
first does not remove the need to answer the second.
DiffResult is the worked example. Its local override algorithm moved out
(checker_policy.apply_policy_file_overrides), so no policy algorithm runs
inside checker_types.py any more — but its methods still call
checker_policy and reclassify, both unclassified public-root leaves, so
the model -> policy-shaped dependency is separated in code without being
removed.
The fix: give the implementation behind each entry a real owner and keep
only the compatibility adapter at the public path; canonical internal
callers use the owner. This is not to be implemented by caching
policy-resolved values on today's mutable DiffResult — callers and tests
change policy after construction, and Phase 2's own lesson rejects that
cache. Preserve the mutation behavior at the supported legacy boundary and
move new internal processing onto completed, explicitly evaluated results.
Completion test: every module reachable through public_root_surfaces
names an owning layer; no canonical internal caller reaches an
implementation through the compatibility route; the supported import paths
still resolve, pinned by facade tests.
Closure status (2026-09-10). Of the eleven public_root_surfaces
entries this gap named, six moved their real implementation to a named
owning layer behind a thin, delegation-only flat facade, with every
physically-migrated internal caller switched to import the owner directly:
api_types -> workflows.request_inputs/workflows.contracts (plus
model.header_ast_frontends for the one constant extract-classified
buildsource/header_compile_context.py also needs); checker_policy ->
policy.classification/policy.evidence_status; contract_coverage_ledger
-> policy.coverage_ledger; qualified_name_segments ->
compare.qualified_name_normalization (the diffing-decision half) and
storage.closure_identity (the snapshot-normalization half, sharing
storage's own qualified_name_segments_walk.py leaf). Two entries
(errors, dumper_contract) already had a real owner recorded in that
owning layer's legacy_paths — their public_root_surfaces membership was
simply stale bookkeeping once the redundant migrated-caller exemption was
confirmed unused, not an open gap; both were dropped from the list.
Three entries — checker_policy, contract_gating, reclassify — stay
deliberately unclassified, confirmed (not merely documented) to be the "no
single layer" leaf this gap's own text anticipates: abicheck/checker_types.py
(model-owned, and model's ADR-061 imports are the standard library only)
imports each directly, so giving any of the three a policy home would turn
that pre-existing import into a real, gate-checked direction violation.
checker_policy's real implementation still moved to policy.classification/
policy.evidence_status — the fix is "give the implementation a real owner,"
not "every facade becomes classified" — contract_gating's to
policy.contract_finding_relevance, and reclassify's to policy.reclassify;
only the thin flat facade stays off every layer's legacy_paths so the one
caller that structurally cannot depend on policy keeps resolving through
it, unchanged.
A twelfth entry was added during this closure, not removed: contract_evidence
(needed by the new policy.coverage_ledger's own TYPE_CHECKING import, but
also read by frontends-classified cli_compare_receipt.py/
cli_scan_receipt.py, so it cannot become policy-classified without
breaking those) — the identical "no single layer" shape, discovered rather
than pre-existing.
abicheck.schemas was investigated and, despite looking like a clean model
move at first (a dependency-free leaf of JSON Schema documents and version
constants), turned out to be this gap's sharpest example of why the two
questions must stay separate: its own current() reaches forward into
report and workflows to answer "what version does abicheck emit for
artifact X" from one place, while workflows/report/frontends callers
reach into it for the reverse direction — a package that is both a callee
and a caller of the same two layers has no ADR-061 layer that doesn't close a
real cycle. Confirmed by attempting the model reclassification directly
against scripts/check_architecture.py, not by inspection alone. Stays
public_root_surfaces-listed.
serialization's own gap-E item is now closed (closure package 6, see that
gap's own entry): its ~1500-line codec proper moved to a real
storage-classified home (storage/snapshot_codec.py and four siblings).
The flat serialization.py facade itself stays public_root_surfaces-listed
— confirmed, the same way this gap already confirmed checker_policy/
contract_gating/reclassify, to be a genuine "no single layer" leaf: it is
the one legal route through which the workflows/policy steps
storage.snapshot_codec.decode_snapshot/finalize_snapshot cannot
themselves run still execute.
header_only_dump is unchanged and is not a gap-B instance at all on closer
reading: its own module docstring already states, and this closure
re-confirmed, that its flat placement is load-bearing — it bridges to
dumper.py/dumper_manifest.py, both still unclassified, and a migrated
package importing either would trip unclassified-import the moment
header_only_dump itself moved into one. Its real fix is classifying
dumper.py, a separate, much larger migration this gap does not attempt.
Net: eight of the eleven original entries now name a real owner (six moved,
two already had one); four (checker_policy, contract_gating,
reclassify, and — closure package 6 — serialization) are confirmed, not
merely asserted, "no single layer" leaves, serialization's own former
migrate gap now closed the same way; public_root_surfaces itself shrank
from eleven to seven entries (checker_policy, contract_evidence,
contract_gating, header_only_dump, reclassify, schemas,
serialization) — every remaining one a reviewed exception with a stated
reason, not an unclassified default.
C. One result, one document, several projections¶
Phase 2 gave every format a real fact-vs-formatting boundary and a
ReportDocument of its own. What it did not establish is the invariant the
definition of done states:
one completed evaluation
-> one completed semantic report document
-> JSON / Markdown / HTML / SARIF / JUnit
service_render.render_output() still dispatches the DiffResult,
snapshots, severity configuration, and presentation options down six
independent paths. This is not evidence that today's formats disagree; it
means that wrapping their separately-built projections in the same immutable
type cannot prove they can't. It matters directly to
ADR-068's "one
analysis, several artifacts" experience, which also makes the canonical
report the replacement for the separate scan schema.
The fix: finalize compatibility, assurance, scope, dispositions,
consumer impact, and the exit decision before format selection, and build
the shared report representation from that completed result. Format-specific
builders may still arrange presentation; they may not be independent
authorities for those decisions. Keep Phase 2's immutable containers and
projection tests and build on them. Coordinate with ADR-063 and ADR-068
rather than starting a second reporting-convergence project — the design is
duplication-and-convergence-assessment.md
Phase 4's ReportEnvelope.
Completion test: render the same completed document repeatedly and in different format orders; the semantic content is identical every time, and no renderer re-runs extraction, policy evaluation, or gate resolution.
Closed by closure package 3 — see the gap C status section under Phase 2
above for what landed and how it was verified. abicheck/report/envelope.py
holds the ReportEnvelope, report/build.build_report_envelope builds it
once above format selection, and service_render.render_envelope projects
it. The completion test is executable in
tests/unit/report/test_build_report_document.py's
TestRendererOrderIndependence (order-independent byte-identity across
formats and repeated renders, plus a call-count spy proving every decision
function runs once at envelope-construction time and never during a
projection).
D. Typed request/plan and operand convergence¶
Re-measured, closure package 4 (this section was stale — it described a 2026-09 snapshot of the tree that several PRs have since moved past; see this ADR's own "A blocker recorded once goes stale" lesson under Phase 3). The gap's two original halves are now in different states:
- The gate-shape half is closed.
gate.py's two callers (compare'sResolvedCompareConfigand the release fan-out'sGateOptions) no longer fold onto independent shapes:policy/effective_gate.py'sEffectiveGate/GateSeverityState/ScopedGateSelectionis the one converged runtime object both resolve to (ResolvedCompareConfig. effective_gate/GateOptions.effective_gate, andworkflows.gate.effective_gate_for_resolved_compare_configfor the former's own no-growth-capped module), covering severity, exit-code scheme,require_complete_analysis, and ADR-043 scoped-gate selection — seetests/test_effective_gate.py's own characterization/completion split. This is deliberately narrower than the plan's fullEffectiveEvaluationConfig(policy/contract/assurance/surface/evidence/ suppressions namespaces beyond gate) — see this ADR's link todocs/contribute/plans/duplication-and-convergence-assessment.md's P0 section, which still names that wider object as not yet attempted. - The scope/inventory/acquisition-state half is landed for the release
fan-out's own resolution, but not yet reachable from a typed Python
entry point.
workflows/release_scope.py'sReleaseScopePlan/ReleaseScopeResult(a realRequest -> ResolvedPlan -> Resultpair for ADR-065's scope model) replaced the four independent locals (old_map/new_map/matched_keys/inventory_evidence)cli_compare_release.pyused to thread by hand, and is the actual execution-authoritative input, not a DTO computed alongside it (seetests/test_release_scope_plan.py's completion tests, includingTestScopePlanIsExecutionAuthoritative). This closure package's own next slice widened that to the entire pre-execution resolution — input discovery, inventory evidence, the scope plan, the resolvedGateOptions, and each side's stored-degraded markers — as one typed request/plan pair,frontends.cli.release_compare_request.ReleaseCompareRequest/ReleaseComparePlan/resolve_release_compare_plan: selection, inventory, and acquisition state are now fields on that shared plan, constructible and resolvable with one ordinary Python function call, not frontend locals (tests/test_release_compare_request.py'sTestReleaseCompareRequestParityis the completion test: a CLI-shaped invocation and a direct typed-request-shaped call resolve to the same scope, gate configuration, and degraded-member markers).
Closed (closure package 4, final slice). The account this section
carried — that resolve_release_compare_plan was reachable from Python with
no Click context but not from abicheck.service, because the functions it
must call (_prepare_compare_release_inputs, frontends.cli.
release_variant_operand._resolve_release_package_side and their siblings:
input discovery, package extraction, stored-variant resolution) were
classified frontends/flat cli_* and at least one raised a real
click.UsageError from inside the resolution — described the state before
that migration slice landed. It has now landed, as this ADR's own rules
require rather than as a wrapper:
- that whole call chain is
abicheck/workflows/release_inputs.py, and everyclick.UsageError/click.ClickExceptionit raised is the typederrors.ReleaseOperandError(aValidationErrorsubclass, so existing usage-error translation already covers it); - the request/plan pair itself is
abicheck/workflows/release_request.py, andabicheck/frontends/cli/release_compare_request.pyis a delegation-only facade whose one remaining job is translating that typed error back into aclick.UsageError— so a user's message and exit64are unchanged; abicheck.serviceexposes it asresolve_release_compare, alongsideReleaseCompareRequest/ReleaseComparePlan/cleanup_release_compare_plan.
tests/test_release_request_parity.py is the completion test for this half,
and it is behavioural on both sides: a real compare CLI run and a direct
service.resolve_release_compare call resolve the same scope, gate and
markers (parametrized over cardinality), and the same malformed operand
raises a typed error for the API caller while still exiting 64 with the
identical message for the CLI one. It deliberately asserts none of this
from file text: an import click grep would pass while the two surfaces
behaved differently.
Also still open (unchanged by this slice): the wider
EffectiveEvaluationConfig namespaces beyond gate (above), and the release
fan-out's execution half (per-library dump/compare dispatch, matrix/probe
expansion, bundle-facts writing, output rendering) — this closure package
covers only the resolution half, matching workflows/artifact/contracts.py's
own Milestone A/B precedent of resolving before executing. The missing
dump fan-out that should own baseline capture rather than
compare --bundle-facts-out is recorded in docs/contribute/known-gaps.md;
it was blocked on exactly the migration above, and is now unblocked.
Completion test: equivalent CLI and typed-API inputs produce equal
resolved scope, configuration, acquisition records, and outcomes across
live and stored operands; selection, inventory, and acquisition state are
fields on the shared request/plan, not frontend locals. Met for the
release fan-out's own pre-execution resolution, reachable as one direct
Python call and from abicheck.service with typed errors -- see the
completion test named above. Not yet met for the execution half, which is
the remaining scope above.
E. Storage and model ownership, not file placement¶
Three unresolved owners, each a value-versus-persistence-versus-
orchestration conflation of the shape Phase 5 already solved for
*_metadata.py:
serialization.pyhas taken real decomposition (platform blocks, severalstorage/codecs). Closed (closure package 5, slice 1):snapshot_from_dict's legacy backfill call topython_ext.detect_python_extension()— evidence derivation inside the loader, whichstorage'smay_import: [model]could not admit — now runs as an explicit post-load step inworkflows/snapshot_load.py, reached throughserialization.snapshot_from_dict's unchanged public signature. The siblingsnapshot_platform_blocks.py's ownstorage -> extractedge (its_xxx_from_dicthelpers importing dataclasses from the flat parser modules) closed the same slice by switching those ~10 imports to each dataclass's canonicalmodel/*_facts.pyhome;snapshot_platform_blocks.pyis now classifiedstorage.serialization.pyitself stays unclassified (thepublic_root_surfacescompatibility-facade treatment) — its own ~1500 remaining lines of codec logic are a separate, not-yet-attempted classification.bundle_facts.pyand its serialization/store siblings were classifiedworkflows, conflating theBundleFactsvalue, its persistence, and capture/comparison orchestration. Closed: applied the same split Phase 5 proved for*_metadata.py. The value type and its construction invariant (require_degraded_members_known, applied at every construction/import choke point) moved tomodel/bundle_facts.py. Persistence split three ways instorage/:bundle_facts_codec.py(JSON (de)serialization, moved out of the flatbundle_facts_serialization.py),bundle_facts_archive.py(the G40 content-addressed zip archive, moved out ofbundle_facts.py's own G40 section), andbundle_facts_package.py(the multi-artifactProjectSnapshotpackage adapter, moved out ofbundle_facts_store.pyand reclassifiedstorage— its own historical docstring had explained why it couldn't bestorageyet, precisely becauseBundleFactshad no settled layer; once it did, every dependency this module actually has turned out to already bestorage-legal). Capture and comparison orchestration moved toworkflows/:bundle_facts_capture.py(capture_bundle_facts/bundle_snapshot_from_facts) andbundle_facts_compare.py(compare_bundle_from_facts). The flatbundle_facts.py/bundle_facts_serialization.py/bundle_facts_store.pymodules are now delegation-only compatibility facades (added tomodules.yaml'sfacadeslist, each under the 150-line facade cap, each with an explicit__all__) re-exporting the same public names, so the documented Python API path (docs/use/multi-binary.md) and every existing internal/test call site are unaffected. The one real, unavoidableserialization.py <-> storage.bundle_facts_codectwo-file cycle (bundle_facts_codec.pyneedssnapshot_to_dict/snapshot_from_dictfromserialization.py;serialization.py's own back-compatbundle_facts_to_dict/etc. wrappers need the codec) is kept dynamic viaimportlib.import_module, same shape as before the move, just retargeted — not a new bridge, and not anIMPORT_CYCLE_ALLOWLISTentry. Every other internal caller (bundle_multibuild.py,bundle_side_input.py,cli_compare_release_helpers.py,storage/import_bundle_facts.py,storage/variant_composition.py,workflows/bundle_compare_operand.py,workflows/bundle_stored_pair_compare.py,workflows/release_scope.py) now imports the canonical owner directly, per D6. Legacy-reader behavior (schema-version fallback, thedegraded_membersmarker, the G40 archive format, resource-limit hardening) is pinned by the existingtests/test_bundle_facts*.pysuite, exercised unchanged before and after the move; two of those files'monkeypatch-based tests were updated to patch the real owner module (storage.bundle_facts_archive/storage.bundle_facts_package) instead of the facade, per D10 ("unit tests patch the owner module").probe_harness.py(compare) neededsnapshot_to_dict/snapshot_from_dictto serialize its own probe matrix. Comparison logic must not become the owner of persistence because a probe workflow needs serialized inputs;compare'smay_import: [model]offers no facade route. Closed (closure package 5, slice 2): the conversion functions (ProbeResult.to_dict/MatrixSnapshot.to_json/.from_dict, and the module-levelwrite_matrix_snapshot/load_matrix_snapshot) moved out ofprobe_harness.pyentirely intoabicheck/workflows/findings.py— already this ADR's own documentedworkflowsre-export surface for "the probe matrix" — which legally imports bothcompare(for theProbeResult/MatrixSnapshotdataclasses) andabicheck.serialization(thepublic_root_surfacesfacade).ProbeResult/MatrixSnapshotare now pure value objects with nostorageimport;probe_harness.pyitself still ownsload_probe_spec/run_probe_matrix(YAML parsing and compile orchestration), unchanged.probe_harness.pyis documented Python API (docs/use/probe-harness.md) — the compatibility question was never whether the module is public, only whether the removed serialization helpers specifically needed a shim at their old path; they didn't, since the documented workflow's supported path is that doc page, which now imports the moved names fromworkflows.findingsinstead (verified round-tripping) — every internal caller (CLI, tests) was switched to the new home in the same slice.serialization.py's own ~1500-line codec proper — the last open item this gap and gap B's own closure status both named. Closed (closure package 6): the encode direction (snapshot_to_dict/snapshot_to_json/snapshot_content_digest), the schema-version history/thresholds, the declarations decode (functions/variables/types/enums/typedefs and their entity-id sidecars), the seven*_facts_reliableflag computations, the platform-block/provenance decode, and the finalAbiSnapshot(...)assembly all moved to a realstorage-classified home —storage/snapshot_codec.pyplus four siblings (snapshot_schema_versions.py,snapshot_encode.py,snapshot_decode_declarations.py,snapshot_reliability_flags.py), split purely to keep each file under the ADR-061 new-file production line ceiling (mechanical extraction, verified againstcheck_architecture.pydirectly — a brand-new file has no adoption-debt exemption available, so each sibling had to clear 800 lines on its own merits, not via a baseline).serialization.pyitself shrank to a thin orchestration-only facade and now joinschecker_policy/contract_gating/reclassifyas a confirmed (not merely documented) "no single layer" leaf, for the identical structural reason gap B's own closure status names for those three: it is the one legal route through which two genuinely storage-illegal steps —workflows.snapshot_load.backfill_python_ext_from_evidence(real evidence-derived extraction, not a fact lookup) andpolicy.analysis_assurance_degraded_facts.degraded_reliability_facts(an assurance judgement over an already-decoded snapshot) — run in betweenstorage.snapshot_codec.decode_snapshotandstorage.snapshot_codec.finalize_snapshot. Reclassifying the facade itself asstoragewould turn those two into real, gate-checkedstorage -> workflows/storage -> policydirection violations, the same way gap B's own investigation found forchecker_policy/contract_gating/reclassify.serializationtherefore stays inarchitecture/modules.yaml'spublic_root_surfaces— it cannot move tofacadeseither (that list's 150-line cap and delegation-only shape don't admit the warning/backfill orchestrationsnapshot_from_dictstill does). The one realserialization.py <-> storage.bundle_facts_codeccycle is unchanged, still resolved dynamically viaimportlib.import_module, not a newIMPORT_CYCLE_ALLOWLISTentry.docs/contribute/known-gaps.md's entry andarchitecture/debt.yaml'sabicheck/serialization.pybaseline were both updated to record the closure.
The owners to establish are: model for snapshot/bundle value types and
their invariants; storage for codecs, schemas, persistence, and schema
migration; extract for evidence derivation; workflows for capture, load,
enrichment, and comparison orchestration. For legacy loading, decide
explicitly where evidence-derived backfill runs and preserve supported
reader behavior — do not remove or relocate it without auditing every direct
snapshot_from_dict() caller. Coordinate with
ADR-062; this is not a competing
storage redesign.
Completion test: bundle values, persistence, evidence backfill, and
orchestration each name one owner; legacy-reader behavior is pinned by tests
written before the move (Phase 3's rule); no compare-classified module
owns a persistence operation.
F. A repository-wide completion track¶
The definition of done requires more than the six phase labels: explicit
ownership for the remaining root modules, legal imports, no responsibility
cycles, shared workflow contracts, pure report projections, bounded facades,
and debt either retired or explicitly accepted. The live ownership map and
architecture/debt.yaml still hold substantial legacy implementation and
migration debt.
This is not "move every file now" — this ADR rejects mass mechanical relocation (see "Split every oversized module immediately" under Alternatives). What it requires is that every remaining legacy area carries a disposition: migrate through a named responsibility slice, retain as a genuinely supported public module, or accept a specific architectural exception with a reason. "Unclassified for now" and "see the ownership map" are not dispositions.
Architectural debt is also tracked independently of line count. A short
module can hold a forbidden dependency; a large, well-owned parser can be a
legitimate reviewed exception. Several ownership questions in this ADR never
got a debt.yaml entry only because the file was not oversized.
Completion test: every unclassified first-party module under abicheck/
has one of the three dispositions recorded; debt.yaml holds only accepted
exceptions, described as exceptions rather than as migration work.
Closure sequence¶
This replaces the open-ended Phase 4 narrative. Packages 3-5 may run as
bounded parallel slices where they do not compete for the same
result/request contracts. Package 6 must not race ahead of ADR-068: deleting
a scan surface before its capabilities and consumers migrate is the
failure that plan's phase ordering exists to prevent.
| Order | Work package | Completion condition |
|---|---|---|
| 1 | Reconcile ADR scope and tracking | Status, acceptance gaps, and the links to ADR-062/063/068 agree; scan no longer appears as a future architecture example |
| 2 | Close enforcement escapes and the rendering back-edge (gap A) | A literal importlib.import_module call naming a first-party module is visible to check_architecture.py's direction check as a real edge; workflows/render.py retires; the remaining workflows -> frontends/frontends -> extract edges this re-measurement found are real, static, visible, and recorded as reviewed dependency_direction_exceptions (not silently dynamic) pending the further migration each names; the public path stays, composed at an outer adapter |
| 3 | Converge completed results and report projections (gap C) | One evaluated result supplies every format; no renderer derives a competing gate, disposition, or assurance decision |
| 4 | Finish typed request/plan and operand convergence (gap D) | Equivalent CLI/API inputs produce equivalent resolved scope, configuration, acquisition records, and outcomes; coordinated with ADR-068's shared-driver work |
| 5 | Close the storage/model splits (gap E) | Bundle values, persistence, evidence backfill, and orchestration have explicit owners; legacy-reader behavior is tested |
| 6 | Finish the facade and legacy retirement (gaps B and F) | Internal callers use canonical owners; unnecessary private shims are gone; scan-related surfaces retire only after capability migration; every remaining exception is explicit |
How closure is validated. Exercise the architecture's promises through real public paths, not through internal detectors: equivalent CLI/API resolution; live and stored operands; selected versus missing members; stripped-binary and header-only tasks; suppressed findings that remain visible; consumer impact that leaves the global gate intact; and repeated multi-format rendering with no re-evaluation. Assert compatibility, assurance, scope, gate, and process exit separately — a passing exit code is not proof of any of the other four.
Migration rules for every phase¶
Each migration PR must be a vertical, behavior-preserving slice and must:
- identify the old owner, new owner, supported public paths, and debt entry;
- move implementation and its unit tests together;
- switch internal callers to the new implementation module in the same PR;
- leave only necessary public delegation in the old module;
- add or update compatibility-import tests for retained public paths;
- prove no new package imports the old facade;
- preserve output/schema behavior unless the PR separately declares and tests a product change;
- update
modules.yamlcoverage and shrink/removedebt.yamlentries; and - run the canonical PR verification profile.
Line-count reduction without ownership transfer does not satisfy a phase.
Alternatives considered¶
Keep the flat package and enforce only a lower line limit¶
Rejected. It encourages more prefix siblings and mechanical splits while leaving ownership and dependency direction undefined. The repository would have smaller files with the same coupling graph.
Split every oversized module immediately¶
Rejected. A mass move creates review noise, import churn, and compatibility risk before target contracts are enforceable. Incremental vertical slices allow parity tests and facade decisions per responsibility.
Preserve every old private import and monkeypatch location¶
Rejected. That makes incidental implementation paths permanent and forces new packages to import through legacy owners. Only documented public paths receive compatibility treatment; internal tests move to the real owner.
Use one broad core package¶
Rejected. core would reproduce the current ambiguity inside a directory.
The eight packages are based on decisions and data transformations, not a
generic notion of importance.
Allow dependency cycles during migration¶
Rejected as a stable policy. Explicit, expiring debt records can describe a temporary edge, but the target graph remains acyclic and no permanent allowlist is created.
Create the full destination tree up front¶
Rejected. Empty directories communicate false progress and create package surfaces with no owner. A package appears when implementation and tests move.
Put all architectural checks into check_ai_readiness.py¶
Rejected. Architecture validation is one focused concern with its own configuration and tests. The verification orchestrator should invoke it, not absorb its implementation.
Consequences¶
Positive¶
- The filesystem answers where new behavior belongs.
- Cross-package dependencies become reviewable and machine-checkable.
- Typed stage boundaries reduce duplicate resolution and frontend drift.
- Compatibility obligations are explicit rather than inferred from every historical internal import.
- Report formats cannot silently disagree about findings, verdicts, or gates.
- Agent guidance becomes shorter because ownership moves into the tree and scoped package contracts.
- Debt is visible as temporary data with owners and review dates rather than normalized by an ever-increasing maximum file size.
Costs and risks¶
- Migration temporarily increases the number of facade and target modules.
- Import-path churn can disrupt tests and external users if public/private boundaries are not explicitly audited.
- An over-eager dependency checker can misclassify dynamic or optional imports; its tests and errors must distinguish unsupported edges from parser limitations.
ReportDocumentmigration may expose format-specific decisions that have accidentally diverged and require deliberate reconciliation.- Reducing root instructions requires careful relocation, not deletion, of unique operational knowledge.
- Until all debt entries are retired, contributors must understand both the target architecture and explicitly recorded legacy exceptions.
Definition of done¶
The repository-wide migration is complete when:
abicheck/root contains only entry points, supported facades/public modules, and responsibility packages.- No new root
cli_*,service_*,reporter_*,diff_*, or equivalent pseudo-package sibling exists. - No ordinary new production module exceeds 800 lines.
- Existing files above 800 lines cannot grow without an explicit reviewed debt-baseline change.
- Every cross-package import follows
modules.yaml. - No responsibility-package dependency cycle exists.
- Root
cli.pyandservice.pyare delegation-only facades below 150 lines — reached by reducing coupling, and checked last. A facade that hits the number by trimming documentation, or that keeps private shims with no external contract, does not satisfy this. - Every major operation follows
Request -> ResolvedPlan -> Result, and the plan carries selection, inventory, and acquisition state rather than leaving them to a frontend. - Dry-run renders the actual resolved plan.
- Every output format consumes one immutable
ReportDocumentbuilt once per completed evaluation — not one document per format. - Extraction cannot import policy, report, workflows, or frontends.
- Compare cannot decide suppression, severity, or exit status.
- Renderers cannot alter findings, verdicts, or gate state.
- Root
AGENTS.mdis a stable routing contract below 350 lines and scoped package instructions exist for every responsibility package. - A contributor adding an ELF fact, detector, policy rule, CLI flag, or report field can identify its owner from one routing table without first opening a legacy monolith.
architecture/debt.yamlis empty or contains only explicitly accepted exceptions that are no longer described as migration work, and every remaining unclassified first-party module carries one of the three dispositions in gap F.- No first-party dependency is hidden behind a dynamic import to keep it
out of the architecture checks (D6); the checks resolve literal
importlib.import_moduleedges. - Every module reachable through
public_root_surfacesnames an owning layer, and no canonical internal caller reaches an implementation through the compatibility route. - Persistence, evidence derivation, value types, and orchestration have distinct owners for snapshots and bundle facts alike.
Items 1-16 are the original criteria; 17-19 make explicit the guarantees the remaining acceptance gaps found were not covered by a passing check.
The immediate deliverable after acceptance is Phase 0: establish ownership, contracts, scoped guidance, and no-growth enforcement. Splitting another large file before those constraints exist is not progress toward this ADR by itself.
Amendment (2026-09-13): the delegation-only facades are retired, not preserved¶
This ADR's Phase 1 acceptance criterion reads "public imports covered by
contract tests continue to work", and its extraction rule (mirrored in
AGENTS.md and abicheck/AGENTS.md) says to "preserve only documented
public imports through a thin explicit façade". Twelve such facades had
accumulated under that rule. They are now deleted outright, and the
rule no longer applies to abicheck's own historical Python import paths:
abicheck.aggregate, aggregate_findings, aggregate_manifest,
api_types, bundle_facts, bundle_facts_serialization,
bundle_facts_store, contract_coverage_ledger,
qualified_name_segments, and buildsource.{entity_identity,
entity_resolver, source_graph_query}. Importing any of them now raises
ModuleNotFoundError; import the owning module named in each deletion's
changelog entry, or — for the typed request/result types — the supported
abicheck.service surface.
Why the preservation rule is retired rather than applied. It was written to keep a migration from breaking callers while the migration was in flight, and it did that job. What it did not anticipate is that the facades would become permanent: each one is a second name for a thing that already has an owner, so every subsequent reader has two places to look, every new contributor has two import paths to choose between, and the "which do I import?" question has to be answered in a docstring on each facade — which is why several of them are longer than the rule's own 150-line ceiling would suggest is possible for a pure re-export. ADR-043 already reset the CLI surface on the basis that abicheck is pre-1.0; the typed API's historical import paths sit on exactly the same footing. An accepted document describing an interface as public is a reason to record its retirement, not on its own a reason to preserve it forever.
This is compatibility of abicheck's own Python interfaces, and it must not be confused with the ABI/API compatibility the product analyses for its users — which is unchanged, and which this batch validated end to end against real compiled binaries across every operand shape.
What replaces the criterion. Phase 1's acceptance now reads: semantic
results and JSON are exactly compatible; internal imports use the new
owner; no reverse facade import or duplicated decision exists; the relevant
debt entries shrink or disappear; and every retired import path is named in
a changelog fragment with its owner. tests/test_adr061_gap_b_facades.py
still holds the remaining facades to their delegation-only contract —
checker_policy, contract_gating and reclassify, which stay because
model-owned checker_types.py imports them and model cannot depend on
policy. Those three are a real dependency-direction constraint, not a
compatibility promise, and they go when DiffResult's policy lookups move
out of model.