Aggregate Reports: Folding a CI Matrix into One Gate¶
abicheck aggregate is a report fan-in, not another way to compare
binaries. It never parses a .so/.dll/.dylib and never runs a header
scan — it reads a directory of already-produced compare/scan JSON
reports (one per CI matrix leg) and reconciles them into one gate decision
and, when the reports came from more than one compiler/build profile, one
reconciled finding matrix.
Three commands, three different jobs¶
It's easy to conflate aggregate with compare's own multi-library mode or
with project — all three touch "more than one thing at once," but they
answer different questions:
| Command | Operand | Answers |
|---|---|---|
compare OLD NEW (directory/package inputs) |
Real binaries — several DSOs in one release bundle | "Does this one release, built once, stay compatible?" Fans a directory/package compare out per library through the same Tier-2 service.run_compare chokepoint a single-pair compare uses, and can additionally check cross-DSO relationships (a release-wide bundle/dependency analysis). |
project validate / project plan |
.abicheck.yml's targets:/profiles:/checks: |
"What should CI check, and how do the pieces fit together?" Validates the declared topology and generates the run plan a CI matrix executes — it doesn't read or produce a compare report itself. |
aggregate REPORTS_DIR |
A directory of already-produced *.json reports |
"Given what the matrix actually reported, does the whole run pass?" Never analyzes a binary — it reconciles reports against the target set the matrix was supposed to build. |
Put differently: compare's directory/package mode fans out within one
toolchain/platform leg (comparing several libraries built the same way in
one process); aggregate fans in across separate CI legs — different
platforms and/or different compiler profiles for the same target(s), each
of which produced its own report independently, possibly using compare's
directory/package mode itself for that one leg.
aggregate replaces the hand-written post-matrix for path in glob('*.json')
heredoc some projects grew organically — that loop silently drops any target
whose build failed before uploading its report, passing green while a
required platform was never analyzed. aggregate's one invariant is the
fix: an expected target with no report is unavailable (unknown), never
folded into the result as compatible.
The five orthogonal axes¶
Every aggregate run answers five independent questions. compatibility is
reporting-only — the worst ABI verdict, shown for context, never itself an
input to the exit code. The exit code is the worst contribution across the
other four (gate/coverage/contract_coverage/analysis_assurance,
max, never additive):
flowchart TD
R["Per-target compare/scan<br/>JSON reports"] --> A["compatibility<br/>(worst verdict — reporting only)"]
R --> B["gate<br/>(each report's own severity/scan<br/>gate, combined — never recomputed)"]
R --> C["coverage<br/>(did every REQUIRED target<br/>report at all?)"]
R --> D["contract_coverage<br/>(for a target that DID report,<br/>was its own contract evidence complete?)"]
R --> F["analysis_assurance<br/>(for a target that DID report,<br/>was its own evidence complete<br/>under assurance.require_complete?)"]
A -.->|context only| E["exit code = max(...)"]
B --> E
C --> E
D --> E
F --> E
- compatibility — the worst ABI verdict over the analyzed targets.
Reported for context; it does not by itself decide the exit code, since a
policy can make a
COMPATIBLEreport block (addition=error) or aBREAKINGreport pass (a demoted severity preset). - gate — each report already carries its own gate decision
(
severity.{exit_code,blocking,blocking_categories}, or ascanreport's own top-levelexit_code).aggregatecombines those — it never recomputes a gate from the compatibility verdict. Reading is fail-closed: a report whose gate block is present but corrupt makes that target unavailable, never silently reverting to the legacy path. - coverage — did every required expected target actually report at
all? A required target with no report is a coverage gap, exit
1— never promoted to a fake ABI-break exit4. - contract_coverage (
abicheck.workflows.aggregate.AGGREGATE_SCHEMA_VERSIONis the versioned fact owner) — for a target that did report, was its own selected--contractdomain's evidence complete? Read back from that report's owncontract_coverage_exit_contributionand folded withmax, exactly the waycompare/scan --againstfold theirs. This is a different question from plaincoverage: a required target can report successfully (no coverage gap) while its own contract-evidence domain was still incomplete (a contract-coverage gap) — both can independently produce exit1, for unrelated reasons, and the JSON output records which targets caused which. - analysis_assurance (P0.4, aggregate schema 1.5) — for a target that
did report, was its own evidence complete under
.abicheck.yml'sassurance.require_complete: true? Read back from that report's ownanalysis_assurance_exit_contributionand folded withmax, exactly the waycompare/scan --againstfold theirs. The exact sibling ofcontract_coverageabove, for a different question: a required target can report successfully with a closed--contractdomain while its own broader evidence (depth, TU/export accounting, header-context drift, ...) was still incomplete — both axes can independently produce exit1, and the JSON output records which targets caused which.
An illustrative report (aggregate_schema_version is omitted here since
abicheck.workflows.aggregate.AGGREGATE_SCHEMA_VERSION is its fact owner, not this
example — every real report carries the field):
{
"status": "fail",
"compatibility": {"verdict": "BREAKING", "analyzed_targets": 2},
"coverage": {
"status": "partial",
"required_targets": 3,
"analyzed_required_targets": 2,
"missing_required_targets": ["windows-x86_64"],
"blocking": true
},
"gate": {
"passed": false,
"exit_code": 4,
"blocking_targets": ["linux-x86_64"],
"coverage_blocking": true
},
"contract_coverage": {
"exit_contribution": 0,
"incomplete_targets": []
},
"analysis_assurance": {
"exit_contribution": 0,
"incomplete_targets": []
},
"targets": [
{
"target_id": "linux-x86_64",
"required": true,
"state": "analyzed",
"compatibility_verdict": "BREAKING",
"gate": {"exit_code": 4, "blocking": true, "blocking_categories": ["abi_breaking"], "from_report": true},
"contract_coverage_exit": 0,
"analysis_assurance_exit": 0
},
{
"target_id": "macos-arm64",
"required": true,
"state": "analyzed",
"compatibility_verdict": "COMPATIBLE",
"gate": {"exit_code": 0, "blocking": false, "blocking_categories": [], "from_report": true},
"contract_coverage_exit": 0,
"analysis_assurance_exit": 0
},
{
"target_id": "windows-x86_64",
"required": true,
"state": "unavailable",
"compatibility_verdict": null,
"gate": null,
"contract_coverage_exit": 0,
"analysis_assurance_exit": 0,
"reason": "no report was produced for this expected target"
}
],
"unexpected_targets": []
}
contract_coverage is present in every -o json=... output, with
exit_contribution: 0 and an empty incomplete_targets list when no
target's report used --contract — it is never omitted. analysis_assurance
is the exact sibling, present the same way, with exit_contribution: 0 and
an empty incomplete_targets list when no target's report used
assurance.require_complete: true.
Declaring the expected-target set¶
Exactly one of these is required (a bare aggregate reports/ with none of
them is a usage error, exit 64 — with no declared target set the gate
cannot tell a missing required target from an intentionally absent one):
--manifest abi-targets.json—{"targets": [{"id": "linux-x86_64", "required": true}, ...]}. Recommended: generate it once in the plan job and feed the same file to both the matrix and the gate.--manifest run-plan.json— the same flag also takes aproject planrun-plan and projects it internally into the manifest shape. Which shape a document is comes from its own content (a run-plan declaresschema: abicheck.run-plan/vN), never from its filename, so either artifact can arrive under any name the CI job gives it.--discovered-only— aggregate whatever reports are present with no required-target coverage gate (a missing target is simply not counted, never a coverage failure). This disables only thecoverageaxis — thecontract_coverage/analysis_assuranceaxes are unaffected: a report that is present with an incomplete contract-evidence or analysis-assurance domain still floors the exit at1, the same as in the declared-target-set modes.
The gate policy for a coverage gap or an unexpected report is the manifest's
(or run-plan-projected manifest's) own gate block (CLI cleanup phase two,
PR 2 — replacing the former --on-missing-required/--on-unexpected-target
CLI flags):
{"aggregate_manifest_version": "2.0",
"targets": [{"id": "linux-x86_64", "required": true}],
"gate": {"missing_required": "warn", "unexpected_target": "fail"}}
missing_required: warn downgrades a coverage gap to advisory.
unexpected_target (include/warn/fail/ignore, default include)
controls a report whose target isn't in the expected set. Omitting gate
keeps the same defaults (missing_required: fail, unexpected_target:
include) this command always had; the resolved policy and its source
(manifest/run-plan/default) are reported back in
effective_policy. A report written by an older version may also carry
explicit (a since-removed Python-API override); the schema still accepts
it, but it is no longer emitted.
Reconciling findings across compiler/build profiles¶
When report ids follow the shape target@profile#channel@depth (produced by
project plan's matrix — see Project Targets
Schema), aggregate groups them
back into two additional, reporting-only blocks that don't affect the
exit code: profile_matrix (one entry per logical target, across profiles)
and finding_matrix (one entry per distinct finding, reconciled across
profiles). This is the part of aggregate worth understanding on its own —
it answers "is this break universal, or specific to one compiler/platform?"
How two findings from different profiles become one entry¶
A finding's identity for reconciliation is the same tiered
canonical/normalized/reduced identity diff_filtering.py already uses as
its own cross-detector dedup key — computed from kind,
symbol, description, old_value/new_value, source_location, and
affected_symbols, read back off each report's changes[] entries. Two
reports naming the same symbol removal on GCC and on Clang reconcile to
one finding_matrix entry with affected_profiles naming both.
A separate, narrower mechanism handles cross-ABI mangling: a
cross_abi_declaration (an Itanium/MSVC-mangling-independent qualified
name, e.g. both _ZN3lib3addEii and ?add@lib@@YAHHH@Z reduce to
lib::add) links declarations across mangling schemes for display, but
is deliberately never used to merge two findings' identities — it can
only make an ambiguous case withhold a "clean" verdict, never claim
provable sameness it can't back up. A true spelling-equivalence merge (e.g.
macOS's extra leading underscore on an otherwise-identical mangled symbol)
is handled separately, and only when the two spellings are exactly
equivalent.
The scope field¶
Every finding_matrix entry gets exactly one scope, resolved by strict
precedence — undetermined always wins if it applies, since an unclear
profile can never be reported as either affected or clean:
scope |
Meaning | Precedence |
|---|---|---|
undetermined |
At least one profile's report was incomplete for this finding (missing, unreadable, not-comparable, or a report format — like a compare-release bundle — that's never complete). |
Highest — wins over everything else. |
all_profiles |
Every profile carries this finding; no profile is confirmed clean of it. | |
partial |
Two or more profiles are affected, and at least one other profile is confirmed clean. | |
profile_specific |
Exactly one profile is affected, and at least one other is confirmed clean. | Lowest. |
The critical, unsuppressible invariant: unaffected_profiles requires
a profile's report to be fully known (every check for that profile
complete=True) — a positive "checked and clean" claim. A profile that
can't clear that bar goes to undetermined_profiles instead, never
unaffected_profiles. Only completeness can clear a profile of a
finding; a finding — even from an incomplete report — can always convict
one.
Four worked outcomes¶
- One removal, both GCC and Clang report it →
scope: all_profiles,affected_profiles: ["linux-gcc14", "linux-clang20"],unaffected_profiles: []. - A layout break only under MSVC (GCC/Clang unaffected and their
reports are complete) →
scope: profile_specific,affected_profiles: ["windows-msvc"],unaffected_profiles: ["linux-gcc14", "linux-clang20"]. - GCC reports
func_removed; a stripped Clang lane reports the narrowerfunc_removed_elf_only— these reconcile to one finding entry (same underlying identity, different evidence tier surfaced the same fact differently) rather than appearing as two unrelated rows. - The Windows leg's report was never produced (build failed before
upload) → that profile lands in
undetermined_profilesfor every finding in the matrix, never inunaffected_profiles— a report that was never produced proves nothing was clean.
Per-profile contract decisions (profile_contract)¶
scope/affected_profiles answer whether a profile has a finding, not
why compatibility policy did or didn't act on it — a question the contract-relevance model (see Compatibility Evaluation
Config) already answers
per finding on a single-pair compare, and that a matrix can disagree on
just as easily as it can disagree on the finding itself. Two profiles
compared under different --contract domains can both report the same
type_size_changed finding while reaching opposite conclusions about it —
one's evidence proves the type is in the public contract (IN_CONTRACT,
gating), the other's evidence can't resolve it either way
(UNKNOWN_UNRESOLVED, not gating) — and scope: all_profiles alone would
present that as one uniformly-understood break.
When at least one affected profile's report ran --contract,
its finding_matrix entry carries a profile_contract array — one entry
per affected profile, in the same order as affected_profiles (a
profile confirmed clean of the finding has no contract decision about it
to report), each with that profile's own contract_relevance,
compatibility_evaluation_status, compatibility_decision, and
gate_contribution, read back verbatim from that profile's own report. A
comparison where no profile ever evaluated a contract omits the field
entirely, so finding_matrix for a contract-unaware CI matrix renders
exactly as before:
{
"kinds": ["type_size_changed"],
"symbol": "Foo",
"scope": "all_profiles",
"affected_profiles": ["linux-clang20", "linux-gcc14"],
"profile_contract": [
{
"profile": "linux-clang20",
"contract_relevance": "UNKNOWN_UNRESOLVED",
"compatibility_evaluation_status": "NOT_EVALUATED",
"compatibility_decision": null,
"gate_contribution": 0
},
{
"profile": "linux-gcc14",
"contract_relevance": "IN_CONTRACT",
"compatibility_evaluation_status": "EVALUATED",
"compatibility_decision": "BREAKING",
"gate_contribution": 1
}
]
}
Neither profile_matrix nor finding_matrix changes the exit code — they
are reporting views over the same gate/coverage/contract_coverage axes
above. The full field list for both blocks is in
aggregate_report.schema.json.
CLI reference¶
| Flag | Default | Notes |
|---|---|---|
--manifest PATH |
— | The single source of truth for the expected-target set; its own gate block sets the missing-required/unexpected-target policy. |
--manifest PATH (run-plan) |
— | The same flag takes a project plan run-plan.json, recognized by its own schema; its projected manifest can carry the same gate block. |
--discovered-only |
— | No required-target coverage gate (contract coverage still applies). |
-o FORMAT=DESTINATION |
text=- |
text | json; - is stdout, repeatable. |
Full generated flag reference, pulled from the live --help output: CLI
Reference.
See Exit Codes → abicheck aggregate
for the exhaustive exit-code matrix, and GitHub Action:
Recipes for the full fan-out/fan-in CI workflow
(matrix job → per-leg report upload → gate job).
See also¶
- Exit Codes — the canonical per-axis exit-code contract
- Project Targets Schema —
profiles:/checks:, the source oftarget@profile#channel@depthreport ids - GitHub Action: Recipes — the worked matrix + gate workflow
- Compatibility Evaluation Config — what feeds a report's
contract_coverage_exit_contribution - Reporting on fork pull requests — publishing an aggregate document as a PR comment from a trusted
workflow_runjob (its per-target member reports must sit in the same directory as the document)