Skip to content

Exit Codes

abicheck uses different exit codes for each command family.

Why they differ: compare is the native interface — 0/2/4 by verdict (or 0/1/2/4 severity-aware), with invalid invocations exiting 64 so a usage error is never mistaken for an ABI verdict. deps has its own narrower contract, documented below. scan had one too, but it was retired outright — see that section's warning before depending on any of its historical codes. compat (the ABICC drop-in, with ABICC's 0/1/2 and its own 3–11) was removed too; abicheck compat exits 64.

Contract relevance decides what the gate sees

Under --contract, each finding's contract relevance is classified before compatibility policy runs, and policy then scores only the EVALUATED findings — those whose relevance is IN_CONTRACT or NOT_APPLICABLE. A PROVEN_OUT_OF_CONTRACT, UNKNOWN_UNPROVEN or UNKNOWN_UNRESOLVED finding is NOT_EVALUATED: its compatibility_decision is JSON null and its gate_contribution is 0, so it moves neither the verdict nor the exit code.

This is a real change to what the compatibility verdict scores, but only in one direction: since policy now scores the EVALUATED subset of what it used to score in full, the verdict can only stay the same or get less severe — a change proven outside the declared contract stops blocking, and a change one domain cannot resolve stops gating as an ABI break. What it never does is make either disappear — an excluded finding keeps its ChangeKind, stays in changes and in the audit ledgers, and is rendered with the relevance and reason code that say why it did not gate. The other direction — the overall process exit getting worse — comes from the separate, independently-orthogonal axis below (missing evidence contributing its own exit 1), not from relevance itself; that's what stops missing evidence from being the cheapest way to pass.

Without --contract no finding carries a relevance, so every finding is scored exactly as before and every exit code below is unchanged.

Contract-coverage contribution

compare carries an orthogonal contract-coverage axis under --contract (the retired scan --against carried the same one). Complete coverage of the mode-selected evidence domain contributes 0; missing, partial, stale, failed, contradictory, or identity-incomplete required domain evidence is recorded as a CoverageFailure in the run-level contract_coverage_failures ledger and contributes 1. Unrelated provider failures stay advisory. This is a run-level ledger, not a per-finding field: a coverage gap does not by itself force every finding to UNKNOWN_UNRESOLVED — an observed root or a kind that's NOT_APPLICABLE regardless of evidence can still resolve normally even while the ledger records incomplete evidence elsewhere.

The axis is folded with max, so it raises a clean 0 to 1 and never lowers a gate's 2/4 — missing coverage cannot demote a real ABI break to "warnings only". It never rewrites a finding's compatibility decision or its gate contribution either; it is a floor on the exit status alone. Both commands fold it identically.

A directory/package compare (the per-library release fan-out) applies the same flag per library, then maxs every library's own contribution into the release's exit code — one library's incomplete coverage still raises a clean release exit 0 to 1, and the release JSON summary states the aggregate in the same contract_coverage_exit_contribution field. --fail-on-removed- library's exit 8 is checked ahead of the coverage-only fallback, so a removed library's own signal is never masked by an unrelated coverage gap (and a real verdict-based 2/4 still wins outright over both, unchanged).

The completeness axis (directory/package compare only)

Two more orthogonal 0/1 contributions, folded with max exactly like the contract-coverage axis above (a clean 0 becomes 1; a 2/4 is never lowered; no finding's decision is rewritten):

  • incomplete_scope — a selected, expected member of the comparison scope never reached a completed comparison: it had no counterpart on the other side and that side's inventory is not proven complete (not_supplied), this build cannot analyze it (unsupported), or its extraction failed (failed). Contributes 0 under the default .abicheck.yml scope.on_incomplete: warn and 1 under scope.on_incomplete: block. Under warn the run still reports the scope as incomplete: the JSON run_outcome.scope reads incomplete, the comparison_scope block names every unchecked member and why, and the Markdown/PR-comment views say the verdict covers the compared members only.
  • no_comparison_completed — the selected scope produced no valid comparison at all (zero matched pairs, or every selected member failed/unsupported). Contributes 1 under either setting: a permissive policy can downgrade missing members, never "nothing compared".

Both appear in the report's exit block (incomplete_scope_contribution, no_comparison_completed_contribution) and, when they decide the code, in exit.reasons. A scalar compare never sets either.

Migration note (exit 8). Before the completeness axis landed, the (now-removed) --fail-on-removed-library flag (today .abicheck.yml's gate.fail_on_removed_library: true) exited 8 on the raw old-minus-new filename set difference, so a partial local build compared against a full baseline read as "N libraries removed". Exit 8 now requires the removal to be proven: the NEW side's inventory must be proven complete. Two things prove it, and nothing else does — a stored ProjectSnapshot package or bundle-facts document whose capture asserted inventory_complete, or a package archive operand, whose extractor unpacks the container in full or fails, so the components found in the extracted tree are the whole set the package ships. A stored package without the assertion and a live directory still cannot prove absence: a directory may simply have been populated partially. An unmatched library under an unproven inventory is reported as an incomplete scope instead — exit 0 under warn, 1 under block — and the JSON key unmatched_old keeps listing it. When NEW is named as a single file (not a directory or archive that happens to hold one member) with exactly one OLD counterpart and NEW's inventory is unproven, the run is a current-artifact comparison (D9): the other OLD members are out_of_scope and the scope is complete. A one-member NEW directory is not narrowed: its unmatched OLD members stay unchecked, so block still gates — discovered cardinality is never read as intent.

Migration note (exit 8, second step). S3 makes a package archive pair a completeness proof, so gate.fail_on_removed_library: true reaches exit 8 again for abicheck compare old.rpm new.rpm (or .deb/.tar.*/ .whl/.conda) where S2 had left it reachable only for a stored snapshot. This is the correction S2's note anticipated, not a reversal of it: a directory pair still exits 0, because a directory proves nothing. If a release genuinely ships fewer components on purpose, that is what --support-promise declared reports as a finding (support_promise_component_retired) rather than as a bare exit code.

Without --contract there is no selected domain, so the contribution is always 0 and every other exit code below is unchanged.

Ordinary change suppressions cannot clear a provider/domain coverage failure — a coverage failure is not a finding, so the suppression machinery structurally cannot reach one. To accept incomplete contract assurance, set contract.unresolved: warn (for example via a kind: contract pack). That zeroes this contribution and changes nothing else: the failures remain listed in contract_coverage_failures — but only -o json=... carries that field at all; markdown, review, HTML, SARIF, and JUnit output surface the same information as a stderr diagnostic, a SARIF notification, or a JUnit error suite instead.

Reports state the applied number in contract_coverage_exit_contribution, which distinguishes contract coverage exit 1 from severity or aggregate required-target coverage. The composite GitHub Action reads the same field and publishes verdict: COVERAGE_INCOMPLETE rather than labelling the axis a severity-policy or operational failure; on compare, where exit 1 is shared, it uses the report's pre-fold severity.exit_code to tell the two apart. The configured GateDecision independently contributes 0/1/2/4: a compatible addition can block, and a breaking finding can be demoted. Only legacy output without a gate block falls back from compatibility verdict to 2/4. Existing command-specific 5, 8, and 64 behavior is as documented below.

compare --no-baseline (single artifact)

abicheck compare --no-baseline NEW declares that no prior surface exists for NEW and runs a candidate-side audit instead of a comparison — the first replacement for scan's audit-only mode (no --against) and integration scenario S5. --no-baseline is an explicit declaration, never inferred from argument count: compare NEW (one operand, no flag) and compare --no-baseline OLD NEW (the flag plus two operands) are both usage errors, exit 64. A directory of libraries, a package archive, or a multi-artifact stored package is audited per member instead — see compare --no-baseline DIR below.

The OLD side is recorded with the declared_absent MemberAcquisition state — distinct from not_supplied (an unproven absence that can leave the scope reading incomplete): the user declared there is no prior surface, so run_outcome.scope reads complete and the comparison_scope block names the one declared_absent member. The run never emits an addition, a removal, or a compatibility verdict — run_outcome.compatibility and the top-level verdict are JSON null, and changes is always []. The compatibility axis therefore always contributes 0 to the exit code.

The audit's own content is the candidate-side finding set — the eleven cross-source hygiene checks and the pattern/preprocessor pre-scan, reported under findings[] (not changes[], which stays empty by construction). These stay advisory (RISK/API_BREAK, never BREAKING), so a hygiene finding never gates on its own — an audit that reports several findings still exits 0 unless one of the orthogonal axes below fires. Every finding carries its evolution state, which with a declared_absent OLD is always persistent or not_evaluated, never introduced.

Four orthogonal axes still apply exactly as they would for a two-sided run, folded with the same max discipline:

Axis Contributes When
Audit gate 3 --severity-preset (any value except info-only) opted the run into gating, and at least one candidate-side finding is BREAKING/API_BREAK-classified
Analysis assurance 1 .abicheck.yml's assurance.require_complete: true and analysis_assurance.status is not complete. On a directory/package (release) compare, max over every compared member's own contribution
Contract coverage 1 --contract and the selected domain's required evidence is incomplete
Evidence contract 7 --depth build/--depth source pinned, but this run's live extraction did not reach it

The contract-coverage row was inert until 2026-09-09 — --contract was parsed and documented on this path but never forwarded, so the ledger it folds from was never populated and a compare --no-baseline NEW --contract public run against a headerless candidate exited 0 where the two-sided equivalent exited 1. It is now wired through the same resolve_contract_evaluation/resolve_contract_domain resolvers the two-sided path uses, so an equivalent one-sided and two-sided invocation activate identically. Without --contract the contribution is still always 0, so every pre-existing invocation is unchanged.

The evidence-contract row was added in the same pass, for the same reason: --depth build with no --sources/--build-info silently degraded to symbols-only evidence and reported a clean audit. It now records the exit-7 axis, matching the two-sided compare path exactly. A stored snapshot candidate (.abi.json) is exempt — this run never extracted it, so it cannot have fallen short of a pinned depth.

compare --no-baseline never emits 2/4: an audit reports no compatibility verdict, so the compatibility family contributes nothing and only the orthogonal axes above can raise its exit code.

The audit-gate axis. Legacy scan's audit mode derived a verdict from its own findings and gated at 2 on a hygiene finding whose kind is BREAKING/API_BREAK-classified. compare --no-baseline cannot reproduce that at 2/4 (D2 forbids it), so it reproduces the same partition through its own code, 3, and the axis is opt-in: every invocation that omits --severity-preset stays exactly as it always was (exit 0 for a hygiene finding, unaffected by this axis's existence). Passing --severity-preset default (or strict) opts the run into gating; --severity-preset info-only is the explicit "don't gate" request and stays 0. A scan-based gating CI job migrating to compare --no-baseline needs exactly one addition to keep gating: --severity-preset default. The axis's own rule reads BREAKING_KINDS/API_BREAK_KINDS membership directly (the same sets checker_policy.py's registry derives), not --severity-preset's own per-category severity mapping — that mapping's potential_breaking bucket merges API_BREAK_KINDS with RISK_KINDS, which would gate an advisory- only RISK finding the same as an API_BREAK one; the preset only activates the axis, it does not decide which findings gate. See policy/audit_gate_exit.py for the full account.

On compare --no-baseline only, --dry-run reports the same condition ahead of any analysis: against a live candidate, a pinned --depth build/ --depth source with no evidence input is a dry-run blocker (exit 1), never a clean preview of a run that would exit 7. Two-sided compare --dry-run does not preview its own floor today — it still exits 0 on the same pinned-but-unsatisfiable depth (recorded in known-gaps.md; giving it the same preview is the natural follow-up, and would make it exit nonzero there too).

compare --no-baseline DIR (N-library audit)

When the operand is a set of libraries — a directory, a package archive (.deb/.rpm/.whl/.tar.*/...), or a multi-artifact or degraded stored ProjectSnapshot package, classified exactly as two-sided compare classifies it — each selected member is audited exactly as compare --no-baseline <member> would audit it, and one audit_set JSON document holds every member's own audit report. Members are selected with --select/--select-required (a usage error, exit 64, for a single artifact), and .abicheck.yml's scope.on_incomplete, release.dso_only and release.include_private_dso apply as they do to a two-sided directory/package compare.

The exit code is the max over:

Axis Contributes When
Audit gate, contract coverage, analysis assurance, evidence contract 3/1/1/7 The max of every audited member's own contribution (table above)
Operational error 4 A selected member's audit failed (an unreadable or unparseable artifact) — the same contribution a directory/package compare folds for a failed member. The member is listed with its reason, never dropped
Incomplete scope 1 A selected member was not audited (expected_not_produced, failed, unsupported) and scope.on_incomplete: block; 0 under the default warn
No audit completed 1 No selected member reached a completed audit — an empty selection, or every member failing. Never a clean pass, under either scope.on_incomplete setting

An artifact this build cannot analyze (a snapshot newer than the reader) is unsupported — the scope axis, not the operational one; a misconfiguration that would be a usage error for one artifact aborts the whole run with exit 64. OLD is declared absent for every member, so no member is ever reported added or removed — whatever the candidate's inventory proves — and the compatibility family never contributes. A directory operand with no readable member exits 1 with the same error a two-sided compare gives. Formats: json, markdown (the default) and oneline; the others are usage errors.

Analysis-assurance contribution (P0.4)

compare always computes and reports analysis_assurance — a third, orthogonal axis alongside the compatibility verdict and the policy/severity gate, answering "how complete and trustworthy was the evidence behind this comparison" independently of whatever the verdict says. Its status field is one of complete, partial, failed, not_comparable, or not_requested, and it is always present in -o json=... output (analysis_assurance — a top-level key on compare's report, nested under diff on scan's), regardless of any flag.

By itself this changes nothing about any exit code — analysis_assurance is purely informational until a caller opts in. Setting .abicheck.yml's assurance.require_complete: true (config-only -- no CLI flag; the former compare --require-complete-analysis flag was demoted here entirely) makes compare additionally contribute exit 1 whenever analysis_assurance.status is not complete, folded with the same max discipline the contract-coverage axis above uses: it raises a clean 0 to 1 and never lowers a 2/4/5/6 — incomplete assurance cannot demote a real ABI break to "warnings only", and it never rewrites the compatibility verdict, any finding, or the severity gate's own contribution.

assurance.require_complete: true applies at any cardinality. A directory/package (release) compare folds each member's own 0/1 contribution with max() — the identical contribution a single-pair compare of that member computes — so a release of one library and that library compared on its own reach the same exit code, and one member short of complete analysis floors the release regardless of how many complete siblings it has. It used to be rejected outright for a directory/package operand ("no single analysis_assurance result to gate on"), which made the setting's meaning depend on the input shape.

A stored BundleFacts OLD side folds identically. The assurance being folded belongs to each comparison, not to either operand: every member of a stored-baseline release is still compared against a live NEW artifact and still gets its own AnalysisAssurance. What a shallow stored side cannot do is manufacture evidence it never captured — it reports partial, and the gate then names it rather than ignoring it.

The release JSON carries an analysis_assurance block naming the members that fell short and why, a top-level analysis_assurance_exit_contribution (the key abicheck aggregate and the Action's deferred gate read), and a per-libraries[] analysis_assurance_status; a non-JSON format gets the same facts as a one-line stderr notice.

This axis stays orthogonal to the completeness axis, and both can apply to the same run: that one asks whether every selected member was compared at all (an inventory question), this one asks whether the comparisons that ran had complete evidence. A release can be scope-complete with a partial analysis, or scope-incomplete with a complete analysis over what ran; on a tie both are named in reasons.

Without assurance.require_complete: true every pre-existing invocation's exit code is unchanged, exactly as --contract's own coverage axis requires no opt-in flag change either.

The composite GitHub Action's own dedicated require-complete-analysis input is retired alongside the CLI flag it forwarded to (rulings.py deferred-option followup) -- see docs/reference/github-action-inputs.md.

The exit report field (CLI cleanup phase two, PR G1 / PR E)

Every real-verdict compare -o json=... report (full/leaf/root-cause modes) carries a top-level exit object (introduced at report schema 2.41; schema 2.42 added its crosscheck_promotion_contribution field — see abicheck/schemas/__init__.py's REPORT_SCHEMA_VERSION docstring for that and later additive fields, e.g. annotations at 2.43) stating the already-resolved decision behind the axes above as one explainable value, rather than requiring a reader to separately combine severity.exit_code/verdict, contract_coverage_exit_contribution, and analysis_assurance_exit_contribution themselves. A stored report from the retired scan --against carries the identical object (scan schema 1.18), nested at diff.exit rather than at the top level — matching where its own constituent contribution fields lived:

"exit": {
  "code": 1,
  "reasons": ["analysis_assurance"],
  "compatibility_contribution": 0,
  "contract_coverage_contribution": 0,
  "analysis_assurance_contribution": 1,
  "crosscheck_promotion_contribution": 0
}

code is exactly max() over the four contributions — the identical number the real process exits with for a single-pair compare. reasons names every axis whose own contribution equals code (a lower, non-winning contribution is excluded, since it did not determine the result); ["clean"] when code is 0.

crosscheck_promotion_contribution (schema 2.42) is always 0 on a compare report — it never had meaning outside the retired scan --against's own maintainer-promoted --crosscheck KEY=error finding (scan_engine._promote_published_gate), which reconstructs the whole diff.exit block through the same resolver whenever the crosscheck contributes anything positive, so reasons can carry promoted_crosscheck even when the crosscheck only ties — rather than exceeds — the baseline comparison's own exit code.

--used-by/--required-symbol(s) scoping does not change compatibility_contribution/reasons at all (workstream D-S1, docs/contribute/plans/vision-api-abi-evolution.md "D. Optional prebuilt-consumer lifecycle"): they always describe the full-library compatibility gate, exactly as an unscoped run's would. A supplied consumer's own confirmed/potential/unresolved assessment is reported separately (used_by/required_symbol_contract/consumer_scope in the JSON report — see the section below) and never feeds into this field.

This field is additive and does not itself change any exit code — it is a persisted view of a resolution every axis above already performs. It does not yet cover not_comparable or a release's removed-required-library policy, each raised through a different code path today; see abicheck/policy/exit_decision.py's own module docstring for that scope boundary.

evidence_contract_error_contribution/budget_overflow_contribution (schema 3.3, docs/contribute/plans/one-comparison-product.md P3). Native compare carries the evidence-contract-error (7) and budget-overflow (5) ExitDecision axes scan --against originated — resolve_compare_exit_decision folds DiffResult.evidence_contract_error/.budget_overflow through the identical precedence rule scan used (exit_decision_precedence.resolve_scan_exit_decision), reused rather than re-derived, so the two commands can never disagree on which axis wins when both apply. Phase 2d gave the evidence-contract axis its first CLI-reachable trigger on compare: compare --abi3 VERSION against a candidate that is not a recognisable CPython extension module exits 7, exactly as scan --abi3 does, since the stable-ABI audit the flag asks for cannot be performed at all (the comparison's own findings are left as they were and still reported). budget_overflow_contribution has no compare trigger yet — there is still no --budget flag — so it stays 0 on every compare report, prerequisite plumbing for the plan's Phase 7.

Commands removed in the pre-1.0 CLI reset

appcompat and plugin-check are gone as standalone commands; their scoping folded into compare itself — see Application- and plugin-scoped comparisons below. baseline (the push/pull/list/delete registry), debian-symbols, collect, merge, inputs validate, and inputs compact were removed outright with no CLI replacement — validating a build-emitted abicheck_inputs/ pack now happens automatically whenever the pack is consumed, and the debian-symbols/collect/merge library functions remain available for programmatic (Python API) use only. None of these have their own exit codes in the current CLI, so they no longer appear in the tables below.


abicheck compare

Legacy exit codes (default, no --severity-* flags)

Exit code Meaning
0 NO_CHANGE, COMPATIBLE, or COMPATIBLE_WITH_RISK — no binary ABI break
2 API_BREAK — source-level API break — recompilation required
4 BREAKING — binary ABI break
5 Budget overflow — the run's wall-clock --budget guard was exceeded. compare has its own --budget flag now (frontends/cli/commands/compare.py); exit 5 applies to both compare and scan on overflow.
7 Evidence-contract error — the analysis a pinned input asked for could not be performed at all, so it was not silently downgraded. Reachable through --abi3 VERSION against a candidate that is not a recognisable CPython extension module, and through a pinned --depth build/--depth source whose evidence doesn't reach it (compare shares this floor with scan/dump now, verified live). The depth trigger applies only when at least one operand is a live extraction; comparing two already-serialized snapshots is exempted from it even when neither embeds L3/L4 evidence (verified live — see Evidence Depth's "pinned depth is a contract" warning), so a clean both-snapshot compare is not proof the pinned depth was reached. On compare, the top-level verdict field is not "EVIDENCE_CONTRACT_ERROR" for either trigger — verified live: it stays whatever the (otherwise-unaffected) compatibility comparison produced (e.g. "NO_CHANGE"). The error is carried in the exit block instead: exit.code: 7 and exit.reasons: ["evidence_contract_error"]. A JSON consumer must check exit, not verdict, to detect this. (scan's own dedicated exit path does set verdict: "EVIDENCE_CONTRACT_ERROR" — see its row below — so the two commands differ here despite sharing the same exit code.) compare folds this through the same ExitDecision precedence rule scan uses (exit_decision_precedence.resolve_scan_exit_decision), so the two commands can never disagree on which axis wins, only on how the JSON surfaces it. dump's own, separately implemented floor for the depth trigger raises DumpDepthNotSatisfiedError and exits 1 instead — no snapshot is written — since dump has no verdict to fall back to.
16 not_comparable — OLD and NEW were not extracted under a comparable profile/scope contract, so no verdict was produced (verdict: null in -o json=-, with a reason object). Pass --diagnostic-comparison to force a tentative diff instead.
64 Invalid invocation — bad arguments/options or an unreadable/unrecognised input, deliberately outside the 0/2/4 verdict space

compare --dry-run does not preview the exit-7 depth floor. Pinning an unsatisfiable --depth build/--depth source under --dry-run still exits 0 and reports 0 TU(s) for the affected layers, where the equivalent real run now exits 7 — see Evidence Depth.

⚠️ Exit 0 covers NO_CHANGE, COMPATIBLE, and COMPATIBLE_WITH_RISK. If your pipeline needs to distinguish them (e.g. warn on deployment risk), use -o json=... and read the verdict field — exit code alone is not sufficient.

Severity-aware exit codes (with any --severity-* flag)

When any --severity-preset or --severity-* option is provided, the exit code is computed from the severity configuration rather than the verdict:

Exit code Meaning
0 No error-level findings
1 Error-level findings in addition or quality_issues only
2 Error-level findings in potential_breaking (but not abi_breaking)
4 Error-level findings in abi_breaking
5 Budget overflow — as in the legacy table above; the run's wall-clock --budget guard was exceeded before any severity classification could run, so it is scheme-independent.
7 Evidence-contract error — as in the legacy table above; raised before severity classification runs, so it is scheme-independent.
16 not_comparable — the comparability gate hard-fails before severity classification ever runs, identical to the legacy scheme's 16.

The highest applicable code wins. For example, if both abi_breaking=error and quality_issues=error have findings, the exit code is 4.

ℹ️ The two exit code paths are mutually exclusive. Without --severity-* flags, the legacy verdict-based path runs. With any --severity-* flag, the severity-aware path runs. They never both execute.

Severity presets

Preset abi_breaking potential_breaking quality_issues addition
default error warning warning info
strict error error error error
info-only info info info info

Per-category overrides — .abicheck.yml's severity: block (abi_breaking/potential_breaking/quality_issues/addition) — take precedence over the preset.

CI gate patterns

# Production gate: fail on any break (legacy exit codes)
abicheck compare old.json new.json
ret=$?
[ $ret -eq 4 ] && echo "BREAKING — release blocked" && exit 1
[ $ret -eq 2 ] && echo "API_BREAK — source-level break" && exit 1
echo "OK (NO_CHANGE or COMPATIBLE)"

# Block unexpected API expansion (severity-aware; `severity.addition: error`
# in .abicheck.yml, which compare discovers from the working directory)
abicheck compare old.json new.json
ret=$?
[ $ret -eq 1 ] && echo "ADDITIONS — unexpected API expansion" && exit 1
[ $ret -eq 4 ] && echo "BREAKING — release blocked" && exit 1
[ $ret -eq 2 ] && echo "API_BREAK — source-level break" && exit 1
echo "OK"

# Strict mode: all categories at error level
abicheck compare old.json new.json --severity-preset strict

# Permissive gate: fail only on binary breaks
abicheck compare old.json new.json
ret=$?
[ $ret -eq 4 ] && exit 1   # BREAKING only; API_BREAK (exit 2) allowed
exit 0

# Parse exact verdict from JSON (with severity info)
abicheck compare old.json new.json -o json=- --severity-preset default -o result.json
verdict=$(python3 -c "import json,sys; d=json.load(open('result.json')); print(d['verdict'])" \
  || { echo "ERROR parsing result.json"; exit 1; })
[ "$verdict" = "BREAKING" ] && exit 1

abicheck compare (multi-library / release inputs)

When compare is handed directory or package inputs (RPM/deb/tar/conda/wheel), it fans out to per-library pairs and aggregates the worst per-library verdict across the release — the behaviour formerly exposed as the standalone compare-release command (folded into compare; the GitHub Action's own compare-release/stack-check mode aliases were removed the same way — mode: compare handles directory/package operands directly). By default a set/release comparison uses the verdict-based scheme below, plus a dedicated code for removed libraries:

Exit code Meaning
0 All libraries compatible (no API/ABI break)
2 Worst verdict is API_BREAK
4 Worst verdict is BREAKING, or an operational ERROR (a library failed to dump/extract/compare)
1 No compatibility break, but the completeness axis contributed: .abicheck.yml's scope.on_incomplete: block with an incompletely checked scope, or a run that completed no comparison at all (under either setting). Also the contract-coverage axis's own floor, as for single-pair compare.
8 A library was proven removed between releases (NEW's inventory is proven complete — a stored ProjectSnapshot package or bundle-facts document whose capture asserted inventory_complete) and .abicheck.yml's gate.fail_on_removed_library: true is set. In the legacy scheme this is emitted only when no API/ABI verdict exit 2/4 and no operational ERROR exit 4 already applies; in the severity-aware scheme it takes precedence over 0/1/2/4. An unmatched library under an unproven inventory never exits 8; see the completeness axis above.
16 not_comparable — at least one library's OLD/NEW DSOs were not extracted under a comparable profile/scope contract. Takes precedence over every other outcome in the release, including 8 (removed-library) and a genuine ERROR: a not_comparable result means the comparison couldn't establish what changed at all, so it dominates in both the legacy and severity-aware schemes. Identical code to native compare's own 16.

On the release path the severity-aware code (0/1/2/4) replaces the verdict-based 2/4 mapping only when a severity map is actually in effect — that is, any --severity-* flag is passed or .abicheck.yml carries a severity: block (a preset or per-category levels). There is no manual override any more (CLI cleanup removed --exit-code-scheme and .abicheck.yml's exit_code_scheme: key): the scheme is fully automatic, so with no severity values to apply, the fan-out has nothing to score against and falls back to the legacy verdict mapping — there is no way to pin it to severity without one. Under the legacy mapping, an operational ERROR exit 4 or nonzero API/ABI verdict exit (2/4) wins before the removed-library check; under an effective severity map, removed-library exit 8 wins over the aggregated 0/1/2/4 code. An operational ERROR without a higher-priority removed-library result still floors the severity-aware exit at 4. One consequence worth gating on: with an effective severity map, a release whose worst verdict is BREAKING can still exit 0 if that map downgrades ABI breaks (e.g. abi_breaking: warning) — parse the verdict from JSON output if you need scheme-independent CI behaviour.


abicheck scan (retired)

scan was retired in 0.6 — hard removal, no deprecation window

0.6 reduced the root surface to six verbs and retired scan as a second analysis product. D8 is explicit that the removal is hard: no hidden alias, no shim, no silent ignoring. abicheck scan now exits 64 with No such command, and the error names compare --no-baseline. The whole table below no longer describes any live command — exit 5, 6 and 7 did not become compare codes by inheritance; each moved onto compare's own ExitDecision axes on its own schedule (where one exists yet), and the compare sections above are where a migrated axis is documented. The table is kept below purely as a historical record of scan's own exit-code contract while it existed.

Do not write new CI against it: pin the equivalent compare invocation instead, and where none exists yet, see known gaps for what is still open. The GitHub Action's own mode: scan input has since been retired outright too — see the migration guide.

Everything below this point, to the end of this section, is historical — scan no longer exists and none of it is a live invocation to copy. While it existed, the one-shot source-intelligence scan had its own contract (it could compare ARTIFACT against --against and added a budget guard). --against was the only thing that selected the mode: omit it and scan ran a one-build audit/hygiene/source-consistency scan only; pass it and scan also compared ARTIFACT against it — there was no separate --audit flag:

Exit code Meaning
0 Compatible; or (audit-only, no --against) advisory-only crosscheck findings with no comparison verdict at all
2 Source-level / API break (incl. API_BREAK cross-source findings). On a baseline (--against) scan, a cross-source finding (exported_not_public, unversioned_exported_symbol, ...) is scored exactly like any other compare finding as of 2026-09-09 — it is not advisory-only just because it came from a cross-source check
4 ABI break (from the --against comparison)
5 --budget overflow — the time guard tripped (scope is never silently shrunk)
6 NOT_COMPARABLE — ARTIFACT and --against were not extracted under a comparable profile/scope contract, so the comparison never ran (diff.reason in -o json=-). Distinct from compat check's 9 and native compare's 16 — every command maintains an independent exit-code scheme.
7 Evidence-contract error — a pinned --depth/--source-method whose required source evidence was never collected, or --abi3 targeting a binary that isn't a recognisable CPython extension module. No comparison ever ran (verdict: "EVIDENCE_CONTRACT_ERROR" in -o json=-); this process's own dedicated exit code (cli_scan.py's _EXIT_EVIDENCE_CONTRACT_ERROR), unambiguous regardless of format or whether a JSON report was written.
64 Invalid invocation (bad arguments/options)

Exit 5 is scan-only in practice: --budget 15m fails the run rather than quietly dropping evidence, and native compare has no --budget flag to raise it (see the exit report field section above for the shared, currently-unreachable ExitDecision axis). Use --dry-run to preview the audit checks and (if --against is given) the comparison that would run, plus the projected per-layer cost, without scanning — like every command's --dry-run it only ever exits 0/1/64, never a verdict code; see --dry-run below.

scan --artifact-set (auditing 2+ libraries together) is retired in 0.6 — a documented breaking change with no deprecation window. The capability is not abandoned: 0.6 retires the mode, and it returns as compare --no-baseline DIR over package component inventories, which are the prerequisite for preserving its per-member selection and coverage accounting. Until then, run compare --no-baseline once per library. Exit 1 on a scan was unchanged there and meant a genuine CLI/operational error.

scan --against and severity (retired; mirrored compare)

scan --against accepts the same severity surface as compare — --severity-preset and the hidden per-category --severity-* overrides (plus .abicheck.yml's severity: block) — and, like compare, uses them to compute the 0/2/4 portion of the exit code above from severity.compute_exit_code instead of the raw verdict when the resolved scheme is severity. A BREAKING verdict under --severity-preset info-only can therefore exit 0, exactly as it can with compare.

Under the severity scheme the JSON report's diff block also carries a severity gate object — the same config/categories/exit_code/ blocking/blocking_categories shape compare's own report uses (one shared builder, so the two are comparable field by field), added in scan_schema_version 1.9. It is what makes a non-zero exit on an otherwise compatible diff self-explanatory: severity.addition: error on an additions-only diff exits 1, and blocking_categories: ["addition"] names the cause, distinguishing it from the orthogonal contract-coverage 1 above. The default text output states the same fact in its Baseline comparison block:

Baseline comparison
  breaking=0 api_break=0 risk=0 compatible=1
  severity gate: exit 1 — blocking: addition

Verdict: COMPATIBLE

Both are absent under the default legacy scheme, which runs no severity gate.

The block is the scan's compatibility gate, not the baseline diff's alone: a cross-check the maintainer promoted with --crosscheck KEY=error raises it too, adding a promoted_crosscheck entry to blocking_categories (deliberately outside the four severity categories, since no severity level produced it). The promotion is a floor — it can add a blocking reason but never clear one a severity category already raised.

aggregate reads that diff.severity block as the target's compatibility gate when it is present, exactly as it reads a compare report's own severity block (and with the same fail-closed validation). This is what keeps the orthogonal axes separable for a scan target: a legacy-scheme scan has no native exit 1, so a raw 1 can only be the contract-coverage and/or analysis-assurance contribution (both orthogonal, both readable from their own report fields regardless of scheme) — but a severity-scheme scan also has a native 1 (an error-level addition), and folding all of these to 1 would otherwise be indistinguishable. See abicheck aggregate.

scan --dry-run previews whichever scheme the invocation resolves — the scheme label, the per-category severity levels, and that scheme's exit codes — so the preview matches the run it is predicting.

A gate pack (--pack) folds a gate.* assignment into a scan's severity the same way it does for compare, and cannot override a value that was actually stated — by an explicit --severity-* flag, or by .abicheck.yml. Every flag in this family is a comparison-only flag (rejected as a usage error without --against, exit 64) — see the table above. The budget (5), NOT_COMPARABLE (6), and evidence-contract-error exit codes are unaffected: they are returned before the baseline comparison — and therefore before any severity computation — ever runs.


abicheck aggregate

The multi-target fan-in gate folds the per-target compare/scan JSON reports a CI build matrix produces (one abi-report-<target>.json per leg) into one gate decision. Four axes stay orthogonal, and the exit code is the worst contribution across them:

  • gate — each report already carries its own severity gate decision (severity.{exit_code,blocking,blocking_categories}); aggregate combines those, it never recomputes a gate from the compatibility verdict. So a COMPATIBLE report with an addition=error policy still contributes exit 1, and a BREAKING report under a demoted preset can contribute 0. A scan report is read via its own nested diff.severity gate block when it has one (a severity-scheme scan --against, schema 1.9+ — read through the identical validator a compare block goes through), and otherwise via its top-level exit_code (keyed on scan_schema_version). Reports produced without any gate block fall back to the legacy verdict→exit mapping (0/2/4). Reading is fail-closed: a report whose gate block is present but corrupt (an out-of-range or non-integer exit_code, a blocking flag that contradicts it, non-string categories) makes that target unavailable — never silently reverting to the greener legacy path.
  • coverage — did every required expected target actually report? An incomplete required coverage is a coverage failure at exit 1; it is never promoted to an ABI-break exit 4. This is a different question from contract coverage below: this one asks whether the matrix ran and reported at all.
  • compatibility — the worst verdict over the analyzed targets, reported for context; it does not by itself drive the exit code.
  • contract_coverage — reads back each already-analyzed target's own contract_coverage_exit_contribution (per-report field; see "Contract-coverage contribution" above) and folds it with max, same as the other axes — added to the aggregate schema alongside this axis (abicheck.workflows.aggregate.AGGREGATE_SCHEMA_VERSION is the versioned fact owner); aggregate never recomputes it. This is a different question again from plain coverage: a required target can have reported successfully (no coverage gap) while its own evidence for the selected contract domain was still incomplete (a contract-coverage gap). Both can independently produce exit 1, for unrelated reasons, and aggregate records which one fired rather than merging them into one undifferentiated 1.
  • analysis_assurance — reads back each analyzed target's own analysis_assurance_exit_contribution (assurance.require_complete; see "Analysis-assurance contribution" above) and folds it with max the same way (aggregate schema 1.5); aggregate never recomputes it. A target whose own evidence was incomplete under that flag aggregates to 1 on this axis alone, independently of contract coverage and of the completeness axis below.
  • scope_completeness — reads back each analyzed target's own completeness-axis contributions (the exit block's incomplete_scope_contribution and no_comparison_completed_contribution; see "The completeness axis" above) and folds their max the same way (aggregate schema 1.8); aggregate never recomputes them. A release that exited 1 because a selected member went unchecked under .abicheck.yml's scope.on_incomplete: block, or because it completed no comparison at all, therefore aggregates to 1 as well — its run_outcome.gate and operational axes read none for that case, so without this axis the target read green. Reported as its own scope_completeness block and a per-target scope_completeness_exit, and named per profile as scope_incomplete_profiles, so an exit of 1 stays attributable to exactly one axis.
Exit code Meaning
0 Every required target analyzed, no blocking findings
1 A required target was unavailable while the effective missing_required policy was fail (the default; warn downgrades this to advisory and contributes nothing here); an analyzed target's gate blocks on an addition/quality finding only; a target's own contract-coverage evidence was incomplete under --contract; a target's own analysis assurance was incomplete under .abicheck.yml's assurance.require_complete: true; a release target's comparison scope gated (.abicheck.yml's scope.on_incomplete: block, or no comparison completed); or a non-verdict per-report failure folds here (e.g. a scan report's budget-overflow exit 5) — these axes are independent and any one of them alone is enough to produce 1
2 An analyzed target's gate is a source-level / API break
4 An analyzed target's gate is an ABI break
64 Invalid invocation (bad arguments/options, malformed manifest, duplicate target id, or no expected-target set given)

The highest applicable code wins: a run with both an ABI break and a coverage gap exits 4; a run whose only problem is a missing required target exits 1, never 4.

A required target with no report is unavailable (unknown), never counted as compatible. This is the whole point of the command: a matrix leg that failed before uploading its report is reported as a coverage gap and fails the gate at exit 1 — it is never silently folded into the verdict as an empty, compatible ABI, and a build that simply never ran is never handed an ABI-break exit 4.

Declaring the expected-target set (required — one of):

  • --manifest abi-targets.json (recommended) — the single source of truth for which targets the matrix must produce. Two document shapes are accepted and told apart by their own content, never by filename: an expected-target manifest ({"targets": [{"id": "linux-x86_64", "required": true}, ...]}), or an abicheck project plan run-plan (schema: abicheck.run-plan/vN), projected to the same expected-target shape internally — each check's own check_id becomes the expected target id, matching what check-target writes as every report's target_id. Generate it once in the plan job and feed the same file to both the matrix and this gate so they never drift.
  • --discovered-only — explicitly aggregate whatever reports are present with no required-target coverage gate (a missing target is simply not counted, never a coverage failure — the contract_coverage axis is unaffected and still applies to whatever is present). Required to run without --manifest: with no declared target set the gate cannot tell a missing required target from an intentionally absent one, so a bare aggregate reports/ is a usage error (exit 64), not a silent pass.

The gate policy for these two situations is no longer a pair of CLI flags (CLI cleanup phase two, PR 2) — it's the manifest's (or run-plan-projected manifest's) own versioned gate block:

{"aggregate_manifest_version": "2.0",
 "targets": [{"id": "linux-x86_64", "required": true}],
 "gate": {"missing_required": "warn", "unexpected_target": "fail"}}

missing_required: warn downgrades a coverage gap to advisory — it stops contributing to exit 1 on its own, but the per-target gate decisions, contract-coverage evidence, and per-report failures above remain independent axes that can each still produce 1 for an unrelated reason. unexpected_target (include/warn/fail/ignore, default include) controls a report whose target is not in the expected set: include counts its real findings in the gate but not in required coverage. Omitting gate entirely keeps the same defaults this command always had (missing_required: fail, unexpected_target: include). The resolved policy is reported back in the JSON output's effective_policy block, including which source (manifest/run-plan/explicit/default) it came from — explicit only appears for a direct Python-API caller of aggregate() forcing a value (there is no CLI spelling for it). The -o json=... output is versioned (aggregate_schema_version — see abicheck.workflows.aggregate.AGGREGATE_SCHEMA_VERSION for the current value) and carries the six axes separately under gate / coverage / compatibility / contract_coverage / analysis_assurance / scope_completeness — the last three are {"exit_contribution": 0, "incomplete_targets": []}-shaped and present even when no target used --contract/assurance.require_complete: true or every release target checked its whole scope (an empty incomplete_targets list, not an omitted block).

When targets are checked under several toolchain profiles (report ids of the form target@profile#channel@depth), two additional reporting-only blocks group them back together: profile_matrix (one entry per logical target, with affected_profiles/verdict_by_profile) and finding_matrix (one entry per distinct finding, with affected_profiles / unaffected_profiles / undetermined_profiles and a scope of all_profiles / profile_specific / partial / undetermined). A profile whose report is missing, unreadable, not-comparable, or carries no changes array is undetermined — never reported as clean of a finding it was never checked for. Neither block affects the exit code; the full field list is in aggregate_report.schema.json.


Application- and plugin-scoped comparisons (compare --used-by/--required-symbol)

The standalone appcompat and plugin-check commands are gone. Their scoping now folds into compare itself:

  • compare --used-by APP (repeatable) — folds appcompat. APP is a real application binary; its actual imports/required symbol versions scope the comparison. OLD/NEW may be real library binaries or JSON snapshots that carry binary evidence (a dump of a real library, not headers-only). Mutually exclusive with --required-symbol.
  • compare --required-symbol SYM (repeatable) / --required-symbol @FILE — folds plugin-check. Scopes the comparison to an explicit dlopen/dlsym entrypoint contract instead of the full diff. Mutually exclusive with --used-by.

The full library comparison still runs once, and (workstream D-S1, docs/contribute/plans/vision-api-abi-evolution.md "D. Optional prebuilt-consumer lifecycle") its own verdict/exit code is always what this run reports — exactly the compare codes documented above (legacy 0/2/4, severity-aware 0/1/2/4, 64 for a usage error) — whether or not a consumer was supplied. In particular, an application requiring symbols or ELF version tags absent from the new library does not, by itself, raise this run's own exit code: that fact is real and reported (used_by/ required_symbol_contract/consumer_scope in the JSON report), but it is the supplied consumer's own informational assessment, separate from the library's own compatibility result. A library removal the consumer's own imports happen to name is still exactly as breaking as any other removal, through the ordinary compatibility axis — supplying --used-by/ --required-symbol changes what is reported, never what the library's own exit code is.


abicheck deps tree

Exit code Meaning
0 All dependencies resolved, all required symbols bound
1 Missing dependencies or unresolved symbols (binary would fail to load)
64 Invalid invocation (bad arguments/options)

--dry-run shows the resolved binary path and search order without resolving the dependency tree — see --dry-run below.


abicheck deps compare

Sysroot flags are --old-root/--new-root (default / for each — renamed from the old --baseline/--candidate).

Exit code Verdict Meaning
0 PASS Binary loads and no harmful ABI changes
1 WARN Binary loads but ABI risk detected in dependencies
4 FAIL Load failure or binary ABI break in dependencies
5 — not_comparable — at least one dependency's before/after DSOs were not extracted under a comparable profile/scope contract, so its per-library ABI diff never ran. Dominates 0/1/4, the same "couldn't establish what changed" precedence the gate uses elsewhere.
64 — Invalid invocation (bad arguments/options)

--dry-run shows the old/new roots, resolved binary paths, and search order without running per-library ABI diffs — see --dry-run below.

CI gate patterns

# Full-stack check: fail on FAIL, warn on WARN
abicheck deps compare usr/bin/myapp --old-root /old-root --new-root /new-root
ret=$?
[ $ret -eq 4 ] && echo "FAIL — load failure or ABI break" && exit 1
[ $ret -eq 1 ] && echo "WARN — ABI risk detected" && exit 1
[ $ret -ne 0 ] && echo "ERROR — unexpected non-verdict exit code: $ret" && exit 1
echo "PASS"

# Permissive: only fail on load failure / ABI break
abicheck deps compare usr/bin/myapp --old-root /old-root --new-root /new-root
ret=$?
[ $ret -eq 4 ] && exit 1   # FAIL only; WARN (exit 1) treated as OK
[ $ret -ne 0 ] && [ $ret -ne 1 ] && exit 1   # fail closed on non-verdict errors
exit 0

--dry-run (dump, compare, scan, deps tree, deps compare)

Every one of these five commands accepts --dry-run: it resolves and validates the invocation — classifies inputs, discovers config, and (per command) shows which data layers (L0–L5) are available, the audit checks and comparison that would run, or the resolved binary path/search order — and prints a report without doing the real work. It is cheap and read-only: no compiler invocation, no build-system query, no network access, and it writes nothing — passing -o/--output together with --dry-run is a usage error.

Exit code Meaning
0 Resolved cleanly — ok to proceed
1 Blocked — the invocation would fail once actually run
64 Usage error (e.g. -o/--output passed together with --dry-run)

--dry-run never returns a verdict code. It exits 0/1/64 only — never 2, 4, 5, or 8, even on a command whose real run could produce one of those.


Summary table

Verdict / State compare exit (legacy) compare exit (severity) scan exit deps tree exit deps compare exit
NO_CHANGE / PASS / compatible 0 0 0 0 0
COMPATIBLE 0 0 0‡ — —
COMPATIBLE_WITH_RISK 0 0–2* 0 / 0–2*‡ — —
Additions only 0 0–1* 0 / 0–1*‡ — —
Quality issues only 0 0–1* 0 / 0–1*‡ — —
WARN (ABI risk) — — — — 1
API_BREAK 2 0–2* 2 / 0–2*‡ — —
BREAKING / FAIL 4 0–4* 4 / 0–4*‡ — 4
--budget overflow — — 5 — —
Missing dependencies/symbols — — — 1 —
Load failure — — — — 4
Invalid invocation / tool error 64† 64† 64† 64† 64†

In the scan column, the value left of the / is the legacy (verdict-based) mapping — the default — and the value right of it applies once scan --against resolves the severity scheme (any --severity-preset/ --severity-*, or a config severity: block), where it follows the same compare exit (severity) column; see "scan --against and severity" above.

App/plugin-scoped comparisons (compare --used-by/--required-symbol) reuse the compare columns above — see Application- and plugin-scoped comparisons. aggregate combines each report's own severity gate (0/1/2/4) over its analyzed targets and adds a coverage gate (a required gap exits 1, never 4) — see abicheck aggregate. --dry-run (on dump/compare/scan/deps tree/deps compare) reuses none of these rows — it always exits 0/1/64; see --dry-run above.

* Severity exit codes depend on the configuration, and the range covers the whole configuration space — including demotion of a real break. With severity.addition: error, additions exit 1; with --severity-preset info-only every category is info, so everything exits 0, a BREAKING comparison included. The default preset leaves potential_breaking at warning, so an API_BREAK exits 0 unless --severity-preset strict (or severity.potential_breaking: error) raises it to 2. Read the report's own severity gate block — exit_code/blocking/blocking_categories — rather than inferring the cause from the code.

† Every command exits 64 for an invalid invocation — bad arguments/options or an unreadable/unrecognised input — deliberately outside the verdict/result space so a usage error is never mistaken for a compatibility result. To reliably distinguish verdicts from errors in a script, use -o json=... and read the verdict field where available.

‡ Two schemes, shown as legacy / severity. scan's legacy scheme (the default) collapses every compatible/advisory-only state (no break, deployment risk, additions, quality signals) to exit 0 — read -o json if your pipeline needs to distinguish them. Under a resolved severity scheme (scan --against with any --severity-* flag, or a config severity: block) scan follows the compare exit (severity) column on the same * terms, in both directions: severity.addition: error exits 1 on an additions-only diff, and --severity-preset info-only exits 0 on a BREAKING one. See "scan --against and severity".