ADR-036: Report view-model and canonical report severity¶
Status¶
Accepted — core implemented (Increments 1-2: ReportModel, the consolidated
verdict→vocabulary table, and the cross-channel integrity tests are in place
and route the Markdown/text/JSON/SARIF/JUnit reporters). Increment 3 (routing
html_report.py and pr_comment.py through ReportModel to delete their
remaining local bucketing, e.g. pr_comment._bucket_changes) is still open —
see "Rollout" below.
2026-09-14 presentation amendment¶
The ordinary compare human output now defaults to the bounded terminal
projection rather than the exhaustive Markdown document. This is an
intentional presentation default change only: comparison, classification,
suppression, gating, and exit-code computation are unchanged. Users and
automation that require the detailed prose report must request it explicitly
with -o markdown=- (or write it beside a compact summary with repeated
exports). JSON, SARIF, JUnit, HTML, oneline, and explicit Markdown requests
retain their existing contracts.
The review projection is bounded across public-surface entries and always
discloses omitted counts. Public-surface operation/entity classification comes
from ChangeKindMeta, not kind-name suffixes or effective severity; dependency
imports and environment requirements are consequently not presented as public
API operations.
The completed redesign adds report-schema 5.1's review_groups and
result_counts blocks. A review group is a presentation relationship, never a
semantic deduplication: it retains every member finding ID and finalized gate
contribution. Its key includes library scope and canonical entity identity;
only explicitly supported corroborating families are joined. Count populations
are named by stage (raw_detected, retained, gating, non_gating) and by
independent review dimension (groups, public operations, runtime/dependency
observations, and hygiene lifecycle).
The compact view has deterministic caps for groups, symbols, warnings,
suppression rules, acknowledgments, and modulation rows. Omitted counts and a
valid detailed-export command are always printed. Full Markdown/HTML and
machine exports remain complete. Release/SONAME advice is shown only when the
project explicitly stated a versioning policy; strict_abi alone does not
imply SemVer.
Evidence metrics such as artifact_backed_findings deliberately remain a
provenance subset, not another retained-total spelling: policy overlays and
findings backed only by source/build context are retained but not
artifact-backed. Compact output therefore uses result_counts.retained for
the retained population and never tries to reconcile it to a provenance
subset as if an equality were expected.
Context¶
abicheck renders a comparison result in many formats (JSON, Markdown, text, HTML, SARIF, JUnit, PR comment). Historically each renderer independently:
- re-applied the
--show-onlydisplay filter (apply_show_only(list(result.changes), …)repeated inreporter,sarif,junit,html_report), and - re-derived how to bucket the change set for display.
The bucketing is where the real hazard lay. There are in fact three different classification axes in the codebase, which had been conflated:
| Axis | Question it answers | Home |
|---|---|---|
| Verdict (BREAKING / API_BREAK / RISK / COMPATIBLE) | "does this break the ABI, and does it gate CI?" | checker_policy, severity, DiffResult._effective_verdict_for_change |
| Display severity (HIGH / MEDIUM / LOW) | ABICC-style report colouring | report_classifications |
| Origin (rtti / internal / public) | "is a big breaking count just RTTI/internal churn?" | report_summary |
Because every renderer re-computed the verdict axis on its own (and the PR comment used yet another string-keyed bucket dict), the formats could disagree with each other and, worse, with the gate/exit code.
Decision¶
-
Introduce a
ReportModel(abicheck/report_model.py) — a render-ready value object built once from aDiffResult: the (optionallyshow_only- filtered) change set, the four verdict-axis buckets, and the headline summary. Renderers become thin projections over it. The verdict→vocabulary projections (native label, SARIF override level, breaking boundary) live in a single authoritativeVERDICT_PRESENTATIONtable; the legacyVERDICT_TO_SEVERITY_LABEL/VERDICT_TO_SARIF_LEVELdicts are derived from it so there is exactly one source of truth, asserted internally consistent bytests/test_report_integrity.py. -
Canonical report severity = the verdict axis, specifically each finding's
result._effective_verdict_for_change(c)(which already honours PolicyFile overrides and ADR-027 A4 per-finding modulation). This is the same partition that produces the overall verdict and the process exit code, so a report can never contradict the gate. -
The other two axes are kept as deliberate, separate projections, not collapsed. Display severity (HIGH/MEDIUM/LOW) is an ABICC-compatibility presentation; origin (rtti/internal/public) explains breaking-count composition. They answer different questions from the verdict axis, so merging them would lose information rather than de-duplicate.
ReportModelexposes the verdict axis; the display/origin projections remain available viareport_classifications/report_summary. -
Cycle-safety:
report_modelimports onlychecker_policyandreport_summary. Theshow_onlyfilter (apply_show_only) stays inreporter; callers apply it and pass the filtered list intoReportModel.from_result(result, changes=…), soreport_modelnever importsreporterandreporterdepends on it one-directionally.
Consequences¶
- Single canonical verdict-axis bucketer:
reporter._classify_changes_by_kindis now a thin wrapper overReportModel.classify. New output formats classify via the model instead of re-deriving buckets. - No behaviour change in this first increment — the Markdown/text path was routed through the model and the golden snapshots are byte-identical.
- The verdict axis is documented as canonical, settling the "which severity is authoritative?" question for future renderer work.
Cross-channel invariant (what "unified" actually means)¶
Investigating the renderers showed they are not meant to emit identical vocabulary, and forcing that would be wrong:
- Native channels (JSON, text/Markdown, JUnit) classify on the verdict axis.
- SARIF keeps a finer per-kind level (
policy_for(kind).severity) — e.g. additions are SARIFwarning, notnote(there is a long-standing test for this). On the A4/PolicyFile override path it maps the overridden verdict viaVERDICT_TO_SARIF_LEVEL. - ABICC-compat HTML uses ABICC's own kind-based HIGH/MEDIUM/LOW so ABICC report parsers/diffs keep working — deliberately not the verdict axis.
So "unified" means two concrete, testable guarantees, not identical buckets:
- Breaking-boundary consistency. A finding on the breaking side of the gate
(BREAKING/API_BREAK) reads as error/failure/breaking in every native
channel; one off it never does. (
ReportModel.is_breaking_boundary.) - Override propagation. A PolicyFile/A4 effective-verdict override is honoured by every native channel (the demoted-change case).
tests/test_report_integrity.py asserts both across JSON/SARIF/JUnit and pins
the ABICC-HTML exception as a conscious decision.
Rollout¶
- Increment 1 (done): add
ReportModel; route the Markdown/text reporter and the shared classifier through it; golden output unchanged. - Increment 2 (done): consolidate the previously-duplicated verdict→vocabulary
maps (
reporter._VERDICT_TO_SEVERITY_LABEL,sarif._VERDICT_TO_SARIF_LEVEL) intoreport_modelso they can no longer drift; add the cross-channel integrity tests above. No behaviour change (the maps were identical); the tests are the new guard. - Increment 3 (follow-up): optionally route
html_report(native, non-compat path) andpr_commentmodel construction throughReportModelto delete their remaining local bucketing. Pure cleanup; the integrity invariant already holds.
Alternatives considered¶
- Collapse all three axes into one severity enum. Rejected: display severity and origin are genuinely different questions; collapsing loses the ABICC-compat colouring and the RTTI/internal-churn explanation.
- Put
apply_show_onlyinreport_model. Rejected: it would force areport_model ↔ reporterimport cycle (the readiness gate flags it). Keeping the filter inreporterand passing filtered changes in keeps the dependency one-directional.