How would you use jacoco-report-aggregation as the single source of truth for coverage across a large org's CI, and what are the trade-offs versus per-module enforcement?
answer
- convention plugin provisions :coverage everywhere
- one canonical testCodeCoverageReport XML -> dashboard/gate
- aggregate = honest cross-module number
- but masks weak modules / no ownership
- pair with per-module rules; cost = run-all-tests
basics
~20 sStandardize a convention plugin that adds a code-free aggregation module, have CI run one testCodeCoverageReport and publish its XML to the coverage service. Aggregated gives a true cross-module number but can mask weak modules; pair it with per-module rules for accountability.
solid answer
~40 sTreat aggregation as the org-wide reporting backbone: a shared convention plugin provisions a standard `:coverage` module per repo, applies `jacoco-report-aggregation`, and exposes one canonical `testCodeCoverageReport` task whose XML feeds the coverage dashboard (Codecov/SonarQube) and the PR gate. The advantage is a single honest number that correctly attributes cross-module integration coverage. The trade-off: a single aggregate threshold can hide a poorly-tested module behind well-tested ones, and it can't enforce ownership. So combine layers — aggregation for the headline number and integration-attributed coverage, plus per-module verification rules (delegated to the Coverage Verification Rules concern) for accountability and to fail the offending module's build. Also weigh CI cost: the aggregation task runs all contributing test suites, so on huge repos you may aggregate only on integration/merge pipelines and rely on per-module checks on fast PR pipelines.
code
kotlin · 7 lines// buildSrc/.../coverage-convention.gradle.kts (applied org-wide)
plugins { id("jacoco-report-aggregation") }
// CI then runs one canonical task and publishes its XML:
// ./gradlew :coverage:testCodeCoverageReport
// upload build/reports/jacoco/testCodeCoverageReport/
// testCodeCoverageReport.xml -> coverage servicego deeper
Know the aggregated XML can feed a coverage service for one combined number.
Explain publishing the single report to CI and that an aggregate alone can hide weak modules.
Design the CI flow: one canonical task + XML, aggregation for reporting, per-module rules for enforcement, cache for cost.
Own the org strategy: convention plugin rollout, layered reporting-vs-enforcement, pipeline placement, and governance of policy changes.
## Goal: one number, governed centrally At org scale you want consistent, low-maintenance coverage reporting that's hard to get wrong per-team. The aggregation plugin is the natural backbone, wrapped in **convention**. ## Architecture 1. **Convention plugin** (precompiled script or binary plugin) that, applied to a repo, creates a standardized `:coverage` project, applies `jacoco-report-aggregation`, and adds `jacocoAggregation` entries for the repo's published modules (or wires a settings convention to auto-include them). 2. **Single CI contract**: pipelines invoke `:coverage:testCodeCoverageReport`. The XML (`testCodeCoverageReport.xml`) is uploaded to the coverage service and/or parsed by the PR gate. 3. **Dashboards** read one artifact per repo, so trend lines and org rollups are clean. ## Why aggregation is the *reporting* truth Only the aggregated view correctly credits coverage when tests in module A exercise code in module B — per-module reports systematically under-count integration-style coverage. For an honest 'how much of our code is exercised' figure, you need aggregation. ## Why aggregation is *not* sufficient for enforcement - **Masking**: one 95%-covered module can hide a 10%-covered one under a single 80% aggregate gate. - **No ownership signal**: an aggregate failure doesn't point at a responsible team/module. - **Coarse feedback**: developers want the *offending module's* build to fail, locally and fast. So enforcement belongs to **per-module verification rules** (a separate concern), with aggregation supplying the trustworthy headline and integration-attributed view. The two compose: per-module `violationRules` keep each module honest; the aggregate number drives org dashboards and detects integration-only coverage. ## CI cost trade-offs The aggregation task transitively runs every contributing test suite, so it's as expensive as 'run all tests'. On large monorepos: - Run **per-module fast checks** + changed-module tests on PR pipelines. - Run **full aggregation** on merge/nightly pipelines and publish the canonical number there. - Lean on the build cache so unchanged modules' tests/exec are reused. ## Governance pointers - Version the convention plugin so coverage policy changes roll out centrally. - Standardize the report XML path so tooling is uniform. - Treat a *drop* in the aggregate as a signal, but enforce *floors* per module.
- Why not just set a single aggregate coverage threshold and skip per-module gates?An aggregate can mask a poorly-tested module behind well-tested ones and gives no ownership signal. Per-module floors keep each team accountable; aggregation supplies the honest headline and integration-attributed view.
- What's the CI cost consideration for large monorepos?The aggregation task transitively runs all contributing test suites, so it's as costly as running everything; run full aggregation on merge/nightly pipelines and lean on build cache, while PR pipelines run changed-module checks.
- How do you roll out a coverage-policy change across many repos?Encapsulate it in a versioned convention plugin that provisions the aggregation module and CI contract, then bump the plugin version per repo.
saying these in an interview costs you the question
- Treating a single aggregate threshold as sufficient governance — it masks weak modules.
- Ignoring that aggregation runs all test suites and is therefore expensive on every PR.