How would you design a CI pipeline so unit tests give fast feedback but integration tests still run reliably, using Surefire and Failsafe?
answer
- two loops: fast unit / slow IT
- mvn test then mvn verify
- -DskipITs to gate stages
- forkCount/parallel bound IT time
- deferred fail keeps CI agents clean
basics
~10 sRun unit tests (Surefire) early via mvn test for fast feedback, then run integration tests (Failsafe) in a later stage via mvn verify with -DskipITs control, so slow ITs don't block quick signal.
solid answer
~40 sI separate the runs by lifecycle and CI stage. Stage one runs `mvn test` (Surefire only) for sub-minute feedback on unit failures — fail-fast is desirable here. Stage two runs `mvn verify` so Failsafe executes integration tests with proper setup/teardown; its deferred-failure-to-verify behavior keeps resources clean even on failure. To avoid double-running unit tests I can reuse the build artifact or accept the cost. I gate ITs with `-DskipITs` so feature branches can opt in/out, and use Failsafe's `<groups>`/parallel options (`forkCount`, `parallel`) to bound IT duration. For flaky-IT isolation I tag and quarantine. The principle: unit tests optimize for speed and locality; integration tests optimize for fidelity and reliable cleanup, and the two plugins let me bind each to the right place.
code
bash · 4 lines# PR stage: fast unit feedback
mvn -DskipITs test
# Merge/nightly stage: full integration verification
mvn verifygo deeper
Knows unit tests are fast and run first; ITs run later.
Maps the two stages to mvn test and mvn verify with skip flags.
Adds parallelism, dynamic ports, and report separation for reliability.
Defines the org-wide pipeline template, quarantine policy, and unit/IT governance.
## Goal: two feedback loops - **Fast loop (unit):** developers and CI want a failure in seconds-to-a-minute. Surefire's fail-fast `test` phase is ideal. - **Slow loop (integration):** higher fidelity, needs infra, must clean up. Failsafe's `verify` flow is ideal. ## Pipeline shape ``` Stage 1: mvn -DskipITs test # Surefire unit tests, fast fail Stage 2: mvn verify # Failsafe ITs with setup/teardown ``` Stage 1 surfaces the cheap failures first. Stage 2 only runs after compile/unit pass, so you don't pay for slow ITs on an obviously broken build. ## Controlling cost and flakiness - `-DskipITs` lets non-IT pipelines (e.g. quick PR checks) stay fast; nightly/merge pipelines run full `verify`. - Parallelism: Failsafe supports `<forkCount>` (separate JVMs) and `<parallel>` (threads within a JVM) plus `<threadCount>` to bound wall-clock time. Watch for shared-resource contention (one DB, one port). - Tag-based quarantine: JUnit 5 `@Tag("flaky")` plus `<excludedGroups>flaky</excludedGroups>` in the main run and a separate tolerant job for quarantined tests. - `<rerunFailingTestsCount>` can retry transient failures, but use sparingly — it can mask real flakiness. ## Reliability guarantees that matter on CI - Failsafe's deferred failure means start/stop goals bracket the run; even on failure, `post-integration-test` cleanup runs, so CI agents don't accumulate zombie servers/containers across jobs. - Reserve dynamic ports (`build-helper` `reserve-network-port`) because CI runners are shared and ephemeral. - Publish `target/failsafe-reports` and `target/surefire-reports` separately so dashboards distinguish unit vs IT failures. ## Avoiding double work `mvn verify` re-runs unit tests by default. Options: accept it (simplest), pass `-Dsurefire.skip=true` in the IT stage if the unit stage already certified them (risky if stages drift), or run a single `mvn verify` and split reporting. Most teams just run one `mvn verify` per merge build and a lighter `mvn test` per push. ## Governance As an architect I'd standardize: naming (`*Test` vs `*IT`), the two-stage pipeline template, the skip flags, the report locations, and a flaky-test quarantine policy — so every module behaves identically and the unit/IT boundary stays clean.
- How do you keep flaky integration tests from blocking the pipeline?Tag them (@Tag("flaky")), exclude via <excludedGroups> in the gating run, and run them in a separate non-blocking job; optionally limited rerunFailingTestsCount, used cautiously.
- Why run unit tests in a stage before integration tests rather than together?To get fast failure signal cheaply; there's no point spinning up infrastructure for ITs if unit tests already prove the build is broken.
saying these in an interview costs you the question
- Treating integration tests as a substitute for fast unit tests.
- Relying on rerunFailingTestsCount to hide chronic flakiness.
- Running everything as ITs so every CI run is slow.