A teammate proposes annotating a suspected intermittently-failing JUnit 5 test with @RepeatedTest(50) and leaving it in CI to catch the problem. How do you evaluate that proposal, and what would you do instead?
answer
- Repetition ≠ retry — all repetitions must pass
- Same JVM/thread/order → misses ordering + concurrency flakes
- Fresh instance per repetition, but statics/DB/clock shared
- 50x setup cost; failureThreshold limits noise not runtime
- Fix determinism: fixed Clock, seeded Random, await not sleep; quarantine meanwhile
basics
~20 sRepetition only catches flakes whose cause varies run to run in the same JVM and thread — random data, timing, leaked static or database state. It misses ordering, cross-test and concurrency flakes, multiplies CI time by 50, and adds no retry. Use it temporarily to reproduce and later to prove a fix; fix the root cause.
solid answer
~60 sIt is a useful **diagnostic**, a poor **permanent** fixture. What repetition does catch: nondeterminism inside the test itself — random or time-dependent inputs, hash/iteration-order dependence, state that leaks through statics, singletons, caches or a shared database (each repetition gets a fresh test-class instance but not a fresh world). What it misses: everything that depends on *context*. All 50 repetitions run in the same JVM, on the same thread, in the same class, at the same point in the suite — so ordering flakes, pollution from a *different* test class, machine-load and parallel-execution races are exactly as likely to reproduce as they were once. Costs: 50x the setup and runtime for that test, 50 report nodes, and no retry semantics — `@RepeatedTest` makes the suite fail *more* reliably, it does not make it green. What I'd do: reproduce locally with a high repetition count plus `failureThreshold`, capture a seed and per-repetition diagnostics, then fix the cause — pin the clock and the random seed, remove shared mutable state, replace sleeps with awaits. Keep repetition in CI only briefly as proof the fix holds, or quarantine the test while it is being fixed.
code
java · 17 linesimport org.junit.jupiter.api.RepeatedTest;
import org.junit.jupiter.api.RepetitionInfo;
import java.util.Random;
import static org.junit.jupiter.api.Assertions.assertEquals;
class SchedulerFlakeReproductionTest {
@RepeatedTest(value = 200, failureThreshold = 1)
void schedulesInOrder(RepetitionInfo info) {
long seed = System.nanoTime();
Random random = new Random(seed);
var input = randomJobs(random);
assertEquals(expected(input), scheduler.schedule(input),
() -> "repetition " + info.getCurrentRepetition() + " seed=" + seed);
}
}go deeper
Say that repetition can surface random or leftover-state problems, that every repetition must pass, and that it is not a retry.
Add the lifecycle detail — new instance per repetition but shared statics/DB — and the CI cost multiplier, plus concrete determinism fixes like a fixed Clock and a seeded Random.
Lead with the diagnostic-versus-permanent distinction, enumerate what repetition structurally cannot reproduce (ordering, parallel neighbours, environment), and lay out a reproduce → measure rate → fix → prove sequence with quarantine as the interim.
Frame flakiness as a suite-health policy question: base-rate measurement, quarantine with ownership and deadlines, whether retry extensions are ever allowed, and the cost of the confidence a green repeated run falsely buys.
## Separate the two goals "Catch a flaky test" conflates two different jobs: 1. **Reproduce** the intermittent failure so you can debug it. 2. **Stop it from breaking the build** while it is unfixed. `@RepeatedTest` is a decent tool for (1) and actively counterproductive for (2) — it does not retry, it does not quarantine, it just gives the flake fifty more chances to turn the build red. ## What repetition can and cannot reproduce Each repetition is a fresh test invocation: `@BeforeEach`/`@AfterEach` run again, and under the default `PER_METHOD` lifecycle a new test-class instance is constructed. But everything outside the instance is *shared*: static fields, singletons, framework caches, an embedded server, the database, the filesystem, the clock, the machine. So repetition reproduces: - **Input nondeterminism** — `Math.random()`, `UUID.randomUUID()`, `Instant.now()`, unordered `HashMap`/`HashSet` iteration, locale/timezone-sensitive formatting. - **Self-pollution** — the test leaves rows, files, static counters or cached entries behind, so run 2 sees a world run 1 did not. - **Timing inside the test** — a fixed `Thread.sleep(50)` that is usually but not always enough. It does **not** reproduce: - **Ordering / cross-test pollution** — when the real trigger is "class B runs after class A". Fifty runs in the same slot of the same order tell you nothing. - **Concurrency between tests** — under parallel execution the flake comes from *which other tests* are running alongside; sequential repetition removes that. - **Environment-dependent flakes** — CI agent load, disk pressure, container CPU throttling, DNS. These show up on the CI machine, not from repeating locally. - **Rare-event flakes** — if the failure rate is 1 in 5,000, 50 repetitions have roughly a 1% chance of catching it, and a green run proves nothing. Getting a useful confidence bound requires understanding the base rate first. ## The costs of leaving it in CI - **Time**: repetition multiplies the *whole* per-test cost, including `@BeforeEach`. A test with a 2-second Spring-context-dependent setup becomes a 100-second test. - **Report noise**: 50 nodes per method, and on failure up to 50 stack traces for one defect. `failureThreshold` mitigates the second problem but not the first. - **False confidence**: a green 50-run does not prove stability; it narrows the plausible failure rate, nothing more. Teams routinely read it as "fixed". - **It is code, not configuration**: the repetition count is compiled into the test, so tuning it means a code change and review, and it applies on every developer's machine too, not just CI. ## A better sequence 1. **Measure the rate.** Run the test many times locally (high repetition count, plus `failureThreshold` so a broken run exits fast) and record how often it fails. A rate turns "flaky" into a number you can act on. 2. **Make failures self-describing.** Inject `RepetitionInfo` and put the repetition index, the random seed and relevant state into assertion messages or `TestReporter` entries, so the one failing run out of fifty is diagnosable from the report alone. 3. **Attack determinism at the source.** Inject a fixed `Clock` instead of calling `Instant.now()`; use a seeded `Random` and log the seed; replace sleeps with polling awaits on a real condition; sort or use ordered collections where iteration order leaks into assertions; ensure teardown truly resets shared state (transaction rollback, truncation, cache eviction). 4. **Test the ordering hypothesis separately.** Randomise or reverse execution order, or run the suspect class both alone and after the suite, rather than repeating it in place. 5. **Contain the bleeding.** While it is unfixed, tag the test (for example with a custom `@Tag("flaky")`) and exclude it from the blocking pipeline, with an owner and a deadline. Quarantine is honest; a green build produced by luck is not. 6. **Prove the fix, then remove the scaffolding.** After the fix, a temporary high-count `@RepeatedTest` is exactly the right evidence — and then it comes back out, or drops to a small count if the value is ongoing. ## On retries Jupiter deliberately has no built-in `@Retry`. `@RepeatedTest` is often mistaken for one, but its semantics are the opposite: it requires *all* repetitions to pass. If retry-on-failure semantics are genuinely wanted, that is an extension-level concern (a custom `TestTemplateInvocationContextProvider` or execution-exception handler) — and it should be an explicit, owner-assigned, temporary decision, because a retried test silently converts a real intermittent product bug into a green build. ## The senior framing The strongest answer names the tradeoff: flakiness is a *signal about the system*, and repetition is a magnifying glass, not a cure. Use it to shorten the feedback loop while you find the nondeterminism; never let it become the team's standing answer to "this test is unreliable".
- The repeated test passes 200 times in CI. What have you actually proven?Only that the failure rate is probably low, not that it is zero. Two hundred green runs are consistent with a one-in-a-thousand defect, and they say nothing at all about failure modes that depend on test ordering, parallel neighbours or machine load, since every repetition ran in the same JVM, thread and position in the suite. The evidence supports 'I could not reproduce it this way', never 'it is fixed'.
- Is @RepeatedTest a way to make an unreliable test stop breaking the build?No — it does the opposite. Every repetition must pass, so repeating a test that fails one time in fifty makes the build fail almost every run. Retry-on-failure is not part of Jupiter's core; it would need a custom extension, and it is a deliberate, dangerous choice because it hides genuine intermittent product bugs. The honest short-term move is to quarantine the test with a named owner and a deadline.
- Which flake causes will repetition never surface?Anything that depends on context rather than on the test itself: pollution left by a different test class, dependence on execution order, races against tests running in parallel, and environment effects such as CI agent CPU throttling or disk contention. All repetitions run sequentially in the same JVM at the same point in the suite, so those conditions are held constant rather than varied.
Repeating a test is like re-driving the same stretch of road fifty times to find an intermittent rattle: great for a fault that depends on the car, useless for one that only appears on a different road surface.
saying these in an interview costs you the question
- Calling @RepeatedTest a retry mechanism, or expecting one passing repetition to make the test green.
- Claiming a green repeated run proves the flake is fixed.
- Ignoring that repetitions share statics, singletons, the database and the clock.
- Leaving a high repetition count permanently in CI without accounting for the runtime multiplier.
- Never considering ordering or parallel-neighbour causes, which repetition holds constant by construction.