Your CI pipeline retries each failed browser end-to-end test up to twice and reports the run green if a retry passes. What does that policy buy you, what does it cost, and what would you put around it?
answer
- retries are sampling, not fixing
- green on retry is evidence, not absolution
- intermittent bug looks exactly like a flake
- gate the aggregate flake rate
- no retries on release verification
basics
~20 sRetries buy a usable pipeline while flakes exist, at the cost of signal: they can hide a rising flake rate and mask real intermittent bugs. Keep them only with retry counts recorded as a metric, a flake budget that fails the build, and quarantine with owners.
solid answer
~50 sRetries are a sampling decision, not a fix. What they buy is throughput — a browser suite of any size will sometimes lose to a runner hiccup, and blocking every merge on that is worse than the alternative. What they cost is truth: a passed-on-retry result is evidence of nondeterminism somewhere, and if it is silently rolled into green, the flake rate climbs invisibly and a genuine intermittent product bug is indistinguishable from noise. The guardrails matter more than the retry count. Emit passed-on-retry as a distinct outcome and track it per test; set a suite-level flake budget that fails the build when breached, so the trend is gated even though individual tests are not; cap retries at one or two and never retry non-idempotent tests that leave data behind; separate infrastructure retries from test retries; and disable retries for release verification, where you want the truth.
go deeper
Know that re-running a failed test until it passes is not the same as fixing it, and that a test which only passes sometimes is telling you something either about the test or about the app.
Explain the tradeoff concretely: retries keep a large browser suite usable, but a passed-on-retry result must be recorded, because otherwise the flake rate grows with no visible symptom.
Be ready to design the guardrails — a distinct passed-on-retry outcome, a low cap, no retries for tests with side effects, raw results for release verification — and to argue that some intermittent failures are product bugs that retrying would ship.
Own the policy end to end: who is accountable for the flake rate, what budget gates the build, how quarantine expiry is enforced, how flakes are attributed to teams without inviting coverage deletion, and how deflaking work is funded against feature pressure.
## What a retry actually is Automatic retry says: this test's result is a sample from a distribution, and we will take up to three samples and keep the best. That is a defensible engineering choice for a browser suite, where a runner can lose a container, a network can blip, and a shared environment can stall. It is not a defensible answer to "why is this test nondeterministic", and confusing the two is the failure mode the question is probing. ## What it buys Throughput and morale. Without retries, a suite with a 0.5% per-test flake rate and 400 tests fails most runs by arithmetic alone, and the team learns that red means nothing — which is a far worse outcome than retries, because it destroys the gate's meaning through habit rather than policy. Retries also let you keep a large end-to-end suite on the merge path at all, and they buy time to deflake without stopping delivery. ## What it costs **Signal.** A passed-on-retry rolled into a green run erases the only evidence that anything was wrong. The flake rate can double over a quarter with no visible symptom until it crosses the point where three attempts are not enough. **Real bugs.** A double-submit that fires one time in twenty, a race between two requests, a stale-cache flash — these fail intermittently in CI exactly as flakes do. Retrying them ships them. This is the strongest argument for treating passed-on-retry as an event that must be looked at rather than absorbed. **Money and time.** Retrying a slow browser test three times triples its wall clock on the runs that need it, and a badly flaky suite spends a large fraction of its compute re-running. **Culture.** Once "just re-run it" is the accepted first response to red, it becomes the response to genuine failures too. ## The guardrails - **Make the flake visible.** Report passed-on-retry as its own outcome, per test, into a durable store. The number you manage is the flake rate, not the individual failures. - **Gate the trend, not the test.** Individual flaky tests do not block merges; a suite whose flake rate exceeds its budget does. That inverts the incentive: the team is accountable for the aggregate, and deflaking becomes a scheduled cost rather than heroism. - **Cap the count.** One or two retries. A test that needs three attempts is not flaky, it is broken, and a high cap mostly buys compute. - **Retry only what is safe to retry.** A test that creates data, sends an email, or moves an external resource may fail differently on the second attempt because the first left residue. Retrying such a test produces confusing failures; the precondition for retry is that the test can start from a clean state. - **Separate infra retries from test retries.** Losing a runner is an infrastructure event and should be retried at the job level without counting against the flake budget. A test that failed on its own merits is a different signal, and merging the two hides both. - **Turn retries off where truth matters.** Release verification, smoke tests against production, and any test guarding a payment or auth path should report the raw result. - **Quarantine with an expiry.** Chronically flaky tests come off the gate with a named owner and a date, keep running, and report. An expired quarantine entry should itself fail the build so the decision — fix it or delete it — is forced rather than deferred. ## The organisational question underneath The retry policy is really a question about who owns suite reliability. If nobody does, retries are a slow leak: each team adds one more flaky test, nobody's dashboard turns red, and in a year the suite is advisory. Making the flake rate a visible, owned, budgeted number — with the retry mechanism feeding it rather than hiding it — is what keeps an end-to-end suite worth running. The honest framing in an interview is that you keep retries *and* you make them expensive to rely on.
- If you gate on a suite-level flake budget rather than on individual flaky tests, how do you stop teams from gaming it?Attribute flakes to owning teams so the aggregate is not an anonymous pool, and count quarantined and skipped tests against the budget too — otherwise the cheapest way to meet it is to remove coverage. Publish the trend, review it on a regular cadence, and treat a budget breach like any other broken build: it blocks until someone acts.
- Which end-to-end tests would you exclude from automatic retries entirely?Tests that are not safely repeatable — anything creating external side effects that a second run would duplicate — and tests whose whole purpose is to detect intermittent behaviour, such as concurrency or double-submit guards. Also exclude release-verification and production smoke runs, where you want the unretried truth rather than the best of three attempts.
- How would you tell whether the retry policy is currently masking a real product bug?Mine the passed-on-retry records for patterns rather than treating each as noise: the same test, the same step, the same console error or failed request across many runs points at the application, not at test timing. Any retry whose artifacts show the app in a state a user could reach — duplicate order, stale value, unhandled rejection — should be triaged as a product defect.
saying these in an interview costs you the question
- Says retries are the standard fix for flaky suites
- Counts passed-on-retry as an ordinary green result
- Sets a high retry cap to keep the pipeline green
- Retries tests that leave data or send external side effects
- Keeps retries enabled for release verification runs
- Assumes an intermittent failure is never a product bug