skip to content

Waiting & Flakiness Management

You will learn how to make an e2e test wait on a real signal instead of a stopwatch, and what to do with the flakes that survive anyway. Interviewers phrase it as 'your suite is red ten percent of the time — what now?' and want a triage process, not a retry flag.

on this pageshow

questions

5

In a browser end-to-end test, why is inserting a fixed sleep (such as Playwright's page.waitForTimeout(500) or Cypress's cy.wait(500)) before an assertion a bad way to handle timing, and what do you use instead?

level: juniorimportance: must knowfreq 78%

answer

  1. a sleep asserts nothing about the app
  2. wrong in both directions
  3. slow machine versus wasted seconds
  4. poll a condition, return early
  5. timeout is a bound, not a cost

basics

~20 s

A fixed sleep hard-codes a guess about duration: too short and the test fails on a slow machine, too long and every run pays the delay. Wait on a retrying condition instead, which returns the moment the expected state appears.

solid answer

~50 s

A sleep asserts nothing — it just bets that the app will be ready within N milliseconds. That bet is wrong in both directions: on a loaded CI runner the app takes longer and the test fails for no product reason, and on a fast machine every run burns the full delay, so a few hundred sleeps become minutes of wall clock. The replacement is a wait on an actual condition with a generous timeout: a retrying, "web-first" assertion such as `expect(locator).toBeVisible()` in Playwright, a retried `cy.get(...).should(...)` in Cypress, or an explicit wait on an expected condition in WebDriver. Those poll until the condition holds and return immediately when it does, so the test is both faster and stable. A sleep is defensible only as a commented last resort for something with no observable signal at all.

code

javascript · 8 lines
javascript
import { test, expect } from '@playwright/test';

test('saving shows a confirmation', async ({ page }) => {
  await page.goto('/settings');
  await page.getByRole('button', { name: 'Save' }).click();
  // No sleep: the assertion re-runs until it passes or the timeout expires.
  await expect(page.getByText('Saved')).toBeVisible({ timeout: 10_000 });
});

go deeper

for a junior

Be ready to say plainly that a fixed pause is a guess and that you wait on a condition instead. Naming one retrying assertion you have actually used is enough at this level.

for a middle

Explain the polling mechanics: the wait re-evaluates until true or the timeout expires, returns early on success, and reports which condition failed. Be able to argue why the timeout value is not the same thing as the wait's cost.

for a senior

Show the judgment about what to poll on — a user-visible end state, not an incidental spinner — and be able to say when a sleep is genuinely unavoidable and how you would get rid of it by asking the app for a signal.

for a principal

Own the policy angle: how sleeps accumulate into pipeline minutes across hundreds of tests, how you would detect and ban them in review or lint, and what you would ask product engineers to expose so tests never need one.

## What a sleep actually says A statement like `await page.waitForTimeout(500)` makes no claim about your application. It says "pause this test for half a second", and nothing about whether the save completed, the list re-rendered, or the spinner went away. The test then asserts and hopes. Every sleep in a suite is a hard-coded guess about a duration that depends on machine speed, network latency, CI contention, browser warm-up, and how much data the fixture happened to create. ## Both failure modes are bad A sleep that is too short is the classic flake: it passes on the developer's laptop and fails on a CI runner that is running six workers on four cores, where the same render takes 900 ms instead of 300 ms. Nothing about the product broke, but the build is red, and the usual reaction — bumping 500 to 2000 — makes the second failure mode worse. A sleep that is too long is a tax paid by every single run, including the 99% that were ready in 80 ms. Sleeps compound: two hundred of them at one second each is over three minutes added to every pipeline run, forever, and that cost is invisible in any one test's diff. Worse, a long sleep does not actually remove the flake — it just lowers its probability. The distribution of "how long the app takes" has a tail, and CI runs often enough to find it. ## The replacement: poll a condition Every serious browser automation tool offers the same primitive under different names: evaluate a condition repeatedly until it becomes true or a timeout expires. Playwright's `expect(locator).toBeVisible()` and friends retry internally; Cypress retries the `cy.get(...).should(...)` chain; Selenium/WebDriver uses `WebDriverWait` with an expected condition. The shape matters more than the API: - it returns as soon as the condition holds, so the fast path is fast; - the timeout is an upper bound on failure, not the normal cost; - when it does time out, the failure message names the condition that never became true, which is a far better diagnostic than "expected 1, got 0" after a blind sleep. ```js // Guess: fails on a slow runner, wastes time on a fast one await page.waitForTimeout(500); await expect(page.getByText('Saved')).toBeVisible(); // Condition: returns the instant the text appears, fails after 10s with a clear message await expect(page.getByText('Saved')).toBeVisible({ timeout: 10_000 }); ``` ## What to poll on Poll on the thing the user would look at to decide the app is ready: the confirmation text, the new row, the disappearance of the skeleton, the disabled state clearing. That keeps the wait and the assertion the same statement, so there is no window between "ready" and "checked" for the app to change again. Polling on something incidental — an internal class name, a spinner element that only exists for 40 ms — reintroduces the race in a new disguise. ## Timeouts still exist, and are not the fix A retrying wait still needs an upper bound, and choosing it is a real decision: long enough to absorb a slow CI runner, short enough that a genuine hang fails in reasonable time. Set a sane default globally and override it per assertion where a step is legitimately slow (a report that takes fifteen seconds to generate). What you should not do is respond to a flaky suite by multiplying every timeout — that converts fast failures into slow ones and hides the underlying cause. ## When a sleep is genuinely the last resort There are cases with no observable signal: waiting out a debounce you cannot control, letting a third-party script settle, or working around an animation you cannot disable. If you must sleep, keep it small, keep it local, and leave a comment naming the signal you wished existed. Better still, treat the missing signal as a product gap and ask for one — a `data-*` attribute the app sets when idle, or an event it emits — so the next test can wait on a real condition rather than on the clock.

  • The suite is flaky, so someone raises every timeout from 5 to 30 seconds. Why is that not a fix?
    It lowers the flake probability without removing the cause, and it makes genuine failures take six times longer to report. A test that hangs now burns 30 seconds before telling you anything. Raising a timeout is correct only when you can point at a step that is legitimately slow — a report build, a large upload — and then you raise that one assertion, not the global default.
  • Is there any case where you would still write a fixed sleep in an end-to-end test?
    Yes, as a commented last resort when there is no observable signal at all: riding out a debounce you cannot control, or letting a third-party widget settle. Keep it short and local, and treat it as a bug report against the app — ask for a state attribute or an event to wait on so the next test does not need the sleep.
  • Why is a retrying wait usually faster than the sleep it replaces, not slower?
    Because the timeout is an upper bound on failure, not the normal cost. The poll returns the instant the condition holds, so a step that is ready in 80 ms costs 80 ms even if the timeout is ten seconds. A 500 ms sleep costs 500 ms every run, including the runs where the app was ready immediately.

Sleeping for a fixed time is guessing how long the kettle takes; waiting on a condition is listening for the whistle. The whistle is right on a cold morning and on a hot one.

saying these in an interview costs you the question

  • Says flakiness is fixed by sleeping longer
  • Treats a global timeout increase as the standard remedy
  • Believes a retrying assertion always waits the full timeout
  • Waits on a spinner appearing rather than on the end state
  • Calls the sleep harmless because the suite is small

context

open as a page

Your end-to-end framework auto-waits for an element to be attached, visible and actionable before every click, yet the suite still has timing flakes. What kinds of waiting does that built-in auto-waiting not cover?

level: middleimportance: must knowfreq 58%

basics

~20 s

Auto-waiting covers only the element you are about to touch, one action at a time. It knows nothing about whether the app is still fetching, whether a later re-render will overwrite what you just asserted, and an assertion of absence gives it nothing to wait for.

open as a page

Your browser end-to-end suite fails on about one run in ten, a different test each time, and every one of them passes when re-run. How do you triage that?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Treat it as one reliability problem, not ten separate bugs. Record per-test outcomes so you can rank flakes by how often they block a merge, reproduce from the failing run's trace, video and logs, classify the cause, quarantine the worst offenders, then fix them.

open as a page

In an end-to-end test that submits a form and expects a new row in a list, when should the test wait on the network response, and when should it just assert on the rendered result?

level: middleimportance: should knowfreq 52%

basics

~20 s

Default to asserting on the rendered result: it is what the user perceives and it implicitly covers the request. Wait on the response only when the effect has no visible consequence, or when the test needs the request or response payload itself.

open as a page

Your CI pipeline retries each failed browser end-to-end test up to twice and reports the run green if a retry passes. What does that policy buy you, what does it cost, and what would you put around it?

level: principalimportance: should knowfreq 44%

basics

~20 s

Retries buy a usable pipeline while flakes exist, at the cost of signal: they can hide a rising flake rate and mask real intermittent bugs. Keep them only with retry counts recorded as a metric, a flake budget that fails the build, and quarantine with owners.

open as a page