skip to content

End-to-End Testing Patterns

You will learn the patterns that decide whether an end-to-end suite is trusted or muted: how you select elements, how you wait, how you get logged in, how you structure helpers, and where test data comes from. Interviewers ask about e2e strategy far more often than about any single runner's API.

on this pageshow

questions

26

An end-to-end suite signs in through the login form at the start of every test, and a full run now takes twenty minutes. Why is UI login per test a poor default, and what do end-to-end suites do instead?

level: juniorimportance: must knowfreq 74%

answer

  1. login is setup, not the assertion
  2. pay the cost once per run
  3. capture cookies and origin storage
  4. seed the context, then navigate
  5. one test still drives the real form

basics

~20 s

Driving the login form repeats a slow, flake-prone flow that adds no coverage after the first test. Authenticate once through the app's auth API instead, save the resulting cookies and browser storage as a session, and start every test already signed in.

solid answer

~40 s

UI login is a fixed cost paid by every test — page load, form fill, submit, redirect, token exchange — and it is one of the most common sources of flake, yet after the first test it proves nothing new. The usual fix is to authenticate out of band: call the application's own login endpoint with a test user's credentials, capture the session the server hands back (cookies, plus any token the app keeps in `localStorage`), and persist it. Each test then opens a fresh browser context seeded with that saved state and navigates straight to an authenticated page. Playwright exposes this as `storageState`; Cypress wraps the same idea in `cy.session()`. You keep exactly one test that drives the real sign-in form, so the flow itself is still covered.

code

typescript · 14 lines
typescript
import { chromium } from '@playwright/test';

export async function saveSession(statePath: string): Promise<void> {
  const browser = await chromium.launch();
  const context = await browser.newContext();
  const res = await context.request.post('https://app.example.com/api/login', {
    data: { email: '[email protected]', password: process.env.E2E_PASSWORD },
  });
  if (!res.ok()) {
    throw new Error(`programmatic login failed: ${res.status()}`);
  }
  await context.storageState({ path: statePath });
  await browser.close();
}

go deeper

for a junior

Be able to say plainly that logging in through the form in every test is slow and repeats coverage, and that suites instead sign in once via the auth API and reuse the saved session.

for a middle

Explain what a saved session actually contains — cookies plus origin storage — why a cookie-only restore can leave the app looking logged out, and why the session must match the origin under test.

for a senior

Show judgment about the seam: the session is real and issued by the server, never a stubbed-out auth check, and you deliberately retain one test that drives the real sign-in form so the flow is not orphaned.

for a principal

Own the tradeoff between suite speed and fidelity: which paths must be exercised end to end at least once, what a shared setup step costs in coupling and debuggability, and how the team stops the injected session from quietly becoming an unverified assumption.

## Why per-test UI login hurts Every end-to-end test needs an authenticated browser. The obvious way to get one is to do what a user does: open the login page, type an email and password, submit, wait for the redirect. For the first test that is exactly right — it *is* the test of login. For the other two hundred tests it is overhead, and the overhead is not small. - **Time.** A login round trip is usually several seconds: document load, framework boot, form interaction, network call, redirect, second page load. Multiply by the number of tests and it dominates the suite. Many teams find that half their end-to-end runtime is logging in. - **Flakiness.** The login screen touches everything fragile at once — a form, a network call, a redirect, sometimes a rate limiter or a bot check. A flake there fails a test that was about something else entirely, and the failure report points at login rather than at the feature under test. - **False coupling.** If the login page is redesigned, every test in the suite breaks, even though none of them are about login. That is a sign the step belongs in setup, not in the test body. The principle underneath: **a test should drive the UI only for the behaviour it is asserting.** Everything else — getting into the right state — should be established by the cheapest reliable means available, which is almost always the application's own API. ## What "programmatic login" means Programmatic login means obtaining a real session the way the app's backend issues it, without rendering the login screen: 1. Send a request to the app's real authentication endpoint with a test user's credentials. 2. The server responds with whatever constitutes a session — most commonly a `Set-Cookie` for a session or refresh cookie, sometimes a token in the response body that the app stores in `localStorage`. 3. Capture that browser state and reuse it. The critical word is **real**. You are not bypassing authentication or stubbing out the authorization check; the server still issued the session, the app still sends it, and the backend still enforces permissions on every request. That is what separates this from the anti-pattern of adding a `TEST_MODE` flag to production code that skips auth — which changes what you are testing and can ship to production by accident. ## Saving and restoring the session Most runners have a name for the saved artifact. Playwright serialises a browser context's cookies and per-origin `localStorage` into a *storage state* JSON file, and a new context created with `storageState` starts with all of it in place. Cypress's `cy.session()` caches the cookies, `localStorage` and `sessionStorage` produced by a setup function and restores them for later tests keyed by an id. Different APIs, same shape: run the expensive step once, snapshot the browser's auth state, replay the snapshot. Two details bite people: - **Cookies are not always enough.** If the app keeps an access token in `localStorage`, a cookie-only restore leaves the app looking logged out. Capture the whole origin state, not just cookies. Note also that Playwright's `storageState` covers cookies and `localStorage` but not `sessionStorage`, so an app that keeps auth there needs it seeded explicitly. - **The state must match the origin.** A cookie captured for one host will not be sent to another, so a session recorded against a local dev server is not reusable against staging. ## What you still owe Injecting a session removes the login flow from every test — including from the tests that should have covered it. The deliberate answer is to keep one (or a small handful of) test(s) that drives the real form: correct credentials land on the dashboard, wrong credentials show an error, the post-login redirect returns the user to where they were headed. That test is the only place login is exercised, which is fine: you need to know it works, not to re-prove it two hundred times. A sketch of the setup step: ```ts const context = await browser.newContext(); const res = await context.request.post('https://app.example.com/api/login', { data: { email: user, password: pass }, }); if (!res.ok()) throw new Error(`login failed: ${res.status()}`); await context.storageState({ path: 'state/admin.json' }); ``` Note the explicit failure check. A setup step that silently fails produces a suite where every test fails at the first assertion with a confusing message; failing loudly in setup tells you in one line that authentication broke. ## When UI login is still correct Do not over-apply the rule. If the thing under test *is* the authenticated entry path — remember-me behaviour, session expiry warnings, the redirect after login, or a first-run onboarding wizard triggered by sign-in — the UI flow is the subject and must be driven. The rule is that login is setup *unless it is the assertion*.

  • The app keeps its access token in localStorage rather than in a cookie — does that change the approach?
    Not the approach, only what you capture. The session now lives in origin storage instead of the cookie jar, so the saved state must include the `localStorage` entry and be restored before the app's first render, otherwise the app boots as anonymous. Snapshot mechanisms that serialise whole-origin state handle both; a helper that only sets cookies will silently produce a logged-out app.
  • Where should the programmatic login actually run — once for the whole run, or before each test?
    Perform the real authentication once per run (per role), and restore the cheap saved state per test. Restoring is milliseconds and gives each test a fresh, isolated browser context; re-authenticating per test puts the network call back on the critical path and hammers the auth endpoint, which some environments rate-limit.
  • Is there a case where you should still log in through the UI?
    Yes — whenever login is the subject rather than the setup: valid and invalid credential handling, the redirect back to the originally requested page, remember-me, session-expiry prompts, or onboarding that only triggers on first sign-in. The rule is that the flow under assertion gets driven through the UI; everything else gets injected.

saying these in an interview costs you the question

  • Says every test must log in through the UI to be realistic
  • Adds a test-only flag to app code that skips authentication
  • Restores only cookies for an app that stores tokens in localStorage
  • Deletes all login coverage once sessions are injected
  • Believes injecting a session bypasses server-side authorization

context

open as a page

An end-to-end browser test clicks a button using the CSS selector `.card > div:nth-child(2) > button.btn-primary`. It passes today and fails after a refactor that changed only markup nesting and class names, with no change to what a user can do. What makes that selector fragile, and what should the test anchor on instead?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A structural CSS path couples the test to DOM nesting and styling class names, which change freely without changing behaviour. Anchor on what a user perceives — the element's role and visible name — or on an explicit test-id attribute the component deliberately exposes.

open as a page

In a browser end-to-end test, why is inserting a fixed sleep (such as Playwright's page.waitForTimeout(500) or Cypress's cy.wait(500)) before an assertion a bad way to handle timing, and what do you use instead?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A fixed sleep hard-codes a guess about duration: too short and the test fails on a slow machine, too long and every run pays the delay. Wait on a retrying condition instead, which returns the moment the expected state appears.

open as a page

In an end-to-end browser test suite, what belongs inside a page object and what should stay in the test — and why do experienced teams keep assertions out of page objects?

level: middleimportance: must knowfreq 65%

basics

~20 s

A page object owns a screen's locators and user actions and returns data or locators; the test keeps the assertions and the scenario. Assertions buried in a shared helper hide what each test actually checks and make one test's expectations everyone's.

open as a page

In an end-to-end browser test, why do many teams locate elements by accessible role and visible name — for example a button named "Save changes" — before trying any other strategy, and where does that strategy break down?

level: middleimportance: must knowfreq 64%

basics

~20 s

Role plus name is how a user identifies a control, so it changes only when the interface genuinely changes, and a test that cannot find the element usually reveals a real labelling gap. It breaks on unlabelled icon controls, non-semantic markup, duplicated names, and copy or locale churn.

open as a page

An end-to-end suite registers with the hard-coded address [email protected] and asserts on a fixture order numbered 1001. What breaks once that suite runs on several parallel workers, or simply runs twice in a row, and what changes if each test builds its data from a factory instead?

level: middleimportance: must knowfreq 65%

basics

~20 s

Fixed identifiers collide: a second worker or a second run is rejected because the address is already taken, or two tests mutate the same order and see each other's changes. A factory mints fresh unique values per test, so each test owns data nothing else touches.

open as a page

Your end-to-end framework auto-waits for an element to be attached, visible and actionable before every click, yet the suite still has timing flakes. What kinds of waiting does that built-in auto-waiting not cover?

level: middleimportance: must knowfreq 58%

basics

~20 s

Auto-waiting covers only the element you are about to touch, one action at a time. It knows nothing about whether the app is still fetching, whether a later re-render will overwrite what you just asserted, and an assertion of absence gives it nothing to wait for.

open as a page

An end-to-end suite creates real records through the application's API. A test crashes halfway through and its cleanup step never runs; over weeks the shared test environment fills with orphaned data and the suite starts failing. How would you design the data lifecycle so that any rerun is safe?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Do not rely on cleanup running. Make tests assert only on data they created, tag every record with a run identifier so a scheduled reaper can delete leftovers, register teardown at creation time so it deletes in reverse, and treat resetting the environment as the real backstop.

open as a page

Your browser end-to-end suite fails on about one run in ten, a different test each time, and every one of them passes when re-run. How do you triage that?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Treat it as one reliability problem, not ten separate bugs. Record per-test outcomes so you can rank flakes by how often they block a merge, reproduce from the failing run's trace, video and logs, classify the cause, quarantine the worst offenders, then fix them.

open as a page

In an end-to-end browser test suite, what do you gain and what do you lose by writing element lookups and steps inline in each test instead of extracting them into shared page objects or helpers?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Inline tests read top to bottom and are easy to debug, but repeat the same element knowledge in every test, so a UI change edits many files. Extraction removes that duplication and adds a layer of indirection to read through.

open as a page

An end-to-end test for an order-detail page needs the signed-in account to already have one completed order. Why do teams create that prerequisite data with an API call or a database setup step instead of driving the application's own UI to create it first, and when is UI-driven setup still the right choice?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Creating prerequisite data through the product's API or a database setup step is faster and far less brittle, because the test then depends on no UI it is not asserting about. Drive the UI for setup only when that creation flow is itself the behaviour under test.

open as a page

An end-to-end suite reuses one saved session for a single shared account across all tests, and the tests run in parallel across several workers. What failures does sharing that one identity cause, and how do you keep programmatic login while restoring isolation?

level: middleimportance: should knowfreq 56%

basics

~20 s

Parallel tests on one account fight over account-scoped state and can invalidate each other's session — a logout, password change or settings edit in one worker breaks another. Give each worker and each role its own account and its own saved session.

open as a page

Your team is adding `data-testid` attributes to components so end-to-end tests can find elements. What makes a test id an explicit contract rather than a workaround, and what rules keep them from spreading uncontrolled through the codebase?

level: middleimportance: should knowfreq 56%

basics

~20 s

A test id is a contract when the component deliberately owns it, names the thing rather than the test or its styling, and is changed as consciously as any public API. Keep them scarce: use them only where role and name cannot identify an element or an instance.

open as a page

An end-to-end suite locates elements by their visible text, such as a button reading "Add to cart". The product then ships in five locales and the copy team rewords labels regularly. What breaks, and how do you keep the tests readable without coupling every test to product copy?

level: middleimportance: should knowfreq 48%

basics

~20 s

Every locator matching visible text breaks when a label is reworded or the build is localised, because the string is product copy, not a stable identifier. Pin the suite to one locale, look labels up from the translation catalogue where possible, and let only copy-focused assertions depend on wording.

open as a page

In an end-to-end test that submits a form and expects a new row in a list, when should the test wait on the network response, and when should it just assert on the rendered result?

level: middleimportance: should knowfreq 52%

basics

~20 s

Default to asserting on the rendered result: it is what the user perceives and it implicitly covers the request. Wait on the response only when the effect has no visible consequence, or when the test needs the request or response payload itself.

open as a page

If every end-to-end test starts from an injected session, nothing exercises the real sign-in flow any more. How do you cover login itself — including a third-party OAuth redirect and a multi-factor prompt — without paying for it in every test?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Cover the real flow in a small number of dedicated tests and inject the session everywhere else. Make those tests deterministic: a dedicated identity-provider test account, and a multi-factor code derived in the test from the account's shared TOTP secret rather than read from a phone.

open as a page

An end-to-end suite loads its authenticated session from a state file committed to the repository. It passed for weeks, then every test started failing with a redirect to the login page. What went wrong, and how should that session be produced instead?

level: seniorimportance: should knowfreq 41%

basics

~20 s

The saved session expired. A stored session is a time-limited credential, not a fixture, so committing it guarantees it goes stale — and leaks a credential. Regenerate it by authenticating at the start of every run, write it to a gitignored path, and fail loudly when setup cannot authenticate.

open as a page

An end-to-end test fails with "element not visible" reported from line 12 of a shared checkout helper, and the report gives no indication of which user step actually broke. How does test abstraction produce failures like this, and how do you structure helpers so failures stay diagnosable?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Deep helper chains report failures from shared code instead of from the step that broke. Keep the layers shallow, name methods after user intent, avoid conditionals and swallowed errors inside helpers, and let the assertion that matters live in the test.

open as a page

An end-to-end suite has one page object per URL, but the application is built from reusable components — a cart widget, a data grid, a date picker — that appear on many of those pages, and the page objects have started duplicating each other. When would you scope helpers to a component instead of to a page, and what does that change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Scope the helper to the widget when the same widget appears on several pages: one owner, reused wherever it renders, constructed from the container element it lives in. Page-shaped objects duplicate that widget's knowledge on every page that hosts it.

open as a page

On an order list page where every row renders a "Cancel" button, an end-to-end test's locator matches several elements and ends up cancelling the wrong order. Why do page-wide selectors go wrong on repeated UI, and how do you target the intended instance reliably?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A selector describing a kind of control cannot identify which instance you mean when the page renders many of them, and index-based picking depends on ordering and data that shift. Find the container that identifies the entity first, then query for the control within it.

open as a page

An end-to-end suite runs eight workers in parallel against one shared staging environment. Giving each test its own freshly created records fixed most of the interference, but a handful of tests still fail only when the whole suite runs together. What kinds of state cannot be made unique per test, and what do you do about them?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Singleton state resists per-test uniqueness: application-wide feature flags and settings, limited inventory, rate limits, a shared mailbox, background jobs, caches and the clock. Partition what you can into a per-worker tenant or account, and run the few tests that must mutate genuinely global state serially.

open as a page

You inherit a 600-test end-to-end suite whose selectors are a mix of XPath, hashed CSS class names and ad-hoc test ids, and it breaks on most user-interface pull requests. How would you move it to one selector convention without freezing feature work?

level: principalimportance: should knowfreq 30%

basics

~20 s

Measure which selector styles actually cause breakage, write the convention down, and stop the bleeding first by gating new and touched tests. Then migrate the highest-churn specs deliberately, funded by the app-side work — labels and owned test ids — and track breakage per pull request, not migration percentage.

open as a page

A team proposes running the end-to-end suite against a nightly anonymized copy of the production database instead of against data the tests create for themselves. How would you evaluate that proposal, and what would you put in place either way?

level: principalimportance: should knowfreq 40%

basics

~20 s

Judge it on what it adds versus what it destabilises: real volume and shapes catch bugs synthetic data never will, but a refreshing snapshot makes assertions on specific records unrepeatable and carries personal data into a lower-trust environment. The usual answer is a hybrid.

open as a page

Your CI pipeline retries each failed browser end-to-end test up to twice and reports the run green if a retry passes. What does that policy buy you, what does it cost, and what would you put around it?

level: principalimportance: should knowfreq 44%

basics

~20 s

Retries buy a usable pipeline while flakes exist, at the cost of signal: they can hide a rising flake rate and mask real intermittent bugs. Keep them only with retry counts recorded as a metric, a flake budget that fails the build, and quarantine with owners.

open as a page

You own the end-to-end testing strategy for an application with several permission roles — admin, editor, read-only — running against a shared staging environment used by other teams. How would you design the set of test identities and the way tests obtain their sessions?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Treat test identities as owned infrastructure: one account per role, multiplied per worker where tests mutate account state, provisioned by a setup step, credentials in secrets, sessions minted per run. On shared staging, namespace the accounts and assume nothing about state you did not create.

open as a page

The screenplay (actor and task) style models an end-to-end test as an actor performing composable tasks rather than as calls on page objects. What does it change compared with page objects, and how would you decide whether adopting it is worth it for a growing suite?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Screenplay reuses tasks organised around user goals instead of classes organised around screens, and composes them from small interactions. It buys composability and readable narration at the cost of more indirection, more concepts and a steeper onboarding curve.

open as a page