skip to content

You inherit a 600-test end-to-end suite whose selectors are a mix of XPath, hashed CSS class names and ad-hoc test ids, and it breaks on most user-interface pull requests. How would you move it to one selector convention without freezing feature work?

level: principalimportance: should knowfreq 30%

answer

  1. breakage is deterministic, not flake
  2. evidence before plan: classify CI failures
  3. stop new violations before cleaning old
  4. most of the work is app-side labels and ids
  5. measure pull requests unblocked, not specs migrated

basics

~20 s

Measure which selector styles actually cause breakage, write the convention down, and stop the bleeding first by gating new and touched tests. Then migrate the highest-churn specs deliberately, funded by the app-side work — labels and owned test ids — and track breakage per pull request, not migration percentage.

solid answer

~50 s

A big-bang rewrite of 600 tests is unfundable and would land on a moving target, so sequence it. First get evidence: attribute recent failures to selector style, and find which specs and which screens produce most of the breakage — it is usually a small set. Second, write the ladder down as policy: role and accessible name first, component-owned test id where identity is not user-perceivable, structural selectors only in markup you do not control. Third, stop the growth — new tests and any spec touched by a pull request must comply, enforced by review and, where possible, a lint rule that bans XPath and hashed class selectors. Fourth, migrate the worst offenders in scheduled slices, and remember most of the cost is app-side: labelling controls and adding entity-scoped test ids. Finally, track the metric that matters — the share of UI pull requests that break tests for no behavioural reason — because migration percentage can hit 100% while the pain stays.

go deeper

for a junior

Recall that a suite breaking on markup changes is a selector problem, and that the fix is a single agreed convention — semantic locators first, owned test ids second — rather than patching paths.

for a middle

Explain why these failures are breakage rather than flake, and describe the boy-scout rule plus a lint gate that stops new violations while the backlog is worked down.

for a senior

Show sequencing and evidence: classify CI failures by cause, find the concentrated offenders, migrate screen by screen with the app-side labelling included, and ship in slices rather than one rewrite.

for a principal

Own the outcome and the organisation — fund the app-side work in the product teams' currency, assign durable ownership of the convention, measure pull requests unblocked rather than specs migrated, and say explicitly where the migration stops.

## Frame the problem correctly first This is not flakiness. A flaky test passes and fails on the same code; this suite fails **deterministically** because the app changed shape. That distinction matters because it rules out the tempting non-fix: adding retries. Retrying a selector that no longer matches burns CI time and delays the same red. Say that explicitly — an interviewer is listening for whether you can tell breakage and flake apart. It is also not primarily a test problem. Most of the fix lives in the application: controls that are real, labelled controls, and components that render entity-scoped test ids. A migration plan that assigns all the work to whoever owns the test suite will stall. ## Step 1: get evidence before you get opinions Run a short diagnostic before proposing anything: - Pull failures from the last few weeks of CI and classify them by cause: selector no longer matches, timing, data, genuine regression. - Attribute the selector failures by style (XPath, generated class, text, test id) and by spec and screen. - Count how many pull requests were blocked, and how much developer time went into re-running or patching. The usual finding is a heavy concentration: a minority of specs, hitting a minority of screens, cause most of the breakage — commonly the screens under active redesign. That concentration is the plan. It also gives you the business case, expressed in blocked pull requests rather than in taste. ## Step 2: write the convention down One page, unambiguous, with the ladder and the escape hatch: 1. Accessible role and name for anything a user identifies by what it is and what it says. 2. A component-owned `data-testid` where identity is not user-perceivable, or where it identifies an instance (`order-row-1042`). 3. Structural CSS only inside third-party markup, with a comment saying why. Include naming rules for test ids, the container-scoping pattern, and the locale decision for text matching. Ambiguity in the document becomes inconsistency in the suite. ## Step 3: stop the bleeding before cleaning up Migration fails when the old style keeps arriving. Two gates: - **Mechanical where possible.** A lint rule banning XPath helpers and selectors that look like generated class hashes catches most new violations cheaply and impersonally, which matters more than catching all of them. - **Boy-scout rule.** Any spec a pull request touches must be brought to the convention. This spreads cost across the teams causing the churn, and it naturally prioritises the files that change most. Apply the gates to new code only, so the existing 600 do not block anyone on day one. ## Step 4: migrate in funded slices, app first Take the top offenders from step 1 and work screen by screen, not spec by spec: fix the screen's markup — real buttons, real labels, entity-scoped test ids on list rows — then re-anchor every spec that touches it. Screen-level batches keep the app work coherent and let one component change fix a dozen tests. Ship each slice separately so value lands continuously, and pair with the team that owns the screen so the labels land with product context rather than as a drive-by patch. Expect to delete as well as migrate. A suite that size usually contains duplicates and tests better served at a lower level; migration is a reasonable moment to ask whether a given test earns its runtime at all. Be careful this does not become an excuse to drop coverage — decide by what the test protects, not by how annoying it is to fix. ## Step 5: measure the outcome, not the activity "78% of specs migrated" is an activity metric and can reach 100% while pull requests still break. Track instead: - share of pull requests whose only failures are non-behavioural test breakages, - mean time from red to green on those, - number of new selector-caused failures per week. When those numbers stop moving, the remaining specs are not worth migrating, and you say so and stop. Knowing where to stop is part of the answer. ## The organisational half The convention survives only if ownership is clear: component owners own the test ids and labels their components render, the platform or quality group owns the document and the lint rule, and removing a test id is reviewed like removing public API. Left ownerless, the suite drifts back to XPath within two quarters — usually via a recorder, because recorders emit exactly the brittle style you just removed. If the team uses one, constrain how its output is allowed to be committed.

  • Why not simply enable retries until the migration is done?
    Because these failures are deterministic. The selector does not match, so every retry fails identically while multiplying CI time and delaying feedback. Retries address genuine nondeterminism; using them here would only hide the signal that tells you which screens still need migrating.
  • How do you get product teams to fund the app-side work?
    Show it in their currency: blocked pull requests and hours lost per week, concentrated on the screens they own. Then make the ask concrete and small — label these controls, add row test ids here — and note that labelling controls is work most teams already owe for accessibility, so it lands against an existing obligation.
  • When do you stop migrating rather than finish?
    When the outcome metric flattens. Once selector-caused breakage per pull request has dropped into the noise, the remaining specs are on stable screens that rarely change, and migrating them buys nothing. Declare the ratchet — new and touched specs comply — and leave the tail alone.
  • How do you keep the convention from decaying after you move on?
    Assign ownership explicitly, keep the lint rule in the default pipeline so compliance costs nothing to enforce, treat removing a test id as a breaking change in review, and keep one dashboard for selector-caused failures. Also constrain recorder output, since generated selectors are the usual route back to XPath.

saying these in an interview costs you the question

  • Proposes a big-bang rewrite of all 600 specs
  • Adds retries to a deterministic selector failure
  • Reports migration percentage as the success metric
  • Assigns all the work to the test team, none to component owners
  • Skips measurement and migrates whichever specs are most annoying

context