An end-to-end suite creates real records through the application's API. A test crashes halfway through and its cleanup step never runs; over weeks the shared test environment fills with orphaned data and the suite starts failing. How would you design the data lifecycle so that any rerun is safe?
answer
- teardown is best-effort, not guaranteed
- assert only on what you created
- every row should say who made it
- something outside the test must sweep
- the strongest cleanup is a fresh environment
basics
~20 sDo not rely on cleanup running. Make tests assert only on data they created, tag every record with a run identifier so a scheduled reaper can delete leftovers, register teardown at creation time so it deletes in reverse, and treat resetting the environment as the real backstop.
solid answer
~50 sCleanup that runs at the end of a test is best-effort by definition: a crashed process, a cancelled CI job or a machine that vanishes skips it, so leftovers are inevitable and the design has to survive them. Three layers do that. First, make the suite indifferent to leftovers — never assert on global counts or "the newest record", always on the identifiers this test created, scoped to them. Second, make every created record self-identifying, with a run id in a name, email prefix or tag, so a scheduled reaper can delete anything older than a day and an engineer can tell test traffic from real data. Third, register the delete at the moment of creation and run it in an always-runs hook, deleting in reverse dependency order, with failures logged rather than failing the test. If the environment can be recreated per run — an ephemeral database, a fresh tenant, a restored snapshot — that beats all three, because there is nothing to clean.
code
javascript · 14 linesconst created = [];
async function seedOrder(api, data) {
const order = await api.createOrder(data);
created.push(() => api.deleteOrder(order.id));
return order;
}
async function cleanup() {
const actions = created.splice(0).reverse();
const results = await Promise.allSettled(actions.map((run) => run()));
const failed = results.filter((r) => r.status === 'rejected').length;
if (failed > 0) console.warn(`cleanup: ${failed} of ${results.length} deletions failed`);
}go deeper
Know that data an end-to-end test creates does not disappear on its own, and that a test should clean up what it made and assert only on its own records rather than on how many rows exist.
Explain why a teardown hook is best-effort — crashes and cancelled jobs skip it — and describe registering cleanup at creation time, deleting in reverse dependency order, and logging rather than throwing when a delete fails.
Demonstrate the layered design: assertions scoped to owned data, a run identifier on every record so an out-of-process reaper can sweep, and idempotent seeding so a rerun against a dirty environment behaves identically.
Own the tradeoff between a long-lived shared environment with reaping and per-run ephemeral environments — provisioning cost and complexity against interference and debugging time — and make leftover volume a monitored signal rather than a periodic surprise.
## Why cleanup alone is never the design A teardown hook is code that runs *if* the process survives. It does not run when the worker segfaults, when CI cancels the job, when the runner is preempted, when someone hits stop, or when a network partition eats the delete request. Orphan records are therefore not a bug to be fixed once; they are a permanent property of any suite that writes to a shared environment. A suite is well designed when leftovers are boring, not when they are impossible. ## Layer 1: make the suite indifferent to leftovers This is the cheapest and most valuable layer, and it is mostly about assertions. - **Never assert on a global count.** "The orders list shows three rows" is a statement about the entire environment, and every other worker and every past run can change it. - **Never assert on "the first" or "the newest" item** unless the test created it and can prove it. - **Scope queries to your own data.** Find the record by the unique value you generated, then make every assertion inside that row or that detail page. A suite obeying these rules keeps passing in an environment full of junk. That is the difference between leftovers being an annoyance and leftovers being an outage. ## Layer 2: make every record self-identifying and reapable Give everything the tests create a marker: a run id in the email prefix, a name prefix, a tag or a custom field. ```javascript const RUN_ID = process.env.CI_RUN_ID ?? `local-${Date.now()}`; const email = (name) => `e2e-${RUN_ID}-${name}@example.test`; ``` That single convention buys three things: a scheduled job can delete every record matching the marker that is older than some window; an engineer looking at the environment can immediately tell test data from real data; and a support question about a strange account has an instant answer. The reaper is ordinary infrastructure — a cron job hitting the same delete API — and it is the only cleanup that is guaranteed to run, because it does not live inside the test process. ## Layer 3: register teardown when you create, not when you write the test Cleanup written as a separate block at the bottom of a test drifts from what the test actually created. Register it at creation time instead, and run the registered actions in reverse: ```javascript const created = []; async function seedOrder(api, data) { const order = await api.createOrder(data); created.push(() => api.deleteOrder(order.id)); return order; } ``` Reverse order matters because of dependencies: delete the order before the customer, the membership before the organisation. And cleanup failures should be **logged, not thrown** — a delete that returns 404 because the test already deleted the record must not turn a passing test red. The reaper will catch whatever slips through. ## Idempotent reruns "Idempotent" here means: running the suite against an environment that already contains the previous run's output produces the same result. Two things achieve it. - **Unique-per-run identifiers**, so a second run never collides with the first — the same discipline that makes parallel runs safe makes serial reruns safe. - **Upsert-style seeding for anything that must be fixed**, such as a reference record the suite expects to exist. Create-if-missing rather than create-and-fail. A common wrong answer is "wrap it in a transaction and roll back". That works for a server-side integration test running in the same process as the database. It does not work for a browser-driven test: the application commits its own transactions in its own process, and the test cannot hold them open. Ephemeral state has to come from the environment, not from a transaction boundary the test does not control. ## The strongest option: no shared world at all If your platform can give each run its own database, its own tenant, or an environment restored from a snapshot, take it. Cleanup becomes deletion of the whole environment, which cannot half-fail, and cross-run interference stops existing. The cost is provisioning time and infrastructure complexity, which is why the layered approach above is what most teams actually run — and why the layers still matter even where ephemeral environments exist, since a long-lived staging environment usually survives somewhere. ## Monitoring the mess Measure it. A count of test-tagged records older than the reaper window is a one-line dashboard that tells you whether the strategy is holding. When it climbs, something changed — a new entity nobody registered for cleanup, or a reaper that lost permission — and you learn it before the suite starts failing.
- Why not just wrap each test in a database transaction and roll it back?Because the application commits its own transactions in its own process. A browser-driven test drives the app from outside and has no transaction to roll back. That technique belongs to server-side tests running in-process with the database; end-to-end tests need environment-level or record-level cleanup instead.
- Should a failing cleanup step fail the test?No. Log it and move on. A delete returning 404 because the record was already removed would otherwise turn a passing test red and train people to ignore failures. Cleanup health belongs on a dashboard — leftover count over time — not in the pass/fail signal of a feature test.
- How do you order deletions when records depend on each other?Register each cleanup at the moment of creation and run them in reverse, so children go before parents: order before customer, membership before organisation. That mirrors creation order automatically and avoids a hand-maintained teardown block that drifts from what the test really made.
- What single metric tells you whether the strategy is working?The number of test-tagged records older than the reaper's window. If it stays flat, cleanup and reaping are keeping up. If it climbs, something new is being created that nobody registered, or the reaper has lost access — and you find out before the suite starts failing.
saying these in an interview costs you the question
- Assumes an after-each hook always runs
- Proposes rolling back a transaction around a browser-driven test
- Fails the test when a delete call errors
- Asserts on total record counts and blames flakiness
- Solves leftovers by wiping the shared database by hand