skip to content

A 27-minute test suite has 213 end-to-end tests and 46 narrow tests: how do you rebalance the levels?

level: seniorimportance: should knowfreq 46%

answer

  1. Measure before you reshape
  2. Freeze the top level first
  3. Ask what each case uniquely proves
  4. Push the rule down, then delete
  5. Judge by escapes and run time

basics

~20 s

Rebalance by risk, not by ratio. Freeze new top-level tests, ask of each existing one what it uniquely proves, push duplicated checks down to the level where the rule lives, and delete a top-level case only once something cheaper covers it.

solid answer

~50 s

Start by measuring rather than reshaping: per level, the test count, the share of the 27 minutes, and how many genuine defects each level has actually caught. Then triage the top layer into three buckets — cases that duplicate coverage available lower down, cases whose risk is a single rule that belongs in a narrow or boundary test, and the few genuine journeys that prove wiring, configuration and deployment. **Freeze the top first** so the shape stops worsening, then work the queue driven by real failures and escaped defects: for each, write the cheap test that would have caught it, and retire the expensive case only after its replacement is green. Expect the base to grow by several hundred tests while the tip falls to a couple of dozen journeys, and track the result by run time and by escaped-defect count, not by a target ratio.

code

pseudocode · 10 lines
pseudocode
function test_duplicate_reading_is_stored_once():
    store  = InMemoryReadingStore()
    ingest = IngestService(store)
    reading = Reading(unitId = "TLM-4417", sequence = 91823, odometerKm = 12.7)

    ingest.accept(reading)
    ingest.accept(reading)          # the transport retried the same delivery

    assert store.countFor("TLM-4417") == 1
    assert store.odometerFor("TLM-4417") == 12.7

go deeper

for a junior

Know that a suite dominated by assembled-system tests is slow and hard to diagnose, and that the fix is to move checks down to the level where the rule lives rather than to delete tests outright.

for a middle

Be able to triage cases into duplicates, misplaced rules and genuine journeys, and to describe the push-down-then-delete order so coverage is never dropped before its replacement is green.

for a senior

Demonstrate the measurement discipline: per-level run time and defect yield before and after, an escape used as the diagnostic, and a design finding accepted when a rule proves untestable narrowly.

for a principal

Own the sequencing and the cost story: what a freeze buys immediately, how much of an engineering quarter a conversion warrants, and what evidence you would show a stakeholder to justify the spend.

## Framing A fleet telematics ingest service accepts position and odometer readings from vehicle units, deduplicates them and writes them to a store. Its suite is 213 assembled-system tests and 46 narrow tests, running 27 minutes. The trigger for the conversation is usually not the run time but an escape: a retried delivery caused **a duplicated side effect** — one reading counted twice, so a vehicle's daily distance was overstated — and none of the 213 tests noticed. That escape is the diagnostic. It says the problem is not the number of tests but *where the checking sits*. The idempotency rule lives in one component, and no test addressed it directly; the 213 journeys all delivered each reading once, because that is what a journey does. ## Step 1 — measure before reshaping Collect three numbers per level: test count, share of wall-clock time, and defect yield — how many genuine defects that level has caught in, say, the last six months. Suites like this usually show something stark: the tip consumes 24 of the 27 minutes, and the majority of its red runs were nondeterministic rather than real. Without these numbers, any reshaping is doctrine; with them, it is an argument anyone can check. ## Step 2 — freeze the top Agree one rule before touching anything: new checking enters at the lowest level that can catch the risk. This costs nothing, needs no budget, and stops the drift on day one. Every later step is optional; this one is not. ## Step 3 — triage the 213 Ask of each top-level case: *what does this uniquely prove?* Three buckets fall out. - **Duplicates.** Cases that walk the same happy path with different data. Twenty variations of accepting a valid reading prove one thing about wiring and nineteen things about field validation that a narrow test proves in milliseconds. Typically the largest bucket. - **Misplaced rules.** Cases whose real subject is a single rule — a rejection, a rounding, a retry, an ordering guarantee. These belong one or two levels down. The duplicate-delivery rule is the archetype: it needs a narrow test asserting that accepting the same reading twice leaves one stored record, plus one boundary test proving the store's uniqueness constraint actually holds. - **Genuine journeys.** Cases that prove the assembled system is wired, configured and deployable — a unit connects, a reading traverses the pipeline, the resulting figure appears where a client reads it. Few in number, and worth every second. ## Step 4 — push down, then delete Work the buckets in the order real failures and escaped defects hand them to you rather than top to bottom. For each item: write the cheap test, watch it fail against the old behaviour if you can arrange that, make it green, *then* retire the expensive case. Deleting before the replacement exists trades coverage for a nicer picture. Where a rule turns out to be untestable narrowly, that is a design finding — the rule is tangled with wiring and should be extracted, which is a benefit of the exercise rather than an obstacle to it. A realistic landing point for this suite: 213 journeys down to roughly 24, the narrow layer up from 46 to a few hundred, a boundary layer of perhaps 60 tests against a real store, and a run measured in single-digit minutes. ## Step 5 — verify with escapes, not ratios The success measure is not that the picture looks like a triangle. It is: did the suite get faster, did failures start naming a cause, and did defect escapes fall? Re-run the defect-yield analysis a quarter later. If a class of defect keeps escaping, the shape is still wrong somewhere specific — and that is far more useful information than a percentage. ## What to watch out for - **Padding the base to hit a ratio.** Hundreds of trivial tests over accessors improve the shape and nothing else. - **Re-running to green.** A tip that is red for nondeterministic reasons is a signal about scope; automatic retries hide the signal. - **Reshaping in one sprint.** A big-bang rewrite of a suite loses the accumulated knowledge encoded in cases nobody remembers writing. Incremental, failure-driven conversion keeps that knowledge. - **Confusing narrowness with isolation everywhere.** Some risks genuinely only exist at a boundary; pushing those down produces a fast test that proves a mock behaves like a mock.

  • Which top-level cases would you keep, and how do you justify keeping them?
    Keep the ones that prove something no lower level can see: that components are wired together, that configuration and credentials resolve in a deployed environment, and that a complete journey produces the figure a client reads. Justify each by naming the risk it uniquely covers; a case that cannot answer that question is a duplicate.
  • How do you know the rebalancing worked rather than just looking tidier?
    Compare the same three measures before and after: wall-clock run time, the share of red runs that were genuine, and escaped defects per period. A faster suite whose failures name a cause and whose escape rate fell is a real improvement; a triangle-shaped picture with the same escape rate is not.
  • What would make you refuse to push a check down a level?
    When the risk only exists at the boundary — serialization, a uniqueness constraint, transaction behaviour, a query the store actually executes. Pushing those down produces a fast test that asserts a stand-in behaves as it was configured to behave, which proves nothing about the real dependency.

saying these in an interview costs you the question

  • Deletes top-level tests before the cheaper replacement exists
  • Targets a fixed ratio instead of measuring escapes
  • Pads the base with trivial tests to reshape the picture
  • Proposes a big-bang rewrite of the whole suite
  • Ignores that a rule may be untestable because of design
  • Retries failing runs automatically to keep the tip green

context