Why do three-row fixtures hide data-access defects that only appear once a table holds realistic row volume?
answer
- three rows multiply by three
- no threshold is ever crossed
- a scan wins on a tiny table
- assert scaling, not an absolute cost
- same statement count at ten times the rows
basics
~20 sVolume is the multiplier. At three rows a per-row statement loop costs three statements and milliseconds, no page boundary or batch limit is crossed, and a full scan is the fastest access path anyway, so the defect is real but produces no visible symptom.
solid answer
~40 sMost read-path defects are **per-row costs that only matter when rows are many**. A loop that issues one statement per row is three statements in the test and three thousand in production; a fetch plan that multiplies rows through nested links produces a handful of rows at fixture size; a missing index is irrelevant when a scan of three rows is the cheapest plan; page-size, batch-size and result-materialisation thresholds are never crossed. Small fixtures are still right for correctness — they are fast and readable. The fix is not to grow every fixture but to add a few **volume-shaped** tests where the assertion is about how cost scales: run the same operation at two row counts and require that the number of statements does not grow with the rows.
go deeper
Understand the multiplier: a defect that costs one extra statement per row costs three statements at fixture size. Small fixtures check that the data is right, not that fetching it is affordable.
Explain which defect classes need volume — per-row loops, multiplying fetch plans, page and batch thresholds, plan choice — and describe the two-row-count comparison that turns scaling into an assertion.
Show judgment about where volume tests go and how you keep them cheap: seed rows in bulk, cover the few read paths that matter, and state plainly which effects you leave to production observation.
Own the budget question. Volume in the suite buys one class of signal at a direct cost in feedback time; decide how much you buy, where, and what production signal covers the rest instead of letting every escape add another slow test.
A fixture of two or three rows per table is the default in most suites, for good reasons: it is fast, it is readable, and it makes a failure easy to diagnose. It also makes an entire family of data-access defects produce no symptom at all, because that family's symptom **is** volume. ## Why the defect is present but silent - **A per-row statement loop scales with rows.** Reading a list and then touching a link on each item emits one statement per item. At three rows that is three statements and no measurable time; at three thousand it is the slowest thing in the request. The code path is identical in both runs. - **A multiplying fetch plan needs rows to multiply.** Joining several to-many links in one query produces a row for each combination. Three parents with two children each is six rows; realistic parents with realistic children is where the result set and the de-duplication cost explode. - **A missing index costs nothing on a tiny table.** For a handful of rows a full scan is genuinely the cheapest plan, and a planner will choose it even when the index exists. The test cannot distinguish a well-indexed query from an unindexed one. - **Thresholds are never crossed.** Page sizes, batch sizes, chunked writes, streaming versus materialising the whole result, prepared-statement caches — each has a limit, and small fixtures stay far below all of them. - **Ordering luck.** With three rows an unordered query often returns them in insert order, so a missing sort clause passes; with many rows and a different access path the order changes and the assertion — or worse, the pagination — breaks. - **Collision windows stay shut.** Unique-key clashes, key-block exhaustion and retry paths need enough rows and enough concurrent work to be hit at all. ## What each fixture size can tell you | Fixture | Good at | Blind to | |---|---|---| | A few rows | Mapping correctness, constraints, the shape of the result | Cost that scales with rows; thresholds; plan choice | | Tens of rows | Crossing one page or batch boundary; ordering stability | Plan choice, which usually needs far more rows | | Thousands of rows | Statement counts that grow, materialisation cost, index usefulness | Real data distribution and concurrency | ## What to assert instead of "it returned the right objects" The useful move is to stop asserting a cost *level* and start asserting how cost *scales*: 1. **Run the operation at two different row counts** — say a handful and ten times that — and require the number of statements it emits to be the same. A count that tracks the row count is a per-row loop, whatever the absolute number is. 2. **Keep the volume-shaped tests few and deliberate.** One per read path that matters, not one per test. They are slower to set up and slower to run, and a suite full of them stops being run. 3. **Generate the rows rather than writing them out.** A loop that inserts many rows with plain statements is cheaper to write and much cheaper to run than a hand-authored fixture, and it keeps the intent — "many" — visible in the test. 4. **Insert the volume the cheap way.** Setting up thousands of rows through the mapper's own write path is slow and tests something else; a bulk insert of the seed data keeps the test about the read path. 5. **Assert result-shape invariants that volume breaks**, such as a page containing exactly the page size, a total count matching the seeded number, and a stable order under an explicit sort. ## Keeping the suite honest without making it slow Volume tests are a small, high-value minority. The bulk of the suite should stay at a few rows and stay fast, because most of what it verifies is mapping and behaviour, not cost. The mistake in both directions is real: a suite with no volume-shaped test cannot see scaling defects at all, and a suite that grew every fixture to thousands of rows becomes slow enough that people stop running it and start marking it as skipped. There is also an honest limit. Even a large synthetic fixture is uniformly distributed, freshly written and unfragmented — it will not reproduce a plan that degrades because one value covers most of the table, or one that changes as a table grows over months. Those belong to production observation, not to the suite. What the suite can own is the scaling *shape* of the code path: whether the work it does is proportional to the rows it returns or to the rows it touches.
- Why assert that a statement count is unchanged across two row counts rather than below a fixed number?An absolute threshold has to be re-tuned whenever the query legitimately changes, and it passes at fixture size for a loop that would explode in production. Comparing two row counts targets the actual property you care about: the work must be proportional to the query, not to the rows.
- Why not just grow every fixture to a realistic size?Cost. Setup dominates a large suite's runtime, and a slow suite gets run less, skipped, or run only in the pipeline. Keep most tests at a few rows for mapping and behaviour, and spend volume on the few read paths where scaling is the risk.
- Does a large fixture make an index test meaningful?Partly. Enough rows will make a scan visibly worse than a lookup, so a plan inspection becomes informative. What synthetic volume still misses is distribution — a column where one value dominates, or a range the real data clusters in — which is what usually flips a plan in production.
Testing a stairwell evacuation with three people in the building: everything works, and nothing about the queue at the door is learned.
saying these in an interview costs you the question
- Says the query is fine because the test finished quickly
- Believes a passing suite proves the read path scales
- Grows every fixture to thousands of rows to be safe
- Assumes an unordered query keeps returning rows in insert order
- Treats a missing index as detectable at any fixture size