skip to content

How do you decide which tests may run on an engine substitute and which must use the real engine?

level: principalimportance: should knowfreq 44%

answer

  1. Ask what each case is evidence for
  2. Placement belongs to a tier, not an author
  3. Write down the feedback budget
  4. Replay the same cases against both engines
  5. Never reshape production code to suit the substitute

basics

~20 s

Split by what each case is evidence for. Logic cases may use a substitute; cases asserting a query, index, constraint, migration or concurrency need the real engine. Hold the split with a stated feedback budget and a periodic replay against both.

solid answer

~50 s

Make the split a property of a **tier**, not a per-author choice. Classify each case by its subject: logic, mapping and orchestration cases merely need somewhere to put a row and may use a substitute; anything whose assertion is about engine behaviour — a non-trivial query, an index, a constraint, ordering, isolation, or the migration history — belongs on the engine. Set an explicit feedback budget for the pre-merge tier and let the engine cases run in a slower tier that still gates the deploy. Two mechanisms keep the split honest: a periodic **replay of the same cases against both engines**, the only check that the substitute is still a fair stand-in; and a standing rule that production code is never reshaped to fit the substitute's supported subset. Revisit on evidence — escaped defects traceable to the gap, and the measured cost of the engine tier.

go deeper

for a junior

Focus on the classification, not the policy: be able to say that a case about a calculation can use a substitute while a case about a query, an index or a constraint needs the engine production runs.

for a middle

Explain how tiers implement the split — what runs on every push, what runs before merge — and why the fast tier's assertion set decides its engine rather than the other way round.

for a senior

Show you would operate it: an explicit feedback budget, a rule for admitting a case to the engine tier, and a scheduled replay against both engines whose findings you act on.

for a principal

Own the tradeoff and its reversal conditions. Name the costs — two schemas, dialect-patched setup, false confidence, and production code reshaped to suit the test engine — and state the evidence that would make you retire the substitute entirely.

## The decision, stated properly The question is not "substitute or real engine" but "which claim does each case make, and which engine can support that claim". Framed that way the split is mostly mechanical. **Cases a substitute may serve.** Their subject is application behaviour and the store is incidental: branching, validation, calculation, mapping, orchestration between components, error handling. Swapping the engine underneath such a case does not change what it proves. **Cases that require the real engine.** Their assertion *is* engine behaviour: a query with non-trivial filtering, grouping or ordering; anything relying on an index; uniqueness, cascade, check and deferred-constraint behaviour; collation and case sensitivity; numeric and fractional-second precision; isolation, locking and deadlock handling; and the migration history. Add one pragmatic rule that catches the rest: any statement using an engine-specific feature is not permitted on the substitute, because the substitute's version of that feature is a reimplementation. ## Make it a tier, not a preference If the choice is left to each author, the codebase ends up with both engines used arbitrarily and nobody able to say what the suite proves. Encode it instead: - one **fast tier** on the substitute, running on every push, with a stated budget — for example, the tier must finish inside its 92nd-percentile target of seven minutes forty seconds, and a change that pushes it past that must buy the time back; - one **engine tier** on the real engine, running before merge or on a short schedule, deliberately smaller and deliberately not a mirror of the fast tier; - a placement rule in the contribution guide that an author can apply without asking, and a review prompt when a case in the fast tier touches a query. The budget matters more than it looks. Almost every argument for keeping a substitute is really an argument about feedback latency, and once the number is written down the argument becomes measurable rather than a matter of taste. ## The check that keeps the substitute honest A substitute is a claim: *behaviour relevant to these cases is the same on both engines*. Nothing in a green fast tier tests that claim. The check that does is a **replay** — run the same cases against both engines and compare. Any divergence is a finding: either the production engine behaves in a way the code did not anticipate, or the substitute has drifted and some case is no longer trustworthy on it. Run the replay on a schedule rather than per commit, because its value is trend, not gating. Record what it finds; a replay that has produced nothing for months is either evidence the substitute is fine or evidence the cases in it never touched the engine. Both are useful conclusions and they are distinguishable by looking at what the replayed cases actually assert. ## Costs to hold in view - **Two schemas.** If the substitute's schema comes from anywhere other than the production migration path, drift is guaranteed and silent. - **Reverse pressure on production code.** The failure mode to forbid outright: dropping an index, rewriting a set-based statement as a loop, or avoiding a feature so the substitute still passes. At that point the test engine is designing the system, and the cost is paid forever in production performance. - **Two dialects in test setup.** Setup fixtures written for one engine and patched for the other are an ongoing tax and a source of cases that pass for the wrong reason. - **False confidence.** The most expensive cost, and the hardest to see, because its symptom is an absence of failures. ## The worked case A subscription renewal job is the most engine-dependent flow in a system: it selects a due set at a boundary, updates rows under concurrency, and depends on an index to be viable at all. Placing its cases in the fast tier would be the classic error — every one of its assertions is about the engine. In one such split, the renewal boundary query was left on the substitute, and a comparison that behaved as a strict inequality on the production engine's truncated timestamps produced an off-by-one at the boundary: 118 subscriptions renewed on two consecutive runs. The engine tier had no case for the boundary because the boundary case already existed and was green. The lesson is about placement, not about the assertion. ## Revisiting the decision Treat it as reversible and set the evidence that would reverse it: escaped defects attributable to the gap over a period, the measured wall-clock cost of moving a tier onto the real engine, and how much of the fast tier genuinely touches the store at all. Often the third number is the surprise — a large share of substitute-backed cases turn out not to need a store, and moving them off it shrinks the problem instead of solving it. Be honest that the wider question is contested. One camp argues substitutes are legacy now that a disposable real engine is routine; another points to constrained build agents, licensing, start-up cost and offline work. The defensible position is not a slogan but the split above plus a replay that keeps checking whether the split still holds.

  • How do you stop the real-engine tier from slowly becoming a full mirror of the fast tier?
    Give it an explicit admission rule — a case belongs there only if its assertion is about engine behaviour — and review additions against it. Track its wall-clock time as a first-class number so growth is visible. Duplicated cases are also detectable: if removing a case from the engine tier changes no coverage of engine behaviour, it was a mirror and should go.
  • What do you do when a replay against both engines finds a divergence?
    Treat it as a defect report, not noise. Establish which engine is right for the intended behaviour, fix the code or the assertion, and then decide placement: a case that diverged has demonstrated it is evidence about the engine, so it moves permanently to the engine tier. Record the divergence family, because repeats in one family are the argument for retiring the substitute.
  • A team asks to move the whole suite onto the real engine. What evidence would persuade you?
    Measured escaped-defect counts attributable to the substitute gap, the wall-clock cost of the move against the stated feedback budget, and whether the build agents can actually run the engine reliably. If the escaped defects are real and the budget survives the move, take it — the substitute has no independent value beyond speed and portability.
  • How would you handle a case that must be fast but genuinely depends on engine semantics?
    Do not compromise the engine dependency for speed. Reduce the cost instead: share one engine instance across the tier, isolate by rollback rather than rebuild, shrink the data to what the assertion needs, and run the tier in parallel. If it still exceeds the budget, move it out of the pre-merge gate rather than moving it onto an engine that cannot support its claim.

saying these in an interview costs you the question

  • Leaves the choice to each test author's preference
  • Treats the substitute tier as proof queries work
  • Mirrors every case onto both engines regardless of subject
  • Changes production queries so the substitute keeps passing
  • Argues the split with taste instead of a stated budget
  • Never checks whether the substitute has drifted

context