Every lifted function is tested and the whole job runs green in one process - which defect classes stay invisible?
answer
- meaning yes, arrangement no
- nothing moves, nothing dies
- small and evenly shaped input
- the function never left the process
- write the uncovered list down
basics
~20 sFour: anything appearing only when records are redistributed between workers; anything appearing only when a worker is lost and its share recomputed; anything appearing only when one key holds most of the records or the input is large; and anything appearing only when the function crosses a process boundary.
solid answer
~50 sAn in-process run is the whole job executed inside the test's own process on one machine, with no network and no second worker - and not every runtime even offers one, while some that do take a different code path than the distributed one. Four classes are structurally out of its reach. Redistribution: a key that does not spread as assumed, a value type that will not travel, the memory cost on the fetching side. Worker loss and recovery: an external write repeated because a unit of work was recomputed, or a non-deterministic rule producing different values the second time. Volume and distribution: one dominant key, spill, per-worker memory limits. And the process boundary: captures, per-worker initialisation, concurrency inside a worker. Say the list out loud, then buy back what you can with a small multi-worker run.
go deeper
Recall that a test running the whole job on one machine still has one worker, a tiny input and nothing that ever crashes, so whole families of production problems have no way to appear in it.
Explain each class by mechanism: no exchange between workers, no unit of work ever recomputed, no uneven or large input, no process boundary for the function to cross.
Demonstrate that you keep the uncovered list written down and buy classes back deliberately - a two-worker run, an idempotence test on the write path, a shaped real run - rather than adding assertions where coverage cannot grow.
The call is how much verification each risk deserves. Some teams rightly accept the recovery class and cover it with idempotent writes and monitoring instead of test infrastructure; make that a stated decision, not an omission.
## Why a green one-process run proves less than it looks An **in-process run** means the whole job executed inside the test's own process on one machine, with no network and no second worker. Two cautions before the list. Not every runtime in this class offers one; where it exists, it may take a partly different code path from the distributed one, so *green* can mean green on a path production never takes. The lifted-function tests are sound, but they only ever proved that one record becomes the right other record. Four classes remain, and an interviewer at this level is listening for exactly these. ## 1. Everything that only exists once records move between workers A step that must compare records held by different machines forces a **redistribution**: every worker writes out the records it holds, addressed by key, and every worker fetches the ones addressed to it. In one process there is no such exchange. What hides there: - a grouping key that does not spread the way you assumed - or that is derived from an object whose equality or hash is not stable across processes, so records that should meet do not; - key or value types that cannot be transported, discovered only when something tries to send them; - the memory cost on the fetching side, where one worker may now hold far more than it read; - an ordering you relied on inside a group without noticing, which the exchange does not preserve. ## 2. Worker loss and recovery A single-process run never kills anything, so it never exercises the second execution of the same work. What hides: - **a repeated external effect**: a rule that writes to something outside the job runs again when its unit of work is recomputed, and the destination now has the row twice; - **non-determinism under recompute**: a rule using randomness, the current time or an unstable ordering returns different values the second time, so recomputed output disagrees with output already read downstream; - **a write that is not idempotent**, which is the same thing seen from the destination. Runtimes recover differently and it matters: some recompute the lost piece of the input from its sources, some re-run the failed unit of work, some restore a saved picture of the job - a consistent copy of everything a running job holds, written periodically - and resume from it. A candidate who names one of these as *the* recovery model has described one engine, not the class. ## 3. Volume and distribution Test input is small and, almost always, evenly shaped. What hides: - **one dominant key**: in production one key holds most of the records and its worker runs for an hour while four hundred others finish in seconds; - **spill** - writing part of a working set to local disk when it no longer fits in memory - which never happens at a thousand rows; - per-worker memory limits, and the per-record cost of a resource that was free at a thousand records and dominant at a billion; - elapsed time itself, and therefore everything that only degrades slowly. ## 4. The process boundary and the environment In one process the function never travels. On a cluster, captures must cross, resources must be built once per worker, several units of work may run concurrently inside one worker, and the installed dependency set is the cluster's rather than the test's. Storage semantics differ too - an operation that is atomic on a local filesystem may not be on object storage, the shared place every worker reads input from and writes output to. ## What to do instead of pretending 1. **Write the list down next to the suite.** The practice most candidates omit is the honest negative statement; a list that names these four classes is worth more than a hundred extra assertions on the same one-process run. 2. **Buy back the cheapest class first.** A run with two real worker processes, even on one machine, catches captures, per-worker initialisation and non-global caches for almost nothing. 3. **Prove idempotence directly.** Call the write path twice over the same input in a test and assert the destination is unchanged. That is provable without losing a worker. 4. **Run on real machines over a cut-down input** before anything ships. Shaping that input so it keeps the dominant key and the rare shapes is a separate subject (Representative Samples), and judging output too large to read is another (Checking the Answer). 5. **Diff two versions side by side** before a change to a pipeline people depend on (Altering a Live Pipeline). | class | why one process cannot see it | cheapest way to buy some of it back | |---|---|---| | redistribution | no exchange between workers exists | a two-worker run over an input with several keys | | worker loss and recovery | nothing is ever lost or recomputed | testing the write path twice; deliberate worker kill in a scheduled run | | volume and distribution | input is tiny and evenly shaped | a real run over a shaped cut of production input | | process boundary | the function never leaves the process | two separate worker processes on one machine | The short interview answer is that a one-process run proves meaning and cannot prove arrangement, failure or scale - and that the value of saying so is that it tells your team what is still uncovered.
- Which of the four classes can you buy back most cheaply, and how?The process boundary. A run with two real worker processes, even on a single machine, forces the function and its captures to travel, builds resources per worker, and makes a per-worker cache visibly not global. It costs minutes to set up and removes a family of failures that otherwise appear only in production.
- How do you test that a repeated write is safe without losing a worker?Call the write path twice over the same input in an ordinary test and assert the destination is unchanged. Idempotence is a property of the write, not of the cluster, so it is provable outside one. That does not prove the rule is deterministic under recompute, which is a separate check on the rule itself.
- Why is a green one-process run a weak answer on its own in an interview?Because it answers a question nobody asked: whether the meaning of one record is right. The interviewer is probing whether you know it cannot touch redistribution, worker loss, scale or the process boundary. Naming those four and saying what you do about each is the answer; claiming coverage you do not have is the failure mode being tested for.
A read-through where one actor speaks every part proves the lines are right. It proves nothing about the cues between people, the actor who does not turn up, or the scene where three hundred extras arrive at once.
saying these in an interview costs you the question
- Says a green one-process run means the job is ready for the cluster.
- Believes a one-process run exercises redistribution because grouping still happens.
- Claims more assertions on the same one-process run close the gap.
- Assumes every runtime recovers a lost worker in the same way.
- Treats a repeated external write after a recompute as impossible.
- Expects a dominant key to show up in a thousand-row test input.