A clean run from empty reproduced the result, but a colleague's run does not — what did the clean run not prove?
answer
- the restart discards the process only
- anything written down survives it
- caches, intermediates, paths, versions
- warm machine against a cold one
- state the scope with the claim
basics
~20 sA clean run from empty proves only that no value living inside the previous process was holding the result up. Anything the work left outside the process — written files, caches, inserted records, local paths, installed versions — survives the restart untouched.
solid answer
~50 sA **clean run from empty** means starting a new process with nothing bound and executing every block once in the order the file lists them. What it discards is the process, and therefore every binding in it. What it does not discard is everything the earlier runs wrote down: an intermediate file a later block reads, a cached extract on disk, records inserted into a store, an output location never cleaned out. Those are still there, so a file that silently depends on one of them reproduces perfectly on your machine and fails on anyone else's. Add to that what the file never states — a path only you have, a moving source read as 'today', a version installed months ago, an unrecorded seed — and the honest claim is narrow: the in-process state was not load-bearing. That is worth proving, and it is not the same as 'it reproduces'.
go deeper
Recall the shape of the claim: restarting throws away the values the process was holding, and nothing else. Files, caches and settings are all still there afterwards.
List the outside state a restart leaves alone — derived files, caches, inserted records, local paths, installed versions, a moving source — and explain why each one makes a file pass here and fail elsewhere.
Demonstrate the method: clear the derived artefacts, run in a fresh location, pin the input by identity, and treat a colleague's failing run as the experiment that found the dependency rather than as their problem.
Decide what a shared result must carry before anyone acts on it, and who owns producing that evidence — because the difference between 'it reproduces' and a bounded, checkable claim is a standing cost somebody pays.
## What the clean run actually proves An **interactive session** is a long-lived process you send blocks of code to one at a time, keeping every value those blocks ever bound. A **clean run from empty** starts a new process with nothing bound and executes every **block** once, in the order the file lists them. Killing the process destroys every binding in it, so a matching result establishes exactly one thing: **no value that lived only in the old process was holding the answer up.** That is genuinely worth establishing — it is the difference between a file and a souvenir. But it is a claim about the inside of one process, and people routinely report it as though it were a claim about the world. ## The state that survives a restart Restarting discards bindings. It discards nothing you wrote down. Everything in this list is still exactly where the last run left it: - **A derived file an earlier run wrote** that a later block reads. The block that produced it may have been deleted, commented out, or moved behind a condition that is now false — and the read still succeeds. - **A cache.** An extract saved to avoid a slow fetch, keyed on something that has not changed. On your machine it hits; on a fresh machine it misses and does the real thing, which may not give the same answer. - **Records a block inserted** into a store somewhere, which a later block then reads back. - **An output location never cleaned.** The file writes into a directory and the review reads from it, so stale outputs from a run three days ago are indistinguishable from today's. - **Paths, locations and settings** that exist only here — a mount, a working location, a credential, a locale that changes how text sorts or how a date is read. - **Installed versions.** The file names no versions, so a colleague gets whatever they have; behaviour that differs between versions differs silently. - **A moving source.** A fetch with no fixed cut-off, or a query that means 'as of now', gives a different answer on Tuesday. The file is identical and the input is not. - **An unrecorded seed.** Anything random reproduces across your two runs only by luck, and across two machines not at all. ## Two runs, two machines | | Your clean run | Their first run | |---|---|---| | In-process bindings | discarded | never existed | | Files earlier runs wrote | present | absent | | Caches | warm | cold | | Paths and installed versions | yours | theirs | | Moving source | as of your run | as of their run | The clean run controls exactly one row of that table. Every other row is a live difference, and the colleague's failing run is the experiment that found one of them. ## Making the claim stronger 1. **Clear the derived artefacts first,** then run from empty. Remove the intermediate files, empty the output location, invalidate the cache. Now the run is exercising the file rather than the machine's history. 2. **Run in a fresh location** rather than the directory you have been working in all week, so a path you never noticed you depended on fails loudly instead of quietly resolving. 3. **Pin the input by identity, not by recency.** A fixed cut-off, a stated snapshot, a recorded content hash — anything that makes 'the same input' checkable rather than assumed. 4. **Record the seed** beside the result for anything random, as part of the same evidence. 5. **Have someone else run it.** This is the only step that tests the rows the first four cannot, and it is why a colleague's failure is a finding rather than an annoyance. ## One more thing the restart does not settle Session designs differ. The common design executes whatever you send it and holds no relationship between blocks. Dataflow-style designs track which blocks read which bindings and re-execute the stale ones, so the visible state stays consistent with the current code. That is a real guarantee and it is a different guarantee: it says the shown values match the code as written, not that a run from nothing would produce them. On such a design it is easy to conclude that reproducibility is handled; the external state in the list above is untouched by it either way. ## What to say instead of 'it reproduces' Replace the bare claim with the claim plus its scope. *'Runs clean from empty, with the intermediate directory emptied first, against the snapshot dated the ninth, seed recorded, on the versions listed at the top of the file.'* That sentence is longer, is checkable, and names the things a colleague's machine will differ on. The bare version is not a weaker claim — it is an unbounded one, and reproducing is in any case not the same as being right: a wrong transform reproduces its wrong answer perfectly.
- How would you find which piece of outside state the file is depending on?Bisect the environment rather than the code. Run from empty in a fresh location with the derived files removed and see whether it still passes; if it does, restore one class of artefact at a time — the intermediate directory, then the cache, then the local paths. The first restoration that changes the outcome names the dependency. A colleague's clean machine does the same test in one step.
- Does a clean run from empty tell you the result is correct?No, and conflating the two is common. Reproducibility says the same inputs and the same file give the same output; correctness says that output is the right answer. A transform with a wrong condition reproduces its wrong answer perfectly on every machine. The checks that speak to correctness are separate work written beside the steps themselves.
saying these in an interview costs you the question
- It ran clean from empty, so it reproduces anywhere.
- Restarting the process also clears what earlier runs wrote to disk.
- A cached intermediate is fine because it came from this same file.
- Reproducing twice on one machine means the result is correct.
- If a path were wrong the run would have failed.