skip to content

Why does reading a configuration setting fresh at each point of use, rather than once at startup, make a run hard to reproduce?

level: seniorimportance: nice to knowfreq 26%

answer

  1. No moment at which the inputs are fixed
  2. The value can differ between two reads
  3. Nothing complete exists to print
  4. Resolve, validate, freeze, hand it down

basics

~20 s

Nothing pins the value. Reads spread across a run see different answers if a layer changes underneath them, no moment exists at which the full input set could be recorded, and parallel cases can act on different values.

solid answer

~50 s

Reproducing a run means stating its inputs and re-supplying them. With one resolution point the inputs are a single frozen object: printable, diffable and constant for the run's duration. With point-of-use reads the inputs are a sequence of reads spread over time, so three things break at once. Nothing pins a value, so anything that alters a layer mid-run — a job step exporting a variable, a case writing to a shared holder — makes reads before and after disagree. No complete list of inputs exists, because the set consumed depends on which code paths ran, which means there is no moment at which the configuration could be echoed. And execution order leaks into results, since two parallel cases can observe different values for the same key. The repair is to resolve once, validate, freeze, and hand the object down as a parameter.

code

pseudocode · 8 lines
pseudocode
# point of use: two reads in one run can disagree, and nothing records that
function open_batch():
    size = read_process_env("BATCH_SIZE")     # read again on every call
    return new_batch(size)

# resolved once: fixed for the whole run, printable, and the same for every worker
function open_batch(config):
    return new_batch(config.batch_size)       # frozen at startup, echoed at startup

go deeper

for a junior

Be ready to say that settings should be read once at the start and then passed to whatever needs them, and that code scattered across a suite reaching out for its own values is what makes a run hard to explain.

for a middle

Explain the mechanics: without one resolution point there is no instant at which the complete input set exists, so nothing can be validated or printed, and a value may differ between two reads in the same run.

for a senior

Demonstrate the diagnosis. Show how parallel workers or a changed layer mid-run produce results that depend on execution order, and describe resolve-validate-freeze-hand-down as the repair, plus the rule that only the entry point touches a layer.

for a principal

Own the boundary between given inputs and produced facts, and the estate-wide rule that no code below the entry point reads a layer directly. Be ready to say how you would migrate a large suite that already reaches outward in dozens of places.

## What a point-of-use read is A point-of-use read is any code that fetches a setting at the moment it needs it, rather than receiving a value that was resolved once when the run began. A helper that asks the process environment for a batch size every time it opens a batch; a page-of-results routine that re-reads a file each call; a shared holder that any code may write to as well as read. The suite still "has configuration" in the sense that values exist. What it lacks is a **moment at which those values became fixed**, and that moment is the whole basis of reproducibility. Reproducing a run means being able to state its inputs and re-supply them. With one resolution point, the inputs are one object: printable, diffable, and unchanged for the run's duration. With point-of-use reads, the inputs are a *sequence of reads spread across time*, and the honest description of what the run used is a timeline rather than a value. ## Why that destroys reproducibility 1. **Nothing pins the value.** If anything alters a layer while the run is in flight — a job step that exports a variable, a case that writes to a shared holder, a file rewritten by a parallel process — reads before and after see different answers. The run is internally inconsistent and nothing records that it was. 2. **There is no complete list of inputs.** The set of settings a run consumed is whatever the code paths it happened to take asked for. A different selection of cases consumes a different set. So the run cannot be echoed, because at no instant does the full configuration exist in one place to echo. 3. **Order and concurrency leak into results.** Two cases executing in parallel, or the same two cases in a different order, can observe different values for the same key. The suite acquires a dependency on execution order that no case declares and no reader can see. Each of those alone makes a run unexplainable. Together they produce the familiar conversation where a run passed on one machine, failed on another, and nobody can identify a single difference — because the differences were never captured. ## Resolve, validate, freeze, hand it down The repair is four steps and one habit: - **Resolve** every layer once, at startup, into a single object. - **Validate** the whole object immediately: required keys present, declared shapes parsed. - **Freeze** it, so that later code physically cannot write to it and no case can change another case's world. - **Hand it down** as a parameter to whatever needs it, rather than letting code reach outward for a mechanism of its own. The habit is that **no code below the entry point reads a layer directly.** One place touches files and process variables; everything else touches the frozen object. That single rule is what makes the echo trustworthy, because the thing printed is the same thing every case read. ## The comparison | | Read at the point of use | Resolved once, frozen | |---|---|---| | Value for a key during a run | may differ per read | one, for the whole run | | Complete list of inputs | only after the fact, per code path | known before the first case | | Can be echoed | not meaningfully | fully, in one block | | Missing required value | surfaces mid-run, as a case failure | surfaces at startup, as a clear abort | | Two parallel cases | may observe different values | observe the same values | | Reproducing the run | replay a timeline | re-supply one object | ## Values that genuinely arrive during the run Not everything a run knows can be resolved at startup, and the distinction is worth stating plainly, because it is the honest objection to freezing. A generated run identifier, an address allocated when a stand-in is started, a record created by a case: these are **facts the run produces**, not configuration it was given. They belong to the run's own working state, and they are reproducible in a different sense — you reproduce them by re-running with the same configuration, not by re-supplying them. Keeping the two categories apart is what lets the configuration object be frozen at all. ## Signals you have this problem - A search of the suite finds direct reads of process variables in more than one place. - Nobody can produce a list of the settings the suite understands without reading all of it. - Changing a value requires knowing which code path will read it and when. - The same run behaves differently depending on how many workers it used, with no case that mentions concurrency. - Diagnosing a difference between two runs starts with comparing two machines rather than two blocks of output.

  • A shared mutable holder is filled once at startup and read everywhere. Is that good enough?
    Better than reading layers directly, but the mutability is still the defect: any code may write to it, so one case can change another case's world and the echo printed at startup stops being true. Freeze it after resolution, and prefer handing it down as a parameter, so the code that needs a value declares that it does rather than reaching outward for it.
  • Some values genuinely are not known until the run starts, such as an allocated address. Does freezing forbid those?
    No, because those are facts the run produces rather than configuration it was given. Keep them in the run's own working state, separate from the frozen inputs. You reproduce them by re-running with the same configuration, not by re-supplying them. Keeping the two categories apart is exactly what allows the input object to be frozen and echoed at all.

Photographing the departures board once and quoting the photograph all day is reproducible; glancing up afresh each time someone asks is not, even though both are honestly reading the same board.

saying these in an interview costs you the question

  • Says reading fresh keeps the run more up to date
  • Treats a mutable global holder as equivalent to freezing
  • Assumes nothing can change a layer during a run
  • Cannot list which settings the suite understands
  • Blames machine differences before comparing recorded inputs