skip to content

questions

5

A mortgage underwriting scorer is retrained from an unchanged code revision a month later and comes out different - why?

level: juniorimportance: must knowfreq 52%

answer

  1. instructions pinned, inputs not
  2. the training window is relative
  3. libraries resolve by range
  4. seed and config arrive at run time
  5. record inputs at run start

basics

~20 s

A code revision pins the training instructions, not the values they read: the applications the training query returns, the resolved library versions, the run-time configuration and the seed all moved. Rebuilding a model version means pinning those inputs too.

solid answer

~50 s

The code revision fixes the model family, the feature transformations, the loss and any literal written into the script. It does not fix what the script reads while it runs. Four things moved in that month: the **training data**, because the query selects a relative window and a month of new and corrected applications landed in it; the **environment**, because library versions and the base image resolved differently; the **run-time configuration**, including hyperparameters supplied by the job definition and a seed the script never set; and **nondeterminism inside the run** - thread count, data order, accelerator reduction order. So the training run is a function whose arguments mostly live outside the repository. To rebuild a model version you record those arguments at run time: an immutable training-data snapshot identity, the resolved config, an environment digest, and the artifact digest the run produced.

go deeper

for a junior

Recall the split: the code revision fixes the instructions, while the data, the environment, the run-time config and the seed are inputs that arrive from outside it.

for a middle

Explain how each of the four drifts - a relative training window, a version range resolving upward, config supplied by the job, unset seeds - and what recording each one looks like at run start.

for a senior

Show that you rebuild on a schedule rather than on demand, and that you resolve a relative window into a snapshot identity before the fit begins rather than after it.

for a principal

Frame the trade-off: pinning costs storage and build discipline on every run, including runs that never ship, and it is the only cost that cannot be paid retroactively when a decision is disputed.

## What the code revision actually fixes A **code revision** is an immutable name for one state of the training source: the model family the script builds, the feature transformations it applies, the loss it optimises, the evaluation it runs, and any value written into the script as a literal. Checking out that revision guarantees the same *instructions* execute. It guarantees nothing about the values those instructions read while they execute - and a training job reads a great deal from outside its own source tree. That gap is the whole subject. A training run is a function whose arguments mostly do not live in the repository. Reproducing the run means reproducing the arguments; a rebuild that pins only the revision has pinned the function without its inputs. ## The four inputs that moved 1. **The training data.** The script almost certainly selects examples by a *relative* window - the last eighteen months of applications, everything up to today. A month later the same selection returns a different set: new applications arrived, some rows were corrected or backfilled, and applications whose outcome was still unknown now carry a label. The query text is pinned; the result set is not. 2. **The environment.** The script declares its numerical, data-handling and model libraries by version *range*, inside a base image that is rebuilt on a schedule. A minor or patch upgrade can change a default argument, a tie-breaking rule, an initialisation scheme or a summation order without changing a single call the script makes. 3. **The configuration read at run time.** Hyperparameters supplied by a job definition, an environment variable or a tuning sweep are not in the revision unless they are literals in it. The random seed is the same story: a script that does not set one lets each run draw its own. 4. **Nondeterminism inside the run.** Thread counts, data-loading order and accelerator kernels that reduce in whatever order threads finish all move floating-point results in the last bits, and an iterative fit amplifies those differences. | Input | Fixed by the code revision alone? | |---|---| | Model family, loss, feature transformations | Yes | | Hyperparameters written as literals in the script | Yes | | Hyperparameters supplied by the job definition | No | | The set of applications the training query returned | No | | Resolved library versions and the base image | No | | Random seed, thread count, reduction order | No | ## Why this is an underwriting problem, not a tidiness problem An applicant disputes a decline eleven months after it was issued. A reviewer asks which model version produced it and what that version was built from. "We re-ran the training script from the same revision and got a slightly different scorer" is not an answer - it concedes that the team cannot demonstrate what scored the application. The defensible position is the opposite: every live model version resolves to a recorded set of inputs, and re-running those inputs reproduces the version to a tolerance the team stated in advance. ## Turning the run into a function of recorded inputs - Resolve the relative training window into an **immutable snapshot identity** at run start, and record that identity rather than the query text. - Record the **resolved** configuration - the hyperparameter set as the run actually saw it, including a seed the run sets rather than inherits. - Record the **environment as a digest**: the base image the job ran in plus the fully resolved dependency versions, never the ranges the script declares. - Record the **artifact digest** the run produced, so a later rebuild has something to compare against. - Treat the rebuild as an **assertion**: rebuild a sampled model version on a schedule, so the first attempt is not the one made under dispute. ## What still will not be identical Even with all of that pinned, a rebuild often differs in the last bits. That is expected, and by itself it is not a failure. The audit question is whether the rebuilt version assigns the disputed application the same score and the same decision within a stated tolerance. Bit-for-bit identity is a stronger property that costs real engineering - single-threaded determinism, deterministic kernels, frozen dependency digests - and it is worth buying only where the decision boundary is genuinely that sharp. ## Why the pinning cannot wait A team that discovers this during a dispute usually discovers two failures at once: the model version cannot be rebuilt, and the data behind it no longer exists in the form it had. They compound, because reconstructing the snapshot after the fact means re-deriving it from upstream systems that have themselves moved. Pinning is cheap while the run is executing and impossible afterwards - which is why it happens on every run, not only on the ones expected to reach production.

  • The team pins the training query text instead of a snapshot identity. Why is that not enough?
    The query text is an instruction, not a result. Re-executing it a month later reads a table that has gained rows, lost rows to corrections and gained labels for applications whose outcome was previously unknown. Pinning the text reproduces the selection logic; pinning a snapshot identity reproduces the actual set of applications the fit consumed, which is the thing the model was a function of.
  • If the script sets a fixed seed, does the rerun now reproduce the earlier model version?
    Not on its own. A seed fixes the pseudo-random draws the script controls - initialisation, shuffling, subsampling - and nothing outside them. The data selected by a relative window still moved, the resolved library versions still moved, and thread-level reduction order is still free. A seed removes one source of variation out of four; it is necessary and nowhere near sufficient.
  • Which of these four inputs is cheapest to pin, and which is most often left unpinned?
    The configuration is cheapest: record the resolved hyperparameter map the run actually used, which the job already holds in memory. The environment is most often left unpinned, because the training script looks pinned - it declares its dependencies - while the resolution happens in the image build and is never written down. Recording a base-image digest and a resolved dependency digest closes it.

saying these in an interview costs you the question

  • Believes checking out the same revision reproduces the model
  • Thinks storing the trained artifact is the same as being able to rebuild it
  • Assumes the training table is immutable because rows are only appended
  • Says a fixed seed alone makes a training run reproducible
  • Treats the environment as pinned because dependencies are declared
open as a page

What must a rebuild record pin, besides the code revision, to reconstruct the underwriting model version that declined an application?

level: middleimportance: must knowfreq 68%

basics

~20 s

Five pins: the code revision, the training-data snapshot identity as a content digest, the fully resolved hyperparameters and seed, an environment digest covering base image and dependency versions, and the artifact digest the run produced, all tied to one training-run identifier.

open as a page

Given only the model version stamped on an eleven-month-old underwriting decline, how does the platform reach its training run?

level: seniorimportance: should knowfreq 43%

basics

~20 s

Through a two-hop chain: the stamped version resolves to an immutable artifact digest, and a lineage entry written when the run finished maps that digest to its training-run identifier and pin set. A moving alias breaks the first hop.

open as a page

Why can a rebuilt underwriting model version match the original's metrics yet differ bit-for-bit from the stored artifact?

level: seniorimportance: should knowfreq 58%

basics

~10 s

Floating-point addition is not associative, so thread count, data order and accelerator reduction order shift the last bits even with every input pinned. Reproducible-in-metric is the achievable claim; reproducible-in-bits costs deliberate determinism work.

open as a page

In an underwriting platform, what makes a retired model version un-rebuildable while its artifact still sits in object storage?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

Rebuildability depends on the pins outliving the artifact. Most often the training-data snapshot was deleted under a shorter retention policy; a withdrawn base image, an unavailable dependency version or a rewritten feature definition do the same.

open as a page