skip to content

Why can a claim-severity training run with its parameters, metrics and code commit recorded still fail to reproduce?

level: middleimportance: should knowfreq 57%

answer

  1. your code is not the only code
  2. a range resolves forward, a lock does not
  3. defaults move between library versions
  4. record the environment digest and seeds
  5. bitwise versus statistical reproduction

basics

~20 s

Because the execution environment is an unrecorded input. A commit pins your code, not the code you depend on: dependency ranges resolve forward, defaults move between library versions, and unpinned seeds and thread counts change results. Record a resolved dependency lock and an environment digest with the run.

solid answer

~50 s

A code commit pins the training job's own source, and nothing else it ran on. The second attempt resolves dependency ranges to newer versions whose defaults differ, starts from a rebuilt base environment, and may run on different hardware with a different number of threads. So the run record has to carry the environment as an input: a **resolved** dependency set (exact versions including transitive ones, not ranges), a digest of the built execution environment, the random seeds, and any value the job read as of run time rather than by snapshot identifier - a rate table or a holiday calendar read live is a hidden input exactly like an unpinned dataset. It is also worth stating which reproduction you are promising: the same evaluation error within a published tolerance is achievable and cheap; byte-identical artifacts are a much stronger and costlier claim.

go deeper

for a junior

Recall that a training run depends on more than the code you wrote: the libraries it calls, the environment it ran in, and the random seed all change the result, so all of them belong on the run record.

for a middle

Explain how an unpinned dependency changes results through moved defaults rather than through any parameter you set, and why a resolved lock covering transitive dependencies is different from declaring version ranges.

for a senior

Demonstrate the operational form: an environment digest recorded per run, a published tolerance for what reproduction means, and a scheduled rebuild of a promoted run so the claim is tested before an auditor tests it.

for a principal

Decide how far up the pinning ladder to buy. Full determinism costs training throughput and fleet flexibility; the usual answer is a reproducible environment plus retained artifacts, with bitwise guarantees reserved for models whose exact output may be disputed.

## The commit is one input of several A training run consumes four kinds of input, and a code commit covers only the first: 1. **Your code** - pinned by the commit. 2. **Everything your code calls** - the training library, the numeric stack beneath it, the data access layer, and their transitive dependencies. A declared range such as *at least this minor version* resolves to whatever exists on the day the environment is built. 3. **The execution environment itself** - the base image, the system libraries, the accelerator runtime, the thread and process counts. 4. **Data read at run time** - the training snapshot (which should be pinned by identifier), but also the small live lookups a job quietly does: a rate table, a region mapping, a holiday calendar, a currency conversion. Any of these that is not written into the run record is a hidden input, and a hidden input is the reason a rebuild disagrees with the record. ## Why an unpinned dependency changes the answer The usual assumption is that a library upgrade changes performance, not results. It routinely changes results, because the thing that moved is a **default**: - a default regularisation, split rule or early-stopping criterion changes between minor versions; - a default handling of missing values changes, and a claim-severity dataset is full of them; - a categorical encoding changes its ordering, which shifts every tree split downstream; - a numerical routine changes its summation or reduction order, which moves the last digits and, through early stopping, the model. None of these appear as a parameter change on the run record, because from the job's point of view nothing was changed at all. ## Pinning, and what each level buys | what you record | what a rebuild reproduces | what it costs | |---|---|---| | code commit only | the job's own logic | nothing | | + direct dependency versions | most behaviour, until a transitive one moves | a lock file to maintain | | + fully resolved lock, transitive included | the library stack as it was | rebuild discipline, periodic upgrade work | | + digest of the built environment | system libraries and runtime too | storing or rebuilding the environment | | + seeds, thread counts, deterministic settings | run-to-run variation on the same hardware | measurably slower training | | + identical hardware class | the last remaining numeric differences | fleet inflexibility | Most teams should stop at the environment digest plus seeds. The rows below it are bought only when a dispute over exact bytes is foreseeable - and note that **retaining the artifact** answers *what did we serve* immediately, without reproducing anything at all. ## Two different promises called reproducibility - **Statistical reproduction**: rebuilding from the record yields a model whose evaluation error on the same evaluation set is within a stated tolerance of the recorded value. This is the practical promise, and the tolerance has to be written down or it will be argued about during an incident. - **Bitwise reproduction**: rebuilding yields an artifact with the same content digest. This requires fixed seeds, deterministic numeric kernels, fixed thread and shard counts, and a fixed data ordering - and it can still be defeated by different hardware. Confusing the two produces a false incident: someone rebuilds a quarter-old claim-severity run, sees an error of 0.186 against a recorded 0.184, and declares the pipeline broken, when the honest reading is that it reproduced fine and nobody had stated the band. ## Seeds are necessary and not sufficient Recording the seed removes sampling and initialisation as explanations for a difference - that is its whole value. It does **not** guarantee identical results across machines: a reduction split across a different number of threads sums floating-point values in a different order, and any downstream branch on that value can diverge. Record the seed, and record the parallelism settings next to it, so that a difference has a shortlist of suspects rather than a shrug. ## What to do about it 1. **Resolve, then record.** Let the build resolve versions once, write the resolved set and its hash onto the run record, and reuse that hash rather than re-resolving. 2. **Make live reads explicit.** Any table the job reads without a snapshot identifier is either promoted to a pinned input or written into the record with the value it returned. 3. **Publish the tolerance.** State what a successful reproduction means for a claim-severity model - the same evaluation metric within a band - and treat a result outside it as an incident with a cause, not as normal variance. 4. **Exercise it.** Rebuild one promoted run per quarter from its record alone. A reproducibility claim that has never been executed is an assumption.

  • What tolerance should reproduced mean for a quarterly claim-severity model?
    Publish a band on the evaluation metric: a rebuild from the record must land within it on the same evaluation set, and anything outside is an incident with a named cause. Byte-identical output is a stronger promise that most pipelines do not need, because a retained artifact under a verified digest already answers which exact model was served.
  • Which is cheaper to guarantee, rebuilding the artifact or keeping it?
    Keeping it, by a wide margin. The retained bytes plus their digest answer what ran and make a rollback a pointer flip; the environment lock answers a different question - can we change this model and still trust the comparison. A pipeline needs both, but retention is the one that buys minutes during an incident.
  • Is recording the base environment's name enough instead of a digest?
    No. A name resolves to different contents over time as the environment is rebuilt with patched system libraries, so it describes the environment rather than pinning it. A digest of the built environment is content-addressed: it either resolves to the same bytes or it fails loudly, which is the behaviour you want a year later.

saying these in an interview costs you the question

  • Believes the code commit alone pins every library the run used
  • Pins direct dependencies but lets transitive ones float
  • Assumes an identical seed guarantees identical results on any hardware
  • Says reproducible without stating whether it means bitwise
  • Reads a rate table live rather than by snapshot identifier
  • Treats a library upgrade as a performance change only