Why can a claim-severity training run with its parameters, metrics and code commit recorded still fail to reproduce?
answer
- your code is not the only code
- a range resolves forward, a lock does not
- defaults move between library versions
- record the environment digest and seeds
- bitwise versus statistical reproduction
basics
~20 sBecause the execution environment is an unrecorded input. A commit pins your code, not the code you depend on: dependency ranges resolve forward, defaults move between library versions, and unpinned seeds and thread counts change results. Record a resolved dependency lock and an environment digest with the run.
solid answer
~50 sA code commit pins the training job's own source, and nothing else it ran on. The second attempt resolves dependency ranges to newer versions whose defaults differ, starts from a rebuilt base environment, and may run on different hardware with a different number of threads. So the run record has to carry the environment as an input: a **resolved** dependency set (exact versions including transitive ones, not ranges), a digest of the built execution environment, the random seeds, and any value the job read as of run time rather than by snapshot identifier - a rate table or a holiday calendar read live is a hidden input exactly like an unpinned dataset. It is also worth stating which reproduction you are promising: the same evaluation error within a published tolerance is achievable and cheap; byte-identical artifacts are a much stronger and costlier claim.
go deeper
Recall that a training run depends on more than the code you wrote: the libraries it calls, the environment it ran in, and the random seed all change the result, so all of them belong on the run record.
Explain how an unpinned dependency changes results through moved defaults rather than through any parameter you set, and why a resolved lock covering transitive dependencies is different from declaring version ranges.
Demonstrate the operational form: an environment digest recorded per run, a published tolerance for what reproduction means, and a scheduled rebuild of a promoted run so the claim is tested before an auditor tests it.
Decide how far up the pinning ladder to buy. Full determinism costs training throughput and fleet flexibility; the usual answer is a reproducible environment plus retained artifacts, with bitwise guarantees reserved for models whose exact output may be disputed.
## The commit is one input of several A training run consumes four kinds of input, and a code commit covers only the first: 1. **Your code** - pinned by the commit. 2. **Everything your code calls** - the training library, the numeric stack beneath it, the data access layer, and their transitive dependencies. A declared range such as *at least this minor version* resolves to whatever exists on the day the environment is built. 3. **The execution environment itself** - the base image, the system libraries, the accelerator runtime, the thread and process counts. 4. **Data read at run time** - the training snapshot (which should be pinned by identifier), but also the small live lookups a job quietly does: a rate table, a region mapping, a holiday calendar, a currency conversion. Any of these that is not written into the run record is a hidden input, and a hidden input is the reason a rebuild disagrees with the record. ## Why an unpinned dependency changes the answer The usual assumption is that a library upgrade changes performance, not results. It routinely changes results, because the thing that moved is a **default**: - a default regularisation, split rule or early-stopping criterion changes between minor versions; - a default handling of missing values changes, and a claim-severity dataset is full of them; - a categorical encoding changes its ordering, which shifts every tree split downstream; - a numerical routine changes its summation or reduction order, which moves the last digits and, through early stopping, the model. None of these appear as a parameter change on the run record, because from the job's point of view nothing was changed at all. ## Pinning, and what each level buys | what you record | what a rebuild reproduces | what it costs | |---|---|---| | code commit only | the job's own logic | nothing | | + direct dependency versions | most behaviour, until a transitive one moves | a lock file to maintain | | + fully resolved lock, transitive included | the library stack as it was | rebuild discipline, periodic upgrade work | | + digest of the built environment | system libraries and runtime too | storing or rebuilding the environment | | + seeds, thread counts, deterministic settings | run-to-run variation on the same hardware | measurably slower training | | + identical hardware class | the last remaining numeric differences | fleet inflexibility | Most teams should stop at the environment digest plus seeds. The rows below it are bought only when a dispute over exact bytes is foreseeable - and note that **retaining the artifact** answers *what did we serve* immediately, without reproducing anything at all. ## Two different promises called reproducibility - **Statistical reproduction**: rebuilding from the record yields a model whose evaluation error on the same evaluation set is within a stated tolerance of the recorded value. This is the practical promise, and the tolerance has to be written down or it will be argued about during an incident. - **Bitwise reproduction**: rebuilding yields an artifact with the same content digest. This requires fixed seeds, deterministic numeric kernels, fixed thread and shard counts, and a fixed data ordering - and it can still be defeated by different hardware. Confusing the two produces a false incident: someone rebuilds a quarter-old claim-severity run, sees an error of 0.186 against a recorded 0.184, and declares the pipeline broken, when the honest reading is that it reproduced fine and nobody had stated the band. ## Seeds are necessary and not sufficient Recording the seed removes sampling and initialisation as explanations for a difference - that is its whole value. It does **not** guarantee identical results across machines: a reduction split across a different number of threads sums floating-point values in a different order, and any downstream branch on that value can diverge. Record the seed, and record the parallelism settings next to it, so that a difference has a shortlist of suspects rather than a shrug. ## What to do about it 1. **Resolve, then record.** Let the build resolve versions once, write the resolved set and its hash onto the run record, and reuse that hash rather than re-resolving. 2. **Make live reads explicit.** Any table the job reads without a snapshot identifier is either promoted to a pinned input or written into the record with the value it returned. 3. **Publish the tolerance.** State what a successful reproduction means for a claim-severity model - the same evaluation metric within a band - and treat a result outside it as an incident with a cause, not as normal variance. 4. **Exercise it.** Rebuild one promoted run per quarter from its record alone. A reproducibility claim that has never been executed is an assumption.
- What tolerance should reproduced mean for a quarterly claim-severity model?Publish a band on the evaluation metric: a rebuild from the record must land within it on the same evaluation set, and anything outside is an incident with a named cause. Byte-identical output is a stronger promise that most pipelines do not need, because a retained artifact under a verified digest already answers which exact model was served.
- Which is cheaper to guarantee, rebuilding the artifact or keeping it?Keeping it, by a wide margin. The retained bytes plus their digest answer what ran and make a rollback a pointer flip; the environment lock answers a different question - can we change this model and still trust the comparison. A pipeline needs both, but retention is the one that buys minutes during an incident.
- Is recording the base environment's name enough instead of a digest?No. A name resolves to different contents over time as the environment is rebuilt with patched system libraries, so it describes the environment rather than pinning it. A digest of the built environment is content-addressed: it either resolves to the same bytes or it fails loudly, which is the behaviour you want a year later.
saying these in an interview costs you the question
- Believes the code commit alone pins every library the run used
- Pins direct dependencies but lets transitive ones float
- Assumes an identical seed guarantees identical results on any hardware
- Says reproducible without stating whether it means bitwise
- Reads a rate table live rather than by snapshot identifier
- Treats a library upgrade as a performance change only