skip to content

Why can a rebuilt underwriting model version match the original's metrics yet differ bit-for-bit from the stored artifact?

level: seniorimportance: should knowfreq 58%

answer

  1. two claims, not one
  2. floating-point addition is not associative
  3. seed covers the random stream only
  4. zero decision flips, bounded score delta
  5. buy bit-identity only near the threshold

basics

~10 s

Floating-point addition is not associative, so thread count, data order and accelerator reduction order shift the last bits even with every input pinned. Reproducible-in-metric is the achievable claim; reproducible-in-bits costs deliberate determinism work.

solid answer

~50 s

These are two different claims. **Reproducible in bits** means the rebuilt artifact hashes to the recorded digest. **Reproducible in metric** means it scores an audit sample to within a stated tolerance and returns the same decisions. Pinning the code revision, snapshot digest, config and environment gets you the second reliably and the first only with extra work, because floating-point addition is not associative: change the thread count, the data-loading order or the accelerator's reduction order and partial sums combine differently, and an iterative fit amplifies the difference. A seed fixes the pseudo-random draws the script controls and nothing outside them. So state the guarantee you actually offer - typically "the rebuild reproduces every decision on the audit sample and every score within tolerance" - and buy bit-identity only where a decision sits close enough to the threshold that last-bit differences could flip it.

code

pseudocode · 18 lines
pseudocode
rebuilt = train(pins.codeRevision, pins.snapshotDigest,
                pins.hyperparameters, pins.environmentDigest)

if digest(rebuilt.artifact) == pins.artifactDigest:
    return "bit-identical"

maxDelta = 0
flipped = 0
for each case in auditSample:
    newScore = score(rebuilt, case.storedInputs)
    maxDelta = max(maxDelta, abs(newScore - case.storedScore))
    if decision(newScore) != case.storedDecision:
        flipped = flipped + 1

if flipped == 0 and maxDelta <= 0.0005:
    return "metric-equivalent: the disputed decline stands"

return "not reproducible: escalate before answering the applicant"

go deeper

for a junior

Remember that two runs with identical inputs can still produce different bits, and that matching behaviour is a separate and weaker claim than matching bytes.

for a middle

Explain the mechanism - non-associative floating-point addition combined with parallel reduction order - and name what a seed covers and what it leaves open.

for a senior

Show that you define the tolerance before the dispute, check zero decision flips against a stored audit sample, and run the rebuild on a schedule rather than on demand.

for a principal

Decide which versions justify the cost of deterministic execution, and make sure the guarantee the organisation publishes is the one the platform can actually meet.

## Two different claims When a team says a model version is "reproducible", they are usually making one of two claims, and confusing them costs an audit dearly. - **Reproducible in bits.** Re-executing the pinned run produces an artifact whose digest equals the recorded one. It is a single comparison, and it is either true or false. - **Reproducible in metric.** The rebuilt version scores a stored audit sample within a stated tolerance and returns the same decision on every case. It is a statement about behaviour, and it needs a tolerance defined in advance. | | Reproducible in bits | Reproducible in metric | |---|---|---| | What is compared | Artifact digest | Scores and decisions on an audit sample | | Needs a tolerance | No | Yes, stated in advance | | Cost to guarantee | High - determinism engineering | Moderate - complete pins | | Answers "what scored this application" | Yes | Yes, for the cases checked | | Survives an accelerator or library change | No | Usually | ## Why the bits move even with everything pinned Floating-point addition is **not associative**: `(a + b) + c` and `a + (b + c)` can differ in the last bits. Any computation that sums many values in a parallel order therefore depends on the order in which partial results arrive. Sources of that variation, none of which a seed touches: 1. **Thread and worker count.** A reduction split across eight workers combines partial sums in a different grouping than one split across sixteen. 2. **Accelerator kernels.** Some kernels reduce in whatever order their threads finish, or select an algorithm based on the hardware they find; the same operation on the same inputs can differ between two runs on the same device. 3. **Data-loading order.** Asynchronous loaders deliver batches in completion order unless explicitly ordered, changing the sequence of gradient updates. 4. **Library minor and patch versions.** A new release may change a default, a tie-break, an initialisation scheme or the internal summation strategy while presenting an identical interface. 5. **Hardware.** Different accelerator generations or instruction sets take different code paths for the same operation. In an iterative fit, a last-bit difference in one update changes the next update's starting point, so tiny divergence compounds across the run. The end artifact differs; the model it encodes usually does not, in any way a decision can see. ## What a seed does and does not fix A seed fixes the pseudo-random draws the script controls: parameter initialisation, shuffling, subsampling, dropout masks. That is genuinely valuable and it should always be set and recorded. It does not fix anything outside the random number stream - not reduction order, not library defaults, not worker count. Teams that believe a seed alone buys reproducibility are usually one library upgrade away from finding out otherwise. ## Choosing and stating the tolerance A metric claim without a tolerance is not a claim. State it as two conditions, checked against a stored audit sample of past applications: - **Decision equivalence.** Every case in the sample receives the same approve, decline or refer outcome. This is the condition that matters to a disputed decline, and it should be *zero* flips, not a percentage. - **Score tolerance.** The maximum absolute difference in the score stays under a small bound. This catches a version that agrees on decisions by luck because no case sat near the threshold. The second condition earns its place because decision equivalence alone is fragile: a sample where every case sits far from the cut-off would pass even after a meaningful change. ## When bit-identity is worth buying Bit-identity is achievable - single-threaded or fixed-worker execution, deterministic kernel selection, ordered data loading, fully digest-pinned environments - and it is slower and more constraining than the alternative. It is worth buying when the decision boundary is sharp enough that last-bit differences could flip an outcome, when an external reviewer demands byte-level evidence, or for a small set of versions that carry unusual exposure. For most versions, complete pins plus a stated metric tolerance is both the honest guarantee and the affordable one. ## Saying it out loud The governance failure is not that the bits differ. It is a team that promised bit-identity, discovers during a dispute that it never held, and now has to explain the gap under time pressure. Write the guarantee into the platform's own documentation in the form it can actually meet - "a rebuild reproduces every decision on the audit sample and every score within tolerance" - and test it on a schedule, so the claim and the behaviour never drift apart.

  • The rebuild flips one decision out of a two-thousand-case audit sample. What do you conclude?
    That the rebuild is not metric-equivalent, and the flipped case gets investigated individually rather than averaged away. Check where that case sat relative to the decision threshold: a case a hair from the cut-off flipping on last-bit noise is a different finding from one flipping by a wide margin, which points at a genuinely different model - usually a wrong snapshot or an unpinned environment.
  • Why include a score tolerance when decision equivalence is what the dispute turns on?
    Because decision equivalence can pass for the wrong reason. If no case in the sample sits near the threshold, a materially different model still returns identical outcomes and the check says nothing. The score bound detects that the rebuilt version computes the same numbers, not merely the same side of a cut-off, so it catches the divergence before some future application lands near the boundary.

Re-baking a cake from the same recipe, the same weighed ingredients and the same oven gives a cake that tastes identical but is not the same crumb for crumb. The audit asks whether it tastes the same, not whether the crumbs match.

saying these in an interview costs you the question

  • Claims a fixed seed alone makes a training run bit-reproducible
  • Treats any bit difference as proof the rebuild failed
  • States a metric-reproducibility claim with no tolerance defined
  • Assumes a patch-level library upgrade cannot change results
  • Reports average score difference instead of checking decision flips