skip to content

Why can a model trained from 'the latest transcript export' not be rebuilt by rerunning the same code?

level: juniorimportance: must knowfreq 62%

answer

  1. a name that keeps moving
  2. same code, different rows
  3. seed fixes draws, not inputs
  4. pin content, not a query
  5. snapshot id plus code plus config

basics

~20 s

'Latest' is a moving pointer, not a version. Between the two runs utterances were appended, transcripts corrected and recordings withdrawn, so the second job trains on a different set. Rebuilding needs an immutable snapshot identifier that resolves to fixed bytes.

solid answer

~40 s

A fixed seed pins what the trainer does *given* an input; it says nothing about what the input was. `latest` is resolved at read time against a corpus that keeps moving - new utterances land every cycle, a reviewer corrects a transcript, a contributor withdraws a recording, shards get re-cut - so resolving the same name a month later yields a different row set and therefore different weights. A rebuildable run pins three coordinates: an immutable **dataset snapshot id**, the **code version** that read and transformed it, and the **configuration** (hyperparameters, schedule, seed). The snapshot id has to name content rather than a location, so resolving it either returns exactly the utterances that trained the model or fails loudly - never a silently different corpus.

go deeper

for a junior

Recall the one-liner: 'latest' is a pointer, not a version, and a seed does not pin the data. Name the three things a rebuildable run records - snapshot id, code version, configuration.

for a middle

Explain the mechanics: which edits move the pointer (appends, corrections, withdrawals, re-sharding), and why read order alone can change the model even when the row set is identical.

for a senior

Show you have debugged this: an unexplained metric move attributed to a code change that was really a data change, and the discipline of failing loudly when a snapshot id does not resolve.

for a principal

Frame the tradeoff: pinning ids keeps every referenced byte alive, so reproducibility, storage cost and retention policy are one decision rather than three.

## What a training run has to pin A trained model is a function of three things: **which examples were read**, **which code turned them into batches and weights**, and **which configuration** that code ran under. A result is reproducible only when all three are named by something that cannot change afterwards. Recording the commit and the seed pins the second and third and leaves the first - the largest and most volatile of them - described by a name that means something different every day. ## Why 'latest' moves under you A speech corpus is edited continuously, and each kind of edit changes the training set without changing the name used to request it: - **Appends** - each collection cycle adds new utterances, which a later read includes and an earlier one did not. - **In-place corrections** - a reviewer fixes a mis-typed transcript, so a row that was read one way is read another way later. - **Withdrawals** - a contributor asks for their recordings to be removed, and rows silently disappear from the set. - **Re-processing** - audio is re-encoded or re-segmented upstream, changing the bytes behind an unchanged identifier. - **Re-partitioning** - shards are merged or split, which changes the order rows arrive in even when the set is identical. The first three change *which* examples the trainer sees. The last two can change the model even when the set is unchanged, because read order feeds shuffling and batching. ## The seed is not a version A seed makes the pseudo-random draws repeatable: the shuffle, the initialisation, any stochastic augmentation of the audio. Every one of those draws is applied **to whatever rows arrived**. Same seed with different rows gives different weights, and nothing in the job errors - the run succeeds, the metrics move, and the difference is attributed to the last code change. This is why 'we fixed the seed, so it is reproducible' is the classic first-screen wrong answer. ## The three coordinates | coordinate | what it pins | what breaks when it is missing | |---|---|---| | dataset snapshot id | the exact utterances and transcripts read | a silently different training set; unexplained metric moves | | code version | the transforms, the trainer, the evaluation | the same rows produce different features and different scores | | configuration | hyperparameters, schedule, seed | the same rows and code produce different weights | All three belong to the model version as recorded facts. Two of them are small; the third is a pointer, which is the only affordable way to reference tens of terabytes. ## What makes a snapshot id trustworthy 1. **Content-derived** - the id is computed from the data it names (a hash over a manifest that itself hashes each shard), so equal ids imply equal bytes. 2. **Immutable** - publishing a changed corpus mints a **new** id; an existing id is never re-pointed. 3. **Resolvable** - a catalogue maps the id to the shard list and the shards to storage, so a reader can get the bytes back years later. 4. **Verified on read** - the reader re-hashes what it loaded and fails if it does not match, rather than training on quietly substituted data. 5. **Recorded by the reader** - the job that resolved the id writes that fact down; nothing in the produced artifact reveals it afterwards. ## The near misses, and why each fails | candidate 'version' | why it is not one | |---|---| | a cutoff timestamp on the source table | late-arriving rows and in-place corrections land *behind* the cutoff, so the same predicate returns a different set later | | a dated folder of copied files | mutable, unverified, and it costs the whole corpus per experiment | | the training job's start time | names when the read happened, not what was read | | the model artifact's own hash | identifies the output; says nothing about the inputs | ## In a design round The expected answer is short: name the moving pointer as the defect, pin an immutable content-derived snapshot id, and say that the model version records the snapshot id, the code version and the configuration together. The stronger follow-through is admitting what pinning buys you - the ability to say *this* model came from *these* utterances - and what it costs, which is a catalogue, a retention policy and a garbage-collection story for the bytes those ids keep alive.

  • Does filtering the source corpus on a cutoff timestamp make the read reproducible?
    Only if the source is strictly append-only with an immutable event time and no updates or deletions - which a transcript corpus is not. Corrections and late-arriving utterances land behind the cutoff, so the same predicate returns a different set next month. The cutoff describes an intent; the snapshot id records the result.
  • What is the minimum you must record to rebuild a model months later?
    The dataset snapshot id, the code version that read it, and the full configuration including the seed, plus enough environment detail to reconstruct the runtime. The seed alone is worthless, and the model artifact carries none of it unless the training job wrote it down at the time.
  • Two runs pin the same snapshot id and the same code, but the weights differ slightly. Is that a versioning failure?
    No - that is residual non-determinism in the execution itself: accumulation order on parallel hardware, non-deterministic kernels, or a data loader whose worker count changes the order. It is a real problem, but it is bounded and diagnosable precisely because the inputs are pinned.

A recipe that says 'whatever is in the fridge' is not a recipe; it is a description of a meal someone once made. Naming the exact ingredients is what makes it repeatable.

saying these in an interview costs you the question

  • Thinks a fixed random seed by itself makes a training run reproducible
  • Believes rerunning the same query against the corpus returns the same rows
  • Treats a date cutoff on a mutable table as a version identifier
  • Assumes appended utterances are harmless because older rows were untouched
  • Says the saved model artifact is the record, so nothing else needs storing