skip to content

What must a rebuild record pin, besides the code revision, to reconstruct the underwriting model version that declined an application?

level: middleimportance: must knowfreq 68%

answer

  1. identity, not description
  2. five pins on one run id
  3. resolved, never declared
  4. written by the run, during the run
  5. the cold-machine test

basics

~20 s

Five pins: the code revision, the training-data snapshot identity as a content digest, the fully resolved hyperparameters and seed, an environment digest covering base image and dependency versions, and the artifact digest the run produced, all tied to one training-run identifier.

solid answer

~50 s

A rebuild record - the pin set - has to name every argument the training run consumed, by **identity rather than description**. That means the code revision; the training-data snapshot as a content digest, not a table name and date; the hyperparameter map as the run resolved it, including the seed it set; an environment digest covering the base image and the fully resolved dependency versions; and the digest of the artifact the run produced, so a rebuild has a target to compare against. All five hang off one training-run identifier, and the live model version resolves to that identifier. The test of a pin set is blunt: hand it to a machine that has never seen this project, and it must be able to start the run with no further lookups. Anything the rebuild would have to ask a human or a moving table for is a missing pin.

code

json · 19 lines
json
{
  "modelVersion": "underwriting-scorer/4.2.0",
  "trainingRunId": "run-2025-11-03-0417",
  "codeRevision": "a41d9c7e2b6f",
  "trainingSnapshot": {
    "id": "applications-window-18m",
    "digest": "sha256:20be4f..."
  },
  "hyperparameters": {
    "maxDepth": 6,
    "learningRate": 0.05,
    "seed": 1337
  },
  "environment": {
    "baseImageDigest": "sha256:c07a91...",
    "resolvedDependenciesDigest": "sha256:e3b0c4..."
  },
  "artifactDigest": "sha256:9f1c8d..."
}

go deeper

for a junior

Memorise the five pins - code revision, data snapshot digest, resolved config and seed, environment digest, artifact digest - and that they hang off one training-run identifier.

for a middle

Be able to explain why each pin is recorded as an identity rather than a description, and why the resolved value differs from the declared one for both data and dependencies.

for a senior

Demonstrate the operating discipline: the run writes its own pins before the fit, a scheduled rebuild proves the set is complete, and decay is caught long before a dispute forces the first attempt.

for a principal

Weigh what the pin set costs on every run against the class of question it makes answerable, and decide which versions deserve the stronger guarantees and for how long.

## What a pin set is A **pin set** is the complete argument list of one training run, recorded while the run happens and stored beside the model version it produced. It exists to answer one question months later: *what exactly was this version built from?* The pin set is not documentation and not a description of the run - it is the input to re-executing it. The governing rule is **identity, not description**. A description names something that may move; an identity names bytes that cannot. "The applications table, as of the third of November" is a description - the table may have been corrected since, and the meaning of "as of" depends on which timestamp column the reader picks. A content digest of the snapshot is an identity: it either matches or it does not. ## The five pins 1. **Code revision** - an immutable revision identifier for the training source, never a branch or tag name that can be repointed. 2. **Training-data snapshot identity** - the resolved snapshot the fit consumed, referenced by a content digest such as a SHA-256 over the snapshot's manifest. Resolved at run *start*, before the fit reads anything. 3. **Resolved configuration** - the hyperparameter map as the run saw it after defaults, sweep values and environment overrides were applied, plus the seed the run set. Not the defaults the script declares. 4. **Environment digest** - the base image digest plus a digest over the fully resolved dependency versions. Declared ranges are not a pin; the resolution is. 5. **Artifact digest** - the digest of the model artifact the run emitted, which gives a later rebuild something to compare its output against. These five hang off a single **training-run identifier**, and the model version served in production resolves to that identifier. One indirection, both directions: version to run, run to pins. ## Identity against description | Recorded as a description | Recorded as an identity | |---|---| | "latest" or a branch name | An immutable code revision identifier | | Table name plus a date | A content digest over the snapshot manifest | | A dependency range in the script | A digest over the resolved dependency versions | | "the production image" | The base image digest | | "the champion model" | The artifact digest | Every row on the left resolves to something different a year later; every row on the right resolves to the same bytes or fails loudly. ## When the pin set is written It is written **by the run, during the run**, not assembled afterwards by whoever promotes the version. Two reasons. First, several pins only exist during execution: the resolved config, the resolved dependency set, the snapshot identity a relative window collapsed to. Second, a pin set assembled at promotion time records what the promoter believes rather than what the job did, and the two diverge exactly when it matters. A useful discipline is that the pin set is written before the fit starts, and the artifact digest is appended when the run finishes. A crashed run then still leaves a complete record of what it attempted. ## Testing the pin set A pin set that has never been used is a guess. Two cheap tests keep it honest: - **The cold-machine test.** Hand the pin set to a worker with no project state and no human available. If it can start the run without a single additional lookup, the set is complete; if it has to ask which snapshot "the November data" means, it is not. - **The scheduled rebuild.** Periodically pick a live model version, rebuild it from its pins and compare against the recorded artifact digest and a stored audit sample. This turns rebuildability from a claim into a monitored property, and it catches decay - an expired image, a deleted snapshot - while there is still time to react. ## The gaps that show up in practice - The **environment** is the pin most often missing, because the training script looks pinned when it declares dependencies; the resolution happens elsewhere and is never written down. - The **seed** is recorded as "the default" when no seed was set at all. - The snapshot is pinned by a **name that was later repointed**, which is worse than no pin, because the rebuild silently succeeds on the wrong data. - The pin set is complete but nothing **links the live model version to it**, so an auditor holding a version stamp has no path to the run. ## What the pin set is not It is not the record of what one prediction did - the inputs, score and outcome of a single application belong to the decision record, which serves a different question and has its own retention. And it is not a substitute for the trained artifact: storing the artifact makes the version *servable*, while storing the pins makes it *rebuildable*. An audit that can only serve the old version can show what it does today; only a rebuild shows where it came from.

  • Why record the artifact digest in the pin set when the artifact itself is already stored?
    Because it is the rebuild's acceptance target. Without a recorded digest, a rebuild can only be compared against whatever artifact currently sits at the version's path, which assumes that path was never overwritten. The digest also detects storage corruption and makes the bit-identity check a one-line comparison instead of a byte-by-byte diff of two large files.
  • The team pins a dependency lockfile instead of a digest over resolved versions. Is that equivalent?
    Nearly, with one gap. A lockfile pins the packages the script asked for, but not the system libraries, compiler or driver stack underneath them, which live in the image. Pinning the lockfile and the base image digest together covers both layers. A digest over the lockfile is also cheaper to compare than the file itself, which is why the pin set stores a digest rather than the text.
  • Who writes the pin set - the training job or the promotion step?
    The training job, while it runs. The resolved config, the resolved dependency set and the snapshot identity a relative window collapsed to only exist during execution. A pin set assembled at promotion time records what the promoter believes about the run rather than what the run did, and those two versions of the truth diverge precisely in the cases an audit cares about.

saying these in an interview costs you the question

  • Pins the training data by table name and date rather than a content digest
  • Records declared dependency ranges instead of the resolved versions
  • Assembles the pin set at promotion time from what the promoter remembers
  • Treats storing the trained artifact as equivalent to being able to rebuild it
  • Pins a branch or tag name that can later be repointed to another revision
  • Leaves the live model version with no link back to its training run