skip to content

questions

5

What must a training run's record hold before two quarterly claim-severity runs can be compared at all?

level: juniorimportance: must knowfreq 68%

answer

  1. a number without its inputs
  2. record the yardstick, not just the score
  3. snapshot identifier, commit, parameters
  4. artifact digest ties bytes to the run
  5. same evaluation set, same metric definition

basics

~20 s

A run record must pin every input and the yardstick: the dataset snapshot identifier, the code commit, the full parameter set, a content digest of the model artifact, and the evaluation set identifier with its exact metric definition. Two runs compare only when that evaluation set and metric are identical.

solid answer

~50 s

A run record is comparable when it pins every input and the yardstick, not just the score. For a claim-severity model that means the `dataset snapshot identifier` the run read, the code commit of the training job, the complete parameter set including values that came from defaults, a content digest of the model artifact the run produced, and the evaluation set identifier plus the exact metric definition and its value. Two runs are then comparable only when the evaluation set and the metric definition are identical. If one candidate was scored on the prior quarter's held-out claims and another on this quarter's, the difference in error mixes a model change with a change in the claims themselves, and the ranking means nothing. Anything a run read that is not in the record is a hidden input.

code

json · 19 lines
json
{
  "run_id": "claim-severity-2026Q3-017",
  "code_commit": "9f2c1ab",
  "environment_lock": "sha256:71d0c4...",
  "dataset_snapshot_id": "claims-severity@2026-07-01T00:00:00Z",
  "params": {
    "objective": "tweedie",
    "tweedie_power": 1.6,
    "max_depth": 8,
    "seed": 20260701
  },
  "evaluation": {
    "set_id": "holdout-2026Q2",
    "metric": "weighted_absolute_percentage_error",
    "value": 0.184
  },
  "artifact_digest": "sha256:0b3e91...",
  "started_at": "2026-07-03T09:14:22Z"
}

go deeper

for a junior

Be able to list what one training run must record: the dataset snapshot identifier, the code commit, the parameters, the artifact digest, and the evaluation set with its metric. A score on its own is not a result.

for a middle

Explain why an identical metric name does not make two runs comparable, and why the evaluation set has to be held fixed across a quarter's candidates before any error difference can be attributed to the model.

for a senior

Show how you would enforce it: the job writes the record, a missing mandatory field fails the run, and an artifact with no run behind it cannot be registered. Recording is a pipeline property, not a discipline.

for a principal

The trade-off is how much of the record is mandatory. Every required field is friction on experimentation; every optional one is the field that turns out to be missing from the run someone must defend two years later.

## What a run record is A **run record** is the durable row a tracking system writes for one execution of the training job. It is not the job's console output and not the orchestrator's task history: those describe what the machine did, while the run record describes what the experiment **was**. One run, one record, written by the job itself as it resolves each input and produces each output. An insurer retraining a claim-severity model every quarter will produce a dozen candidate runs before one is promoted: different objectives, different feature sets, different regularisation. The record is what lets someone say, months later, which of those runs produced the model that is currently deciding reserve amounts, and on what evidence it was chosen. ## The five things a comparable record must pin 1. **The dataset snapshot identifier** - not the table name, the identifier of the exact frozen set the run read. Without it, two runs reporting different error may simply have read different claims. 2. **The code version** - the commit of the training job, so the transform that produced the model is recoverable rather than remembered. 3. **The complete parameter set** - every value the run used, *including the ones that came from defaults*. A default that changes in a later version is a silent parameter change, and a record that only lists what someone typed will not show it. 4. **The artifact digest** - a content hash (SHA-256, for example) of the model bytes the run produced, so the run and the thing it produced are tied together by content rather than by a path that can be rewritten. 5. **The evaluation set identifier and the metric definition** - which held-out claims were scored, by exactly which formula (weighted absolute percentage error over the held-out quarter, say), and the resulting value. ## Why comparability is a stricter bar than logging | what the record has | what it answers | what it still cannot answer | |---|---|---| | a metric value | how this run scored | whether another run's score means anything beside it | | value + metric definition | what was measured | whether it was measured on the same claims | | + evaluation set identifier | comparable **within** one evaluation set | which input change moved the number | | + snapshot identifier, commit, parameters | which change moved the number | which bytes are actually serving | | + artifact digest | the whole chain, run to served model | - | Each row is a strictly stronger claim than the one above it. Most teams stop at row one or two and are then surprised that their quarterly leaderboard cannot be defended. ## The quarterly trap The most common comparability mistake in a quarterly retrain is to score each quarter's candidate on **that quarter's own** held-out claims and then compare the error to the number recorded three months ago. Claim severity is not stationary: a hail season, a change in repair costs or a reserving-policy change moves the error on its own. The comparison therefore confounds the model change with a change in the underlying claims, and it moves in whichever direction the data moved. Two honest forms exist, and they answer different questions: - **Which candidate is better?** Score every candidate of this quarter on one fixed evaluation set with one metric definition. The set is held constant; only the model varies. - **Is the model getting worse over time?** Re-score the *currently promoted* model on the new quarter's evaluation set, alongside the candidate. Now the set varies and the model is held constant for one of the two readings, so the drift and the improvement are separable. ## What a partial record costs - **No snapshot identifier**: an error change cannot be attributed to the model rather than the data. - **No artifact digest**: you can say which run looked best, but not prove which bytes it produced. - **No parameters, or only the ones typed by hand**: the run cannot be re-run, and a moved default looks like noise. - **No evaluation set identifier**: the leaderboard is a list of numbers measured on unknown ground. - **A record a person filled in afterwards**: it captures intent, not what ran, and silently misses the value someone changed for one rerun. ## Making it hold in practice - The **job** writes the record, at the moment it resolves each input, and a mandatory field that cannot be resolved **fails the run** instead of logging a warning. - The record is written before an artifact can be registered: an artifact with no run behind it is unpromotable by construction. - Metric definitions are named and versioned, so one label cannot quietly mean two formulas in two quarters. - The record is queryable, not just readable. The questions it has to answer later - which registered versions read a given snapshot, which runs used a given commit - are queries, and a per-run text file cannot serve them.

  • A data contribution used in past training is later found untrustworthy - what in the run records finds every affected model version?
    Query the run records for every run whose dataset snapshot identifier covers that contribution, follow each run to the artifact digest it produced, and then to any registered model version pointing at that digest. That set is the blast radius, and the retained earlier versions are what you can fall back to. Without the snapshot identifier on the run, the same question becomes a forensic guess.
  • Should the record hold metrics measured on the training data as well as on the evaluation set?
    Yes, clearly labelled as separate fields. The training-fit metric is what later tells you whether a poor evaluation number came from underfitting or from the evaluation set moving underneath you. Only the evaluation metric on a fixed set is comparison ground, though - an unlabelled single metric field is how a run's number becomes unusable a quarter later.
  • Who writes the run record, the training job or the engineer?
    The job, automatically, as it resolves each input and writes each output. A record filled in afterwards records what someone believed they ran; it misses the parameter changed for one rerun and the default that moved with a dependency upgrade. Make the job fail when a mandatory field cannot be resolved, so an unrecorded run never becomes a candidate.

A lab notebook entry that records only the reading is useless: the next person needs the sample, the instrument settings and the scale it was read on before that number means anything.

saying these in an interview costs you the question

  • Treats a metric value alone as enough to compare two runs
  • Records hyperparameters but not which dataset snapshot the run read
  • Compares this quarter's error against last quarter's recorded number
  • Expects to find the right artifact later by filename or timestamp
  • Calls the orchestrator's console log a run record
  • Logs only the parameters someone typed, not the resolved defaults
open as a page

When a claim-severity model must be rolled back to the previously promoted version, what must the registry already hold for that to take minutes?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Rollback is fast only when the previous version is still fully resolvable: retained artifact bytes under a verified digest, the feature-definition and input-schema version it was trained against, its environment lock, and a promotion pointer that serving resolves at load - so the act is a pointer flip, not a rebuild.

open as a page

Why can a claim-severity training run with its parameters, metrics and code commit recorded still fail to reproduce?

level: middleimportance: should knowfreq 57%

basics

~20 s

Because the execution environment is an unrecorded input. A commit pins your code, not the code you depend on: dependency ranges resolve forward, defaults move between library versions, and unpinned seeds and thread counts change results. Record a resolved dependency lock and an environment digest with the run.

open as a page

A rerun overwrote the stored artifact of an already-promoted claim-severity model in place - what breaks?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Every claim about that promoted version becomes false. Its recorded metrics and its approval now describe bytes that no longer exist, the version you would roll back to may be the one overwritten, and two serving instances loading the same version can run different models with no version change to explain it.

open as a page

In an insurer's model registry, who should hold promotion authority for a claim-severity model, and what makes that approval a real control?

level: principalimportance: should knowfreq 38%

basics

~20 s

Authority belongs to whoever is accountable for the outcome, and never to the author of the run. The approval is a control only when it binds one immutable version, is made against a stated basis for refusal, is recorded with identity and time, and is reversible in minutes.

open as a page