skip to content

questions

25

What must a training run's record hold before two quarterly claim-severity runs can be compared at all?

level: juniorimportance: must knowfreq 68%

answer

  1. a number without its inputs
  2. record the yardstick, not just the score
  3. snapshot identifier, commit, parameters
  4. artifact digest ties bytes to the run
  5. same evaluation set, same metric definition

basics

~20 s

A run record must pin every input and the yardstick: the dataset snapshot identifier, the code commit, the full parameter set, a content digest of the model artifact, and the evaluation set identifier with its exact metric definition. Two runs compare only when that evaluation set and metric are identical.

solid answer

~50 s

A run record is comparable when it pins every input and the yardstick, not just the score. For a claim-severity model that means the `dataset snapshot identifier` the run read, the code commit of the training job, the complete parameter set including values that came from defaults, a content digest of the model artifact the run produced, and the evaluation set identifier plus the exact metric definition and its value. Two runs are then comparable only when the evaluation set and the metric definition are identical. If one candidate was scored on the prior quarter's held-out claims and another on this quarter's, the difference in error mixes a model change with a change in the claims themselves, and the ranking means nothing. Anything a run read that is not in the record is a hidden input.

code

json · 19 lines
json
{
  "run_id": "claim-severity-2026Q3-017",
  "code_commit": "9f2c1ab",
  "environment_lock": "sha256:71d0c4...",
  "dataset_snapshot_id": "claims-severity@2026-07-01T00:00:00Z",
  "params": {
    "objective": "tweedie",
    "tweedie_power": 1.6,
    "max_depth": 8,
    "seed": 20260701
  },
  "evaluation": {
    "set_id": "holdout-2026Q2",
    "metric": "weighted_absolute_percentage_error",
    "value": 0.184
  },
  "artifact_digest": "sha256:0b3e91...",
  "started_at": "2026-07-03T09:14:22Z"
}

go deeper

for a junior

Be able to list what one training run must record: the dataset snapshot identifier, the code commit, the parameters, the artifact digest, and the evaluation set with its metric. A score on its own is not a result.

for a middle

Explain why an identical metric name does not make two runs comparable, and why the evaluation set has to be held fixed across a quarter's candidates before any error difference can be attributed to the model.

for a senior

Show how you would enforce it: the job writes the record, a missing mandatory field fails the run, and an artifact with no run behind it cannot be registered. Recording is a pipeline property, not a discipline.

for a principal

The trade-off is how much of the record is mandatory. Every required field is friction on experimentation; every optional one is the field that turns out to be missing from the run someone must defend two years later.

## What a run record is A **run record** is the durable row a tracking system writes for one execution of the training job. It is not the job's console output and not the orchestrator's task history: those describe what the machine did, while the run record describes what the experiment **was**. One run, one record, written by the job itself as it resolves each input and produces each output. An insurer retraining a claim-severity model every quarter will produce a dozen candidate runs before one is promoted: different objectives, different feature sets, different regularisation. The record is what lets someone say, months later, which of those runs produced the model that is currently deciding reserve amounts, and on what evidence it was chosen. ## The five things a comparable record must pin 1. **The dataset snapshot identifier** - not the table name, the identifier of the exact frozen set the run read. Without it, two runs reporting different error may simply have read different claims. 2. **The code version** - the commit of the training job, so the transform that produced the model is recoverable rather than remembered. 3. **The complete parameter set** - every value the run used, *including the ones that came from defaults*. A default that changes in a later version is a silent parameter change, and a record that only lists what someone typed will not show it. 4. **The artifact digest** - a content hash (SHA-256, for example) of the model bytes the run produced, so the run and the thing it produced are tied together by content rather than by a path that can be rewritten. 5. **The evaluation set identifier and the metric definition** - which held-out claims were scored, by exactly which formula (weighted absolute percentage error over the held-out quarter, say), and the resulting value. ## Why comparability is a stricter bar than logging | what the record has | what it answers | what it still cannot answer | |---|---|---| | a metric value | how this run scored | whether another run's score means anything beside it | | value + metric definition | what was measured | whether it was measured on the same claims | | + evaluation set identifier | comparable **within** one evaluation set | which input change moved the number | | + snapshot identifier, commit, parameters | which change moved the number | which bytes are actually serving | | + artifact digest | the whole chain, run to served model | - | Each row is a strictly stronger claim than the one above it. Most teams stop at row one or two and are then surprised that their quarterly leaderboard cannot be defended. ## The quarterly trap The most common comparability mistake in a quarterly retrain is to score each quarter's candidate on **that quarter's own** held-out claims and then compare the error to the number recorded three months ago. Claim severity is not stationary: a hail season, a change in repair costs or a reserving-policy change moves the error on its own. The comparison therefore confounds the model change with a change in the underlying claims, and it moves in whichever direction the data moved. Two honest forms exist, and they answer different questions: - **Which candidate is better?** Score every candidate of this quarter on one fixed evaluation set with one metric definition. The set is held constant; only the model varies. - **Is the model getting worse over time?** Re-score the *currently promoted* model on the new quarter's evaluation set, alongside the candidate. Now the set varies and the model is held constant for one of the two readings, so the drift and the improvement are separable. ## What a partial record costs - **No snapshot identifier**: an error change cannot be attributed to the model rather than the data. - **No artifact digest**: you can say which run looked best, but not prove which bytes it produced. - **No parameters, or only the ones typed by hand**: the run cannot be re-run, and a moved default looks like noise. - **No evaluation set identifier**: the leaderboard is a list of numbers measured on unknown ground. - **A record a person filled in afterwards**: it captures intent, not what ran, and silently misses the value someone changed for one rerun. ## Making it hold in practice - The **job** writes the record, at the moment it resolves each input, and a mandatory field that cannot be resolved **fails the run** instead of logging a warning. - The record is written before an artifact can be registered: an artifact with no run behind it is unpromotable by construction. - Metric definitions are named and versioned, so one label cannot quietly mean two formulas in two quarters. - The record is queryable, not just readable. The questions it has to answer later - which registered versions read a given snapshot, which runs used a given commit - are queries, and a per-run text file cannot serve them.

  • A data contribution used in past training is later found untrustworthy - what in the run records finds every affected model version?
    Query the run records for every run whose dataset snapshot identifier covers that contribution, follow each run to the artifact digest it produced, and then to any registered model version pointing at that digest. That set is the blast radius, and the retained earlier versions are what you can fall back to. Without the snapshot identifier on the run, the same question becomes a forensic guess.
  • Should the record hold metrics measured on the training data as well as on the evaluation set?
    Yes, clearly labelled as separate fields. The training-fit metric is what later tells you whether a poor evaluation number came from underfitting or from the evaluation set moving underneath you. Only the evaluation metric on a fixed set is comparison ground, though - an unlabelled single metric field is how a run's number becomes unusable a quarter later.
  • Who writes the run record, the training job or the engineer?
    The job, automatically, as it resolves each input and writes each output. A record filled in afterwards records what someone believed they ran; it misses the parameter changed for one rerun and the default that moved with a dependency upgrade. Make the job fail when a mandatory field cannot be resolved, so an unrecorded run never becomes a candidate.

A lab notebook entry that records only the reading is useless: the next person needs the sample, the instrument settings and the scale it was read on before that number means anything.

saying these in an interview costs you the question

  • Treats a metric value alone as enough to compare two runs
  • Records hyperparameters but not which dataset snapshot the run read
  • Compares this quarter's error against last quarter's recorded number
  • Expects to find the right artifact later by filename or timestamp
  • Calls the orchestrator's console log a run record
  • Logs only the parameters someone typed, not the resolved defaults
open as a page

In a content-moderation labeling operation, why is each reported post judged by three annotators instead of one?

level: juniorimportance: must knowfreq 60%

basics

~20 s

One verdict on a moderated post carries no error signal of its own. Three verdicts make disagreement visible, route the contested post into an adjudication path, and turn agreement into a running health check on the written guideline.

open as a page

Why can a model trained from 'the latest transcript export' not be rebuilt by rerunning the same code?

level: juniorimportance: must knowfreq 62%

basics

~20 s

'Latest' is a moving pointer, not a version. Between the two runs utterances were appended, transcripts corrected and recordings withdrawn, so the second job trains on a different set. Rebuilding needs an immutable snapshot identifier that resolves to fixed bytes.

open as a page

A nightly ad job keeps every click and one non-click in a hundred — what must ship with that dataset to keep scores calibrated?

level: middleimportance: must knowfreq 62%

basics

~20 s

The negative sampling rate, recorded per stratum and per build. Thinning non-clicks a hundredfold multiplies the odds by about a hundred, and without that factor nobody can convert the model's scores back into real click probabilities.

open as a page

A retraining graph's aggregate step is retried after a timeout and one region's daily units double — what went wrong?

level: middleimportance: must knowfreq 68%

basics

~20 s

The step appends. Scheduled steps run at least once, so a retry re-executed work whose rows were already written, and the partition now holds two copies. A step must replace its declared output, not add to it.

open as a page

A click can arrive a day after the impression — when may the nightly job label that impression a non-click?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Only after the impression's attribution window has fully closed. Until then a click can still attach, so the assembly job must lag the log by at least the window length or it labels recent impressions as non-clicks that simply have not been clicked yet.

open as a page

When a claim-severity model must be rolled back to the previously promoted version, what must the registry already hold for that to take minutes?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Rollback is fast only when the previous version is still fully resolvable: retained artifact bytes under a verified digest, the feature-definition and input-schema version it was trained against, its environment lock, and a promotion pointer that serving resolves at load - so the act is a pointer flip, not a rebuild.

open as a page

A moderation team can hand-label 20,000 posts a week while the platform receives 4,000,000. Which posts should reach a human?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Humans can touch 0.5 percent of the volume, so spend it where it buys most: high-precision programmatic rules decide the bulk, uncertainty sampling sends the genuinely ambiguous band to annotators, and a small random slice re-reviews the automatically decided stream.

open as a page

A six-hour fit step in a nightly forecasting graph dies near hour five — what must its checkpoints record to resume?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Enough to continue rather than restart: the model and optimiser state, the exact position in the input, the random-number state, and the identity of the inputs and parameters being fitted — each published atomically and retained alongside the previous good checkpoint.

open as a page

Recordings from one contributor are found corrupted months later - what lineage lets you name every model trained on them?

level: seniorimportance: must knowfreq 58%

basics

~20 s

You need edges recorded at write time: utterance to ingest batch to shard, shard to snapshot manifest, snapshot to model version, model version to deployment window. Walk them backwards from the bad batch and the affected models enumerate themselves; without them the blast radius is 'everything since March'.

open as a page

Why does a click-prediction model need the ad server's impression log and not only its click log?

level: juniorimportance: should knowfreq 50%

basics

~20 s

A click log holds positives only. The negative examples are the ads that were shown and not clicked, so the impression log is what decides which training rows exist at all and what the base click rate is.

open as a page

Why is a nightly retraining run for store demand forecasts built as a graph of declared steps rather than one long script?

level: juniorimportance: should knowfreq 50%

basics

~20 s

A graph makes each step's inputs and outputs explicit, so the run gains restart points: a failed step reruns alone, unchanged branches are skipped, independent steps fan out in parallel, and progress is visible per step instead of per run.

open as a page

Why can a claim-severity training run with its parameters, metrics and code commit recorded still fail to reproduce?

level: middleimportance: should knowfreq 57%

basics

~20 s

Because the execution environment is an unrecorded input. A commit pins your code, not the code you depend on: dependency ranges resolve forward, defaults move between library versions, and unpinned seeds and thread counts change results. Record a resolved dependency lock and an environment digest with the run.

open as a page

In a moderation annotation queue, what do seeded gold questions with known answers measure that an agreement score cannot?

level: middleimportance: should knowfreq 42%

basics

~20 s

Gold questions score one annotator against a verdict already known to be right. An agreement score only compares annotators with each other, so a pool that has all drifted onto the same misreading still scores as agreeing.

open as a page

In a 40 TB transcription corpus, why does a dataset snapshot record a manifest of content hashes instead of copying the files?

level: middleimportance: should knowfreq 50%

basics

~20 s

Copying forty terabytes per experiment is unaffordable and unverifiable. A snapshot is instead a manifest naming each shard by the hash of its bytes, so unchanged shards are stored once, a new version costs a manifest rather than a copy, and any reader can verify what it loaded.

open as a page

You backfill 90 days of ad training rows with a 7-day click count computed from today's full log — what breaks?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Every historical row's count includes clicks that happened after its own impression, so the feature is computed after the outcome it predicts. Offline metrics jump and the gain vanishes online, because serving can only see the past.

open as a page

A rerun overwrote the stored artifact of an already-promoted claim-severity model in place - what breaks?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Every claim about that promoted version becomes false. Its recorded metrics and its approval now describe bytes that no longer exist, the version you would roll back to may be the one overwritten, and two serving instances loading the same version can run different models with no version change to explain it.

open as a page

Inter-annotator agreement on a moderation queue collapses the week after a policy edit. What does that drop indicate?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Suspect the guideline before the people: a freshly edited clause that two careful annotators can read differently shows up immediately as disagreement. The diagnostic is locality, whether the drop sits only in the categories the edit touched.

open as a page

A retraining graph rerun after a pricing fix reused a cached aggregate step, so the fix never reached the model — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Step caching skips a step whose declared inputs and code are unchanged. The aggregate step read the pricing table without declaring it, so the fix was invisible to the cache key and the branch was reused instead of recomputed.

open as a page

In an insurer's model registry, who should hold promotion authority for a claim-severity model, and what makes that approval a real control?

level: principalimportance: should knowfreq 38%

basics

~20 s

Authority belongs to whoever is accountable for the outcome, and never to the author of the run. The approval is a control only when it binds one immutable version, is made against a stated basis for refusal, is recorded with identity and time, and is reversible in minutes.

open as a page

A speaker withdraws consent after three published models trained on their recordings - how do you honour the erasure?

level: principalimportance: should knowfreq 36%

basics

~20 s

You cannot delete the bytes and keep every frozen snapshot byte-rebuildable. Remove the recordings, mint new snapshots without them, and keep the old ids as tombstoned manifests plus a counted suppression list - so a rebuild is exact except for a named, recorded exclusion.

open as a page

For a multi-billion-row impression log, when does reservoir sampling beat stratified sampling for an assembly job?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

When you need a faithful, unbiased miniature of the log and cannot pre-count or pre-key it — a single streaming pass of unknown length. It is the wrong pick when a rare class must be over-represented, which is the usual case for click data.

open as a page

Why do user reports and appeal reversals, used as moderation training labels, teach a model only one side of its errors?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Each signal exists on only one side of the decision. An appeal can follow a removal, so reversals expose wrongly removed posts, while a correctly approved post and a wrongly approved post both generate exactly the same silence.

open as a page

How do you run a ninety-day forecast backfill on the same batch capacity as tonight's retraining graph without starving it?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Fan the backfill out into independent per-day, per-region tasks shaped like the nightly steps, run them under a concurrency cap sized from capacity the nightly graph does not need, and keep per-task completion durable so it resumes where it stopped.

open as a page

A new per-utterance field and a changed audio sample rate land mid-corpus - what keeps last quarter's snapshot readable?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Each snapshot declares its own schema version in its manifest, and readers resolve the schema from the snapshot rather than from current code. Added optional fields stay backward-readable as explicit absences; a changed meaning under an unchanged name is not evolution and needs a new field.

open as a page