skip to content

Recordings from one contributor are found corrupted months later - what lineage lets you name every model trained on them?

level: seniorimportance: must knowfreq 58%

answer

  1. walk the graph backwards
  2. rows, shards, snapshots, models, deployments
  3. edges written at read time
  4. granularity decides the blast radius
  5. derived artifacts are a second hop

basics

~20 s

You need edges recorded at write time: utterance to ingest batch to shard, shard to snapshot manifest, snapshot to model version, model version to deployment window. Walk them backwards from the bad batch and the affected models enumerate themselves; without them the blast radius is 'everything since March'.

solid answer

~50 s

Lineage is a graph, and containment is a **reverse** traversal of it. The forward edges each job writes are: the ingest run that produced a shard and the contributor and batch attributes carried on its rows; the manifest that lists that shard hash; the model version that resolved that snapshot id; and the deployment records that served that model version, with their start and end times. Given a suspect batch you resolve its shard hashes, find every manifest referencing them, find every model version built from those manifests, and intersect with what is still serving or still consumed downstream. Two properties decide whether this works: the edges must be written **by the job that read the data**, because the resolution of a name to bytes is not recoverable from the artifact afterwards; and the provenance must be recorded at the **granularity you will be asked about** - per contributor and per ingest batch - since an attribute you never stored cannot be re-derived.

code

sql · 7 lines
sql
SELECT m.model_version, d.environment, d.served_from, d.served_to
FROM   shard         s
JOIN   snapshot_shard ss ON ss.shard_hash   = s.shard_hash
JOIN   model_version  m  ON m.snapshot_id   = ss.snapshot_id
LEFT JOIN deployment  d  ON d.model_version = m.model_version
WHERE  s.ingest_batch = :suspect_batch
ORDER  BY d.served_from;

go deeper

for a junior

Recall the chain in order - utterance, ingest batch, shard, snapshot, model, deployment - and that it must be readable in both directions.

for a middle

Explain why the reverse traversal is the one that matters, and why a date range is a poor substitute that both over- and under-scopes.

for a senior

Demonstrate incident sequencing: freeze the input first, enumerate from the edges, intersect with what is live, then measure impact on a clean evaluation set.

for a principal

Set the policy: which attributes every row must carry forever, who is accountable for writing edges, and what containment latency the organisation commits to.

## Containment is a reverse query Every data-quality incident - a supplier delivering misaligned audio, a mis-configured re-encoder, deliberately tampered rows - asks the same question in the same direction: *given these bad rows, which models ate them, and which of those are live?* Lineage is only useful if it can be walked **backwards**. A system that can say 'this model was trained on snapshot X' but cannot say 'snapshot X is one of the forty that contain shard Y' answers the wrong direction and leaves you scoping by date. ## The edges, and who writes each one | edge | written by | without it | |---|---|---| | utterance to ingest batch and contributor | the ingest job, as row attributes | you cannot select the suspect rows at all | | ingest batch to shard hashes | the shard writer | you know the rows are bad but not which shards hold them | | shard hash to snapshot id | the manifest, on publish | you cannot find the datasets that included them | | snapshot id to model version | the training job, when it resolves the id | the models look unrelated to any dataset | | model version to deployment window | the release path | you know which models are tainted but not which are live | The critical property is that each edge is written **at the moment the fact is true**. A training job resolves a name to bytes, a branch to a commit, a mount to a storage location; by the time an auditor asks, the artifact records none of it. Reconstructed lineage is a guess dressed as a record. ## Granularity decides the blast radius Provenance recorded at corpus level answers only 'yes, the corpus contained it'. Since almost every model reads almost the whole corpus, every model then looks affected, and containment over-scopes to the entire period. Recording the ingest batch and the contributor as **row attributes**, and indexing shards by the batches they contain, is what lets you *exclude* models - which is the expensive half of an incident. Per-utterance provenance is cheap as a column and expensive as an index; batch-level indexes with row-level attributes is the usual middle. ## Two hops people forget 1. **Derived datasets.** A pronunciation lexicon, a normalisation table or a distilled student model built from the tainted model's outputs are downstream nodes. If derivation edges are not recorded, the contamination outlives the rollback of the direct consumers. 2. **Evaluation sets.** If the bad rows also reached a held-out set, the measurements that would tell you the impact are themselves suspect, and the first honest step is re-measuring on a clean set. ## Ordering the response 1. **Freeze the input** - stop new snapshots from including the suspect batch before anything else, or the blast radius keeps growing while you investigate. 2. **Enumerate** - reverse-walk the edges to the full candidate list of model versions. 3. **Intersect with reality** - which of those are serving now, which feed a downstream artifact, which are merely archived. 4. **Quantify** - measure the affected models on a clean evaluation set; a small corrupted fraction may be immaterial, and the graph tells you *which* models to measure, not *whether* they are hurt. 5. **Record the incident against the snapshot ids** so the next reader of those ids sees the finding. ## What good and weak answers sound like A weak answer reaches for logs and dates: 'we would find when it arrived and retrain everything after that'. It over-scopes, and it still misses jobs that read older shards containing the same batch. A good answer names the edges, insists they are written by the reader at read time, and picks the granularity deliberately - then adds that retraining a replacement does **not** by itself take the affected model out of service; that is a separate, explicit step in the serving path.

  • How fine does provenance have to be - per utterance, per ingest batch, or per corpus?
    Fine enough to answer the questions you are obliged to answer, which in a contributed corpus means at least contributor and ingest batch, carried as row attributes so any future selection is possible. Index shards by the batches they contain; per-utterance indexes are usually reserved for the subject-level lookups that consent and erasure require.
  • Why must lineage edges be written by the job that reads the snapshot rather than inferred afterwards?
    Because name-to-bytes resolution happens only at read time: which id a name pointed at, which commit was checked out, which mount was used. None of that survives in the produced model, so a later reconstruction is inference over timestamps, and it fails exactly when the timeline is contested.
  • The corrupted batch also reached the evaluation sets. What changes?
    The measurements you would use to size the impact are compromised, so impact assessment moves first: rebuild a clean evaluation set from unaffected shards and re-measure the candidate models on it. Until then, any claim that the contamination was immaterial rests on numbers computed from the contamination.

saying these in an interview costs you the question

  • Believes the training run's timestamp is enough to identify its data
  • Plans to reconstruct lineage after the incident from logs
  • Records provenance only at corpus level, so every model looks affected
  • Stops the graph at the model and never reaches what is serving
  • Assumes training a clean replacement removes the affected model from service
  • Forgets artifacts derived from the tainted model's own outputs