skip to content

A new per-utterance field and a changed audio sample rate land mid-corpus - what keeps last quarter's snapshot readable?

level: seniorimportance: nice to knowfreq 31%

answer

  1. the snapshot declares its own schema
  2. add is safe, rename is major
  3. same type, changed meaning, no alarm
  4. absent is not zero
  5. per-shard statistics catch silent shifts

basics

~20 s

Each snapshot declares its own schema version in its manifest, and readers resolve the schema from the snapshot rather than from current code. Added optional fields stay backward-readable as explicit absences; a changed meaning under an unchanged name is not evolution and needs a new field.

solid answer

~50 s

Two changes arrive together and they are not the same kind of change. The **added field** is additive: declare it optional, bump the schema version the manifest records, and let readers of older snapshots represent it as **explicitly absent** rather than defaulted to zero - a zero is a real value that the trainer will learn from. The **changed sample rate** keeps the field's name and type while changing what the bytes mean, so no mechanical compatibility check fires, nothing errors, and a model trained across the boundary silently mixes two distributions; the symptom appears later as degraded word error rate on the affected population. Semantic changes have to become a new field or a declared per-shard property. And no change may be backfilled into an existing snapshot: rewriting a frozen set destroys the ability to rebuild the model that trained on it.

go deeper

for a junior

Recall the two kinds of change: adding an optional field is safe, and changing what an existing field means is not, even when the type is identical.

for a middle

Explain how the manifest's schema version makes an old snapshot readable, and why a missing field must be represented as absent rather than as zero.

for a senior

Show the detection story: per-shard statistics recorded at ingest are the only thing that catches an undeclared semantic change, and readers must fail rather than drop fields.

for a principal

Decide the contract: which changes are allowed without a major version, how long old schema versions stay supported, and who approves a boundary in the corpus.

## The schema belongs to the snapshot, not to the code The rule that makes old snapshots readable is a single inversion: the reader does not apply *its* schema to the data, it resolves the schema **the snapshot declares**. The manifest records a schema version, the schema registry holds every version ever published, and a reader loading a two-year-old snapshot interprets it exactly as it was written. The alternative - one current schema that all data must conform to - forces either backfills or breakage, and backfilling a frozen snapshot is the one thing versioning exists to prevent. ## Three kinds of change, only two of which a checker can see | change | mechanically detectable | what old snapshots need | |---|---|---| | add an optional field | yes, trivially | nothing; readers see an explicit absence | | rename, remove or re-type a field | yes | a new major schema version; old snapshots stay on the old one | | same name and type, different meaning | **no** | a new field or a declared per-shard property, plus a recorded boundary | The third row is the dangerous one, and the sample-rate change is its textbook instance. The field still holds audio, still has the same type, still passes every compatibility gate - and the model now sees two populations under one name. ## Missing is not zero When a reader meets a snapshot written before a field existed, it has three options and only one of them is honest: - **Represent it as absent** and let the trainer decide - exclude those rows, carry an explicit missingness indicator, or use a model that handles absence. Honest. - **Default it to a neutral-looking value** such as zero or an empty string. Dishonest: the trainer cannot distinguish 'no value existed' from 'the value was zero', and it learns the difference as signal. - **Backfill it into the old shards.** Worst of the three: it changes the content behind published hashes, so either the ids stop resolving or, if the ids are re-pointed, every past claim about what trained a model becomes false. ## Catching an undeclared change A change of meaning is invisible to schema comparison, so the defence is **statistical and written at ingest**: record per-shard properties - sample rate, channel count, bit depth, duration distribution, the proportion of rows carrying each optional field - in the manifest at write time, and compare them across the shards of one snapshot and across consecutive snapshots. A property that shifts *inside* a single snapshot is the alarm, and it is only visible because the writer recorded it; nothing can be recovered later from bytes that never declared what they were. ## Derived values are versioned too Normalisation statistics, the vocabulary, the feature-extraction window - anything computed **from** a snapshot inherits the snapshot's version and must be recorded alongside it rather than recomputed from the current corpus at read time. Recomputing quietly imports today's distribution into an old snapshot's rebuild, which is the same defect as a backfill wearing different clothes. ## Failing loudly A reader that encounters a schema major it does not understand must **refuse**, not drop the fields it does not recognise. Silent field-dropping is how a training job quietly loses an input it was supposed to have and produces a worse model with no error anywhere. The corresponding rule on the write side is that a new schema version is published before the first shard that uses it, so no shard ever references a schema nobody can resolve. ## The design-round answer Schema version per snapshot, resolved from the snapshot; additive fields optional and absent rather than defaulted; semantic changes get a new name and a recorded boundary; nothing is ever backfilled into a frozen snapshot; per-shard statistics written at ingest are what catch what the schema check cannot see.

  • An older snapshot lacks the new field - do you train with it absent or backfill it?
    Neither silently. Declare the field optional and let the trainer handle absence explicitly, by excluding those rows or carrying a missingness indicator. Backfilling rewrites a frozen set, so either the published id stops resolving or it now resolves to data no model was ever trained on.
  • How would you catch a sample-rate change that nothing declared?
    Record per-shard audio properties in the manifest at ingest, then compare them across the shards of one snapshot and between consecutive snapshots. A property that shifts inside a single snapshot is the signal. A schema-compatibility check never fires here, because the name and the type are unchanged.

saying these in an interview costs you the question

  • Backfills a new field into older snapshots so everything looks uniform
  • Calls a units or sample-rate change compatible because the type is unchanged
  • Fills a missing field with zero and trains on it as a real value
  • Keeps one current schema and expects old snapshots to conform to it
  • Assumes a mechanical compatibility check catches a change of meaning
  • Lets a reader drop fields it does not recognise instead of failing