skip to content

In a 40 TB transcription corpus, why does a dataset snapshot record a manifest of content hashes instead of copying the files?

level: middleimportance: should knowfreq 50%

answer

  1. name the bytes, not the place
  2. pointers, not forty terabytes
  3. hash per shard, hash the list
  4. unchanged shards shared across versions
  5. Merkle DAG; verify on read

basics

~20 s

Copying forty terabytes per experiment is unaffordable and unverifiable. A snapshot is instead a manifest naming each shard by the hash of its bytes, so unchanged shards are stored once, a new version costs a manifest rather than a copy, and any reader can verify what it loaded.

solid answer

~40 s

Each shard is addressed by a hash of its contents - SHA-256 over the bytes - and a snapshot is an ordered manifest of those hashes with per-shard row counts and sizes. Hashing the manifest yields the snapshot id, so the whole structure is a Merkle DAG: one changed utterance changes one shard hash, which changes the manifest, which changes the snapshot id, and every other shard is shared unchanged with every earlier snapshot. That gives three things a copy cannot: **storage proportional to what changed** rather than to the corpus, **verification on read** because the reader re-hashes and compares, and **a cheap diff** between two versions as a set difference over hashes. The price is a catalogue, a reference-counted garbage collector, and the discipline that an id is never re-pointed.

code

json · 10 lines
json
{
  "snapshotId": "sha256:9f2c7ae1...",
  "createdAt": "2026-09-18T04:11:00Z",
  "schemaId": "utterance/3",
  "shards": [
    { "hash": "sha256:1a0b44d9...", "utterances": 20000, "bytes": 4187593216 },
    { "hash": "sha256:7c41e802...", "utterances": 19873, "bytes": 4102338048 }
  ],
  "totals": { "utterances": 39873, "bytes": 8289931264 }
}

go deeper

for a junior

Recall the shape: a snapshot is a list of hashes plus counts, not a copy of the corpus, and the hash names the bytes rather than the place they sit.

for a middle

Explain the Merkle structure - shard hash, manifest, snapshot id - and what it buys: dedup across versions, verification on read, and a diff that is a set difference.

for a senior

Show the operational half: re-hashing on read and failing loudly, reference-counted reclamation, and the fact that re-sharding mints a new id over unchanged rows.

for a principal

Own the consequence: every retained id pins storage, so retention, reproducibility windows and cost are decided together rather than discovered in conflict.

## Content addressing in one paragraph Addressing by **location** means a name points at a place - a path, a bucket key, a table - and the bytes at that place can change without the name changing. Addressing by **content** inverts it: the name *is* a cryptographic hash of the bytes, so a given name can only ever denote one sequence of bytes. Change anything and you get a different name. Applied to a training corpus, the shard is the unit: hash each shard, and a set of utterances becomes an unambiguous list of hashes. ## What the manifest carries A snapshot manifest is a small document listing, per shard: the **content hash**, the **utterance count**, the **byte size**, and the **schema version** the shard was written under, plus snapshot-level totals and a creation timestamp in ISO 8601. Hashing that document canonically produces the **snapshot id**. The structure is a Merkle DAG - a tree of hashes where each level commits to the level below - which is what makes the id a proof rather than a label: if any shard's bytes were substituted, re-hashing on read fails the comparison. ## Why copies fail at this size | property | copy the data per experiment | manifest of content hashes | |---|---|---| | cost of a new version | the whole corpus, every time | one document, plus only the shards that changed | | shared storage | none; every copy is independent | identical shards stored once and referenced by many snapshots | | verification | compare files, or trust the copy | re-hash on read and fail on mismatch | | diff between versions | a full scan of both copies | a set difference over hash lists | | time to cut a version | hours of transfer | seconds | The cost argument alone usually decides it, but the verification argument is the one that matters after an incident: you can prove which bytes trained a model, not merely assert it. ## Five things content addressing does *not* give you - **Correctness.** A hash proves identity, not quality. A corrupted shard hashes perfectly well. - **Meaning.** Identical bytes under a changed schema interpretation are still a different dataset. - **Set equality.** The manifest commits to layout as well as content, so re-sharding or re-ordering the same rows yields a different snapshot id. Different ids do **not** prove different utterances; to compare sets you diff the hash lists or compare canonical per-row digests. - **Deletion.** Immutability is precisely the property that makes erasure and retention hard work rather than an update. - **Free storage.** Every retained id keeps its shards alive. Without a retention policy the corpus only ever grows. ## Reclaiming bytes Because shards are shared, no single job can decide a shard is finished with. Removal is a **reference question**: a shard's bytes may go when no retained manifest still references that hash, which in practice is a reference count maintained on publish and delete, or a periodic mark-and-sweep over the live manifests. The policy input is which snapshots stay retained - typically any snapshot referenced by a model that is still serving or still legally required to be explainable - and that policy, not the storage layer, is what actually determines the corpus's footprint. ## What this looks like in a design round Say the unit (a shard), the addressing scheme (hash of the bytes), the document (a manifest of hashes with counts), and the identity (a hash of the manifest). Then say what it buys - dedup, verification, cheap diffs, instant version cuts - and what it obliges you to build - a catalogue, reference counting, and a rule that an id is minted, never edited. A candidate who adds that re-sharding changes the id without changing the data has understood what the hash actually commits to.

  • Two snapshots have different ids - does that prove they contain different utterances?
    No. It proves the manifests differ. Re-sharding, re-ordering the shard list, or writing under a new schema version all change the manifest without changing the row set. To compare sets you diff the hash lists, or compare canonical per-row digests when the shard boundaries have moved.
  • When can the bytes behind a shard hash actually be deleted?
    When no retained manifest references that hash any more - a reference count kept on publish and delete, or a mark-and-sweep over live manifests. Which snapshots stay retained is a policy decision driven by which models must remain explainable; erasure is a separate and stronger requirement that overrides it.
  • Does the manifest have to hash the audio bytes, or is hashing the transcript enough?
    Everything the training job reads has to be covered, audio and transcript alike, or the id stops being a proof: re-encoding the audio under an unchanged transcript would leave the snapshot id identical while the model's inputs changed. Hash each artifact and commit to the whole set in the manifest.

saying these in an interview costs you the question

  • Thinks a dataset snapshot must be a physical copy of the data
  • Treats a timestamped folder name as a version identifier
  • Assumes hashing the file path identifies the content behind it
  • Expects the same rows re-sharded differently to hash identically
  • Believes a matching content hash proves the data is correct
  • Deletes shards per experiment without checking other references