skip to content

In a cloud-drive service, why are files stored as chunks keyed by the SHA-256 hash of their content rather than as whole files?

level: juniorimportance: must knowfreq 58%

answer

  1. name derived from the bytes
  2. ordered list per file version
  3. ask which hashes are missing
  4. immutable, so retries are safe

basics

~10 s

Chunking plus content addressing stores identical bytes once, lets an edited file re-upload only the chunks whose hashes changed, and turns each chunk's key into an integrity check on read.

solid answer

~40 s

The client splits each file into chunks and computes the SHA-256 of each one; that digest is the chunk's storage key, so identical bytes always map to the same key. A file version becomes a **manifest**: an ordered list of chunk hashes plus metadata. That gives three wins. **Dedup**: a chunk already stored is never stored twice. **Delta sync**: after an edit the client sends the new hash list, the server says which hashes it lacks, and only those chunks travel. **Integrity**: re-hashing a chunk on read proves the bytes are the ones the manifest names. Because the key is derived from the bytes, chunks are immutable, so uploads are idempotent and safe to retry. The cost is bookkeeping: shared chunks need reference counting before anything can be deleted.

code

json · 10 lines
json
{
  "path": "/reports/q3-model.bin",
  "version": 7,
  "size": 2147483648,
  "chunkHashes": [
    "sha256:9f2c...a41b",
    "sha256:03de...77c0",
    "sha256:e81a...5d92"
  ]
}

go deeper

for a junior

Remember the three pieces: chunks, a SHA-256 key per chunk, and a manifest listing the keys in order. Be able to say how that stores duplicates once and uploads only changed chunks.

for a middle

Walk through the delta-sync handshake step by step: chunk and hash, ask which hashes are missing, upload those, then commit the manifest. Explain why chunks are immutable and uploads idempotent.

for a senior

Show the operational consequences: server-side hash verification, the commit-after-upload ordering, and why shared chunks force reference counting and delayed garbage collection.

for a principal

Frame content addressing as a trade: large bandwidth and storage savings bought with a hash index that scales with chunk count, reference-count upkeep, and a dedup-scope decision with privacy implications.

## Why not store whole files A cloud-drive service keeps users' files in sync across devices. The naive design stores each file as one object and re-uploads the whole object whenever it changes. That fails on three fronts: - **Bandwidth**: a user who edits one paragraph of a 2 GiB file re-sends 2 GiB. - **Storage**: the same installer, photo or shared document saved by many devices or versions is stored many times. - **Integrity**: a single checksum over a 2 GiB object tells you *that* something is corrupt, not *where*. ## Chunks, hashes and manifests The standard fix splits every file into **chunks** — contiguous byte ranges, commonly a few KiB to a few MiB each — and names each chunk by a cryptographic digest of its bytes, typically **SHA-256**. This is **content-addressed storage**: the key is derived from the value, so identical bytes always get the same key, and different bytes in practice never do. A file version is then described by a **manifest**: an ordered list of chunk hashes plus metadata such as the total size. Reading the file means fetching each chunk by its hash and concatenating them in order. | Concept | What it is | Mutable? | |---|---|---| | Chunk | a byte range kept in an object store | no — its key is its content | | Chunk key | `SHA-256(chunk bytes)` | no | | Manifest | ordered list of chunk keys for one file version | no — each version gets a new one | | Metadata record | path, owner, pointer to the current manifest | yes | Because a chunk's key is its content, a chunk is **immutable**: there is no such thing as updating chunk X, only writing a new chunk under a new key. Immutability makes uploads **idempotent** — retrying a chunk upload writes the same bytes under the same key — and lets chunks be cached anywhere indefinitely. ## Delta sync: uploading only what changed When a user saves an edited file, the client: 1. Splits the new version into chunks and hashes each one. 2. Sends the ordered hash list to the server and asks which hashes it does not already hold. 3. Uploads only the missing chunks; the server recomputes each chunk's SHA-256 and rejects any mismatch. 4. Commits the new manifest once every referenced chunk is durably stored. With a chunker whose boundaries stay stable across edits (content-defined chunking), a paragraph edit changes one or two chunks. If the 2 GiB file averages 4 MiB per chunk it has about 512 chunks, and the sync uploads roughly 4-8 MiB instead of 2 GiB. The order in step 4 matters: a manifest committed before its chunks exist describes a file nobody can read. ## What content addressing buys beyond bandwidth - **Deduplication**: a chunk already present in the dedup scope is never stored twice, whether the duplicate comes from an older version of the same file, a copy in another folder, or — if the scope allows it — another user. - **Verification on read**: re-hashing a fetched chunk proves it is exactly the bytes the manifest names, so silent corruption is detected per chunk and can be repaired from another replica. - **Cheap versioning**: keeping the previous version costs only the chunks that differ, because both manifests share the rest. - **Resumable sync logic**: a half-finished sync can pick up again by asking which hashes are still missing. ## The costs you take on - **Shared chunks cannot be deleted naively.** Deleting a file must not delete chunks another manifest still uses, which is why deduplicating stores keep **reference counts** and run garbage collection with a safety delay. - **An index lookup per chunk.** Every has-check and every read goes through a hash-to-location index that grows with the number of chunks. - **Trusting the key.** The server must verify the digest of uploaded bytes itself; storing bytes under a client-supplied key without checking lets one faulty client corrupt every file that references that key. - **Dedup scope is a design decision.** Deduplicating across users saves more storage but can open a privacy side channel, so many designs limit what the client can observe. ## Summary Chunking turns a file into many small immutable pieces; content addressing gives each piece a name derived from its bytes; the manifest ties them back together in order. Together they make edits cheap to sync, duplicates cheap to store and corruption detectable — at the price of reference counting and an index that must scale with the chunk count.

  • Why must the client commit the new manifest only after every missing chunk is confirmed stored?
    A manifest is the only thing a reader follows. If it is committed first and a chunk upload then fails, the current version of the file references bytes that do not exist, and every device that syncs it gets an unreadable file. So chunks are stored durably first, and the server verifies at commit time that every hash in the manifest is present before accepting it.
  • Why does the server recompute each chunk's SHA-256 instead of trusting the hash the client sent?
    Every manifest that contains a hash trusts that the bytes stored under it match. A buggy or malicious client that stores wrong bytes under a popular hash would silently corrupt every file referencing it, including other versions and, under a wider dedup scope, other users' files. Recomputing on receipt and rejecting mismatches keeps the key honest.

It is like a recipe book that lists ingredients by catalogue number instead of reprinting each ingredient's description: identical ingredients share one catalogue entry, and a revised recipe only adds entries for new ingredients.

saying these in an interview costs you the question

  • Chunking exists mainly so an interrupted upload can resume.
  • The server can store chunks under client-supplied hashes without checking them.
  • Deleting a file can immediately delete all of its chunks.
  • A single whole-file hash is enough to sync only the changed parts.
  • Updating a file means overwriting its existing chunks in place.