In a resumable multipart upload, how do checksums prove the assembled object matches the file the client read from disk?
answer
- TLS covers the wire only
- two levels of verification
- resend just the bad part
- hash of hashes
- part boundaries change the value
basics
~20 sThe client hashes each part and the whole file; the store verifies every part on arrival and rejects mismatches, and after assembly the backend checks the whole-file hash before committing, so corruption between disk and store is caught.
solid answer
~50 sTLS protects bytes on the wire, but not mistakes before encryption or after decryption: a wrong read offset, a reused buffer, a file edited mid-upload, or a faulty intermediary that terminates TLS. So integrity is checked at two levels. Each part carries its own checksum in a header, either the standard `Content-Digest` field from HTTP Digest Fields or a store-specific one; the store recomputes it and rejects a mismatch, and the client resends only that part. End to end, the client computes a whole-file `SHA-256` while reading and records it in the pending upload record; after assembly the backend verifies it before committing. The catch is that many stores report a **composite** checksum, a hash of the part hashes, which depends on part boundaries and never equals the full-file `SHA-256`, so compare like with like.
code
json · 9 lines{
"uploadId": "up_81f3",
"status": "pending",
"sizeBytes": 4294967296,
"partSizeBytes": 8388608,
"partCount": 512,
"fileSha256": "3b1f...",
"compositeSha256": "a90c...-512"
}go deeper
Recall that a checksum is a fingerprint of the bytes, and that uploads check one for each part and one for the whole file.
Explain what each level catches and what happens on a mismatch, and why the part token returned by the store ties completion to verified bytes.
Show that you know the traps: TLS does not cover client bugs or a file edited mid-upload, and a composite checksum cannot be compared with a full-file hash.
Weigh the cost of recomputing full hashes server-side for every object against trusting client-computed composites, and decide where in the pipeline integrity is enforced.
## Where corruption actually comes from A common assumption is that **TLS** makes integrity checks redundant. TLS does authenticate every record on the wire, so a bit flipped between the phone and the store is detected and the connection fails. But an upload has many places outside that protected tunnel: - **Client bugs**: reading the wrong offset for a part, reusing a buffer, or truncating the last part. - **A file that changes during upload**: the user edits or re-saves a video while parts are in flight, so early and late parts come from different versions. - **Intermediaries that terminate TLS**, such as corporate proxies, which decrypt and re-encrypt the bytes. - **Server-side faults** during storage or assembly. A byte count alone does not catch any of these: corrupted data usually has exactly the expected size. ## Two levels of checksum | Level | Computed by | Checked when | Catches | On mismatch | |---|---|---|---|---| | **Per-part** | client, per part | as the part arrives | a bad part in transit or in the client buffer | store rejects; client resends that part | | **Whole-file** | client, over the full file | after assembly, before commit | wrong offsets, missing or swapped parts, file changed mid-upload | backend refuses to commit; upload restarts | **Per-part checksums** make failures cheap: the store computes the hash of what it received and compares it with the value sent in a header. HTTP has a standard for this, the Digest Fields headers (`Content-Digest` and `Repr-Digest`), and object stores also define their own headers. The part token returned on success then binds exactly those verified bytes, and the complete call references those tokens. **Whole-file checksums** close the end-to-end gap. The client streams the file through `SHA-256` as it reads parts and sends the digest when it starts the upload, so the backend stores it in the `pending` record. When the upload completes, the backend compares it with the stored object before flipping the record to `committed`. ## The composite checksum trap Many object stores do not compute a plain hash of an assembled multipart object. Instead they report a **composite checksum**: 1. Hash each part: h1, h2, ..., hn. 2. Concatenate those digests. 3. Hash the concatenation, often tagging the result with the part count. This value depends on **where the part boundaries fall**. The same 4 GiB file uploaded with 8 MiB parts (512 parts) and with 16 MiB parts (256 parts) produces two different composite values, and neither equals the `SHA-256` of the whole file. Two ways to compare correctly: - **Match the method**: the client computes the composite the same way, with the same part size, and records that too. - **Recompute server-side**: a processing worker streams the assembled object and computes the full-file `SHA-256`, at the cost of reading the whole object once. ```json { "uploadId": "up_81f3", "status": "pending", "sizeBytes": 4294967296, "partSizeBytes": 8388608, "partCount": 512, "fileSha256": "3b1f...", "compositeSha256": "a90c...-512" } ``` Here `sizeBytes` is 4 GiB and `partSizeBytes` is 8 MiB, which gives exactly 512 parts. ## Detecting a file that changed mid-upload Per-part checksums cannot catch this case alone: every part is internally consistent, it just came from a different version of the file. Useful defences: - Record size, modification time and the whole-file hash at the start; re-check them before sending the complete call. - For files the app does not own, copy or snapshot them before uploading, so the source cannot change underneath. - If anything changed, abort the session and start again rather than committing a mixed object. ## Cost and when it matters Hashing is cheap next to network time: a streaming `SHA-256` pass over a few gigabytes typically takes seconds to tens of seconds on a modern phone. Integrity checking is still a differentiator rather than a screening topic, and interviewers who ask about it are usually probing whether the candidate knows that TLS is not end-to-end for application bugs, and whether they know composite checksums are not full-file hashes.
- Why can't the backend compare the store's composite multipart checksum with the client's full-file SHA-256?A composite value is a hash over the list of part digests, so it depends on where the part boundaries fall; the same bytes split differently give a different composite, and neither equals the hash of the concatenated bytes. Compare like with like: have the client compute the composite with the same part size, or have a worker stream the assembled object and compute the full-file hash.
- What should happen when the file on the phone is modified while its parts are uploading?Committing would produce an object that matches neither version, since early and late parts come from different contents. The client should record size, modification time and hash at the start, re-check them before completing, and abort and restart the session if anything changed. Snapshotting the file before upload avoids the race entirely.
saying these in an interview costs you the question
- TLS already guarantees the stored object equals the file on disk.
- The composite checksum equals the SHA-256 of the whole file.
- Per-part checksums are redundant if the whole file is hashed.
- A matching byte count proves the upload is intact.
- Per-part checks also catch a file edited during the upload.