skip to content

In a cloud-drive service, why should an upload's metadata move from pending to committed only after the object store confirms the complete file?

level: seniorimportance: should knowfreq 50%

answer

  1. no transaction spans both stores
  2. what does a crash leave
  3. invisible but findable state
  4. conditional status transition
  5. sweeper plus expiry

basics

~20 s

Committing first lists files whose bytes are missing; storing bytes with no record leaves untracked orphans. A pending record before upload, then a verified, conditional flip to committed afterwards, keeps listings truthful and leftovers findable.

solid answer

~50 s

The object store and the metadata database share no transaction, so the order of writes decides what a crash leaves behind. At grant time the backend writes a `pending` row: owner, server-chosen object key, expected size and hash, and an expiry; it is hidden from listings but can reserve quota. The bytes land. On a completion signal the backend reads the stored object, checks size and checksum, and runs `UPDATE ... SET status = 'committed' WHERE status = 'pending'`, so duplicate or late signals change nothing. Only committed rows appear in listings, shares and sync. Every failure then leaves a pending row with missing, partial or unverified bytes, which a sweeper can find: it expires stale rows, aborts unfinished multipart sessions and deletes their objects. Replacing a file writes a fresh key and switches the pointer at commit, so readers never see a half-written version.

code

sql · 6 lines
sql
UPDATE uploads
SET status = 'committed', committed_at = CURRENT_TIMESTAMP
WHERE upload_id = :upload_id
  AND status = 'pending'
  AND expected_sha256 = :observed_sha256;
-- 0 rows updated: already committed, expired, or checksum mismatch

go deeper

for a junior

Remember the rule of thumb: create a hidden pending record before the upload, and make the file visible only after the storage confirms every byte arrived.

for a middle

Explain what each alternative ordering leaves behind after a crash, and why a pending row plus a verified commit is the only one where leftovers are both invisible and findable.

for a senior

Demonstrate operational depth: idempotent conditional commits, verification against the store rather than the message, a sweeper with store-side backstops, and version replacement through a fresh key.

for a principal

Discuss where quota is reserved, how long pending uploads may live, and how the commit feeds downstream processing reliably, for example through an outbox in the same transaction.

## Two stores, no shared transaction A cloud-drive service keeps **bytes** in an object store and **metadata** (name, folder, owner, size, version, sharing) in a database. A user expects one thing: a file is either in their drive and fully readable, or it is not there. But no transaction spans both systems, and the upload itself happens between them over minutes. The design question is therefore: *in what order do we write, so that every crash leaves a state we can detect and repair?* ## The failure matrix | Ordering | What a crash can leave | Consequence | |---|---|---| | Metadata committed first, then bytes | a visible file with no or partial bytes | users list, share or sync a file that fails to open | | Bytes first, metadata written only at the end | bytes with no record at all | orphans nobody tracks; storage cost leaks; no quota or ownership | | **Pending row first, bytes, then verified commit** | a pending row with missing, partial or unverified bytes | invisible to users; findable and cleanable by a sweeper | Only the third ordering makes **every** intermediate state both harmless and discoverable. ## The pending -> committed lifecycle 1. **Grant.** The API authorizes the request and inserts a row with `status = 'pending'`, the owner, a **server-chosen object key**, the expected size and `SHA-256`, and an `expires_at`. Quota can be reserved here so a burst of uploads cannot overrun it. 2. **Upload.** The client sends bytes directly to the object store, often as a multipart session. 3. **Completion signal.** The client calls a `complete` endpoint and/or the store emits a notification. 4. **Verify.** The backend reads the stored object's actual size and checksum from the store. It never trusts the size or hash inside the message alone. 5. **Commit.** A conditional update moves the row to `committed`. Listings, shares, search and multi-device sync read only committed rows. 6. **Process.** The commit (or an outbox row written in the same transaction) triggers scanning, thumbnails and other downstream work. ```sql UPDATE uploads SET status = 'committed', committed_at = CURRENT_TIMESTAMP WHERE upload_id = :upload_id AND status = 'pending' AND expected_sha256 = :observed_sha256; -- 0 rows updated: already committed, expired, or checksum mismatch ``` ## Idempotent completion Completion signals are messy in practice: - The client's `complete` call and the store's notification can both arrive, in either order. - Notifications are commonly delivered **at least once**, so duplicates happen. - A signal can arrive after the sweeper has already expired the row. The conditional `WHERE status = 'pending'` handles all three: the first valid signal wins, and every later one updates zero rows and simply returns the current state. Because verification reads the store rather than the message, a stale or duplicated message cannot commit bad data. ## Cleaning up abandoned uploads Users close apps, lose signal and never come back. Without cleanup, stored parts and objects accumulate and cost money forever. A **sweeper** job runs periodically: - Select `pending` rows whose `expires_at` has passed. - Abort the associated multipart session so its stored parts are released. - Delete any object already written at the key. - Mark the row `expired` (or delete it) and release reserved quota. Two backstops make this robust: a **store-side lifecycle rule** that removes incomplete multipart sessions after a few days, in case the sweeper misses one, and an occasional **reverse audit** that lists objects under the incoming prefix older than the longest allowed upload and deletes any with no matching row. The sweeper and a late completion can race; the conditional update resolves it, because only one of `pending -> committed` and `pending -> expired` can succeed. ## Replacing an existing file Overwriting the live object in place would let readers observe a half-written or mixed file during the upload, and a failed upload would destroy the old version. Instead: 1. Upload the new content to a **fresh key** under its own pending row. 2. On commit, switch the file's current-version pointer to the new key in one database update. 3. Garbage-collect the old key later, according to version-retention rules. Readers therefore always see either the complete old version or the complete new one. ## What interviewers listen for That the candidate notices the missing transaction, picks the ordering whose failures are invisible and recoverable, makes the commit idempotent, verifies against the store, and budgets for cleanup rather than assuming uploads always finish.

  • How do you clean up cloud-drive uploads that are started but never completed?
    A periodic sweeper selects `pending` rows past their expiry, aborts the multipart session, deletes any object at the key, marks the row expired and releases reserved quota. A store-side lifecycle rule that drops incomplete sessions after a few days is a backstop, and an occasional audit deletes incoming objects that have no matching row.
  • What happens if the store's completion notification arrives twice, or after the client's own complete call?
    Nothing extra: the commit is a conditional update from `pending`, so the first valid signal wins and every later one updates zero rows and returns the current state. Each handler re-reads the object's size and checksum from the store, so a duplicated or stale message can never commit unverified data.
  • Why not overwrite the existing object in place when a user uploads a new version?
    Readers could see a partially written or mixed file during the upload, and a failed upload would destroy the previous version. Writing to a fresh key and switching the current-version pointer at commit means readers always get a complete version, old or new, and the old key is garbage-collected later.

saying these in an interview costs you the question

  • Insert the file row as visible first; the bytes will follow shortly.
  • Write metadata only after the upload; nothing is needed beforehand.
  • Wrap the object write and the database insert in one transaction.
  • Trust the size and hash in the completion message without checking the store.
  • Overwrite the live object in place when a new version arrives.
  • Abandoned uploads cost nothing, so a cleanup job is optional.