skip to content

In a cloud-drive service, why is client-side deduplication across all users' files a privacy risk?

level: seniorimportance: should knowfreq 35%

answer

  1. a skipped upload is observable
  2. does anyone else have this file
  3. templated form, guessed field
  4. hash is not a secret

basics

~10 s

If the client skips uploading content any user already stored, the skipped upload reveals that someone else holds that exact file, letting an attacker confirm a document exists or brute-force its secret fields.

solid answer

~50 s

Client-side cross-user dedup means the server answers 'already have it' for a chunk hash no matter who stored it, and the client skips the upload. That answer — or the visible absence of upload traffic and the faster completion — is a **side channel**: an attacker uploads a candidate file and learns whether any other account holds identical content. That enables *confirmation of a file* (does anyone have this leaked document?) and *learning the remaining information* (upload every variant of a templated form with a guessed PIN until one is skipped). A related hazard: if presenting a hash grants access, the hash becomes a bearer token for the content. Mitigations: keep client-visible dedup within one user or account; dedup across users only server-side after a full upload; charge quota by logical size; require proof of ownership; or make the skip unreliable with a random threshold. Per-user scope still keeps most edit-sync savings.

go deeper

for a junior

Remember that whether an upload was skipped is visible to the uploader, and that visibility tells them the content already exists somewhere in the service.

for a middle

Distinguish dedup scope from where dedup runs, and explain why only the client-side, cross-user combination exposes the upload-skip signal.

for a senior

Describe the two classic attacks, the hash-as-token hazard with proof of ownership, and the leaks that survive naive fixes, such as quota accounting and timing.

for a principal

Treat dedup scope as a security decision: quantify the storage saved by cross-user dedup, compare it with bandwidth saved per user, and choose a hybrid you can defend in a threat review.

## Dedup scope and where dedup happens A cloud-drive service that stores chunks under their SHA-256 hashes must decide its **dedup scope**: the set of stored chunks a new upload is compared against. - **Per-user (or per-account) scope**: a chunk deduplicates only against chunks the same user already stores. Old versions, copies and renamed files dedup; other users' identical files do not. - **Cross-user (global) scope**: a chunk stored by anyone satisfies the check. Popular content — the same installer, the same widely shared video — is stored once for everyone. Separately, dedup can be **client-side** — the client sends hashes first and skips uploading chunks the server reports as present — or **server-side** — the client always uploads the bytes and the server discards duplicates after receipt. Client-side dedup saves bandwidth and storage; server-side dedup saves only storage. ## The side channel Client-side dedup with cross-user scope leaks information, because the uploader can observe whether bytes were actually sent: 1. The attacker creates a file whose content they want to test. 2. Their client asks the server which chunk hashes are missing. 3. If the server reports them present, the client sends nothing — visible in client behaviour, in network traffic and in how quickly the upload completes. 4. The attacker concludes that **some other account already stores exactly this content**. Hiding the answer inside the official client does not help: an attacker can run a modified client or simply count the bytes leaving the machine. ## The two classic attacks | Attack | How it works | What leaks | |---|---|---| | Confirmation of a file | upload a copy of a known document and watch for the skip | whether anyone holds that document | | Learning the remaining information | upload many variants of a templated file, each with a different guess for a secret field, until one is skipped | the secret value, such as a PIN or salary on a form letter | The second attack works because a templated document is mostly identical from person to person; only a small field changes, and when that field has a small value space the attacker can enumerate it. ## The hash as an access token A related mistake is letting a client add content to its account merely by presenting its hash. A hash is not a secret — it can appear in logs, shared manifests, published checksums or a leaked database — so anyone holding it could claim the content and then download it. Designs that accept hashes need a **proof of ownership**: the server challenges the client with something only the real bytes can answer, such as a keyed hash over the content with a fresh random nonce, or the bytes at randomly chosen offsets. ## Mitigations and their costs | Mitigation | What it fixes | What it costs | |---|---|---| | Per-user client-side dedup | removes the cross-user signal | no bandwidth savings across users | | Server-side cross-user dedup | uploads look the same either way | full bandwidth for every duplicate | | Hybrid: client-side within a user, server-side across users | keeps edit-sync savings and hides the signal | full upload the first time a user stores popular content | | Random threshold before client-side skipping | makes the signal unreliable for rarely held files | fewer savings; popular files still reveal that they are popular | | Proof of ownership | stops hash-as-token theft | an extra round trip and some computation | Two details decide whether a mitigation really holds: - **Quota accounting**: if a user is charged only for unique bytes, the quota figure itself reveals dedup. Charge the logical file size. - **Timing**: server-side dedup must not finish noticeably faster for duplicates, or timing restores the signal. Encryption does not remove the problem by itself. Encryption at rest with a service-held key leaves the has-check unchanged. Client-side encryption with per-user keys makes cross-user dedup impossible. Encryption keyed by the content's own hash (**convergent encryption**) keeps dedup working precisely because equal plaintexts still produce equal ciphertexts — so equality still leaks. ## Choosing a scope For typical cloud-drive workloads, much of the bandwidth saving comes from a user's own edits, versions and copies, which per-user scope already captures. Cross-user dedup mainly saves storage on popular content, and that saving can be taken server-side without exposing the signal. A reasonable default is client-side dedup within a user or account and server-side dedup across users only where measured storage savings justify the extra machinery. Whatever the choice, write it down as a security decision, not just a storage optimisation.

  • Why doesn't server-side cross-user dedup leak in the same way?
    The client always uploads the full bytes, so traffic and completion look identical whether or not the content is a duplicate; the server discards the copy afterwards. The leak can return through side doors, though: charging quota only for unique bytes, or finishing duplicate uploads visibly faster. Charge logical size and keep response behaviour uniform.
  • Why is it unsafe to let a client claim a stored chunk just by presenting its hash?
    Hashes are not secrets; they appear in logs, shared manifests and published checksums. Anyone holding one could attach the content to their own account and download it. A proof of ownership fixes this: the server issues a fresh challenge, such as a keyed hash over the content with a random nonce, that only a client holding the actual bytes can answer.

saying these in an interview costs you the question

  • Cross-user dedup is safe because attackers only ever see hashes.
  • Encrypting chunks at rest removes the dedup side channel.
  • Knowing a chunk's hash should be enough to download it.
  • Limiting dedup to one user throws away all sync bandwidth savings.
  • Client-side and server-side dedup have the same privacy properties.