skip to content

questions

5

In a cloud-drive service, why are file bytes kept in an object store while paths, versions and permissions live in a separate metadata database?

level: juniorimportance: must knowfreq 70%

answer

  1. opposite data shapes
  2. big and immutable vs small and mutable
  3. version row points at blob
  4. opaque ID, not the path
  5. orphans need a collector

basics

~20 s

File contents are large, immutable blobs read whole, while names, versions and permissions are tiny records that change often and need transactions and indexed queries. Splitting them lets each store scale and be tuned for its own workload.

solid answer

~50 s

The two kinds of data have opposite shapes. Bytes are big, written once and read sequentially, so they belong in an object store built for cheap capacity and durability. Metadata - path, owner, ACLs, version list, checksum - is small, mutable and queried constantly, so it belongs in a database with transactions and indexes. Each blob is stored under an opaque ID (a random ID or a `SHA-256` content hash), and a version row in the metadata database points at it. That means renames, moves, shares and version restores are small row updates that never touch the bytes, and each store scales independently. The cost is that the two stores cannot share one transaction: the design needs a garbage collector for orphaned blobs and must avoid metadata rows that point at bytes that were never stored.

go deeper

for a junior

Remember the core contrast: bytes are large and written once, metadata is small and edited constantly. Be ready to say a version row points at a blob by ID.

for a middle

Explain why a rename, share or version restore never touches the bytes, and why blob keys are opaque IDs rather than user paths.

for a senior

Show you have handled the gap between two stores: orphaned blobs, dangling pointers, garbage collection with a grace period, and backing up the metadata store as the irreplaceable part.

for a principal

Frame the split as an ownership and scaling boundary: capacity and query load grow independently, and durability, tiering and deduplication policies can evolve on the byte side without reshaping metadata.

## Two very different kinds of data A **cloud-drive service** stores two things for every file. The first is the **content** - the bytes of a photo, a spreadsheet or a 4 GB video. The second is the **metadata** - the file's name, the folder it sits in, its owner, who it is shared with, its size, its checksum and its list of past **versions**. These two kinds of data behave so differently that nearly every large file store keeps them in separate systems: an **object store** (a flat key-to-bytes service, sometimes called a blob store) for content, and a **metadata database** for everything else. | Property | Content (blobs) | Metadata | |---|---|---| | Size per item | Kilobytes to many gigabytes | Hundreds of bytes | | Mutability | Written once, never edited in place | Changes constantly (renames, shares, moves) | | Access pattern | Read or written whole, sequentially | Point lookups, range scans, joins | | What it needs | Cheap capacity, high durability, throughput | Transactions, secondary indexes, low latency | | Share of total bytes | Nearly all | A tiny fraction | ## How the two stores connect Each blob is stored under an **opaque identifier** - a random ID or a content hash such as `SHA-256` of the bytes - never under the user's path. The metadata row for a file version carries that identifier as a pointer: 1. A client asks for `/team/plan.pdf`. 2. The service resolves the path in the metadata database, checks permissions, and finds the current version's `blob_id`. 3. The bytes are fetched from the object store by `blob_id`, often by handing the client a short-lived download URL. Because blobs are **immutable**, editing a file never overwrites bytes: the new content becomes a new blob, and the metadata database adds a new version row pointing at it. Version history is therefore a metadata feature built on top of an append-only byte store. ## What the split buys you - **Metadata operations never touch bytes.** Renaming, moving, sharing or restoring an old version is a small database write, however large the file. - **Each store scales on its own axis.** Capacity grows by adding storage nodes; query load grows by scaling the database. Neither forces the other to grow. - **Right tool for each job.** Databases are poor at holding multi-gigabyte values: huge rows bloat backups, replication streams and caches. Object stores, in turn, typically offer little beyond put, get, delete and list-by-prefix - usually no multi-object transactions, no rich secondary indexes, and often no rename. - **Independent durability schemes.** Bytes can be protected by replication or erasure coding tuned for cost, while metadata gets a transactional, strongly consistent setup. - **Room to grow.** With content addressed by ID, identical bytes can later be shared and cold bytes moved to cheaper tiers without changing the metadata model. ## What it costs Two stores cannot be updated in one atomic transaction, so the design must tolerate the gap between them: - **Orphaned blobs.** If bytes are written but the metadata commit never happens (a crashed client, a failed request), the object store holds data nothing points at. A background **garbage collector** that removes unreferenced blobs after a grace period is part of the design. - **Dangling pointers.** The reverse - a metadata row pointing at a missing blob - is worse, because users see a file they cannot open. Designs avoid it by making a version visible only after its bytes are confirmed stored. - **Two hops per read.** Every download is a metadata lookup plus a blob fetch, so the metadata path must be fast and well cached. - **Metadata is the crown jewel.** Blobs keyed by opaque IDs are meaningless without the database that names them, so the metadata store needs backups and replication at least as careful as the bytes get. ## A useful mental check Ask of any operation: does it change *what the bytes are*, or only *what we know about them*? Uploading new content changes the bytes and creates a blob. Everything else a user does most of the day - browsing, renaming, sharing, moving, restoring a version - changes only knowledge, and so lands entirely in the metadata database. That asymmetry in traffic is itself a reason for the split: the metadata store sees far more operations, the object store far more bytes. ## In an interview Lead with the shape difference (large and immutable versus small and mutable), name the pointer that joins them (an opaque blob ID on a version row), give one concrete payoff (a rename is a row update), and name one cost (orphan cleanup). That covers what most interviewers are listening for.

  • Why key blobs by an opaque ID instead of by the user's file path?
    A path is mutable metadata; a blob key should never change. With opaque IDs a rename or move is a metadata update only, whereas path-keyed blobs would need a copy and delete for every object, since object stores usually lack rename. Opaque keys also avoid leaking names in storage, sidestep collisions between users, and allow content-hash keys for deduplication.
  • If the metadata database is lost but the object store survives, what can you recover?
    Very little that users recognise: the blobs are opaque IDs with no names, owners, folders, permissions or version order. That is why the metadata store gets rigorous backups and replication. Some designs also write a small descriptive record alongside each blob so a disaster-recovery job could rebuild at least ownership and names.
  • Why can sequential or timestamp-prefixed blob keys hurt an object store?
    Some object stores partition their key space by lexicographic range. Monotonically increasing keys all sort to the end, so every new write lands on the same partition and creates a hotspot until it splits. Random IDs or a hashed prefix spread writes evenly. Other stores hash keys internally and are less sensitive, so the risk depends on the system.

A library keeps books on warehouse shelves identified only by a stock number, while the catalogue records title, location, borrowers and editions; re-cataloguing a book never moves the book.

saying these in an interview costs you the question

  • Store every file as a binary column in the main database; it scales fine.
  • The object store can serve folder listings and permission checks by itself.
  • Renaming a file means copying its bytes to a new key.
  • The two stores stay consistent automatically, so no cleanup job is needed.
  • Metadata is small, so it needs less backup care than the bytes.
open as a page

In a cloud-drive service, how should the API list a folder that holds millions of entries without timing out?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Return the folder's children in capped pages through a cursor: read a sorted index on parent ID and name, and resume after the last name returned. Never load the whole folder, use deep offsets, or read blobs to build a listing.

open as a page

In a cloud-drive metadata database, what does renaming a folder with a million descendants cost under a full-path model versus a parent-id model?

level: middleimportance: should knowfreq 50%

basics

~20 s

Storing full paths on every entry makes a folder rename rewrite all million descendant rows, a huge write that is hard to make atomic. Storing parent-id pointers makes it a one-row update, but path resolution then costs one lookup per level.

open as a page

For a cloud-drive blob store, how do three-way replication and Reed-Solomon erasure coding compare on storage overhead and failure tolerance?

level: middleimportance: should knowfreq 50%

basics

~20 s

Three-way replication costs 3x raw storage and survives two lost copies, with simple reads and repairs. A Reed-Solomon (6,3) code costs 1.5x and survives any three lost fragments, but reads and repairs must touch several nodes and decode.

open as a page

For a cloud-drive blob store where most files are never opened after their first month, how would you design a storage-tiering and lifecycle policy?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Drive tiering from access data in the metadata database: move blobs unread for a measured period to cheaper, slower tiers, expire old versions and trash on schedule, and weigh per-object, minimum-duration and retrieval costs against the savings.

open as a page