skip to content

For a cloud-drive blob store where most files are never opened after their first month, how would you design a storage-tiering and lifecycle policy?

level: principalimportance: nice to knowfreq 30%

answer

  1. access is heavily skewed
  2. last read, not creation date
  3. tiny files, per-object fees
  4. archive means slow restore
  5. versions and trash go first

basics

~20 s

Drive tiering from access data in the metadata database: move blobs unread for a measured period to cheaper, slower tiers, expire old versions and trash on schedule, and weigh per-object, minimum-duration and retrieval costs against the savings.

solid answer

~40 s

Start from measured access curves, not a fixed age. The metadata database already knows last access, version state, size and holds, so it drives the policy and records each blob's tier. Keep current, recently read files hot; move files unread for a threshold to a warm tier that still reads instantly; reserve archive tiers with slow restores for data users do not open on demand, such as old noncurrent versions, or expose a visible restore step. Expire noncurrent versions by count and age, purge trash after a stated window, and let legal holds override deletion. Watch the traps: per-object charges make moving tiny files a loss, some cold tiers bill a minimum duration, and retrieval costs money and time. Throttle transitions, add hysteresis so files do not bounce between tiers, and re-measure the savings.

go deeper

for a junior

Recall that storage comes in tiers that trade price against read speed, and that old, unread data can move to cheaper ones.

for a middle

Explain which signals mark a blob cold, why the metadata database drives the decision, and why it must record the current tier.

for a senior

Show the operational traps: per-object charges on small files, minimum-duration billing, retrieval latency, throttled transitions and tier bouncing.

for a principal

Own the trade-off between savings and user experience, justify thresholds from measured access curves, and set how the policy is reviewed as usage changes.

## Why tiering exists A **cloud-drive blob store** grows forever, yet access is heavily skewed: files are opened often in the days after upload, then rarely or never. Keeping every byte on fast, replicated storage wastes money. **Storage tiering** moves bytes between classes of storage with different price and speed, and a **lifecycle policy** is the set of rules that decides when bytes move or are deleted. Designing one is a judgment call - the right thresholds depend on access data, product promises and cost structure - which is why it is a lead-level question. ## A typical tier ladder | Tier | Durability scheme | Read latency | Relative cost (illustrative) | |---|---|---|---| | Hot | Replicated, fast disks | Milliseconds | Highest | | Warm | Erasure coded, dense disks | Milliseconds to tens of ms | Lower | | Cold or archive | Erasure coded, offline or spun-down media | Minutes to hours to restore | Lowest | Systems differ in exactly which tiers they offer and how retrieval works; the shape above is a common pattern, not any specific offering. ## Signals that should drive the policy The **metadata database** is the natural place to decide, because it already knows every file's history: - **Last access time** - the strongest single signal of coldness; age since creation alone misclassifies old files people still open. - **Version state** - noncurrent versions and items in the trash are read far less than current files and are the safest first candidates. - **Object size** - tiny files dominate object counts but not bytes. - **Holds and retention** - legal holds or retention promises override any deletion rule. - **Account plan** - a product may promise different retention to different customers. The metadata row should also record **which tier** a blob is in, so reads route correctly and the client can show a "restoring" state instead of an error. ## Costs that can erase the savings A naive "move everything older than 30 days to the cheapest tier" rule can raise the bill: 1. **Per-object charges.** Transitions and per-object overhead are often priced per object, so moving millions of tiny files can cost more than their few bytes save. Exclude small objects, or pack them into larger containers first. 2. **Minimum storage durations.** Some cold classes bill a minimum retention period; data deleted early still pays for the full period. 3. **Retrieval fees and latency.** Reading from archive may cost per gigabyte and take hours - unacceptable for a file a user just clicked. 4. **Transition load.** Moving petabytes is a large background job; it must be throttled so it does not starve user traffic. ## A worked estimate Assume, purely for illustration, a 1,000 TB store where 800 TB has not been read in 90 days, and a cold tier priced at one quarter of hot per GB-month. Before tiering the monthly bill is proportional to 1,000 units. After moving the cold 800 TB, it is 200 + 800 x 0.25 = 400 units - a **60%** cut, before subtracting transition, retrieval and early-deletion costs. The estimate is only as good as the access data behind it, which is why the policy should be measured and revised rather than set once. ## A defensible policy sketch - Keep current versions of recently accessed files **hot**. - Move files not read for a threshold period, chosen from the measured access curve, to **warm**, which still serves reads instantly. - Reserve **archive** for data users do not expect to open on demand - old noncurrent versions, backups, compliance copies - or offer it explicitly with a visible restore step. - Expire noncurrent versions by **count and age**, purge trash after a stated window, and never delete a blob while any metadata row still references it. - Leave small objects alone, or pack them before moving. - Promote a blob back to hot on repeated access, with **hysteresis** so files do not bounce between tiers. ## Running it safely - Run transitions as throttled batch jobs driven by metadata queries, and make them idempotent so a crashed batch can resume. - Update the blob's tier in metadata only after the move is confirmed, so reads never route to a tier that does not yet hold the bytes. - Track the realised savings against transition and retrieval spend, and adjust thresholds when the access curve shifts. ## In an interview There is no single right threshold. Show that the policy is **data-driven** (access curves from metadata), **cost-aware** (per-object, minimum-duration and retrieval charges) and **product-aware** (what a user experiences when opening a cold file). Stating those three lenses matters more than any particular number.

  • What should happen when a user opens a file whose blob sits in an archive tier with slow retrieval?
    Ideally that never surprises the user: user-openable current files stay in tiers with instant reads. If archive is used for them, the metadata records the tier, the client shows a restoring state, a restore job runs asynchronously and notifies on completion, and restores are rate-limited. Keeping a small hot preview lets the user recognise the file meanwhile.
  • How do file versions and deletion fit into the lifecycle policy?
    Noncurrent versions are the best cold candidates, expired by both count and age so heavy editors do not accumulate unbounded history. Trash is purged after a stated window. A blob is deleted only when no metadata row references it any more, and legal holds or retention promises suspend deletion regardless of age.

saying these in an interview costs you the question

  • Move everything older than 30 days to the cheapest archive tier.
  • File age since creation is a good enough signal; reads do not matter.
  • Moving tiny files to a cold tier always saves money.
  • The object store can pick tiers alone; metadata need not record them.
  • Archive retrieval delays are fine for files users open on demand.