skip to content

A claim-check pipeline has been running for six months and the object storage bucket backing it has grown to 40 TB, far more than the traffic volume would suggest. What's the likely root cause, and how would you diagnose and fix it?

level: seniorimportance: must knowfreq 58%

answer

  1. orphaned objects = no lifecycle owner
  2. happy-path-only deletion leaves failures behind
  3. fan-out means no single consumer can safely delete
  4. bucket lifecycle policy as the backstop
  5. delete-on-success as optimization, not the only mechanism

basics

~20 s

Old files in storage are probably never getting deleted after being processed -- orphaned blobs pile up forever. Fix by adding a cleanup step or an automatic expiry rule on the storage bucket so old, unused files get removed.

solid answer

~50 s

This is classic orphaned-object accumulation: every payload written for a claim check needs an explicit lifecycle -- someone has to delete it, or it lives forever. Common causes: a consumer that only deletes on success, silently leaving failed or retried payloads behind; a fan-out topology where no single consumer can safely delete because others haven't processed it yet; or simply no cleanup logic written at all. Diagnose by sampling object ages against processing logs to see whether already-processed payloads are still present. The standard fix is to stop relying on consumer-side deletion as the primary mechanism and instead put a time-based lifecycle policy directly on the storage bucket, expiring objects after N days, so cleanup happens independently of whether consumer logic ran correctly; consumers can still delete early as a cost optimization once no other consumer needs the object.

go deeper

for a junior

Should recognize that storage doesn't clean itself up automatically and that something must explicitly delete old objects.

for a middle

Should identify delete-on-success-only as the typical root cause and know that a storage lifecycle policy is a common fix.

for a senior

Should reason through the fan-out coordination problem, propose a concrete diagnostic approach (age sampling against processing logs), and design a layered fix (lifecycle policy backstop plus optional early deletion).

for a principal

Should design the retention/lifecycle strategy as a first-class part of the pipeline's architecture up front, including monitoring that catches divergence between throughput and storage growth before it becomes costly, and know when cold-storage tiering versus outright deletion is the right call.

## Why the bucket keeps growing Unbounded storage growth in a claim-check pipeline is almost always a symptom of the payload lifecycle never being closed out -- every object written for a claim check is a liability until something explicitly removes it, and if nothing ever does, the bucket grows monotonically for as long as the pipeline runs, regardless of how much traffic is currently flowing through it. This is a distinct failure mode from anything the broker itself would surface, because the broker has no visibility into the storage system at all; a broker dashboard showing healthy, low queue depth gives no signal whatsoever about whether the storage side is quietly accumulating. ## The usual root cause: cleanup designed for the happy path The most common root cause is that cleanup was designed around the happy path only: a consumer deletes the object from storage after it finishes processing successfully, and that's the only deletion path that exists anywhere in the system. Every payload becomes permanently orphaned when it is: - one whose processing failed; - one whose consumer crashed before reaching the delete call; - one whose message was redelivered and processed by a different consumer instance that also forgot to delete; - or one whose consumer was simply never written to delete at all. In a **fan-out topology**, where several independent consumers each need to read the same payload -- for example separate transcoding, thumbnailing, and moderation services all reading the same uploaded video -- this gets structurally worse, because no single consumer can correctly delete the object without knowing whether every other subscriber has finished, and naive implementations often just skip deletion altogether to avoid that coordination problem, guaranteeing every object lives forever by default. ## How to diagnose it Diagnosing this starts with characterizing the growth: sample a set of objects by age and compare their creation timestamps against processing logs or a database table that records when each corresponding message was successfully consumed. - If a large fraction of objects older than the pipeline's expected processing SLA are still present and their corresponding processing did complete, that confirms deletion isn't happening on the success path. - If a smaller but still significant fraction correlates with failed or dead-lettered messages, that points at error-path cleanup being the gap specifically. It's also worth checking whether the bucket has any lifecycle configuration at all -- surprisingly often, teams build the write and fetch sides of a claim-check pipeline carefully and never configure the storage system's own expiry rules, treating cleanup as purely an application-level concern. ## The robust fix The robust fix is to stop treating consumer-side deletion as the primary or only cleanup mechanism and instead configure a **lifecycle policy** directly on the storage bucket or container -- most object stores support rules like 'delete objects older than N days' or 'transition to cheaper cold storage after N days, then delete after M more' natively, entirely independent of application code. - **The backstop.** This guarantees an upper bound on storage growth regardless of whether any consumer's delete call ever runs, which matters because it removes an entire category of failure (crashed consumers, buggy delete logic, forgotten error paths) from the equation. - **The optimization on top.** Application-level deletion can still be layered on top as an optimization -- deleting immediately after successful processing reduces storage cost sooner than waiting for the lifecycle policy's expiry window -- but it should never be the only safety net. - **Fan-out topologies.** For fan-out topologies specifically, the retention window needs to be set generously enough that the slowest expected consumer has time to finish before expiry, or the pipeline needs an explicit completion-tracking mechanism (a counter or a set of consumer IDs that have acknowledged processing) before any deletion, whether time-based or explicit, is allowed to happen. ## What it looks like in practice A concrete real-world pattern: an image-processing pipeline storing uploads in an `S3` bucket configures an S3 Lifecycle rule to expire objects seven days after creation, well past the pipeline's typical processing time of minutes to low hours, giving ample buffer for retries and dead-letter reprocessing. Consumers additionally delete their own objects immediately after successful processing as a cost optimization, since S3 storage cost, while cheap per gigabyte, is not free at scale, and shaving days off the retention of successfully-processed objects meaningfully reduces steady-state storage cost when processing millions of objects a day. The combination -- application-level deletion as the fast path, bucket lifecycle policy as the guaranteed backstop -- is what prevents both premature growth and the silent unbounded accumulation described in the scenario.

  • How would you handle cleanup safely in a fan-out topology where three different consumers all need to read the same payload?
    Either rely purely on a generous time-based lifecycle policy sized to the slowest consumer's expected processing time, or track completion explicitly -- for example a small counter or set of consumer IDs in a database that increments as each consumer finishes, with deletion triggered only once all expected consumers have acknowledged. The former is simpler and usually sufficient; the latter is worth it when storage cost or object sensitivity makes early deletion valuable.
  • Would you ever prefer transitioning objects to cold storage instead of deleting them outright?
    Yes, if there's a compliance or debugging need to retain processed payloads for a longer period, tiering to cheaper cold/archive storage after a short window and deleting only after a much longer retention period balances cost against retainability, which many object stores support as a built-in lifecycle transition rather than requiring custom logic.
  • What monitoring would catch this problem before it grows to 40 TB?
    Track bucket size and object count over time against message throughput, and alert if storage growth rate stays roughly flat or keeps climbing while throughput is steady -- that divergence is the early signal that objects aren't being cleaned up, well before it becomes an expensive surprise.

Like a self-storage unit nobody ever cancels: if the only way to stop paying for it is remembering to call and close it out yourself, forgotten units just keep accumulating and costing money indefinitely -- which is why storage facilities also offer auto-expiring or auto-billed plans as a backstop.

saying these in an interview costs you the question

  • Assumes storage will just 'take care of itself' with no lifecycle policy or deletion logic
  • Relies solely on consumer-side delete-on-success with no backstop for failure paths
  • Doesn't recognize that fan-out consumers can't independently delete without coordination
  • Proposes manually deleting the entire bucket's old contents as a one-off fix without addressing the root cause
  • Has no monitoring that would have caught the growth before it reached 40 TB

context