skip to content

Your S3 bucket's storage bill is far larger than the total size of the objects that ListObjectsV2 returns for it. What is the most likely cause, and how do you fix it permanently?

level: seniorimportance: should knowfreq 44%

answer

  1. billed but not listed
  2. an upload has only two terminal states
  3. list-multipart-uploads, not list-objects
  4. a lifecycle action, not a script
  5. retry loops multiply the leak

basics

~20 s

Abandoned multipart uploads. Parts from uploads that were never completed or aborted are stored and billed but never appear in an object listing. Find them with ListMultipartUploads and stop the bleeding with a lifecycle rule using the AbortIncompleteMultipartUpload action.

solid answer

~40 s

Almost certainly incomplete multipart uploads. When a client dies between `CreateMultipartUpload` and `CompleteMultipartUpload`, the parts it already sent stay in the bucket: billed as storage, invisible to `ListObjectsV2`, and retained forever unless something aborts them. A retry loop that starts a fresh upload on each attempt can leak the whole object size several times over. To confirm, run `aws s3api list-multipart-uploads` on the bucket, or look at the incomplete-multipart-upload metrics in S3 Storage Lens. The permanent fix is a lifecycle configuration with the `AbortIncompleteMultipartUpload` action and a `DaysAfterInitiation` of, say, 7 — S3 then aborts anything still open after a week. Set it on every bucket that receives large uploads, as a default, not as a remediation.

go deeper

for a junior

Know that unfinished multipart uploads keep their parts in the bucket, that those parts cost money, and that a normal object listing does not show them.

for a middle

Explain that an upload only ends by completion or abort, and name the lifecycle action AbortIncompleteMultipartUpload with DaysAfterInitiation as the fix.

for a senior

Diagnose before you act: quantify with ListMultipartUploads or Storage Lens, name noncurrent versions as the other hidden-storage cause, and fix the client's retry path as well as the bucket.

for a principal

Make it policy — the abort rule and a versioning-expiration rule on every new bucket by default, with Storage Lens as the standing detection for buckets that drift.

## The symptom Someone compares the bucket's billed storage against the sum of object sizes from a listing and finds a large, unexplained gap. Sometimes the gap is bigger than the bucket's real contents. Two things produce this in S3, and you should name both before diving in: 1. **Noncurrent object versions** in a versioned bucket — old versions and delete markers are billed but hidden from a default listing. 2. **Incomplete multipart uploads** — stored parts belonging to uploads that were never completed. If the bucket takes large uploads, the second is usually the culprit and is the one people forget entirely. ## Why the parts persist A multipart upload has exactly two terminal states: **completed** or **aborted**. There is no timeout. If a client calls `CreateMultipartUpload`, streams 40 GB of parts, and then the process is killed, the upload simply stays open. S3 keeps every part it received. Those bytes are charged at the storage class the upload was created with, for as long as the upload remains open — which is forever, absent intervention. The reason it hides so well is that the parts are not objects. `ListObjectsV2` reports objects, and no object exists at the key until completion. The parts are reachable only through `ListMultipartUploads` (which uploads are open) and `ListParts` (what a given upload holds). The pathological version is a retry loop. A job uploads a 200 GB file, fails at 90%, and its retry logic starts a *new* multipart upload from scratch rather than resuming the old `UploadId`. Three retries later you are paying for roughly 540 GB of orphaned parts and a 200 GB object, and nothing in the object listing has changed. ## Confirming it ```bash aws s3api list-multipart-uploads --bucket my-bucket # For one open upload, see how much data it is actually holding. aws s3api list-parts --bucket my-bucket --key big.bin --upload-id "$UPLOAD_ID" ``` `ListMultipartUploads` gives you the key, the `UploadId`, the initiator, and the initiation timestamp — the timestamp is what tells you whether an upload is genuinely abandoned or merely slow. `ListParts` gives per-part sizes, which is how you convert "37 open uploads" into a number of gigabytes. For an account-wide view, **S3 Storage Lens** reports incomplete-multipart-upload bytes and object counts per bucket, and it is the right tool when you do not yet know which bucket is leaking. You can abort one immediately: ```bash aws s3api abort-multipart-upload --bucket my-bucket \ --key big.bin --upload-id "$UPLOAD_ID" ``` ## The permanent fix A one-off cleanup script is remediation, not a fix — the leak resumes the next time a job dies. The durable answer is an S3 **lifecycle rule** with the `AbortIncompleteMultipartUpload` action: ```json { "Rules": [{ "ID": "abort-stale-multipart-uploads", "Status": "Enabled", "Filter": { "Prefix": "" }, "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } }] } ``` S3 then aborts any upload still open more than seven days after it was initiated, discarding its parts and their charges. Applied with `aws s3api put-bucket-lifecycle-configuration`. Choosing `DaysAfterInitiation` is a judgment call: it must be comfortably longer than your slowest legitimate upload, including any resume window your clients rely on. A day is fine for a fast ingest pipeline; a week is a safe general default; anything measured in months defeats the purpose. Note that the action is time-based only — there is no way to say "abort when the initiator's job has failed", because S3 has no idea what your job is doing. ## The other half: fix the client The lifecycle rule caps the damage; it does not stop the leak. Two client-side habits matter: - **Abort on failure.** Wrap the upload so that a non-retryable error calls `AbortMultipartUpload`. The SDK transfer managers generally do this; hand-written loops usually do not. - **Resume, do not restart.** Persist the `UploadId`, and on retry call `ListParts` and send only what is missing. This is faster *and* it stops the leak from multiplying. ## What to say in the interview Name both hidden-storage causes, say why parts are invisible to the object listing, show that you would measure with `ListMultipartUploads` or Storage Lens before acting, and land on the lifecycle rule as the durable control — then add that you would also fix the client's retry path, because the rule bounds the cost rather than eliminating the cause. Finally: put that lifecycle rule on new buckets by default. It costs nothing and it is the single most common piece of S3 hygiene that teams discover only via the bill.

  • What other kind of hidden storage inflates an S3 bill in the same way?
    Noncurrent object versions in a versioned bucket. Overwrites and deletes retain previous versions and delete markers, all billed, none shown in a default listing. The lifecycle counterpart is `NoncurrentVersionExpiration`, and `ListObjectVersions` is how you see them. In practice, buckets that leak multipart parts often leak versions too.
  • How do you choose a value for DaysAfterInitiation?
    It has to exceed your slowest legitimate upload plus whatever resume window your clients use, otherwise S3 aborts an upload someone intended to finish. A day suits a fast ingest pipeline, a week is a safe general default. Values in the tens of days keep the cost around long enough to defeat the point.
  • Why does a lifecycle rule alone not fully solve this?
    It bounds how long orphaned parts are billed, but the leak still happens on every failed upload. The real fix is client-side: abort on non-retryable failure, and on retry resume the existing `UploadId` via `ListParts` rather than starting a fresh upload. The rule is a safety net for the cases your client misses.

saying these in an interview costs you the question

  • Looks only at ListObjectsV2 and concludes the bill is wrong
  • Thinks incomplete uploads expire on their own
  • Writes a cleanup cron job instead of a lifecycle rule
  • Confuses object expiration with aborting multipart uploads
  • Never questions the retry path that creates the orphans

context