skip to content

S3

S3 is the default place to put bytes on AWS: buckets, storage classes, lifecycle policies, versioning, access control, and static site hosting. You meet it early because nearly every AWS design puts something in a bucket, and the follow-up is always about access control and cost.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

A nightly job that has always read files with a plain S3 GetObject call starts failing with an InvalidObjectState error, shortly after a lifecycle rule was added to the bucket. What happened, and what are your options for fixing it?

level: seniorimportance: should knowfreq 56%

basics

~20 s

The lifecycle rule transitioned those objects into S3 Glacier Flexible Retrieval or Deep Archive, which cannot be read by GetObject at all. You must first issue a RestoreObject request and wait, or move the data to S3 Glacier Instant Retrieval, which serves GETs directly.

open as a page

A versioning-enabled S3 bucket's storage bill keeps climbing even though the team has a lifecycle rule that expires objects after 30 days. Why is the data not going away, and what would you change?

level: seniorimportance: should knowfreq 42%

basics

~20 s

On a versioned bucket, an Expiration rule only makes the current version noncurrent by adding a delete marker; the bytes stay as noncurrent versions and keep billing. Add NoncurrentVersionExpiration, plus a rule with ExpiredObjectDeleteMarker to sweep the markers.

open as a page

A compliance team requires that audit logs in S3 cannot be deleted by anyone — including an administrator whose credentials are stolen — for seven years. How would you build that on S3, and what are the tradeoffs of the mode you choose?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Enable S3 Object Lock on the versioned bucket with a seven-year default retention in compliance mode, so no principal — including the account root user — can shorten it or delete a protected version. Governance mode is the reversible, testable alternative.

open as a page

When would you choose S3 Intelligent-Tiering for a bucket over hand-written lifecycle transition rules, and when is the hand-written rule the better call?

level: principalimportance: should knowfreq 46%

basics

~20 s

Choose Intelligent-Tiering when access patterns are unknown or changing: it moves objects between tiers on observed access and charges no retrieval fees, for a per-object monitoring fee. Choose explicit lifecycle rules when the pattern is known and deterministic, or when objects are tiny.

open as a page

A team wants to use SSE-C for an Amazon S3 bucket so that AWS never stores their key material. What does S3 keep on its side, what must every request carry, and what breaks operationally?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

With SSE-C you send the key on every request; S3 uses it, discards it, and keeps only a randomly salted HMAC to validate later requests. HTTPS is mandatory, a lost key means a lost object, and clients that cannot send custom headers cannot read it.

open as a page

What does a byte-range GET against an Amazon S3 object do, and when would you use one instead of downloading the whole object?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

A byte-range GET sends a Range header on GetObject so S3 returns only those bytes, with HTTP 206 Partial Content. Use it to download one object in parallel chunks, to resume an interrupted download, or to read a small region such as a file footer.

open as a page

A browser PUT to an S3 presigned URL fails with a CORS error in the developer console, yet the same URL works from curl. What is actually wrong, and what do you configure to fix it?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

Nothing is wrong with the signature — curl proves that. The bucket has no CORS configuration permitting your page's origin and the PUT method, so the browser blocks the cross-origin request. Fix it by putting a CORS configuration on the bucket.

open as a page

An S3 bucket had versioning enabled and later suspended. What does S3 do with new uploads and with the versions already stored, and why can that silently lose data?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Suspending stops S3 from minting new version IDs: every later upload to a key is stored with the literal version ID null and overwrites the previous null version. Versions created while versioning was enabled are kept and still billed.

open as a page

In Amazon S3, what does Replication Time Control (RTC) add to a replication rule, and how would you measure replication lag on a rule that does not use it?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Default S3 replication is best-effort with no time commitment. Replication Time Control adds a service level agreement covering replication of 99.99% of objects within 15 minutes, for an extra per-GB charge, and surfaces replication metrics and events so breaches are visible.

open as a page

An archival lifecycle rule moved a few hundred million small log objects into S3 Glacier Deep Archive, and the monthly S3 bill went up rather than down. Which AWS-specific charges explain that, and what would you do differently?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Archiving is billed per object, not just per gigabyte. Glacier Flexible Retrieval and Deep Archive add fixed per-object metadata overhead, each transition is a paid request, and Deep Archive carries a 180-day minimum duration — so hundreds of millions of tiny objects lose on every axis.

open as a page

A shared data-lake bucket is used by dozens of teams and its bucket policy is approaching the 20 KB limit. What are S3 Access Points, and how do they change the way access to that bucket is delegated?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

An S3 Access Point is a named endpoint attached to a bucket, with its own hostname, its own policy and its own Block Public Access settings. Each consumer gets one, so per-consumer rules live in many small policies instead of one growing bucket policy.

open as a page

You own the S3 landing zone for a data lake that receives a few million small events per day from many producer teams. How do you lay out buckets, prefixes and object sizes, and what does getting it wrong cost you later?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Keep a raw immutable landing zone separate from curated output, partition prefixes by a stable dimension such as source and ingestion date, and aggregate events into files of tens to hundreds of megabytes instead of millions of tiny objects. Reorganising a lake later means copying everything.

open as a page

showing 31–42 of 42