skip to content

S3

S3 is the default place to put bytes on AWS: buckets, storage classes, lifecycle policies, versioning, access control, and static site hosting. You meet it early because nearly every AWS design puts something in a bucket, and the follow-up is always about access control and cost.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

In Amazon S3, how does a bucket policy differ from an IAM identity policy attached to a user or role, and which element does a bucket policy have that an identity policy never has?

level: juniorimportance: must knowfreq 78%

answer

  1. one attaches to whom, one to what
  2. which document names the caller
  3. union inside an account
  4. two ARN shapes for one bucket
  5. only one of them can say "*"

basics

~20 s

A bucket policy is a resource-based policy attached to the S3 bucket itself and must name a Principal. An IAM identity policy attaches to a user, group or role, and never names a Principal because the identity it is attached to is the principal.

solid answer

~40 s

Both are JSON policy documents with `Effect`, `Action`, `Resource` and optional `Condition`, but they hang off opposite ends of the request. An **identity policy** is attached to an IAM user, group or role and says "this principal may do X to these buckets" — it has no `Principal` element, because the attachment point is the principal. A **bucket policy** is attached to one bucket and says "these principals may do X to me", so `Principal` is required. Inside a single AWS account, an `Allow` in *either* one is enough, as long as nothing explicitly denies the call. Across accounts, you need an `Allow` on both sides. A bucket policy is also the only one of the two that can grant anonymous access, via `"Principal": "*"`, which is exactly why S3 Block Public Access exists.

code

json · 19 lines
json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ListTheBucket",
      "Effect": "Allow",
      "Principal": {"AWS": "arn:aws:iam::111122223333:role/ReportReader"},
      "Action": "s3:ListBucket",
      "Resource": "arn:aws:s3:::reports-bucket"
    },
    {
      "Sid": "ReadTheObjects",
      "Effect": "Allow",
      "Principal": {"AWS": "arn:aws:iam::111122223333:role/ReportReader"},
      "Action": "s3:GetObject",
      "Resource": "arn:aws:s3:::reports-bucket/*"
    }
  ]
}

go deeper

for a junior

Be able to say plainly that a bucket policy lives on the bucket and names a Principal, while an identity policy lives on the user or role and does not. Know the two ARN shapes: the bucket for listing, bucket/* for objects.

for a middle

Explain the combination rules: inside one account an allow on either side suffices, across accounts both are required, and an explicit deny anywhere is final. Show why an over-broad bucket policy defeats careful IAM.

for a senior

Demonstrate judgment about where a rule belongs — guardrails and origin conditions on the resource, everyday grants on the identity — and be ready to debug an AccessDenied by naming which of the two documents was silent.

for a principal

Own the standard: which permissions are expressed centrally versus per-team, how bucket policies stay under their size limit as the estate grows, and how you keep a resource-policy sprawl auditable rather than accumulating one statement per consumer.

## Two policies, two ends of the same request Every S3 request has a caller (the principal) and a target (the bucket or object). AWS lets you write permissions from either end, and both kinds of document use the same grammar: a `Version`, a list of `Statement` objects, each with `Effect` (`Allow` or `Deny`), `Action`, `Resource`, and an optional `Condition` block. An **identity policy** is attached to an IAM user, group or role. It is a statement about what that identity may do: ```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::reports-bucket/*" }] } ``` There is no `Principal` element, and adding one is a syntax error — the principal is whoever the policy is attached to. A **bucket policy** is a resource-based policy attached to exactly one bucket. It is a statement about who may touch that bucket, so it *must* name the principal: ```json { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": {"AWS": "arn:aws:iam::111122223333:role/ReportReader"}, "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::reports-bucket/*" }] } ``` ## How the two combine Within one AWS account, S3 unions the allows. If the caller's identity policy allows `s3:GetObject` but the bucket policy is silent, the call succeeds. If the bucket policy allows it and the identity policy is silent, it also succeeds — this is why an over-broad bucket policy is dangerous even when your IAM policies are tight. Across accounts the rule tightens: the caller's own account must allow the action *and* the bucket's policy must allow that caller. Neither side can grant access unilaterally, which is what makes cross-account sharing safe by construction. On top of that sits the rule that decides ties: an explicit `Deny` in *any* applicable policy — identity policy, bucket policy, a Service Control Policy in the organization, a permissions boundary, a session policy — wins over every `Allow`. Deny is final; absence of an allow is merely an implicit deny that another policy can fill in. ## Resource ARNs: the bucket and the objects are different things A very common junior mistake is writing one `Resource` and expecting it to cover everything. Bucket-level actions such as `s3:ListBucket`, `s3:GetBucketLocation` and `s3:PutBucketPolicy` take the bucket ARN, `arn:aws:s3:::my-bucket`. Object-level actions such as `s3:GetObject`, `s3:PutObject` and `s3:DeleteObject` take the object ARN pattern, `arn:aws:s3:::my-bucket/*`. A policy that lists only `arn:aws:s3:::my-bucket` will allow directory-style listing but every object GET will fail; a policy that lists only `arn:aws:s3:::my-bucket/*` will read objects fine while `aws s3 ls` returns AccessDenied. Most real policies contain two statements, one per ARN shape. ## What only a bucket policy can do Three capabilities exist only on the resource side: 1. **Anonymous access.** `"Principal": "*"` with no condition grants the whole internet. There is no identity policy for an anonymous caller, so this can only be expressed on the bucket. It is the mechanism behind almost every "public S3 bucket" headline, and the reason S3 Block Public Access is enabled by default on new buckets. 2. **Cross-account grants.** Naming another account's principal in `Principal` is how you offer access outward. 3. **Conditions on the request's origin.** Keys such as `aws:SourceVpce`, `aws:SourceIp` and `aws:PrincipalOrgID` let the bucket refuse traffic that does not arrive the way you expect, regardless of what the caller's own account permits. ## Practical guidance Default to identity policies: they scale with your principals, live with your roles, and are easy to audit per team. Reach for a bucket policy when the statement is genuinely about the bucket — sharing outward to another account, forbidding a class of request no matter who makes it, or setting a guardrail you want to survive changes to IAM. Keep bucket policies short: they are limited to 20 KB, and a bucket policy that has grown into a directory of every team in the company is a sign you should be delegating instead.

  • If the identity policy allows s3:PutObject and the bucket policy explicitly denies it, what happens?
    The request is denied. An explicit `Deny` in any applicable policy — bucket policy, identity policy, SCP, permissions boundary or session policy — overrides every `Allow`, and there is no ordering or specificity rule that can rescue it. This is why teams put non-negotiable guardrails in a `Deny` statement rather than relying on the absence of an allow.
  • Why do so many S3 policies need two statements for what feels like one permission?
    Because S3 has two resource types. Bucket-level actions such as `s3:ListBucket` are authorized against `arn:aws:s3:::bucket`, while object-level actions such as `s3:GetObject` are authorized against `arn:aws:s3:::bucket/*`. One statement cannot naturally cover both, so the idiomatic policy has one statement per ARN shape.
  • When would you deliberately choose a bucket policy over an identity policy inside a single account?
    When the rule belongs to the data rather than to a person: a guardrail you want to hold no matter which role is created later, a condition on where requests may originate, or a grant that must be visible to auditors when they look at the bucket. Everything else is cleaner as an identity policy attached to the role.

An identity policy is the badge in your pocket listing the doors you may open; a bucket policy is the guest list taped to one door listing who may come in.

saying these in an interview costs you the question

  • Says a bucket policy needs no Principal element
  • Thinks bucket policies replace IAM policies entirely
  • Uses one ARN and expects ListBucket and GetObject both to work
  • Believes an Allow can override an explicit Deny
  • Assumes an allow in only one account is enough cross-account

context

open as a page

Amazon S3 refuses a single PutObject request larger than 5 GB. How does an S3 multipart upload work, and what does it give you beyond raising that ceiling?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Multipart upload splits one object into parts sent independently: CreateMultipartUpload returns an UploadId, UploadPart sends each numbered piece, CompleteMultipartUpload assembles them. Beyond the 5 GB ceiling it adds parallelism, per-part retry, and resumable uploads.

open as a page

What is an Amazon S3 presigned URL, and whose permissions does the person holding one actually exercise?

level: juniorimportance: must knowfreq 78%

basics

~20 s

An S3 presigned URL is an object URL carrying a SigV4 signature and an expiry in its query string. Whoever holds it performs one specific operation on one specific object using the signing principal's permissions, with no AWS credentials of their own.

open as a page

In Amazon S3, what does a bucket's Lifecycle configuration do, and what are the two kinds of action a lifecycle rule can take?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An S3 Lifecycle configuration is a set of bucket-level rules that act on objects automatically as they age. A rule takes one of two kinds of action: transition an object to a cheaper storage class, or expire (delete) it.

open as a page

In an Amazon S3 bucket with versioning enabled, what happens when a client sends a DELETE for an object key without specifying a version ID, and how do you get the object back?

level: juniorimportance: must knowfreq 68%

basics

~20 s

S3 erases nothing: it adds a delete marker as the new current version, so plain GETs return 404 while every earlier version stays stored and billed. You restore by deleting the delete marker by its own version ID.

open as a page

An S3 bucket that should be private is serving objects to anonymous callers. How does S3 decide that a bucket is public, and what does each of the four S3 Block Public Access settings stop?

level: middleimportance: must knowfreq 70%

basics

~20 s

S3 treats a grant to a wildcard principal or to the legacy AllUsers/AuthenticatedUsers groups as public unless conditions pin it to fixed values. Block Public Access has four independent switches that block or ignore public ACLs and public policies, and it applies at both account and bucket level.

open as a page

For an Amazon S3 bucket, compare SSE-S3, SSE-KMS, DSSE-KMS and SSE-C: who holds the key material in each, what each one adds operationally, and when would you pick each?

level: middleimportance: must knowfreq 76%

basics

~20 s

SSE-S3 lets S3 hold the keys: free and invisible. SSE-KMS routes each object key through KMS, adding per-key control and an audit trail at a per-request cost. DSSE-KMS applies two layers for dual-encryption mandates. SSE-C makes you supply the key on every request.

open as a page

An S3 bucket is configured with S3 Event Notifications that invoke a Lambda function on every object upload. What delivery guarantees does S3 give for those notifications, and what does that force you to build into the handler?

level: middleimportance: must knowfreq 62%

basics

~20 s

S3 Event Notifications are at-least-once and not ordered. The same upload can produce two events and events for one key can arrive out of order, so the handler must be idempotent; the record's sequencer field lets you order events for the same key.

open as a page

You add a replication rule to an Amazon S3 bucket that already holds millions of objects. New uploads appear in the destination bucket within minutes, but none of the pre-existing objects ever show up. Why, and how do you copy the backlog?

level: middleimportance: must knowfreq 60%

basics

~20 s

An S3 replication rule only applies to objects written after the rule is saved; it never walks the existing contents of the bucket. To copy the backlog you run S3 Batch Replication, an S3 Batch Operations job that replicates existing object versions using the same rule.

open as a page

Your team wants to cut an S3 bill with a lifecycle rule that transitions every object in a bucket to S3 Standard-IA after 30 days. For which objects does that raise the bill instead of lowering it, and why?

level: middleimportance: must knowfreq 66%

basics

~20 s

Small, short-lived and frequently-read objects. S3 Standard-IA bills a per-object minimum size and a 30-day minimum duration, and charges a per-GB retrieval fee, so tiny objects, objects deleted early, and hot data all cost more there than in S3 Standard.

open as a page

Amazon S3 has provided strong read-after-write consistency since December 2020. What exactly does that guarantee cover, and which S3 behaviours are still not covered by it?

level: middleimportance: must knowfreq 55%

basics

~20 s

Since December 2020, any successful S3 PUT or DELETE is immediately visible to every subsequent GET, HEAD and LIST, in all regions, at no extra cost. Bucket configuration changes and replication to another bucket remain eventually consistent.

open as a page

An upload to a single S3 bucket must trigger three independent consumers: a thumbnailer, an antivirus scan, and an audit writer. How do you wire that with S3 event notifications, and what does routing through EventBridge change compared with S3's native destinations?

level: seniorimportance: must knowfreq 58%

basics

~20 s

S3 refuses overlapping notification configurations for one event type, so three consumers on the same prefix need a fan-out point: either one SNS topic with three subscribers, or EventBridge notifications with three independently filtered rules. EventBridge adds richer filtering, archive and replay, and cross-account routing.

open as a page

You create a new Amazon S3 bucket, set no encryption configuration, and upload objects without sending any encryption header. Is that data encrypted at rest, and what does S3 server-side encryption actually protect you against?

level: juniorimportance: should knowfreq 62%

basics

~20 s

Yes. Since January 2023 S3 encrypts every new object with SSE-S3 (AES-256) automatically and at no charge. That protects the data on AWS's disks; it does nothing against a caller whose IAM permissions already allow GetObject.

open as a page

An S3 bucket receives many kinds of uploads. Using S3 Event Notifications, how do you trigger a function only for .jpg objects written under the uploads/ prefix, and what kinds of matching are not possible?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Attach a Filter with two key FilterRules on the notification configuration: prefix uploads/ and suffix .jpg. S3 only supports these two literal, case-sensitive string rules — there are no wildcards, no regular expressions, and no filtering on tags, size, or metadata.

open as a page

In Amazon S3, what is the difference between Cross-Region Replication (CRR) and Same-Region Replication (SRR), and what does each one solve?

level: juniorimportance: should knowfreq 68%

basics

~20 s

Both are the same S3 replication feature and differ only in where the destination bucket lives. CRR copies objects to a bucket in another AWS Region for disaster recovery, lower read latency, or data-residency rules. SRR copies within one Region, typically across accounts or for log aggregation.

open as a page

Account A owns an S3 bucket and a role in account B must read objects under one prefix. What has to be configured in each account, and why does an Allow in only one of them fail?

level: middleimportance: should knowfreq 66%

basics

~20 s

Cross-account S3 access needs two allows: the bucket policy in account A must permit account B's role, and an identity policy in account B must permit the same actions on that bucket. Either side alone is an implicit deny, so the request fails.

open as a page

You are serving a static site from an S3 bucket. What is the difference between S3's static website endpoint and putting CloudFront with Origin Access Control in front of the bucket's REST endpoint, and which do you choose?

level: middleimportance: should knowfreq 50%

basics

~20 s

The S3 website endpoint is HTTP-only and needs a public bucket, but it maps directory paths to index documents and serves custom error pages. CloudFront with Origin Access Control keeps the bucket fully private, adds HTTPS on your own domain and caching, but does no directory-index mapping below the root.

open as a page

You are uploading 500 GB objects to Amazon S3 with multipart uploads. What constrains the part size you can choose, and how do you decide on one?

level: middleimportance: should knowfreq 48%

basics

~20 s

S3 allows at most 10,000 parts, each at least 5 MiB except the last and at most 5 GiB. A 500 GB object therefore needs parts of roughly 50 MB or more; above that floor, larger parts mean fewer requests but more data re-sent per retry and more memory buffered per concurrent part.

open as a page

Code running in AWS Lambda generates an S3 presigned GET URL with a 24-hour expiry, but recipients start getting 403 responses a few hours later. Why does the URL die early, and what is the real ceiling on a presigned URL's lifetime?

level: middleimportance: should knowfreq 55%

basics

~20 s

Two clocks govern a presigned URL, and the shorter wins. Lambda signs with the execution role's temporary STS credentials, so the URL dies when that session expires, no matter what X-Amz-Expires says. SigV4 itself caps the requested expiry at seven days.

open as a page

You let browsers upload directly to S3. With a presigned PUT URL, what stops a client from writing a key you did not intend or pushing a 5 GB file, and what does a presigned POST policy give you that the PUT does not?

level: middleimportance: should knowfreq 46%

basics

~20 s

A presigned PUT pins the bucket, key and method into the signature, so the client cannot redirect the write — but it cannot bound the body size. A presigned POST signs a policy document whose conditions, including content-length-range, are enforced by S3 on upload.

open as a page

A partner account uploads objects into your S3 bucket and your own account then gets AccessDenied reading some of them. Why does the bucket owner not automatically own uploaded objects, and what do the S3 Object Ownership settings change?

level: seniorimportance: should knowfreq 42%

basics

~20 s

With legacy ACLs enabled, an object is owned by the account that uploaded it, not by the bucket owner, so a cross-account writer's objects can be unreadable by the bucket's own account. Setting Object Ownership to BucketOwnerEnforced disables ACLs and makes the bucket owner own every object.

open as a page

How would you write an S3 bucket policy that denies every request coming from outside your AWS Organization, and what mistake in that policy can lock you out of your own bucket?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Add a Deny statement on all S3 actions when aws:PrincipalOrgID does not equal your organization ID. The trap is that AWS service principals do not carry that key, so log delivery and other service writes break unless you exempt them with aws:PrincipalIsAWSService.

open as a page

A batch job that reads tens of millions of small SSE-KMS-encrypted objects from Amazon S3 starts failing with KMS throttling errors, although the same job ran fine at a tenth of the scale. What is causing it, and how do S3 Bucket Keys change the picture?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Each read of an SSE-KMS object costs a KMS decrypt request, so object throughput becomes KMS throughput and hits the account's per-Region KMS request quota. S3 Bucket Keys let S3 derive object keys from a short-lived bucket-level key, cutting KMS calls sharply.

open as a page

An auditor requires that objects in an Amazon S3 bucket can only be written with SSE-KMS and that nothing in the bucket is ever accessed over plain HTTP. How do you enforce both with a bucket policy, and where do such policies commonly go wrong?

level: seniorimportance: should knowfreq 55%

basics

~10 s

Attach two Deny statements to the bucket policy: deny s3:PutObject unless the s3:x-amz-server-side-encryption condition equals aws:kms, and deny s3:* when aws:SecureTransport is false, listing both the bucket ARN and the object ARN.

open as a page

A Lambda function is triggered by s3:ObjectCreated:* on a bucket and writes its processed output back into that same bucket. What goes wrong, and how do you design around it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The output write fires the same notification, so the function invokes itself in an unbounded loop that burns concurrency, S3 request charges and Lambda duration until someone stops it. Fix it by writing output to a different bucket, or by fencing input and output with disjoint prefix or suffix filters.

open as a page

Your S3 bucket's storage bill is far larger than the total size of the objects that ListObjectsV2 returns for it. What is the most likely cause, and how do you fix it permanently?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Abandoned multipart uploads. Parts from uploads that were never completed or aborted are stored and billed but never appear in an object listing. Find them with ListMultipartUploads and stop the bleeding with a lifecycle rule using the AbortIncompleteMultipartUpload action.

open as a page

A batch job writing thousands of objects per second under one Amazon S3 key prefix starts getting HTTP 503 SlowDown responses. What is being throttled, and what do you change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

S3 scales request rate per key prefix, not per bucket: roughly 3,500 write and 5,500 read requests per second each. Spread keys across many prefixes to multiply that ceiling, and retry 503 SlowDown with exponential backoff plus jitter while S3 repartitions.

open as a page

A presigned S3 GET URL for a private customer document turns up in a third party's access logs and is still being fetched. Why can you not simply revoke that one URL, and what actually cuts off access?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A presigned URL is a bearer token that S3 never recorded: validity is recomputed from the signature, the expiry and the signer's current permissions on every request. With nothing to revoke individually, you must deny the object, remove the signer's permission, kill the signing session, or move the object.

open as a page

An Amazon S3 replication rule targets a bucket in another AWS account. Unencrypted objects replicate fine, but every object encrypted with SSE-KMS shows a replication status of FAILED. What are the likely causes, and what does a correct configuration look like?

level: seniorimportance: should knowfreq 42%

basics

~20 s

SSE-KMS objects are excluded unless the rule opts them in via SourceSelectionCriteria, and the replication role needs kms:Decrypt on the source key plus kms:Encrypt on a destination-Region key named in the rule. Cross-account also needs both key policies and the destination bucket policy to allow that role.

open as a page

Your team proposes relying on Amazon S3 Cross-Region Replication as protection against someone accidentally deleting production objects. Which deletions does S3 replication propagate to the destination, which does it never propagate, and why is replication still a poor substitute for backup?

level: seniorimportance: should knowfreq 50%

basics

~20 s

S3 replication optionally propagates delete markers if the rule enables delete-marker replication, but it never propagates a DELETE that names a specific version ID. Replication mirrors live state rather than preserving history, so it protects against losing a Region or a bucket, not against a person or process destroying data.

open as a page

showing 1–30 of 42