A shared data-lake bucket is used by dozens of teams and its bucket policy is approaching the 20 KB limit. What are S3 Access Points, and how do they change the way access to that bucket is delegated?
answer
- one door per consumer
- the bucket policy delegates once
- its own hostname and its own policy
- can narrow, never widen
- VPC-only as a first-class setting
basics
~20 sAn S3 Access Point is a named endpoint attached to a bucket, with its own hostname, its own policy and its own Block Public Access settings. Each consumer gets one, so per-consumer rules live in many small policies instead of one growing bucket policy.
solid answer
~50 sInstead of one bucket policy that names every consumer, you create one access point per consumer or per use case. Each has its own DNS name and ARN, and callers use that ARN in place of the bucket name in ordinary S3 API calls. Each carries its own policy — typically scoped to one prefix and one set of actions — plus its own Block Public Access settings, and optionally a VPC network origin so it is reachable only from inside a named VPC. The bucket policy shrinks to a single delegation statement that trusts access points owned by the account, using the `s3:DataAccessPointAccount` condition. Requests are still evaluated against the underlying bucket, so an access point can never grant more than the bucket allows — it is a way to subdivide and delegate permission, not to escape it.
code
bash · 9 linesaws s3control create-access-point \
--account-id 111122223333 \
--bucket data-lake \
--name team-fraud-ro \
--vpc-configuration VpcId=vpc-0abc1234
aws s3api list-objects-v2 \
--bucket arn:aws:s3:us-east-1:111122223333:accesspoint/team-fraud-ro \
--prefix fraud/go deeper
Know that an S3 Access Point is a named endpoint on a bucket with its own policy, and that you pass its ARN where a bucket name normally goes.
Explain that requests are evaluated against both the access point policy and the bucket, and that the bucket policy shrinks to one delegation statement conditioned on the access point's owning account.
Design the split in practice: which prefixes get their own access point, which are VPC-origin, what stays as a bucket-level guardrail, and how CloudTrail attribution improves once each consumer has its own ARN.
Own the decision itself — when per-consumer delegation beats a single reviewed document, who administers each access point, how grants are revoked and audited at scale, and when Multi-Region Access Points or Access Grants are the better fit.
## The problem being solved A bucket policy is one document, it is limited to 20 KB, and it is edited by whoever administers the bucket. On a shared data lake this becomes a bottleneck and then a hazard: dozens of statements, each naming a team and a prefix, all in one file that a single careless edit can break for everybody. There is no way to delegate "team B manages its own access" without handing over the whole policy. ## What an access point is An S3 Access Point is a named entry point attached to exactly one bucket. Creating one gives you: - **A distinct hostname and ARN.** The ARN has the form `arn:aws:s3:<region>:<account-id>:accesspoint/<name>`, and it is used in place of the bucket name in normal S3 calls — the SDK and CLI accept it wherever a bucket name goes. - **Its own access point policy**, a resource policy in the same language as a bucket policy, usually scoped to one prefix. - **Its own Block Public Access settings**, independent of the bucket's. - **A network origin**, either Internet or VPC. A VPC-origin access point only accepts requests from the named VPC, which is a much simpler way to express "this data is private-network only" than conditioning on endpoint IDs in a bucket policy. Access point names are unique within an account and Region, so they can encode the consumer: `team-fraud-ro`, `partner-b-ingest`. ## How authorization actually works A request through an access point is evaluated against **both** the access point policy and the underlying bucket — the access point cannot grant anything the bucket does not permit. The standard pattern makes the bucket policy delegate once: ```json { "Effect": "Allow", "Principal": {"AWS": "*"}, "Action": "*", "Resource": [ "arn:aws:s3:::data-lake", "arn:aws:s3:::data-lake/*" ], "Condition": { "StringEquals": {"s3:DataAccessPointAccount": "111122223333"} } } ``` That wildcard principal looks alarming and is not public: the condition means the request must arrive through an access point owned by that account, and the access point's own policy still has to allow the caller. After delegation, the bucket policy stops growing, and each new consumer is a new access point with a small policy of its own. Direct requests to the bucket, bypassing access points, are still evaluated against the bucket policy in the ordinary way — so keep any deny-style guardrails there, where they cover both paths. ## What this buys you organisationally 1. **Blast radius.** Breaking one access point policy affects one consumer, not the lake. 2. **Delegation.** You can let a team manage its own access point policy through IAM permissions on that specific access point ARN, without granting `s3:PutBucketPolicy`. 3. **Legibility.** "What can the fraud team read?" is answered by one small document named after the fraud team, instead of by finding their `Sid` in a 20 KB file. 4. **Network scoping.** VPC-origin access points express private-only access declaratively. 5. **Auditability.** CloudTrail records the access point ARN, so usage attributes cleanly per consumer. ## Where it does not help Access points are not a performance feature and not a cost feature; they do not change durability, storage class or request pricing behaviour. They add an object to manage per consumer, which is overhead you do not want for a bucket with two readers. And they cannot widen permissions: if the bucket policy or the caller's identity policy blocks something, the access point is irrelevant. Two related mechanisms are worth naming so you choose deliberately. **Multi-Region Access Points** front the same data set in several Regions behind one global endpoint — a different problem, latency and failover rather than delegation. **S3 Access Grants** maps identities, including directory identities, to prefixes as grants rather than as policy documents, which fits estates where the unit of access is a user or group rather than a role. ## The judgment call The decision is about how many independent consumers you expect and who should own each grant. Under a handful of stable consumers, a bucket policy is simpler and one document is easier to review. Once consumers are numerous, churn independently, or belong to teams that should self-serve, the per-consumer object is the better shape — and the trigger is usually not the 20 KB limit but the moment when a change for one team requires a review from the team that owns the lake. Separately, keep the non-negotiables — encryption requirements, organisation-only denies — in the bucket policy. Guardrails belong on the resource everyone shares; grants belong on the per-consumer object.
- Does an access point policy let a consumer do something the bucket policy forbids?No. A request through an access point is authorized against the access point policy and the underlying bucket together, so an access point can only narrow. That is what makes the delegation safe: handing a team control of its own access point policy cannot let it escape the guardrails you keep on the bucket.
- The delegation statement uses "Principal": "*". Why is that not a public bucket?Because the condition requires the request to arrive through an access point owned by your account, and that access point's own policy must still allow the caller. No anonymous request satisfies both. It is the AWS-documented delegation idiom, though it is worth pairing with Block Public Access left fully enabled so the intent is unambiguous.
- When would you not bother with access points?When the consumer list is short and stable, or when one team owns both the bucket and every reader. Each access point is another object to create, name, audit and clean up. The trigger to adopt them is organisational — consumers churning independently, or grant changes needing a review from the lake's owners — rather than a raw count of statements.
- Where should guardrails live once access points are in use?On the bucket policy, because direct requests to the bucket are still evaluated there and every access point request is evaluated against the bucket as well. Denies such as organisation-only access or required encryption belong on the shared resource; per-consumer grants belong on the per-consumer access point.
The bucket policy is a single guest list on the main door that everyone must edit; access points are separate side doors, each with its own short list that its own team maintains.
saying these in an interview costs you the question
- Thinks an access point can grant more than the bucket policy allows
- Reads the delegation statement's wildcard principal as public access
- Treats access points as a performance or cost optimisation
- Puts organisation-wide guardrails in each access point policy
- Assumes callers need a new SDK rather than passing the ARN as the bucket