skip to content

Event Notifications & Integrations

S3 can fire a notification to Lambda, SQS, SNS, or EventBridge whenever objects arrive, which is how most AWS ingest pipelines actually start. The same bucket also fronts static sites and serves as the landing zone for a data lake.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

An S3 bucket is configured with S3 Event Notifications that invoke a Lambda function on every object upload. What delivery guarantees does S3 give for those notifications, and what does that force you to build into the handler?

level: middleimportance: must knowfreq 62%

answer

  1. at-least-once, not exactly-once
  2. duplicates and out-of-order arrivals
  3. one field orders same-key events
  4. versionId or ETag as dedupe key
  5. no backfill for pre-existing objects

basics

~20 s

S3 Event Notifications are at-least-once and not ordered. The same upload can produce two events and events for one key can arrive out of order, so the handler must be idempotent; the record's sequencer field lets you order events for the same key.

solid answer

~40 s

S3 promises at-least-once delivery, not exactly-once. A single `PutObject` normally produces one event, but S3 can deliver the same notification twice, so the consumer has to be idempotent — key its work on bucket plus object key plus version ID or ETag, and make a repeat a no-op rather than a second charge or a second row. There is also no ordering guarantee: a create and a delete for the same key can arrive in either order. For that, each record carries an `s3.object.sequencer` string, and comparing sequencers lexicographically tells you which event happened later for that key. Delivery is usually within seconds but is explicitly not bounded, so nothing downstream should assume a deadline. Finally, notifications only fire for events after the configuration was created — existing objects are never backfilled.

code

javascript · 20 lines
javascript
const lastSeen = new Map(); // in production: a durable store keyed per object

exports.handler = async (event) => {
  for (const record of event.Records) {
    const key = decodeURIComponent(record.s3.object.key.replace(/\+/g, ' '));
    const seq = record.s3.object.sequencer;
    const previous = lastSeen.get(key);

    // Pad to equal length, then compare lexicographically.
    if (previous && pad(previous, seq.length) >= pad(seq, previous.length)) {
      continue; // duplicate or stale event for this key
    }
    lastSeen.set(key, seq);
    await process(record.s3.bucket.name, key, record.s3.object.versionId);
  }
};

function pad(value, width) {
  return value.padStart(width, '0');
}

go deeper

for a junior

Know that one upload triggers the notification and that the message carries the bucket name and object key, not the file contents — you fetch the object yourself with a GetObject call.

for a middle

Be ready to state at-least-once and unordered plainly, then show the mechanics: name a concrete idempotency key from the record and explain what the sequencer field is for.

for a senior

Show the production judgment: where you put a durable queue so failures are retryable, how you back-fill existing objects with a batch job, and what metric proves the dedupe layer is working.

for a principal

Own the tradeoff between building idempotency into every consumer and centralising deduplication in one ingest layer, and be clear about which downstream contracts can tolerate an unbounded delivery delay.

## What S3 actually promises When you attach a notification configuration to a bucket, S3 watches for the event types you named (`s3:ObjectCreated:*`, `s3:ObjectRemoved:*`, `s3:LifecycleExpiration:*`, and so on) and pushes a small JSON record to your destination. Two properties of that push decide how you must write the consumer: - **At-least-once delivery.** S3 is designed to deliver each event at least once. It is not exactly-once. Duplicates are rare but real, and they are not a bug you can configure away. - **No ordering guarantee.** Events for the same object key can arrive out of the order in which the writes happened. A quick overwrite-then-delete can surface as delete-then-overwrite at the consumer. There is also no delivery deadline. AWS describes notifications as typically arriving in seconds, but occasionally taking a minute or longer. Any SLA you promise downstream must absorb that. ## The record you receive The notification carries *metadata*, never the object body. A create event looks roughly like this: ```json { "Records": [ { "eventSource": "aws:s3", "awsRegion": "us-east-1", "eventTime": "2026-08-21T12:00:00.000Z", "eventName": "ObjectCreated:Put", "s3": { "bucket": { "name": "ingest-bucket" }, "object": { "key": "uploads/report.csv", "size": 10240, "eTag": "9b2cf5...", "versionId": "3sL4kqtJlcpXro...", "sequencer": "0062F0A1B2C3D4E5F6" } } } ] } ``` Two fields do the heavy lifting. `versionId` (present when the bucket is versioned) identifies exactly which write this is. `sequencer` is an opaque hexadecimal string that increases for successive events on the *same key*: compare two sequencers as strings — left-padding the shorter with zeros — and the larger one is the later event. It is meaningless across different keys. One more trap in the record: the `key` is URL-encoded, so a space arrives as `+` and a slash inside a prefix stays literal. Decode it before you call back into S3. ## Why the handler must be idempotent Idempotent means running the handler twice on the same event leaves the same end state as running it once. Concretely: - **Derive an identity, don't invent one.** `bucket/key/versionId` is a natural idempotency key on a versioned bucket; on an unversioned bucket use `bucket/key/eTag` plus sequencer. - **Make the write conditional.** Writing a derived object to a deterministic destination key (`thumbnails/<same-key>.jpg`) is naturally idempotent — the second run overwrites with identical bytes. Appending a row to a ledger or charging a card is not; those need a conditional insert on the idempotency key. - **Guard ordering separately.** If your state machine cares that a delete beats a create, store the last-seen sequencer per key and drop any event whose sequencer is smaller. ## What happens when delivery fails S3 hands the event to a destination — Lambda, SQS, SNS, or the EventBridge default bus — and the destination's own semantics take over from there. If the destination is misconfigured (a queue policy that does not allow S3 to publish, a deleted function), the event is not parked anywhere for you to replay later. That is the main reason production ingest pipelines put a durable queue between S3 and the worker rather than invoking business logic directly: the queue is where retries and a dead-letter path can live. If you need replay of the notifications themselves, routing through EventBridge and enabling an archive on the bus is the mechanism that gives it to you. ## The silent gap: existing objects Enabling notifications does nothing for objects already in the bucket. There is no backfill. To process history you drive it yourself — generate the object list (S3 Inventory or a listing) and run an S3 Batch Operations job that invokes the same Lambda per object. Because the handler is already idempotent, re-processing an object the live pipeline also caught is harmless, which is exactly the payoff for building it that way. ## How to say it in an interview "At-least-once and unordered" is the whole answer in four words; everything else is what you do about it. The strong version adds the concrete idempotency key, mentions `sequencer` for same-key ordering, and notes that existing objects need a batch backfill.

  • The bucket is not versioned, so the event has no versionId. What do you use as an idempotency key?
    Fall back to bucket, key, ETag and sequencer together. The ETag distinguishes different content at the same key, and the sequencer distinguishes successive writes. It is weaker than a version ID — two writes of byte-identical content collapse to one identity — which is usually acceptable, since reprocessing identical bytes produces identical output. If it is not acceptable, turn versioning on.
  • You enabled notifications on a bucket that already holds ten million objects. How do you process the backlog?
    S3 will not replay them, so you drive the backlog explicitly. Produce an object manifest with S3 Inventory or a listing, then run an S3 Batch Operations job that invokes the same Lambda for each object in the manifest. Because the handler is idempotent, overlap between the backfill and live notifications is harmless, and the job gives you per-object success and failure reporting.
  • How would you detect that duplicate deliveries are actually happening in production?
    Emit a metric from the idempotency check itself: count every event you skip because its identity was already recorded. That distinguishes true S3 duplicates from your own retries only if you also tag the source. Pair it with a count of received events; a persistent nonzero skip rate confirms duplicates and tells you the dedupe layer is earning its keep.

Treat each notification like a postcard that the post office may deliver twice and may deliver out of order: you act on what it says, but you write down which postcards you have already acted on.

saying these in an interview costs you the question

  • Claims S3 notifications are exactly-once
  • Assumes events for one key always arrive in order
  • Ignores duplicates because they are rare
  • Thinks enabling notifications replays existing objects
  • Believes the event payload contains the object body

context

open as a page

An upload to a single S3 bucket must trigger three independent consumers: a thumbnailer, an antivirus scan, and an audit writer. How do you wire that with S3 event notifications, and what does routing through EventBridge change compared with S3's native destinations?

level: seniorimportance: must knowfreq 58%

basics

~20 s

S3 refuses overlapping notification configurations for one event type, so three consumers on the same prefix need a fan-out point: either one SNS topic with three subscribers, or EventBridge notifications with three independently filtered rules. EventBridge adds richer filtering, archive and replay, and cross-account routing.

open as a page

An S3 bucket receives many kinds of uploads. Using S3 Event Notifications, how do you trigger a function only for .jpg objects written under the uploads/ prefix, and what kinds of matching are not possible?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Attach a Filter with two key FilterRules on the notification configuration: prefix uploads/ and suffix .jpg. S3 only supports these two literal, case-sensitive string rules — there are no wildcards, no regular expressions, and no filtering on tags, size, or metadata.

open as a page

You are serving a static site from an S3 bucket. What is the difference between S3's static website endpoint and putting CloudFront with Origin Access Control in front of the bucket's REST endpoint, and which do you choose?

level: middleimportance: should knowfreq 50%

basics

~20 s

The S3 website endpoint is HTTP-only and needs a public bucket, but it maps directory paths to index documents and serves custom error pages. CloudFront with Origin Access Control keeps the bucket fully private, adds HTTPS on your own domain and caching, but does no directory-index mapping below the root.

open as a page

A Lambda function is triggered by s3:ObjectCreated:* on a bucket and writes its processed output back into that same bucket. What goes wrong, and how do you design around it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The output write fires the same notification, so the function invokes itself in an unbounded loop that burns concurrency, S3 request charges and Lambda duration until someone stops it. Fix it by writing output to a different bucket, or by fencing input and output with disjoint prefix or suffix filters.

open as a page

You own the S3 landing zone for a data lake that receives a few million small events per day from many producer teams. How do you lay out buckets, prefixes and object sizes, and what does getting it wrong cost you later?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Keep a raw immutable landing zone separate from curated output, partition prefixes by a stable dimension such as source and ingestion date, and aggregate events into files of tens to hundreds of megabytes instead of millions of tiny objects. Reorganising a lake later means copying everything.

open as a page