Object-storage triggers, like S3 event notifications invoking a function on every object upload, are push-based and deliver events with at-least-once semantics. What does 'at-least-once' mean operationally for a handler, and what other delivery quirks (ordering, timing) does this trigger type have that a handler needs to account for?
answer
- at-least-once, not exactly-once
- duplicates from storage-side retries
- no cross-event ordering guarantee
- key idempotency off version ID/ETag
- re-check current object state, don't trust notification order
basics
~20 sAt-least-once means the same upload event might trigger your function more than once, so it can't assume it only runs exactly one time per file. Events also might not arrive in the exact order files were uploaded, and there can be a short delay before the event fires.
solid answer
~50 sObject-storage event notifications are asynchronous, push-based deliveries from the storage service to the function, and the storage service guarantees each event is delivered at least once — meaning duplicate deliveries of the same event are possible, typically due to internal retries on the storage side, and the handler cannot assume 'one event equals exactly one invocation.' There's no ordering guarantee across events either: if a client rapidly overwrites the same key multiple times, or uploads several related objects in quick succession, the notifications for those actions can arrive out of order. Handlers therefore need to be idempotent — safe to run twice on the same object — typically by keying processing on the object's version ID or ETag and checking whether that specific version was already processed, and shouldn't assume strict causal ordering between related object events without adding their own sequencing logic (e.g., checking the object's actual current state rather than trusting notification order).
go deeper
Should know the basic fact that the same event might arrive more than once and that the handler shouldn't assume it runs exactly once.
Should explain roughly why duplicates happen (delivery retries) and describe a basic idempotency approach like checking if an object was already processed.
Should articulate both the duplication and ordering quirks precisely, name version ID/ETag-based deduplication, and explain why re-checking current object state is more robust than trusting notification sequence.
Should design end-to-end pipelines that treat notifications purely as wake-up signals rather than a source of truth, choosing appropriate intermediaries (ordered queues, sequence numbers, conditional writes) when strict ordering or exactly-once effective behavior is a real business requirement.
## How an object-storage trigger fires An object-storage event trigger, like S3's event notifications, works by having the storage service itself publish a notification — to a Lambda function directly, or via an SNS topic or SQS queue as an intermediary — whenever a configured action occurs on a bucket, such as an object being created, deleted, or restored. Unlike a queue or stream trigger, this is a **push** mechanism: S3 doesn't wait to be polled, it actively delivers the event to the destination the moment (or shortly after) the underlying storage operation completes. Mechanically, the write to storage and the emission of the notification are two separate steps happening in sequence just after each other, connected by the storage service's own internal event-publishing pipeline rather than by any read-time polling. ## Why this trigger type exists This trigger type exists to let object writes drive downstream processing: - thumbnail generation on image upload, - virus scanning, - ETL ingestion, - data-lake cataloging, without the consuming application needing to poll the bucket's listing API to notice new objects, which would be both slow (polling interval delay) and expensive at scale (repeated LIST calls against a bucket). Push notification means processing can start within moments of the write completing, and the storage service handles fan-out to potentially multiple destinations. ## At-least-once delivery The central operational property to internalize is **at-least-once delivery**: the storage service guarantees that a notification for a given event will be delivered one or more times, but never explicitly guarantees exactly one delivery. This isn't a bug or a rare edge case — it's the deliberate trade-off nearly all distributed messaging systems make, because achieving true exactly-once delivery across a distributed system requires either a shared, synchronous transaction between producer and consumer (impractical at this scale) or a complex de-duplication protocol, and at-least-once with idempotent consumers is a simpler, more robust design overall. In practice, duplicates arise from the storage service's own internal retry logic — if its first attempt to deliver a notification to a downstream destination times out or gets an ambiguous response, it retries, and if that first attempt actually succeeded despite the ambiguous response, the destination now gets the same event twice. ## The ordering quirk Alongside duplication, there is no cross-event ordering guarantee. If a client uploads object A, then quickly overwrites it with a new version, then deletes it, the three resulting notifications (created, created, removed) are not guaranteed to arrive at the destination in that same sequence — network paths, internal queuing, and retry timing inside the storage service's notification pipeline can reorder them. AWS explicitly documents that for a sequence of write and delete requests to the same key, event notifications may not necessarily arrive in the same order those requests were made. This means a handler that naively assumes 'the events I receive reflect the true chronological history of this object' can act on stale information — for example, processing a 'created' notification for a version of the object that has already been deleted by the time the handler actually runs. ## The two practical mitigations - **The practical mitigation for duplication** is idempotency keyed on something more specific than the object key alone: most object-storage services attach a version ID or an equivalent unique identifier to each notification, and a handler should check — typically against its own persisted processing log or a conditional write — whether that specific version has already been fully processed before doing any non-idempotent side effect (charging a customer, sending a notification email, appending to an output file) a second time. - **The practical mitigation for ordering** is to treat notifications as a hint to 'go check the current state of this object,' rather than as an authoritative event log — a handler that, upon receiving any notification for a key, re-fetches that key's actual current metadata/state and acts on that observed reality is far more robust to reordering than one that trusts the notification's implied sequence. ## The timing quirk There is also a subtler timing quirk worth knowing: event notification delivery, while typically fast (seconds), is not instantaneous and is not contractually bounded — under load or during service issues, delivery can lag meaningfully behind the underlying write. A handler or downstream system relying on notification latency for any time-sensitive guarantee (e.g., 'the file will definitely be processed within 2 seconds of upload, guaranteed') is building on an assumption the storage service doesn't actually make. ## A concrete scenario A concrete scenario: a photo-sharing app triggers a thumbnail-generation function on every S3 object-created event. If a user re-uploads the exact same photo twice in quick succession (perhaps due to a flaky client retry), the handler could receive two 'created' notifications, potentially in either order relative to which physical object version is actually current in the bucket by the time each invocation runs. A naive handler that always generates a thumbnail from 'the current object at this key' and overwrites the same thumbnail output key is accidentally safe here (idempotent by construction, and self-correcting regardless of notification order), while a handler that tries to append thumbnail-generation records to an audit log without checking the object's version ID would end up with duplicate or out-of-order audit entries.
- How would you make an S3-triggered thumbnail generator idempotent against duplicate notifications for the same object version?Use the object's version ID (or ETag) as a deduplication key — before generating a thumbnail, check a durable store (a database table, or even a conditional check against whether the thumbnail output already exists for that version) to see if that exact version has already been processed, and skip if so. This turns a duplicate invocation into a cheap no-op instead of redundant work or a duplicate side effect.
- Why is polling the bucket's list API instead of using event notifications generally a worse design, despite avoiding the duplicate/ordering quirks?Listing objects to detect new uploads introduces a polling-interval delay before new objects are even noticed, and at any meaningful scale, repeated LIST calls against a large bucket are both slow and can incur significant request cost, especially compared to push notifications that arrive within moments of the write at effectively no extra request cost. It trades away real quirks for a strictly worse latency and cost profile, not eliminating complexity so much as relocating it into diffing logic.
- If a handler needs to guarantee it processes objects in the exact order they were uploaded, how should it actually achieve that given S3 notifications don't guarantee order?It should not rely on notification delivery order at all; instead it should have notifications route into an intermediary that supports ordering (like an SQS FIFO queue keyed on the object prefix, or a stream/log) and derive true sequence from a source of truth the handler can query, such as the object's own last-modified timestamp or an explicit sequence number embedded in metadata at write time, rather than the arrival order of notifications.
It's like a courier who guarantees your package will get delivered — possibly by knocking twice if the first knock seemed to go unanswered, and possibly delivering a follow-up parcel before the first one if routes get crossed. You have to check what's actually on your doorstep rather than trust the order the knocks happened in.
saying these in an interview costs you the question
- Assumes S3-style object notifications guarantee exactly-once delivery
- Believes notifications always arrive in the same order the underlying writes happened
- Doesn't build idempotency into handlers that have non-idempotent side effects
- Thinks a duplicate notification implies a duplicate object was actually written
- Assumes notification delivery is guaranteed within a fixed time bound