skip to content

Messaging & Eventing

How AWS services talk to each other asynchronously: SQS queues with visibility timeouts and dead-letter queues, SNS fan-out, EventBridge rules and buses, and Kinesis for streaming. You learn which one fits a given delivery guarantee, and when choreography beats explicit orchestration.

part ofAWSoverview, primer and where to startread it →
on this pageshow

explore

questions

65 · 7 sections

In Amazon SQS, what does the visibility timeout do after a consumer calls ReceiveMessage, and why must the consumer still call DeleteMessage?

level: juniorimportance: must knowfreq 82%
basics
~20 s

ReceiveMessage hides a message for the visibility timeout instead of removing it, giving one consumer a temporary exclusive lease. Only DeleteMessage removes it; if the lease expires first, the message becomes visible again and is delivered to another consumer.

open as a page

In Amazon SQS, what are the differences between a standard queue and a FIFO queue, and when would you choose each?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Standard SQS queues give near-unlimited throughput with best-effort ordering and possible duplicate deliveries. FIFO queues preserve order within each MessageGroupId and deduplicate messages, but cap throughput. Choose FIFO only when ordering or deduplication is genuinely required.

open as a page

In an Amazon SQS redrive policy, what does maxReceiveCount actually count, and what has to happen before a message lands in the dead-letter queue?

level: middleimportance: must knowfreq 78%
basics
~20 s

maxReceiveCount caps how many times SQS may deliver one message. Every ReceiveMessage delivery increments that message's ApproximateReceiveCount — whether the consumer failed, crashed, or never answered — and once the count passes the limit, SQS routes the message to the dead-letter queue.

open as a page

In an SQS FIFO queue, what does MessageGroupId control, and how does the choice of group key affect consumer parallelism?

level: middleimportance: must knowfreq 64%
basics
~20 s

MessageGroupId is the ordering scope of a FIFO queue: SQS delivers messages sharing a group id strictly in order, one in flight at a time, while different groups proceed independently. The number of distinct group ids therefore sets the maximum consumer parallelism.

open as a page

A worker reading from an Amazon SQS queue takes about 90 seconds per message, and operators notice the same message being processed by several workers at once. What is happening, and how do you fix it?

level: seniorimportance: must knowfreq 62%
basics
~20 s

The visibility timeout is shorter than the handler — with the 30-second default, the lease expires while the first worker is still running, so SQS makes the message visible again and hands it to another worker. Raise the timeout above worst-case processing time, or heartbeat with ChangeMessageVisibility.

open as a page

In an AWS design, why do teams subscribe an SQS queue to an Amazon SNS topic for each consumer instead of subscribing the consumers (Lambda functions, HTTPS endpoints) to the topic directly?

level: middleimportance: must knowfreq 74%
basics
~20 s

Amazon SNS pushes once and eventually drops a message it cannot deliver. Putting an SQS queue in front of each consumer stores that consumer's copy durably, so each one retries, backs off and drains at its own rate without losing events.

open as a page

How do you make one Amazon SNS subscription receive only a subset of a topic's messages, and what is the difference between filtering on message attributes and on the message body?

level: middleimportance: must knowfreq 58%
basics
~20 s

Attach a filter policy to the subscription. It is a JSON document of accepted values that SNS evaluates before delivery, either against the publisher's message attributes (the default) or against the JSON message body when the subscription's filter policy scope is set to the body.

open as a page

What does an Amazon SNS FIFO topic give you that a standard topic does not, and what does it require from publishers and subscribers?

level: middleimportance: should knowfreq 40%
basics
~20 s

A FIFO topic preserves publish order within a message group and deduplicates repeated publishes inside a short window. Publishers must supply a message group id and a deduplication id (or enable content-based deduplication), and the ordering only survives if the subscriber is a FIFO SQS queue.

open as a page

You publish to an Amazon SNS topic and the call succeeds, but the subscribed SQS queue in another AWS account never receives anything. Which permission and configuration layers do you check, and in what order?

level: seniorimportance: should knowfreq 42%
basics
~20 s

Check the topic's access policy, the subscription's confirmation state and filter policy, the destination queue's resource policy allowing the SNS service principal to send, and the KMS key policy if the queue is encrypted. A successful Publish only proves you could write to the topic.

open as a page

An HTTPS endpoint subscribed to an Amazon SNS topic is unavailable for an hour. What does SNS do with the notifications during that time, and how do you stop them from being lost?

level: seniorimportance: should knowfreq 46%
basics
~20 s

SNS retries each failed delivery according to that subscription's delivery policy and then discards the message. To preserve it, attach a redrive policy naming a dead-letter SQS queue on the subscription, and enable delivery-status logging so failures are visible.

open as a page

In Amazon EventBridge, what is the difference between the default event bus, a custom event bus and a partner event bus, and how do you decide which one an application should publish to?

level: middleimportance: must knowfreq 68%
basics
~20 s

Every account and region has one default bus, which is where AWS services emit their own events. Custom buses are ones you create for your application's events. Partner buses receive events from a SaaS provider you have associated. Application events belong on a custom bus.

open as a page

How does an EventBridge rule's event pattern decide whether an event matches, and how would you debug a rule that never fires?

level: middleimportance: must knowfreq 72%
basics
~20 s

An event pattern mirrors the event's JSON structure: every field named in the pattern must match, values inside an array are alternatives, and fields the pattern omits are ignored. String matches are exact and case-sensitive unless you use a content filter such as prefix or wildcard.

open as a page

When an application publishes a custom event to Amazon EventBridge with the PutEvents API, what does an event entry contain, and why can a call that returns HTTP 200 still have lost events?

level: juniorimportance: should knowfreq 52%
basics
~20 s

A PutEvents entry carries Source, DetailType and Detail, plus optional EventBusName, Resources and Time. Because PutEvents is a batch call, a 200 response can still report FailedEntryCount above zero, and those rejected entries are lost unless the publisher resends them.

open as a page

What is Amazon EventBridge Pipes, and when does it replace a glue Lambda function written to move events from one AWS service to another?

level: middleimportance: should knowfreq 50%
basics
~20 s

EventBridge Pipes is a managed point-to-point connector: one source (SQS, Kinesis, DynamoDB Streams, MSK, Amazon MQ) to one target, with optional filtering and enrichment in between. It replaces glue Lambda code whose only job is polling, filtering and forwarding.

open as a page

Why would you use Amazon EventBridge Scheduler instead of a scheduled EventBridge rule (a cron rule on an event bus)?

level: middleimportance: should knowfreq 52%
basics
~20 s

EventBridge Scheduler is a separate service built for scale and per-entity timers: one-time schedules, a far higher schedule quota, a per-schedule IAM role, retry and dead-letter settings per schedule, time-zone support, and flexible time windows that spread invocations.

open as a page

In Amazon Kinesis Data Streams, what is a shard, and what throughput and ordering does a single shard give you?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A shard is Kinesis Data Streams' unit of capacity and ordering: roughly 1 MB/s or 1,000 records per second in, 2 MB/s of shared reads out. Records are placed by partition-key hash and are ordered only within one shard.

open as a page

When would you run a Kinesis Data Streams stream in on-demand capacity mode instead of provisioned, and what do you give up by doing so?

level: middleimportance: should knowfreq 52%
basics
~10 s

On-demand suits unpredictable or new workloads: AWS manages shard count and you pay per stream-hour plus data volume. Provisioned is cheaper at steady, well-understood throughput but makes you size and reshard the stream yourself.

open as a page

In Kinesis Data Streams, how does an enhanced fan-out consumer differ from the default shared-throughput consumer, and when is the extra cost justified?

level: middleimportance: should knowfreq 45%
basics
~20 s

A standard consumer polls GetRecords and shares one shard's 2 MB/s read budget with every other standard consumer. An enhanced fan-out consumer registers with the stream and gets its own 2 MB/s per shard, pushed over HTTP/2, at extra cost.

open as a page

Producers to a Kinesis Data Streams stream are getting ProvisionedThroughputExceededException even though the stream's total IncomingBytes sits far below shard count multiplied by 1 MB/s. What is happening, and how do you fix it?

level: seniorimportance: should knowfreq 50%
basics
~20 s

A hot shard. Kinesis enforces throughput per shard, and a low-cardinality or skewed partition key concentrates writes on one shard's hash-key range while the others idle. Fix the key distribution first; resharding alone only buys time.

open as a page

How does the Kinesis Client Library distribute a stream's shards across worker instances and remember progress, and what operational problems come from that design?

level: seniorimportance: nice to knowfreq 36%
basics
~20 s

The KCL keeps a DynamoDB lease table, one item per shard, holding the owning worker and the last checkpointed sequence number. Workers heartbeat to hold leases and steal stale ones, which makes processing at-least-once and puts a second billable table in the critical path.

open as a page

When would you choose an Amazon Data Firehose delivery stream over a Kinesis Data Streams stream for a high-volume event feed, and what do you give up?

level: middleimportance: must knowfreq 74%
basics
~20 s

Amazon Data Firehose is a managed delivery pipeline: it buffers records and writes them to destinations such as S3, Redshift or OpenSearch with no consumer code and no shards to manage. You give up replay, retention and sub-second latency.

open as a page

A service writes records to an Amazon Data Firehose delivery stream targeting S3, but objects only appear in the bucket every few minutes. Why, and which settings control it?

level: juniorimportance: should knowfreq 58%
basics
~20 s

Firehose batches records before writing. Two buffering hints control the flush — a buffer size in MiB and a buffer interval in seconds — and whichever is reached first triggers delivery, so a low-traffic stream waits for the interval.

open as a page

An Amazon Data Firehose delivery stream writes to Amazon OpenSearch Service and the domain starts rejecting writes. What does Firehose do with those records, and how do you recover?

level: seniorimportance: should knowfreq 40%
basics
~20 s

Firehose retries for the configured retry duration, then writes the still-failing records to the backup S3 bucket under an error output prefix. They are not retried again automatically — recovery means reading those objects and re-ingesting them yourself.

open as a page

How would you configure an Amazon Data Firehose delivery stream so JSON events land in S3 as partitioned Parquet that Athena can query efficiently, and what constrains that setup?

level: seniorimportance: should knowfreq 42%
basics
~20 s

Enable record format conversion, which reads the schema from an AWS Glue Data Catalog table and writes Parquet, and enable dynamic partitioning to build the S3 prefix from fields in each record. Both must be planned at stream creation, and partition keys must be low cardinality.

open as a page

You attach a Lambda transformation to an Amazon Data Firehose delivery stream. What must the function return for each record, and what happens to records it cannot process?

level: middleimportance: nice to knowfreq 38%
basics
~20 s

The function must return a records array with one entry per input record, each carrying the same recordId, a result of Ok, Dropped or ProcessingFailed, and base64-encoded data for Ok records. ProcessingFailed records are written to the delivery stream's S3 error output prefix.

open as a page

In an AWS Step Functions Task state, how do the Retry and Catch fields interact, and which retrier fields control the wait between attempts?

level: middleimportance: must knowfreq 72%
basics
~20 s

Step Functions matches a failure against the Retry array first and re-runs the state until that retrier's MaxAttempts is used up. Only then is the error offered to Catch. IntervalSeconds, BackoffRate and MaxAttempts define the wait between attempts.

open as a page

In an AWS Step Functions Task state, what is the difference between the default request-response integration, the .sync pattern, and .waitForTaskToken?

level: middleimportance: must knowfreq 66%
basics
~20 s

Request-response calls the API and moves on as soon as it returns. The .sync suffix makes Step Functions wait until the underlying job reaches a terminal state and returns its result. The .waitForTaskToken suffix pauses the execution until something calls SendTaskSuccess or SendTaskFailure with the injected token.

open as a page

In AWS Step Functions, how do Standard and Express workflows differ, and how would you choose between them?

level: middleimportance: must knowfreq 78%
basics
~20 s

Standard workflows are durable and auditable: they run up to a year, execute exactly once, keep a queryable execution history, and bill per state transition. Express workflows run up to five minutes, are at-least-once, log to CloudWatch, and bill by requests plus duration.

open as a page

In AWS Step Functions, what does the error name States.ALL match inside a Retry or Catch ErrorEquals array, and where is it allowed to appear?

level: juniorimportance: should knowfreq 55%
basics
~20 s

States.ALL is the Step Functions wildcard error name that matches nearly every error a state can raise. It must be the only entry in its ErrorEquals array and must sit in the final retrier or catcher, because they are evaluated top to bottom.

open as a page

In AWS Step Functions, what are the required top-level fields of an Amazon States Language definition, and how do StartAt, Next and End control the flow?

level: juniorimportance: should knowfreq 52%
basics
~20 s

An Amazon States Language definition requires two top-level fields: StartAt and States. StartAt names the first state to run, each state's Next names the state that follows it, and End set to true finishes that branch.

open as a page

A brand-new AWS account cannot email arbitrary recipients through Amazon SES. What is the SES sandbox, exactly what does it restrict, and how do you get production access?

level: juniorimportance: must knowfreq 76%
basics
~20 s

The SES sandbox is the default restricted state of every new SES account, per Region. In it you may only send to identities you have verified (plus the SES mailbox simulator), at a low daily cap and send rate. You leave it by requesting production access for that Region.

open as a page

How do you get Amazon SES bounce and complaint notifications into a system you can act on, and why is email feedback forwarding not enough at scale?

level: middleimportance: must knowfreq 62%
basics
~20 s

Attach an SES configuration set with an event destination — SNS, Kinesis Data Firehose, CloudWatch or EventBridge — and subscribe to the Bounce and Complaint event types. Email feedback forwarding just drops human-readable mail into a mailbox nobody parses or monitors.

open as a page

What is the Amazon SES account-level suppression list, and what happens when you send to an address that is on it?

level: juniorimportance: should knowfreq 48%
basics
~20 s

It is a per-account list of addresses SES refuses to deliver to, populated automatically after a hard bounce or a complaint. A send to a suppressed address is still accepted and returns a message ID, but SES never delivers it and reports a bounce event instead.

open as a page

In Amazon SES, what is the difference between Easy DKIM and BYODKIM, and when would you choose BYODKIM?

level: middleimportance: should knowfreq 50%
basics
~20 s

Easy DKIM has SES generate and rotate the signing key pair — you publish three CNAME records and forget it. BYODKIM means you supply your own private key and publish the public-key TXT record yourself, keeping key custody and rotation.

open as a page

In Amazon SES, what is the difference between verifying a domain identity and verifying an email address identity, and which should a production service use?

level: middleimportance: should knowfreq 61%
basics
~20 s

An email identity verifies one exact address by emailing a confirmation link to it. A domain identity is verified by publishing DNS records, and then any address at that domain may be the From. Production services verify the domain, because addresses like no-reply@ have no mailbox to click a link.

open as a page