Messaging & Eventing
How AWS services talk to each other asynchronously: SQS queues with visibility timeouts and dead-letter queues, SNS fan-out, EventBridge rules and buses, and Kinesis for streaming. You learn which one fits a given delivery guarantee, and when choreography beats explicit orchestration.
part ofAWSoverview, primer and where to startread it →on this pageshowhide
explore
- SQS16 questions
- Standard vs FIFO Queues5 questions
- Consuming Messages & Visibility Timeout5 questions
- Dead-Letter Queues & Redrive6 questions
- SNS6 questions
- EventBridge11 questions
- Event Buses, Rules & Targets6 questions
- Pipes & Scheduler5 questions
- Kinesis Data Streams6 questions
- Data Firehose5 questions
- Step Functions11 questions
- State Machines & Service Integrations6 questions
- Error Handling, Retries & Observability5 questions
- SES10 questions
- Sending Email5 questions
- Deliverability & Feedback5 questions
questions
65 · 7 sectionsIn Amazon SQS, what does the visibility timeout do after a consumer calls ReceiveMessage, and why must the consumer still call DeleteMessage?
basics
~20 sReceiveMessage hides a message for the visibility timeout instead of removing it, giving one consumer a temporary exclusive lease. Only DeleteMessage removes it; if the lease expires first, the message becomes visible again and is delivered to another consumer.
In Amazon SQS, what are the differences between a standard queue and a FIFO queue, and when would you choose each?
basics
~20 sStandard SQS queues give near-unlimited throughput with best-effort ordering and possible duplicate deliveries. FIFO queues preserve order within each MessageGroupId and deduplicate messages, but cap throughput. Choose FIFO only when ordering or deduplication is genuinely required.
In an Amazon SQS redrive policy, what does maxReceiveCount actually count, and what has to happen before a message lands in the dead-letter queue?
basics
~20 smaxReceiveCount caps how many times SQS may deliver one message. Every ReceiveMessage delivery increments that message's ApproximateReceiveCount — whether the consumer failed, crashed, or never answered — and once the count passes the limit, SQS routes the message to the dead-letter queue.
In an SQS FIFO queue, what does MessageGroupId control, and how does the choice of group key affect consumer parallelism?
basics
~20 sMessageGroupId is the ordering scope of a FIFO queue: SQS delivers messages sharing a group id strictly in order, one in flight at a time, while different groups proceed independently. The number of distinct group ids therefore sets the maximum consumer parallelism.
A worker reading from an Amazon SQS queue takes about 90 seconds per message, and operators notice the same message being processed by several workers at once. What is happening, and how do you fix it?
basics
~20 sThe visibility timeout is shorter than the handler — with the 30-second default, the lease expires while the first worker is still running, so SQS makes the message visible again and hands it to another worker. Raise the timeout above worst-case processing time, or heartbeat with ChangeMessageVisibility.
In an AWS design, why do teams subscribe an SQS queue to an Amazon SNS topic for each consumer instead of subscribing the consumers (Lambda functions, HTTPS endpoints) to the topic directly?
basics
~20 sAmazon SNS pushes once and eventually drops a message it cannot deliver. Putting an SQS queue in front of each consumer stores that consumer's copy durably, so each one retries, backs off and drains at its own rate without losing events.
How do you make one Amazon SNS subscription receive only a subset of a topic's messages, and what is the difference between filtering on message attributes and on the message body?
basics
~20 sAttach a filter policy to the subscription. It is a JSON document of accepted values that SNS evaluates before delivery, either against the publisher's message attributes (the default) or against the JSON message body when the subscription's filter policy scope is set to the body.
What does an Amazon SNS FIFO topic give you that a standard topic does not, and what does it require from publishers and subscribers?
basics
~20 sA FIFO topic preserves publish order within a message group and deduplicates repeated publishes inside a short window. Publishers must supply a message group id and a deduplication id (or enable content-based deduplication), and the ordering only survives if the subscriber is a FIFO SQS queue.
You publish to an Amazon SNS topic and the call succeeds, but the subscribed SQS queue in another AWS account never receives anything. Which permission and configuration layers do you check, and in what order?
basics
~20 sCheck the topic's access policy, the subscription's confirmation state and filter policy, the destination queue's resource policy allowing the SNS service principal to send, and the KMS key policy if the queue is encrypted. A successful Publish only proves you could write to the topic.
An HTTPS endpoint subscribed to an Amazon SNS topic is unavailable for an hour. What does SNS do with the notifications during that time, and how do you stop them from being lost?
basics
~20 sSNS retries each failed delivery according to that subscription's delivery policy and then discards the message. To preserve it, attach a redrive policy naming a dead-letter SQS queue on the subscription, and enable delivery-status logging so failures are visible.
In Amazon EventBridge, what is the difference between the default event bus, a custom event bus and a partner event bus, and how do you decide which one an application should publish to?
basics
~20 sEvery account and region has one default bus, which is where AWS services emit their own events. Custom buses are ones you create for your application's events. Partner buses receive events from a SaaS provider you have associated. Application events belong on a custom bus.
How does an EventBridge rule's event pattern decide whether an event matches, and how would you debug a rule that never fires?
basics
~20 sAn event pattern mirrors the event's JSON structure: every field named in the pattern must match, values inside an array are alternatives, and fields the pattern omits are ignored. String matches are exact and case-sensitive unless you use a content filter such as prefix or wildcard.
When an application publishes a custom event to Amazon EventBridge with the PutEvents API, what does an event entry contain, and why can a call that returns HTTP 200 still have lost events?
basics
~20 sA PutEvents entry carries Source, DetailType and Detail, plus optional EventBusName, Resources and Time. Because PutEvents is a batch call, a 200 response can still report FailedEntryCount above zero, and those rejected entries are lost unless the publisher resends them.
What is Amazon EventBridge Pipes, and when does it replace a glue Lambda function written to move events from one AWS service to another?
basics
~20 sEventBridge Pipes is a managed point-to-point connector: one source (SQS, Kinesis, DynamoDB Streams, MSK, Amazon MQ) to one target, with optional filtering and enrichment in between. It replaces glue Lambda code whose only job is polling, filtering and forwarding.
Why would you use Amazon EventBridge Scheduler instead of a scheduled EventBridge rule (a cron rule on an event bus)?
basics
~20 sEventBridge Scheduler is a separate service built for scale and per-entity timers: one-time schedules, a far higher schedule quota, a per-schedule IAM role, retry and dead-letter settings per schedule, time-zone support, and flexible time windows that spread invocations.
In Amazon Kinesis Data Streams, what is a shard, and what throughput and ordering does a single shard give you?
basics
~20 sA shard is Kinesis Data Streams' unit of capacity and ordering: roughly 1 MB/s or 1,000 records per second in, 2 MB/s of shared reads out. Records are placed by partition-key hash and are ordered only within one shard.
When would you run a Kinesis Data Streams stream in on-demand capacity mode instead of provisioned, and what do you give up by doing so?
basics
~10 sOn-demand suits unpredictable or new workloads: AWS manages shard count and you pay per stream-hour plus data volume. Provisioned is cheaper at steady, well-understood throughput but makes you size and reshard the stream yourself.
In Kinesis Data Streams, how does an enhanced fan-out consumer differ from the default shared-throughput consumer, and when is the extra cost justified?
basics
~20 sA standard consumer polls GetRecords and shares one shard's 2 MB/s read budget with every other standard consumer. An enhanced fan-out consumer registers with the stream and gets its own 2 MB/s per shard, pushed over HTTP/2, at extra cost.
Producers to a Kinesis Data Streams stream are getting ProvisionedThroughputExceededException even though the stream's total IncomingBytes sits far below shard count multiplied by 1 MB/s. What is happening, and how do you fix it?
basics
~20 sA hot shard. Kinesis enforces throughput per shard, and a low-cardinality or skewed partition key concentrates writes on one shard's hash-key range while the others idle. Fix the key distribution first; resharding alone only buys time.
How does the Kinesis Client Library distribute a stream's shards across worker instances and remember progress, and what operational problems come from that design?
basics
~20 sThe KCL keeps a DynamoDB lease table, one item per shard, holding the owning worker and the last checkpointed sequence number. Workers heartbeat to hold leases and steal stale ones, which makes processing at-least-once and puts a second billable table in the critical path.
When would you choose an Amazon Data Firehose delivery stream over a Kinesis Data Streams stream for a high-volume event feed, and what do you give up?
basics
~20 sAmazon Data Firehose is a managed delivery pipeline: it buffers records and writes them to destinations such as S3, Redshift or OpenSearch with no consumer code and no shards to manage. You give up replay, retention and sub-second latency.
A service writes records to an Amazon Data Firehose delivery stream targeting S3, but objects only appear in the bucket every few minutes. Why, and which settings control it?
basics
~20 sFirehose batches records before writing. Two buffering hints control the flush — a buffer size in MiB and a buffer interval in seconds — and whichever is reached first triggers delivery, so a low-traffic stream waits for the interval.
An Amazon Data Firehose delivery stream writes to Amazon OpenSearch Service and the domain starts rejecting writes. What does Firehose do with those records, and how do you recover?
basics
~20 sFirehose retries for the configured retry duration, then writes the still-failing records to the backup S3 bucket under an error output prefix. They are not retried again automatically — recovery means reading those objects and re-ingesting them yourself.
How would you configure an Amazon Data Firehose delivery stream so JSON events land in S3 as partitioned Parquet that Athena can query efficiently, and what constrains that setup?
basics
~20 sEnable record format conversion, which reads the schema from an AWS Glue Data Catalog table and writes Parquet, and enable dynamic partitioning to build the S3 prefix from fields in each record. Both must be planned at stream creation, and partition keys must be low cardinality.
You attach a Lambda transformation to an Amazon Data Firehose delivery stream. What must the function return for each record, and what happens to records it cannot process?
basics
~20 sThe function must return a records array with one entry per input record, each carrying the same recordId, a result of Ok, Dropped or ProcessingFailed, and base64-encoded data for Ok records. ProcessingFailed records are written to the delivery stream's S3 error output prefix.
In an AWS Step Functions Task state, how do the Retry and Catch fields interact, and which retrier fields control the wait between attempts?
basics
~20 sStep Functions matches a failure against the Retry array first and re-runs the state until that retrier's MaxAttempts is used up. Only then is the error offered to Catch. IntervalSeconds, BackoffRate and MaxAttempts define the wait between attempts.
In an AWS Step Functions Task state, what is the difference between the default request-response integration, the .sync pattern, and .waitForTaskToken?
basics
~20 sRequest-response calls the API and moves on as soon as it returns. The .sync suffix makes Step Functions wait until the underlying job reaches a terminal state and returns its result. The .waitForTaskToken suffix pauses the execution until something calls SendTaskSuccess or SendTaskFailure with the injected token.
In AWS Step Functions, how do Standard and Express workflows differ, and how would you choose between them?
basics
~20 sStandard workflows are durable and auditable: they run up to a year, execute exactly once, keep a queryable execution history, and bill per state transition. Express workflows run up to five minutes, are at-least-once, log to CloudWatch, and bill by requests plus duration.
In AWS Step Functions, what does the error name States.ALL match inside a Retry or Catch ErrorEquals array, and where is it allowed to appear?
basics
~20 sStates.ALL is the Step Functions wildcard error name that matches nearly every error a state can raise. It must be the only entry in its ErrorEquals array and must sit in the final retrier or catcher, because they are evaluated top to bottom.
In AWS Step Functions, what are the required top-level fields of an Amazon States Language definition, and how do StartAt, Next and End control the flow?
basics
~20 sAn Amazon States Language definition requires two top-level fields: StartAt and States. StartAt names the first state to run, each state's Next names the state that follows it, and End set to true finishes that branch.
A brand-new AWS account cannot email arbitrary recipients through Amazon SES. What is the SES sandbox, exactly what does it restrict, and how do you get production access?
basics
~20 sThe SES sandbox is the default restricted state of every new SES account, per Region. In it you may only send to identities you have verified (plus the SES mailbox simulator), at a low daily cap and send rate. You leave it by requesting production access for that Region.
How do you get Amazon SES bounce and complaint notifications into a system you can act on, and why is email feedback forwarding not enough at scale?
basics
~20 sAttach an SES configuration set with an event destination — SNS, Kinesis Data Firehose, CloudWatch or EventBridge — and subscribe to the Bounce and Complaint event types. Email feedback forwarding just drops human-readable mail into a mailbox nobody parses or monitors.
What is the Amazon SES account-level suppression list, and what happens when you send to an address that is on it?
basics
~20 sIt is a per-account list of addresses SES refuses to deliver to, populated automatically after a hard bounce or a complaint. A send to a suppressed address is still accepted and returns a message ID, but SES never delivers it and reports a bounce event instead.
In Amazon SES, what is the difference between Easy DKIM and BYODKIM, and when would you choose BYODKIM?
basics
~20 sEasy DKIM has SES generate and rotate the signing key pair — you publish three CNAME records and forget it. BYODKIM means you supply your own private key and publish the public-key TXT record yourself, keeping key custody and rotation.
In Amazon SES, what is the difference between verifying a domain identity and verifying an email address identity, and which should a production service use?
basics
~20 sAn email identity verifies one exact address by emailing a confirmation link to it. A domain identity is verified by publishing DNS records, and then any address at that domain may be the From. Production services verify the domain, because addresses like no-reply@ have no mailbox to click a link.