You need every line landing in an Amazon CloudWatch Logs log group delivered to another system in near real time. Explain what a subscription filter is, which destinations it can send to, and how you would choose between them.
answer
- push, not pull
- forward-only, as events arrive
- gzip and base64 before you can read it
- the per-group quota is tiny
- retries, but no infinite buffer
basics
~20 sA subscription filter is a standing rule on a log group that pushes matching events, as they arrive, to Lambda, Kinesis Data Streams, Amazon Data Firehose, or a cross-account destination. Firehose suits bulk delivery to storage; Lambda suits per-event transformation.
solid answer
~50 sA subscription filter attaches to a log group and streams matching events out as they are ingested — it is the push counterpart to querying stored logs. Destinations are Lambda, Kinesis Data Streams, Amazon Data Firehose, and a cross-account destination created with `PutDestination` when the receiver is in another account. Choose by what you need downstream: **Firehose** when the target is S3, OpenSearch or a partner sink and you want buffering, retry and no code to run; **Lambda** when each record needs transforming, routing or forwarding to an API; **Kinesis Data Streams** when several independent consumers must read the same feed, or you need ordering and replay. Payloads arrive gzip-compressed and base64-encoded, so a Lambda consumer must decompress `event.awslogs.data` before doing anything. Two constraints shape designs: a small hard quota of subscription filters per log group, and no back-pressure — a chronically undersized destination shows up as gaps, not as queueing.
code
python · 14 linesimport base64
import gzip
import json
def handler(event, context):
raw = base64.b64decode(event["awslogs"]["data"])
payload = json.loads(gzip.decompress(raw))
if payload["messageType"] == "CONTROL_MESSAGE":
return # validation ping sent when the subscription is created
for record in payload["logEvents"]:
print(payload["logGroup"], record["timestamp"], record["message"])go deeper
Know that a subscription filter pushes log events out of a log group in near real time, and that Lambda, Kinesis Data Streams and Amazon Data Firehose are the destinations.
Explain the payload — gzip plus base64 with logEvents inside — the two permission models (a role for Kinesis/Firehose, a resource policy for Lambda), and that filters only see events ingested after they are created.
Justify the destination choice for a given workload, design around the tiny per-log-group quota with a fan-out point, and know that delivery failure appears as silent downstream gaps, so alarms belong on the destination's metrics.
Own the estate-wide shape: an account-level subscription policy so new log groups are captured automatically, a single central logging account destination, and an explicit stance on paying twice for every byte versus shortening retention at the source.
## The mechanism A **subscription filter** is a standing rule on a log group: as events are ingested, those matching its `filterPattern` are delivered to a destination in near real time. Same pattern syntax as a metric filter; entirely different output — raw events rather than a number. An empty pattern means everything. ```bash aws logs put-subscription-filter \ --log-group-name /myapp/api \ --filter-name to-firehose \ --filter-pattern "" \ --destination-arn arn:aws:firehose:us-east-1:111122223333:deliverystream/logs \ --role-arn arn:aws:iam::111122223333:role/CWLtoFirehoseRole ``` Permissions differ by destination and this is a routine stumbling block: for **Kinesis Data Streams and Firehose** you pass a `roleArn` that CloudWatch Logs assumes to write; for **Lambda** you do not pass a role — you add a resource-based permission on the function allowing the `logs.amazonaws.com` principal to invoke it. ## The payload Every destination receives the same shape: a gzip-compressed, base64-encoded JSON document containing `messageType`, `owner`, `logGroup`, `logStream`, `subscriptionFilters` and `logEvents[{id, timestamp, message}]`. A Lambda consumer must decompress before it can read anything, and must tolerate a `messageType` of `CONTROL_MESSAGE` — the test record CloudWatch Logs sends when validating a destination. Code that assumes `DATA_MESSAGE` throws on the very first delivery. ## Choosing a destination **Amazon Data Firehose** — the default for bulk. It buffers by size or time, retries, can back up failed records to S3, and delivers to S3, OpenSearch, Redshift and partner destinations without you running anything. Choose it when the answer is "put these logs somewhere durable and searchable" and you do not want an operational footprint. The trade is latency measured in buffering interval rather than milliseconds. **Lambda** — the default for transformation. Choose it when each record needs redaction, reshaping, enrichment or forwarding to an HTTP endpoint. You now own concurrency, errors and retries: a high-volume log group can drive a lot of invocations, so consider reserved concurrency so it cannot starve the rest of the account, and remember failures here mean lost deliveries, not queued ones. **Kinesis Data Streams** — choose it when more than one independent consumer needs the same feed, when you want ordering within a shard, or when you want a replay window so a downstream outage does not lose data. The cost is that you now size and operate shards. **Cross-account destination** — when the receiving system lives in a security or logging account, the receiver creates a destination with `PutDestination` and a destination policy naming the sending account; the sender then subscribes to that destination ARN. This is the standard shape for a centralised logging account. ## Constraints that actually shape the design **Quota per log group.** Only a very small number of subscription filters may exist on one log group — two, at the time of writing. This is a hard architectural constraint: you cannot let five teams each attach their own tap. The usual resolution is one subscription to a fan-out point (Kinesis Data Streams or an EventBridge-style hub) that everybody else reads from. **Account-level subscriptions.** Rather than configuring every group, `PutAccountPolicy` with a `SUBSCRIPTION_FILTER_POLICY` applies one subscription across the account's log groups with an optional selection criteria — the right tool for a security team that must capture everything without chasing new log groups as teams create them. **No back-pressure.** CloudWatch Logs retries a failing or throttling destination, but it will not buffer indefinitely. An undersized Kinesis stream returning throughput-exceeded errors, or a Lambda hitting a concurrency ceiling, eventually manifests as **missing events downstream while the log group itself is complete** — which is exactly the failure that erodes trust in a log pipeline. Alarm on the destination's own error and throttle metrics, not on the log group. **Cost.** You continue to pay CloudWatch Logs ingestion for every byte, and then pay the destination as well. Streaming everything to a second system is a doubling, not a substitution — if the goal is cheaper long-term storage, pair the subscription with a short retention on the log group. ## What it is not Subscription filters are not the tool for "let me watch this log right now" — that is Live Tail (`StartLiveTail`), an interactive session streaming events to your console or CLI. Nor are they the tool for a one-off bulk extract of history: subscriptions are forward-only, so getting existing data out means an export to S3 (`CreateExportTask`), which is asynchronous, batch, and does not backfill your stream either. Being able to name all three — subscription for continuous, export for historical bulk, Live Tail for interactive — is what a senior answer sounds like.
- Three teams each want their own real-time copy of one log group's events. How do you serve them?Not with three subscription filters — the per-log-group quota is far too small. Subscribe once to a fan-out point, typically a Kinesis Data Stream, and let each team consume it independently with its own iterator position. That also isolates a slow consumer from the others, which parallel subscriptions would not do.
- Downstream is missing events but the log group clearly has them. Where do you look?At the destination's own failure metrics. Subscription delivery retries but does not queue forever, so a throttled Kinesis stream, a Lambda at its concurrency ceiling, or a Firehose delivery failure produces silent gaps. Check the destination's throttle and error metrics and Firehose's S3 error prefix, not the log group, which is complete by definition.
- You need the last month of an existing log group in S3. Does a subscription filter help?No — subscriptions are forward-only and never see history. Use an export task to S3, which is an asynchronous bulk job over a time range, and add the subscription filter alongside it so everything from now on flows continuously. The two mechanisms cover different halves of the timeline.
- Why does a Lambda subscription consumer often fail on its very first delivery?Because CloudWatch Logs sends a validation record with messageType CONTROL_MESSAGE when the subscription is created, and code that assumes every payload is a DATA_MESSAGE with logEvents throws immediately. Decompress, check messageType, and return early on control messages before touching the event list.
saying these in an interview costs you the question
- Thinking a subscription filter replays historical log events
- Reading the Lambda payload without gunzipping and base64-decoding it
- Assuming you can attach many subscription filters to one log group
- Believing CloudWatch Logs buffers indefinitely when the destination throttles
- Expecting to stop paying CloudWatch ingestion once logs stream elsewhere