skip to content

During an incident you need to search a dozen CloudWatch log groups belonging to different services, some of them in a second AWS account, from a single Logs Insights query. How do you do that, and what has to be in place first?

level: middleimportance: should knowfreq 40%

answer

  1. one query, a set of log groups
  2. @log carries account and log group
  3. @logStream is the instance, not the service
  4. cross-account needs a prior link
  5. every added log group is more scanned

basics

~20 s

One Logs Insights query can span many log groups: select them individually or by name prefix, and use the @log field to tell events apart. Reaching another account first requires CloudWatch cross-account observability linking that account to your monitoring account.

solid answer

~50 s

Within one account and Region it is just a selection: in the console you tick several log groups or match them by name prefix, and over the API `StartQuery` accepts `logGroupNames` or `logGroupIdentifiers`. Each returned event then carries `@log`, which is `accountId:logGroupName`, so you can `stats count(*) by @log` to see which service is actually producing the errors — `@logStream` will not tell you that, because it identifies the instance. There is a ceiling on how many log groups one query may span, and every one you add multiplies the data scanned, so a broad prefix is expensive as well as noisy. Crossing accounts needs setup done in advance: CloudWatch cross-account observability, where the monitoring account creates a sink and each source account creates a link to it that shares its Logs data. Once linked, the monitoring account can select the source accounts' log groups — by ARN over the API — in the same query. Without that link there is no query-time credential trick that will work.

go deeper

for a junior

Know that one Logs Insights query can select several log groups at once, and that @log tells you which log group each returned event came from.

for a middle

Explain the difference between @log and @logStream, why adding log groups raises the data scanned, and that StartQuery takes log group names or ARNs.

for a senior

Show that cross-account querying is prior setup — a monitoring-account sink and per-source links sharing Logs — and handle the field-name mismatch across services rather than assuming a common schema.

for a principal

Decide where the monitoring account sits and what each source account shares, and enforce the logging field conventions that make a fleet-wide query meaningful instead of merely possible.

## Many log groups, one query A Logs Insights query is scoped to a *set* of log groups, not to one. In the console you select them from the picker, either individually or by matching a name prefix; over the API, `StartQuery` takes `logGroupNames` (names in the current account) or `logGroupIdentifiers` (names or ARNs, and ARNs are what you need for log groups in another account). The query language itself does not change at all. The field that makes a multi-log-group result usable is `@log`, which holds `accountId:logGroupName` for each event. A first move worth memorising during an incident is: ``` filter @message like /ERROR/ | stats count(*) as errors by @log | sort errors desc ``` That tells you *which service* is generating the noise before you go read individual lines. `@logStream` is not a substitute — a stream is one instance, task or Lambda execution environment, so grouping by it tells you which host is unhappy, not which service. ## What it costs you Every log group added to the selection is more data read, and Logs Insights bills by data scanned. A prefix such as `/aws/lambda/` may look convenient and quietly pull in fifty functions' worth of logs. There is also a cap on how many log groups a single query may span — tens, not thousands — so "search everything" is not an available strategy even when you are willing to pay. Pick the log groups that could plausibly hold the answer, and lean on a short time range to keep the cost of casting a wide net bounded. ## Fields do not line up across services The practical friction in a cross-service query is schema, not syntax. Service A logs JSON with `http.status`, service B logs a plain line, service C names the same thing `statusCode`. A `filter http.status >= 500` silently matches nothing for the services that do not have that field — nulls do not error. Three ways out, in increasing order of goodness: filter on `@message` text, which works everywhere but is crude; write per-service branches as separate queries; or fix the logging convention so a shared set of field names exists across the fleet. Cross-service querying is the moment a logging standard proves its worth. ## Crossing account boundaries Log groups live in an account and a Region, and a query runs against one Region. To read another account's log groups you need **CloudWatch cross-account observability** set up beforehand: - The **monitoring account** creates a *sink* — the object that is willing to receive shared observability data. - Each **source account** creates a *link* to that sink and declares which data types it shares: Logs, Metrics, Traces, and Application Insights resources among them. Sharing Logs is what makes its log groups queryable. - Access is scoped by the link's configuration, and the source account can restrict which log groups are shared. These sink and link resources belong to the Observability Access Manager (`oam`) service. Once the relationship exists, the monitoring account's console shows source-account log groups in the picker with their account labelled, and API callers pass the log group ARNs in `logGroupIdentifiers`. The `@log` field's account-id prefix is what keeps the merged results legible. What this is *not*: it is not a way to assume a role at query time, and it is not something you can arrange mid-incident. The corollary for design is that the cross-account link belongs in the landing-zone setup, alongside the accounts themselves — the middle of an outage is the wrong moment to discover that the payments account was never linked. ## Saving the result Once a cross-service query works, save it. Saved queries live in the account and are visible to anyone with the right permission, which is how a multi-log-group query written painfully at 3am becomes the first thing the next responder runs. A saved query plus a note about which log groups it expects is worth more to an incident runbook than a paragraph of prose.

  • Grouping by @logStream instead of @log gives hundreds of rows. Why?
    A log stream is one producer — an instance, an ECS task, a Lambda execution environment — so a busy service has many of them. `@log` is the account and log group, which is the service-level identity you actually want when comparing across services. Group by `@log` to find the guilty service, then by `@logStream` inside it to find the guilty instance.
  • A cross-service filter on a JSON field returns nothing for half the services. What is going on?
    Those services do not emit that field, and a filter on a null field simply matches no events rather than erroring. Either fall back to matching text in `@message`, or run per-service queries. The durable fix is a shared logging convention so that field names such as status, trace id and tenant mean the same thing everywhere.
  • Can you set up cross-account log querying during an incident?
    Realistically no. The sink in the monitoring account and the link from each source account are provisioned resources, and the link governs which data types are shared. Treat it as landing-zone work done when accounts are created. The incident-time consolation is a subscription pipeline if one already exists, not a new trust relationship.

saying these in an interview costs you the question

  • Thinks a query can only cover one log group
  • Uses @logStream to identify the service
  • Believes assuming a role at query time crosses accounts
  • Selects a broad prefix without considering scan cost
  • Assumes every service exposes the same JSON fields

context