AWS
AWS is the default assumption in most backend and platform interviews, so you are expected to name services without hesitating. This branch walks the pillars — identity, networking, compute, storage and databases, messaging, observability and cost — in the order you would actually assemble a system.
on this pageshowhide
guide
overview
~2 minAWS is the platform most backend, platform and system-design interviews assume unless told otherwise, so you are expected to reach for the right service by name and then defend the choice. Naming is the easy part. What gets probed is whether you know what a service promises and what it leaves to you: who may make a call, what can reach what over the network, what survives an Availability Zone going down, which delivery guarantee a queue gives, and what the design costs at the end of the month. A strong AWS answer picks a service, states the property that made it the right pick, and names the trade it accepts. The hub follows the order in which you would assemble a system. [IAM and security](/topics/cloud-aws-iam) comes first, because every action on AWS is an authorized API call: policies, roles and temporary credentials, organization guardrails, KMS and secrets. [Networking and CDN](/topics/cloud-aws-networking) is the part you draw — VPCs, subnets and routing, load balancers, Route 53 and CloudFront. [Compute](/topics/cloud-aws-compute) is where code runs, from EC2 through ECS, EKS and Fargate to Lambda. [Storage and databases](/topics/cloud-aws-storage) covers S3, block and file volumes, and the managed databases and caches. [Messaging and eventing](/topics/cloud-aws-messaging) connects services asynchronously with SQS, SNS, EventBridge, Kinesis and Step Functions. [Observability and cost](/topics/cloud-aws-observability-cost) closes the loop with CloudWatch, X-Ray, CloudTrail and the bill. Junior rounds check vocabulary and defaults: what makes a subnet public, what a role is, how two queue types differ. Middle rounds become choices between neighbouring services — ALB or NLB, Multi-AZ or a read replica, Lambda or containers. Senior and staff rounds are designs and incidents: a multi-account landing zone, a bill that doubled, an alarm that stayed quiet during an outage, a failover that did not behave the way the diagram promised. Start with IAM and the VPC, because nearly every later answer assumes both. Take compute and storage next, together, since the choice of one constrains the other, and add messaging and observability once you can draw a working system end to end.
primer
### Every action is an authorized API call Consoles, CLIs, SDKs and AWS services themselves all end up sending signed API requests, and each one is judged against policies before anything happens. A **principal** asks to perform an **action** on a **resource**, optionally under **conditions**. Permission can come from the caller's side (identity-based policies) or the target's side (resource-based policies), and it can be capped from above by permissions boundaries and organization policies. With no matching grant, the answer is no. That single model explains cross-account access, why a Lambda function has two separate permission mechanisms, and why least privilege is something you design rather than declare. ### Credentials should be short-lived Workloads and people should get permissions by assuming **roles** and receiving temporary credentials, not by carrying permanent access keys. EC2 instances, containers and functions each have a way to receive a role; engineers sign in once through a central identity source. In an interview, a long-lived key on a server or a laptop is a finding, not a detail. ### Regions and Availability Zones are the failure model A Region is a separate geography; an Availability Zone is an isolated group of data centers inside it. Every resource is global, Regional or zonal in scope, and that scope decides what a single failure takes down. Highly available designs spread across zones, and managed services differ in how much of that spreading they do for you. Crossing zones is also a line on the bill, so availability and cost are argued together. ### The network is private until you route it A VPC is an isolated address space: nothing enters or leaves unless routes, gateways or endpoints you configure allow it. Subnets take their character from their route tables, and traffic is filtered at two layers with different semantics. Most "why can't it connect" questions are answered by walking routes, then filters, then DNS. ### Managed is a spectrum EC2 gives you servers; ECS and EKS run containers on servers or on Fargate; Lambda runs functions and hides the host entirely. Each step along that line hands AWS more of the operational work — and more of the **shared responsibility** for it — in exchange for limits on duration, runtime control and how the cost scales. The choice is argued from traffic pattern, job length and state, not from preference. ### Choose storage by access pattern S3 stores objects addressed by key, EBS gives one instance a block device, EFS shares a file system, RDS and Aurora run relational engines, and DynamoDB serves key-value access at almost any scale if the keys are designed for it. Each has its own durability, consistency and scaling story. Interviewers probe the mismatch: a relational model forced into DynamoDB, a shared file system used as a queue, a database scaled for writes when reads were the bottleneck. ### Asynchronous delivery is usually at-least-once Queues, topics, event buses and streams decouple producers from consumers, but most of them can deliver a message more than once, and some do not preserve order. So consumers are written to be idempotent, failures are parked in dead-letter queues, and ordering is bought explicitly where it matters. Knowing which service is a queue, which is fan-out, which is a routed bus and which is a replayable stream is most of the messaging round. ### Cost and observability are design inputs Every architecture has a bill with a shape: per-request against per-hour pricing, commitments against flexibility, and data transfer nobody drew on the diagram. Metrics, logs, traces and the API audit trail are how you show a system is healthy and find out who changed it. Senior answers name the alarm and the cost driver alongside the service.
- Region
- A separate geographic area containing several Availability Zones. Most resources and service endpoints are Regional, and data stays in its Region unless you copy it out.
- Availability Zone (AZ)
- One or more isolated data centers inside a Region, with independent power and networking. The unit of failure that highly available designs spread across.
- Principal
- The identity making a request, such as a user, a role session, an AWS service or an account. Resource-based policies name principals; identity-based policies are attached to one.
- IAM role
- An identity with permissions but no permanent credentials. Trusted principals assume it and receive temporary credentials, which is how workloads and federated users should get access.
- Resource-based policy
- A policy attached to the target resource, such as a bucket, queue, key or function, that says which principals may act on it, including principals in other accounts.
- Service control policy (SCP)
- An AWS Organizations policy that caps the permissions available in the accounts it applies to. It limits what IAM policies can grant and grants nothing itself.
- KMS key
- A key held by AWS Key Management Service, typically used to protect data keys rather than to encrypt bulk data, with its own key policy deciding who may use it.
- VPC
- A logically isolated virtual network in one Region, with IP ranges you choose, divided into subnets that each live in a single Availability Zone.
- Route table
- The routes associated with a subnet, deciding where its traffic goes: inside the VPC, to an internet or NAT gateway, a peering connection or a transit gateway.
- Security group
- A stateful, allow-only virtual firewall attached to network interfaces. Rules can reference other security groups, which is how application tiers are allowed to talk.
- NAT gateway
- A managed gateway that lets resources in private subnets open outbound connections to the internet while refusing connections started from outside. Billed per hour and per gigabyte processed.
- Execution environment
- The isolated sandbox Lambda creates to run a function. Creating a new one is what a cold start costs; later invocations can reuse it.
- Multi-AZ deployment
- An RDS configuration that keeps a synchronously replicated standby in another Availability Zone and fails over to it automatically. In its classic form it adds availability, not read capacity.
- Partition key
- The DynamoDB attribute whose hash decides which partition stores an item. How its values spread decides whether traffic is shared evenly or concentrated on hot keys.
- Dead-letter queue (DLQ)
- A queue that receives messages a consumer failed to process after a set number of attempts, so they can be inspected and redriven instead of retried forever.
- Shared responsibility model
- AWS's split of security duties: AWS secures the infrastructure services run on; the customer secures data, identities, configuration and, on EC2, the operating system.
- CloudTrail
- The service that records API calls made in an account — who called what, when and from where — for auditing, investigation and change tracking.
- Savings Plan
- A one- or three-year commitment to a steady hourly compute spend, in exchange for a discount on usage that falls within the plan's scope.
Follow one request through a typical web system. Route 53 answers the DNS query, and CloudFront serves what it can from the edge and forwards the rest to a load balancer in public subnets. The load balancer passes the request to targets — EC2 instances in an Auto Scaling group, ECS tasks or a Lambda function — that sit in private subnets and reach the internet, if at all, through a NAT gateway. The code calls other AWS services with credentials from its role, reads and writes a database or S3, and hands slow work to a queue or an event bus for another consumer. CloudWatch collects metrics and logs along the way, X-Ray stitches the hops into a trace, and CloudTrail records the API calls made against the account. The six sections map onto that path, and each leans on the others: - **IAM** is on nearly every arrow, since each call to an AWS API is signed and checked; KMS adds a second permission check wherever data is encrypted under a KMS key. - **Networking** decides which arrows exist at all: a correct policy does not help a call that has no route. - **Compute** fixes how the role arrives (instance profile, task role, execution role) and how capacity scales. - **Storage** fixes consistency, the failover story and often the largest line on the bill. - **Messaging** turns synchronous chains into buffered ones, moving the failure mode from timeouts to retries and duplicates. - **Observability and cost** read the same resources through metrics, logs, audit events and billing data. A small CloudFormation template shows three sections meeting on one function: ```yaml Resources: Worker: Type: AWS::Lambda::Function Properties: Runtime: python3.12 Handler: app.handler Code: { S3Bucket: !Ref ArtifactBucket, S3Key: worker.zip } Role: !GetAtt WorkerRole.Arn # identity: what the code may call VpcConfig: # network: where the code sits SubnetIds: [!Ref PrivateSubnetA, !Ref PrivateSubnetB] SecurityGroupIds: [!Ref WorkerSecurityGroup] JobsTrigger: Type: AWS::Lambda::EventSourceMapping Properties: EventSourceArn: !GetAtt JobsQueue.Arn # messaging: Lambda reads the queue FunctionName: !Ref Worker ``` Each block answers a different section's question. The role decides what the handler may call, and it also has to allow reading and deleting from the queue and creating network interfaces, because both the poller and the VPC attachment act with the function's permissions. The VPC settings place the function in private subnets, so any AWS service the handler calls needs a way out — a NAT gateway or a VPC endpoint. The event source mapping is what connects queue to function here. Leave out any one of the three and the design fails, each time in a different place.
- IAM →
Principals, policies and roles: the permission model every other section assumes whenever one service calls another.
- VPC →
Subnets, route tables, gateways and the two firewall layers: the network every system-design answer ends up drawing.
- Choosing a Compute Model →
How to argue for EC2, containers or Lambda from traffic shape, job duration and state, the usual opening of a design round.
- S3 →
Objects, storage classes, access control and encryption: the store almost every AWS system touches first.
- SQS →
The simplest asynchronous building block, where retries, duplicates, visibility and dead-letter queues are first met.
- CloudWatch →
Metrics, logs and alarms: how you show a design is healthy, and the tool behind most incident questions.
Calling a subnet public because of its name or an auto-assign setting: the route table decides, and it is the first thing to check when an instance cannot get out.
Giving servers, containers or CI jobs long-lived access keys when a role with temporary credentials would do; interviewers treat a stored key as a security finding.
Adding Multi-AZ to a database that is short on read capacity: the classic standby takes no queries, so read load needs a replica — see Multi-AZ vs read replicas.
Expecting an SCP that allows a service to grant it; organization policies only cap what the IAM policies inside each account can grant.
Designing for one Availability Zone, or spreading across zones without pricing the cross-zone traffic and NAT gateway processing that come with it.
Picking Lambda for long-running or steady high-throughput work without mentioning the 15-minute timeout, concurrency limits and per-invocation pricing.
Writing SQS or SNS consumers that assume each message arrives once; standard queues and most event paths are at-least-once, so handlers must be idempotent.
Making an S3 bucket public to serve files instead of putting CloudFront in front of it or handing out short-lived presigned URLs.
Treating a silent metric as a healthy one: an alarm with the wrong missing-data setting can sit idle through an outage — see the missing-data question.
Listing services without reasons; naming five services with no trade-off scores lower than naming two and defending each.
AWS has no single release number, so "which version" means which service behaviour you assume. This guide assumes behaviour current in 2026. Several changes still come up in interviews, because older courses, blog posts and answers describe the earlier state: - **S3 consistency.** Since December 2020, S3 gives strong read-after-write consistency for every object operation, overwrites, deletes and listings included. Designs that work around eventual consistency on S3 describe the old model. - **S3 secure defaults.** Since January 2023 new objects are encrypted with SSE-S3 by default, and since April 2023 new buckets start with Block Public Access on and ACLs disabled. - **Spot pricing.** Spot prices now change slowly, following long-term trends in spare capacity; the bidding model many older explanations describe no longer applies. - **EBS volume types.** gp3 separated performance from volume size, so growing a gp2 volume just to gain IOPS is now a legacy habit. - **Public IPv4.** Since February 2024 every public IPv4 address is billed hourly, attached or idle, which changed the cost argument around public subnets, NAT and IPv6. - **Renames.** Kinesis Data Firehose is now Amazon Data Firehose, and ElastiCache offers Valkey alongside Redis OSS and Memcached. Use the current names, and recognise the old ones when an interviewer uses them. When an answer depends on one of these, say which behaviour you mean.
AWS is one of three hyperscale clouds, alongside Microsoft Azure and Google Cloud. Most core building blocks exist on all three — object storage, managed Kubernetes, serverless functions, queues, managed relational databases — so a well-reasoned AWS design usually translates. What does not translate is the identity model, the networking details and the service limits. Interviewers rarely want a provider comparison; they want to hear that you chose a service for a property, because that reasoning carries across providers. Around AWS sits tooling you will be asked to place. Infrastructure is described as code with CloudFormation or the CDK, which are AWS's own, or with Terraform, which spans providers. Containers run on ECS, which is AWS-specific and simpler, or on EKS, which is standard Kubernetes and more portable at the price of more operations. Several managed services are weighed against open-source systems run yourself or elsewhere: Kinesis against Kafka, ElastiCache against self-run Redis or Valkey, SQS against RabbitMQ. The trade is nearly always operational effort against control and portability. [Choosing a compute model](/topics/cloud-aws-compute-model-selection) and [EKS in the AWS compute menu](/topics/cloud-aws-compute-eks) take the container side further.
explore
- IAM & Security69 questions
- Policy Documents & Conditions6 questions
- Policy Evaluation Logic5 questions
- Permissions Boundaries & Delegated Admin4 questions
- Roles, STS & Temporary Credentials5 questions
- Service Roles & Instance Profiles5 questions
- Identity Center & Federation5 questions
- Cognito Customer Identity5 questions
- Organizations & Service Control Policies6 questions
- KMS Keys & Envelope Encryption6 questions
- Encryption at Rest & In Transit5 questions
- Secrets Manager & Parameter Store5 questions
- CloudTrail & Access Auditing5 questions
- Detection & Compliance Services6 questions
- IAM1 questions
- Networking & CDN88 questions
- Elastic Load Balancing18 questions
- VPC27 questions
- Route 5310 questions
- CloudFront17 questions
- API Gateway16 questions
- Compute78 questions
- Choosing a Compute Model6 questions
- ECS and Fargate6 questions
- EKS in the AWS Compute Menu3 questions
- Elastic Beanstalk, App Runner and Lightsail5 questions
- EC229 questions
- AWS Lambda29 questions
- Storage & Databases82 questions
- Block & File Storage16 questions
- Managed Databases & Caches23 questions
- S342 questions
- Glacier1 questions
- Messaging & Eventing65 questions
- SQS16 questions
- SNS6 questions
- EventBridge11 questions
- Kinesis Data Streams6 questions
- Data Firehose5 questions
- Step Functions11 questions
- SES10 questions
- Observability & Cost60 questions
- X-Ray Distributed Tracing6 questions
- CloudTrail Audit Events5 questions
- Cost Management & Optimization21 questions
- Well-Architected Reviews & Trusted Advisor5 questions
- CloudWatch23 questions
questions
442 · 6 sectionsSigning in to an Amazon Cognito user pool returns three tokens. Name them, and say which one your own backend API should accept as proof that the caller is authorized.
basics
~20 sA Cognito user pool sign-in returns an ID token, an access token and a refresh token. Your API should accept the access token, which carries scopes and group claims; the ID token describes the user to the client, and the refresh token only obtains new tokens.
Your team currently gives every engineer an IAM user with long-lived access keys in each of five AWS accounts. What does AWS IAM Identity Center replace that with, and why is it considered safer?
basics
~20 sIAM Identity Center centralises workforce sign-in: each engineer exists once in one identity source, is assigned permission sets per account, and receives short-lived credentials at login instead of access keys that live on a laptop forever.
In AWS Organizations, what is the difference between the management account and a member account, and what does grouping accounts into organizational units (OUs) give you?
basics
~20 sIn AWS Organizations one management account creates the organization, pays every bill and cannot be restricted by service control policies; the rest are member accounts holding workloads. OUs group accounts so one attached policy applies to all beneath.
An AWS IAM role has the AWS managed policy AdministratorAccess attached as its identity policy and also has a permissions boundary that allows only s3:* and cloudwatch:*. What can the role actually do, and what could it do if the boundary were the only policy attached to it?
basics
~20 sEffective permissions are the intersection of the two, so the role can call only S3 and CloudWatch actions. A permissions boundary never grants anything, so with the boundary alone and no identity policy the role could do nothing.
In AWS IAM, what is the difference between an implicit deny and an explicit Deny, and what does each mean for how you write and troubleshoot policies?
basics
~20 sEvery AWS request starts denied. An implicit deny is simply the absence of any matching Allow, and adding an Allow fixes it. An explicit Deny is a statement with Effect Deny, and no Allow anywhere can override it.
An API Gateway route uses the Lambda proxy (`AWS_PROXY`) integration and every call returns HTTP 502 with `{"message": "Internal server error"}`, yet the function's own CloudWatch logs show it completing without error. What is wrong, and what must a Lambda proxy handler return?
basics
~20 sThe function returned a shape API Gateway cannot map to an HTTP response. A Lambda proxy handler must return an object with a numeric statusCode, an optional headers map, and a body that is already a string — returning a raw object or a bare value produces 502.
In a CloudFront distribution, what is the difference between an origin and a cache behavior, and how does CloudFront decide which origin serves an incoming request?
basics
~20 sAn origin is a backend CloudFront fetches from: an S3 bucket, a load balancer, any HTTP server. A cache behavior maps a URL path pattern to one origin. CloudFront tests behaviors in their configured order and falls back to the default behavior.
On an AWS Application Load Balancer, what can a listener rule match on, and how does the load balancer decide which rule applies to a given request?
basics
~20 sAn ALB listener rule matches on host header, URL path, HTTP header, HTTP method, query string, or source IP. Rules are evaluated in priority order, lowest number first; the first match wins, and unmatched requests fall through to the listener's default action.
In AWS Elastic Load Balancing, what is a target group, and what happens to a registered target once it starts failing that target group's health check?
basics
~20 sA target group is the named set of backends an Elastic Load Balancing listener forwards to, plus the health check run against each member. A target that fails enough consecutive checks is marked unhealthy and stops receiving new requests.
In an AWS VPC, how do security groups and network ACLs differ in what they attach to, what a rule can express, and how return traffic is treated?
basics
~20 sSecurity groups attach to network interfaces, are allow-only and stateful, so return traffic is automatically permitted. Network ACLs attach to subnets, allow or deny in numbered order, and are stateless, so every direction needs an explicit rule.
What is EC2 user data, at what point in an instance's life does the script run, and as which OS user?
basics
~20 sUser data is a text blob you attach to an EC2 instance that cloud-init fetches from the instance metadata service at first boot. A script starting with a shebang runs as root, once per instance — not on every reboot.
In an EC2 Auto Scaling group, what do the minimum, maximum and desired capacity settings each control, and what happens if you manually terminate one of the group's instances?
basics
~20 sDesired capacity is the instance count the group tries to run right now; minimum and maximum are the floor and ceiling that scaling can never cross. Manually terminating an instance leaves the group below desired, so it launches a replacement.
Decode the EC2 instance type name m7g.2xlarge: what does each part of the name tell you, and how many vCPUs and how much memory does that instance have?
basics
~20 sm7g.2xlarge splits into four parts: family m (general purpose, roughly 4 GiB of memory per vCPU), generation 7, processor letter g (AWS Graviton, arm64), and size 2xlarge — 8 vCPUs and 32 GiB of memory.
In EC2, what do you get and what do you give up by running an instance as Spot instead of On-Demand, and how is the Spot price actually determined?
basics
~20 sSpot Instances run on EC2 capacity AWS currently has idle, at a steep discount off On-Demand, but AWS can reclaim them at any moment with a two-minute warning. Spot prices move gradually with supply and demand; there is no bidding auction.
In Amazon ECS, explain what a task definition, a task, and a service each are, and when you would use the RunTask API directly instead of putting the workload behind a service.
basics
~20 sAn ECS task definition is the immutable blueprint: image, CPU and memory, IAM roles, logging. A task is one running copy of it. A service keeps a desired count of tasks alive and replaces failures; RunTask launches one-off jobs.
In AWS, what happens to data on an EC2 instance store volume, compared with data on an attached EBS volume, when the instance is rebooted, stopped and started again, or terminated?
basics
~20 sEC2 instance store data survives a reboot but is lost on stop, hibernate, or terminate, because it lives on disks physically inside the host. EBS volumes are network-attached and persist independently of the instance lifecycle.
An Aurora cluster gives you a cluster endpoint, a reader endpoint, and one endpoint per instance. What does each of those resolve to, and what goes wrong if an application points all of its traffic at the cluster endpoint?
basics
~20 sThe Aurora cluster endpoint always resolves to the current writer and follows failover; the reader endpoint round-robins DNS across available readers; an instance endpoint names one fixed instance. Sending everything to the cluster endpoint puts all reads on the writer while paid-for readers sit idle.
In DynamoDB, what is the difference between an eventually consistent read and a strongly consistent read, and how does that choice change the read capacity a request consumes?
basics
~20 sEventually consistent reads, the DynamoDB default, may return a slightly stale copy of an item and cost half a read unit per 4 KB. Strongly consistent reads always reflect the latest acknowledged write and cost a full read unit per 4 KB.
Amazon ElastiCache lets you run Valkey, Redis OSS, or Memcached as the cache engine. What does each option give you operationally, and for which workload would you actually choose Memcached?
basics
~10 sValkey and Redis OSS on ElastiCache add replicas, automatic failover, snapshots and authentication; Memcached has none of those but is multi-threaded and simple. Choose Memcached only for a plain, disposable, horizontally sharded key-value cache.
In Amazon Kinesis Data Streams, what is a shard, and what throughput and ordering does a single shard give you?
basics
~20 sA shard is Kinesis Data Streams' unit of capacity and ordering: roughly 1 MB/s or 1,000 records per second in, 2 MB/s of shared reads out. Records are placed by partition-key hash and are ordered only within one shard.
A brand-new AWS account cannot email arbitrary recipients through Amazon SES. What is the SES sandbox, exactly what does it restrict, and how do you get production access?
basics
~20 sThe SES sandbox is the default restricted state of every new SES account, per Region. In it you may only send to identities you have verified (plus the SES mailbox simulator), at a low daily cap and send rate. You leave it by requesting production access for that Region.
In Amazon SQS, what does the visibility timeout do after a consumer calls ReceiveMessage, and why must the consumer still call DeleteMessage?
basics
~20 sReceiveMessage hides a message for the visibility timeout instead of removing it, giving one consumer a temporary exclusive lease. Only DeleteMessage removes it; if the lease expires first, the message becomes visible again and is delivered to another consumer.
In Amazon SQS, what are the differences between a standard queue and a FIFO queue, and when would you choose each?
basics
~20 sStandard SQS queues give near-unlimited throughput with best-effort ordering and possible duplicate deliveries. FIFO queues preserve order within each MessageGroupId and deduplicate messages, but cap throughput. Choose FIFO only when ordering or deduplication is genuinely required.
In Amazon EventBridge, what is the difference between the default event bus, a custom event bus and a partner event bus, and how do you decide which one an application should publish to?
basics
~20 sEvery account and region has one default bus, which is where AWS services emit their own events. Custom buses are ones you create for your application's events. Partner buses receive events from a SaaS provider you have associated. Application events belong on a custom bus.
AWS CloudTrail shows recent API activity in the console under Event history without any setup. What does Event history actually cover, and what does creating a trail add?
basics
~20 sCloudTrail Event history is free, always on, and covers only the last 90 days of management events in the current region. A trail is the delivery configuration that writes events to an S3 bucket you own, giving you retention you control and the option to log data events.
What is an Amazon CloudWatch dashboard actually made of, what kinds of widget can it hold, and why does putting a metric on a CloudWatch dashboard not mean anyone will be told when it moves?
basics
~20 sA CloudWatch dashboard is a JSON document of widgets — metric graphs, log tables, alarm-status tiles, text — and each widget carries its own region, statistic and period. Dashboards only draw data; notification comes from CloudWatch alarms, never from a dashboard.
An application on an EC2 instance writes to a log file on disk and you want those lines in Amazon CloudWatch Logs. What are the ways log data gets into CloudWatch Logs, and what does each path require you to install or grant?
basics
~20 sThree paths: the CloudWatch agent tailing files on a host, AWS services delivering logs natively (Lambda, the ECS awslogs driver, VPC Flow Logs), or your code calling the PutLogEvents API. All three need IAM permission to create streams and put events.
Walk through the anatomy of a CloudWatch Logs Insights query — the pipeline of commands and the @-prefixed fields that are always available — and show how you would pull the 20 most recent lines containing ERROR out of one log group.
basics
~20 sA Logs Insights query is a pipeline of commands joined by the pipe character: fields or display picks columns, filter narrows rows, sort orders them, limit caps the output. Every event exposes @timestamp, @message, @logStream and @ingestionTime.
Your AWS bill is up sharply month over month and nobody knows why. Walk through how you would use AWS Cost Explorer — granularity, Group by and filters — to narrow the increase down to a specific cause.
basics
~20 sSet AWS Cost Explorer to Daily granularity to find the day spend jumped, Group by Service to name the culprit, then filter to that service and regroup by Usage Type, Region or Linked Account until the charge line is identified.