What data does Amazon GuardDuty analyse to produce its findings, and why does enabling it not require you to first turn on VPC Flow Logs or CloudTrail S3 data events yourself?
answer
- no pipeline to build first
- service plane, not your log destinations
- CloudTrail, flow logs, DNS
- optional plans cost extra
- one source really does use an agent
basics
~20 sGuardDuty reads AWS-side telemetry directly from the service plane: CloudTrail management and S3 data events, VPC Flow Logs, and Route 53 Resolver DNS query logs. It gets its own independent copy, so you neither enable nor pay for those logs.
solid answer
~50 sGuardDuty's foundational sources are CloudTrail management events, VPC Flow Logs and DNS query logs resolved inside your VPCs, with S3 data events added by S3 Protection. The key point is that it consumes an **independent internal stream** from the service plane, not the log destinations you configured — so it works in a brand-new account with no trail, no flow-log subscription and no Resolver query logging, and you are not billed by CloudTrail or VPC for the volume it reads (GuardDuty prices that analysis itself). Turning your own flow logs off does not blind it. On top of the foundational sources you can enable optional protection plans — EKS audit logs, RDS login activity, Lambda network activity, malware scanning of EBS volumes, and runtime monitoring agents — each priced separately. Findings carry a structured type such as `UnauthorizedAccess:EC2/SSHBruteForce` plus a severity, and are published to EventBridge.
code
json · 8 lines{
"source": ["aws.guardduty"],
"detail-type": ["GuardDuty Finding"],
"detail": {
"severity": [ { "numeric": [">=", 7] } ],
"type": [ { "prefix": "UnauthorizedAccess:" } ]
}
}go deeper
Recall that GuardDuty is the threat-detection service and that switching it on requires no log setup of your own — it reads AWS telemetry directly.
Name the foundational sources (CloudTrail management events, VPC Flow Logs, DNS queries, plus S3 data events with S3 Protection) and explain why the independent internal copy means no extra logging bill.
Discuss the visibility gaps — external DNS resolvers, non-VPC traffic — and how you route findings through EventBridge or Security Hub into a real response, including suppression of known-benign noise.
Be ready to justify which optional protection plans are worth their per-GB or per-agent cost across a fleet, and how GuardDuty coverage is guaranteed in every Region and every new account rather than left to teams.
## The design idea Most security products need you to build a pipeline first: turn on logging, ship it somewhere, pay for storage, then analyse it. GuardDuty deliberately does not. Because AWS operates the control plane, the VPC network fabric and the Route 53 Resolver, it can hand GuardDuty a private copy of that telemetry directly. The consequence — and the thing interviewers actually probe — is that **enabling GuardDuty is a single switch with no data-source prerequisites**. ## Foundational data sources - **CloudTrail management events** — every control-plane API call: role assumptions, key creation, security-group changes, console logins. This is where credential-misuse findings come from. - **VPC Flow Logs** — network flow metadata for ENIs in your VPCs, used for communication with known-malicious IPs, port-scan detection and traffic-volume anomalies. - **DNS query logs** — names resolved through the AWS-provided Route 53 Resolver inside your VPCs. This catches command-and-control domains and crypto-mining pools. Note the limitation: if an instance is configured to use an external resolver such as 8.8.8.8, those queries bypass the AWS Resolver and GuardDuty does not see them. - **S3 data events**, when S3 Protection is enabled — object-level `GetObject`/`PutObject` activity, used for findings about anomalous data access or exfiltration. Because the copy is internal, three things follow. You do not need a trail, a flow-log subscription or Resolver query logging configured. You are not charged by those services for what GuardDuty reads — GuardDuty charges per GB analysed instead. And an attacker who deletes your trail or disables your flow logs to hide does not disable GuardDuty (disabling the detector itself is a separate, auditable API call). ## Optional protection plans Beyond the foundational sources, GuardDuty has add-on plans, each independently enabled and independently billed: - **EKS Protection** — analyses the EKS control-plane audit log. - **RDS Protection** — analyses login activity against supported Aurora and RDS engines to spot credential brute force. - **Lambda Protection** — analyses network activity from Lambda execution environments. - **Malware Protection for EC2** — on a suspicious finding, snapshots attached EBS volumes and scans them for malware without touching the instance. - **Runtime Monitoring** — deploys an agent to EC2, ECS on Fargate or EKS to observe process, file and network behaviour inside the workload. This is the one source that *is* agent-based. ## What a finding looks like A finding has a hierarchical type string of the shape `ThreatPurpose:ResourceType/ThreatFamilyName`, for example `Recon:EC2/PortProbeUnprotectedPort` or `CryptoCurrency:EC2/BitcoinTool.B!DNS`, plus a severity, the affected resource, and the actor's IP and geography. Findings are re-emitted for continuing activity rather than duplicated endlessly. ## Getting findings out Every finding is delivered to EventBridge with a `detail-type` of `GuardDuty Finding`, so routing is a rule away: ```json { "source": ["aws.guardduty"], "detail-type": ["GuardDuty Finding"], "detail": { "severity": [ { "numeric": [">=", 7] } ] } } ``` GuardDuty also pushes findings into Security Hub automatically when both are enabled, which is usually the better fan-in point once you run more than one detection service. ## Scope and rollout GuardDuty is per-account and per-Region. In an Organization you designate a delegated administrator, then auto-enable GuardDuty for existing and new member accounts so nobody has to remember. Member findings surface in the administrator account. ## The traps to name Being dependent on the AWS Resolver for DNS visibility; foundational analysis being priced by volume, so a chatty account costs more; and the fact that Runtime Monitoring is the only source needing an agent — candidates who claim GuardDuty needs agents everywhere have not run it.
- If someone deletes the account's CloudTrail trail to cover their tracks, does GuardDuty stop producing findings?No. GuardDuty consumes an independent internal stream of CloudTrail events, so deleting your trail removes your own audit copy but not GuardDuty's visibility. It would in fact likely raise a finding about the suspicious API activity. Disabling detection requires disabling the GuardDuty detector itself, which is a separate API call and is itself worth alarming on.
- Why might GuardDuty miss command-and-control DNS traffic from an EC2 instance?Its DNS source is queries resolved through the AWS-provided Route 53 Resolver inside the VPC. If the instance is configured with an external resolver, or tunnels DNS over HTTPS to a third party, those lookups never pass through the Resolver and GuardDuty cannot inspect them. Flow-log-based detections may still fire on the destination IP.
- How do you stop a known-benign activity from raising the same finding repeatedly?Use suppression rules in GuardDuty, which auto-archive findings matching criteria such as finding type plus a specific instance or IP, so they never reach your response path. Trusted IP lists work for known-good addresses. Both are preferable to disabling a whole finding type, and Security Hub automation rules can do the equivalent once findings are aggregated.
saying these in an interview costs you the question
- Claims you must enable VPC Flow Logs before GuardDuty works
- Says GuardDuty needs an agent on every instance
- Believes disabling CloudTrail blinds GuardDuty
- Thinks GuardDuty scans EBS volumes or files by default
- Assumes all protection plans are included in the base price