What does the Datadog Agent do on a host, and how do infrastructure metrics, custom metrics and traces each reach it?
answer
- One process, several different intake paths
- Some telemetry is polled, some is pushed
- Custom metrics arrive on a StatsD listener
- Spans come from the in-process tracing library
- DogStatsD on 8125, trace intake on 8126
basics
~20 sThe Datadog Agent is one host process with several separate intake paths: it collects host metrics itself, runs integration checks that poll local services, accepts custom metrics pushed to its DogStatsD listener, and receives spans from in-process tracing libraries.
solid answer
~50 sThe Datadog Agent is one long-running process per host, or one pod per node in Kubernetes, that gathers telemetry locally and forwards it to Datadog. Four producers feed it, and they work in opposite directions. **Core checks** collect the host's own CPU, memory, disk and network with no configuration. **Integration checks** are scheduled jobs inside the Agent that connect out to a local service - a database, a web server, a queue - and turn its statistics into metrics; the Agent pulls here, configured under `conf.d`. **DogStatsD** is a StatsD-compatible listener embedded in the Agent (UDP port 8125 by default) that applications push counters, gauges, histograms and distributions to, fire and forget. **Traces** are produced by the tracing library inside the application, which posts spans to the Agent's trace intake on port 8126. A fifth path, enabled with `logs_enabled`, has the Agent tail files or container output.
code
yaml · 7 lines# datadog.yaml on the ferry-booking web tier
api_key: "<api key>"
site: datadoghq.eu
logs_enabled: true
dogstatsd_port: 8125
apm_config:
enabled: truego deeper
Be ready to say what the Datadog Agent is - one process per host that gathers telemetry and sends it to Datadog - and to name that host metrics, custom metrics, logs and traces all pass through it.
An interviewer expects the mechanics: which paths pull and which are pushed, that custom metrics arrive on the DogStatsD listener while spans arrive on the trace intake, and that integration checks connect out to a local service on a schedule.
Show that you diagnose from the split. Explain what a missing signal tells you about which path failed, why silent loss is expected on the metric path, and how the Agent's status output narrows a single failing check versus a dead Agent.
Own the boundary between application-owned instrumentation and platform-owned collection: who upgrades what, how the Agent gets deployed and configured across a fleet, and what defaults you set centrally so teams do not each invent their own submission path.
## One process, several intake paths The Datadog Agent is a single long-running process installed on a host, or scheduled as one pod per node in a container platform. Its job is narrow: gather telemetry locally and forward it to Datadog's intake over HTTPS using an API key. What confuses people in interviews is that the Agent is not one collector with one input. It is a supervisor around several listeners and schedulers, each fed by a different producer, and the producers move data in **opposite directions**. Knowing which path carries which signal is what makes an incident like *the dashboards went blank but traces kept arriving* diagnosable instead of mysterious. The paths: 1. **Core checks.** The Agent collects the host's own CPU, memory, disk, network and process data on a schedule with no configuration beyond an API key. This is the baseline infrastructure metric stream. 2. **Integration checks.** For a local service - a relational database, a web server, a message broker - the Agent runs a check that connects *out* to that service, reads its statistics, and converts them into metrics. Configuration lives in a per-integration file under the Agent's `conf.d` directory, and in a container platform Autodiscovery generates that configuration from container metadata instead of you writing files. **The Agent pulls here.** 3. **DogStatsD.** A StatsD-compatible listener embedded in the Agent, bound by default to UDP port 8125 (a Unix socket is also supported). Application code submits counters, gauges, histograms, distributions and sets to it. **The application pushes here**, fire and forget, and the Agent aggregates over a flush interval before submitting upstream. 4. **Trace intake.** The tracing library running *inside* the application creates spans, makes the sampling decision, and posts batches to the Agent's trace intake on port 8126. The Agent buffers them, applies its own ingestion controls, computes aggregate metrics from the spans, and forwards. 5. **Logs.** With log collection enabled in `datadog.yaml`, the Agent tails log files or container output and ships records on a separate pipeline, where server-side processing turns them into structured events. ## What consequently runs where | Signal | Producer | Direction | Where the work happens | |---|---|---|---| | Host metrics | Agent core checks | Agent collects locally | Entirely in the Agent | | Service metrics | Agent integration check | Agent polls the service | Agent, on its check schedule | | Custom metrics | Your application code | App pushes to DogStatsD | App emits, Agent aggregates | | Spans | In-process tracing library | Library posts to the Agent | Library instruments, Agent forwards | | Logs | Files or container output | Agent tails them | Agent ships, server side parses | The practical consequence is that **instrumentation and collection are owned by different artefacts**. The tracing library is a dependency of your application, versioned with your build; the Agent is infrastructure, versioned with the host image. Upgrading one does not change the other, and a team that cannot say which one they need to change has usually mistaken the Agent for something that reads their code. ## Failure modes that follow from the split - **No Agent reachable.** DogStatsD submissions over UDP vanish silently - there is no acknowledgement and no error - while the tracing library usually logs a connection failure and drops spans. Host metrics simply stop. Three different symptoms, one cause. - **Wrong address in a container platform.** The application must be told where the node's Agent lives, commonly through the `DD_AGENT_HOST` environment variable or a mounted socket. Nothing about this is automatic just because the Agent is running somewhere in the cluster. - **One failing integration check.** A check whose credentials for the local service are wrong fails on its own while everything else stays green. The Agent's status command reports per-check errors, which is where you look before assuming a platform-wide outage. - **Custom metrics are code.** They originate in your application, not in the Agent's configuration, so changing them means a deploy - and it is why they are the part of the bill an application team controls directly. - **Ports carry different signals.** Metrics on the DogStatsD listener and spans on the trace intake are separate sockets with separate configuration. A network policy that opens one and not the other produces exactly half a working setup. ## How to answer this in an interview Say the shape first: one Agent process, several intake paths, some pull and some push. Then name what runs where - the Agent polls services and collects the host, the application pushes custom metrics and spans - and finish with the diagnostic payoff, that a signal going missing tells you which path broke before you have looked at anything.
- What changes about those intake paths when the Datadog Agent runs as a DaemonSet in Kubernetes?There is one Agent per node rather than per application. Pods must address the node-local Agent explicitly - by node IP through an environment variable, or through a mounted socket - for DogStatsD and span submission to land anywhere. Integration configuration is generated by Autodiscovery from container metadata instead of files you write, and host-level checks now describe the node, not any single workload on it.
- Custom metrics stopped arriving from one service and nothing anywhere reports an error. Where do you look?Start from the fact that DogStatsD submission is fire and forget over UDP, so loss is silent by design. Check the Agent's status output for DogStatsD packet counts and drops, confirm the application resolves and reaches the Agent's address, and confirm the metric name did not change in a recent deploy. A metric that was renamed looks identical to a metric that vanished.
- Why is the tracing library a separate concern from the Agent version?The library is an application dependency that decides what is instrumented, what attributes spans carry, and whether a trace is sampled; the Agent is infrastructure that receives, buffers and forwards. Upgrading the Agent never adds instrumentation, and upgrading the library never changes host metric collection. Teams that conflate the two spend incidents editing the wrong artefact.
The Agent is less a sensor than a loading dock: some deliveries it drives out and collects, others are dropped at the door by the application, and it forwards everything on one truck.
saying these in an interview costs you the question
- Says the Agent scrapes an endpoint the application exposes for custom metrics
- Believes the Agent generates spans itself rather than the in-process library
- Thinks one Agent port carries metrics, spans and logs alike
- Assumes DogStatsD submission is acknowledged, so losses would be visible
- Cannot separate where instrumentation happens from where forwarding happens
- Expects an application in a container to find the Agent with no configuration