An ECS service on Fargate is instrumented with the AWS X-Ray SDK, but no traces appear in the X-Ray console. Walk through how segment data actually reaches X-Ray from a container, and where you would look for the break.
answer
- the app never calls the API itself
- fire-and-forget on a local port
- two hops, two ways to fail
- task role, not execution role
- the sidecar's own logs say why
basics
~20 sThe SDK does not call X-Ray directly: it sends segments over UDP to a local collector — the X-Ray daemon or an ADOT collector sidecar on port 2000 — which batches them and calls PutTraceSegments. Check that the sidecar exists, that the app points at it, and that the task role grants X-Ray write.
solid answer
~50 sSegments travel in two hops. In-process, the SDK serialises each segment and sends it as a **UDP datagram to a local listener on port 2000** — the X-Ray daemon, or an ADOT collector configured with the X-Ray receiver. That listener batches segments and calls `PutTraceSegments` (and `PutTelemetryRecords`) against the X-Ray API. Each hop is a separate failure: no sidecar in the task definition, or the app pointing at the wrong address; a task role without `xray:PutTraceSegments` (the `AWSXRayDaemonWriteAccess` managed policy covers it); no network path to the X-Ray endpoint from a private subnet with no NAT and no VPC endpoint; or sampling saying no, so nothing was ever produced. UDP is the reason the symptom is silence — the application never learns that its segments went nowhere, which is deliberate, since telemetry must not be able to fail a request.
code
json · 23 lines{
"family": "checkout",
"networkMode": "awsvpc",
"taskRoleArn": "arn:aws:iam::111122223333:role/checkout-task-role",
"containerDefinitions": [
{
"name": "app",
"image": "111122223333.dkr.ecr.eu-west-1.amazonaws.com/checkout:1.4.2",
"essential": true,
"environment": [
{ "name": "AWS_XRAY_DAEMON_ADDRESS", "value": "127.0.0.1:2000" }
]
},
{
"name": "xray-daemon",
"image": "public.ecr.aws/xray/aws-xray-daemon:latest",
"essential": false,
"portMappings": [
{ "containerPort": 2000, "protocol": "udp" }
]
}
]
}go deeper
Know that instrumented applications send segments to a local agent on UDP port 2000, and that the agent is what actually uploads them to AWS.
Explain both hops and their distinct failure modes — a missing or unreachable sidecar on the first, and permissions or network egress on the second — and why UDP makes the first one silent.
Drive the investigation in cost order: read the sidecar's logs first, then network mode, then the task role rather than the execution role, then subnet egress, and rule out sampling before rewriting instrumentation.
Decide the fleet-wide collection standard — daemon versus ADOT collector, sidecar versus shared agent — and make the task role permissions and egress path part of a paved-road template so individual teams do not rediscover this failure one service at a time.
## The delivery path, hop by hop A very common misconception is that an instrumented application calls the X-Ray API. It does not. There are two hops: **Hop 1 — process to local collector, over UDP.** The SDK builds a segment document and sends it as a UDP datagram to `127.0.0.1:2000` by default. UDP is chosen deliberately: it is fire-and-forget, it cannot block the request path, and it cannot fail the request. The cost of that choice is that the application gets **no feedback whatsoever** if nothing is listening — which is why the symptom of a broken trace pipeline is silence rather than errors. **Hop 2 — collector to the X-Ray API, over HTTPS.** The listener buffers segments and calls `PutTraceSegments` in batches, plus `PutTelemetryRecords` for its own health counters. This hop is authenticated with the task's credentials and requires a network route. The listener is either the **X-Ray daemon** (a small AWS-provided agent, run as a sidecar container in the ECS task or as a process on an EC2 host) or an **ADOT collector** — AWS's distribution of the OpenTelemetry Collector — configured to receive X-Ray segments and export them to X-Ray. On Lambda with active tracing the equivalent is provided by the platform, which is why Lambda needs no sidecar. ## Where it breaks, in the order worth checking **1. There is no sidecar.** The single most common cause on ECS. The task definition has one container. Nothing is listening on port 2000, the datagrams are discarded by the kernel, and there is no log line anywhere. Fix: add the daemon or ADOT collector as a second container in the same task definition. **2. The app is pointing at the wrong address.** In `awsvpc` network mode all containers in the task share a network namespace, so `127.0.0.1:2000` works. In `bridge` mode it does not — the sidecar is a different network namespace, and the app has to reach it by the link or the host address. Getting this wrong produces exactly the same silence. **3. The task role cannot write to X-Ray.** Whoever calls `PutTraceSegments` — the daemon, or the collector — uses the **task role**, not the execution role. The execution role is what ECS itself uses to pull images and write container logs; the task role is what code inside the task uses. Attaching X-Ray permissions to the wrong one of those two is a classic and produces access-denied errors *in the sidecar's own logs* — which is why step zero of any investigation is reading them. **4. There is no network path.** In a private subnet with no NAT gateway and no VPC endpoint for X-Ray, the HTTPS call to the X-Ray endpoint simply times out. The sidecar logs will show it. Note that this is a general egress problem: if X-Ray cannot be reached, other AWS API calls from the task probably cannot either. **5. Sampling said no.** Before blaming plumbing, confirm that anything was supposed to be produced. If a rule matched with a zero rate, or the incoming trace header said `Sampled=0`, there is nothing wrong — there was nothing to send. Check the effective sampling rules and generate load against a path you know is sampled. **6. The instrumentation never started a segment.** Work happening outside a request context — a startup task, a background poller, a thread the SDK does not know about — has no segment to attach to. Some SDKs raise a "no segment found" style condition here; whether that is loud or silent depends on the SDK's configured context-missing behaviour. ## The diagnostic order that actually works Go outside-in, because the cheap checks eliminate the most causes: 1. **Read the sidecar's logs.** They tell you immediately whether it received segments and whether uploads succeeded, and they distinguish causes 1–4 from each other in one look. 2. **Confirm the sidecar is running and reachable** — is the container present and healthy, is the network mode what you think it is. 3. **Check the task role's policy**, not the execution role's, for `xray:PutTraceSegments` and `xray:PutTelemetryRecords`. 4. **Check egress** — NAT route or VPC endpoint from the subnet the task runs in. 5. **Check sampling** last, by looking at whether *any* service is producing traces or only this one. If other services in the same cluster are producing traces normally, causes 3 and 4 become much less likely and the problem is almost certainly in this task definition. ## Choosing between the daemon and ADOT The X-Ray daemon speaks only X-Ray and does one job well. An ADOT collector can receive the same X-Ray segments and also OTLP telemetry, and can export to more than one destination. Teams standardising on OpenTelemetry instrumentation across clouds usually run the collector; teams whose only backend is X-Ray keep the daemon for its simplicity. Either way the UDP-then-API shape of the delivery path is the same, and so is the failure list above.
- Why does the application see no error at all when the collector is missing?Because the first hop is UDP. The SDK sends a datagram to the local port and never waits for acknowledgement, so a missing listener is indistinguishable from a successful send. That is a deliberate design choice — telemetry must never be able to add latency to or fail a customer request — and the price is that the failure is silent. The evidence lives in the sidecar's logs, or in the sidecar's absence.
- Which IAM role needs the X-Ray permissions on an ECS task, and why is that confusing?The **task role** — the identity code inside the task assumes. The execution role is what ECS itself uses to pull the image and ship container logs, and attaching X-Ray permissions there does nothing for the daemon. Because both are configured in the same task definition and both are called "the role" in conversation, mixing them up is one of the most common ECS misconfigurations, not just for X-Ray.
- When would you run an ADOT collector instead of the X-Ray daemon?When the task also emits OpenTelemetry telemetry, or when you want the same data to reach more than one backend. The collector receives X-Ray segments and OTLP, and exports to X-Ray among other destinations, so it becomes a single sidecar for everything. The daemon remains the simpler choice when X-Ray is the only backend — fewer moving parts and nothing to configure beyond running it.
saying these in an interview costs you the question
- Thinking the X-Ray SDK calls the AWS API directly
- Attaching X-Ray permissions to the ECS execution role
- Expecting an application error when the daemon is missing
- Assuming 127.0.0.1 works in bridge network mode
- Blaming instrumentation before checking sampling produced anything