In AWS X-Ray, what is the difference between a segment and a subsegment, and how does X-Ray turn the segments it receives into the service map shown in the console?
answer
- one service's work versus one step
- every piece shares a trace id
- parent id rebuilds the tree
- map is derived, not stored
- uninstrumented hop simply disappears
basics
~20 sA segment is what one service records for one request; subsegments break that segment into finer steps such as downstream calls. X-Ray links all segments sharing a trace id and aggregates them into the service map graph.
solid answer
~40 sA **segment** is the record one instrumented service (or an AWS service acting on your behalf) emits for its part of a request: a name, a trace id, start and end times, HTTP details, and error/fault/throttle flags. **Subsegments** sit inside a segment and describe finer-grained work — an SDK call to DynamoDB, an outbound HTTP request, a slow block of your own code. Every segment carries the same `trace_id`, plus its own `id` and the `parent_id` of whatever caused it, so X-Ray can reassemble the whole request as a tree. The **service map** is not stored separately: X-Ray aggregates the segments from a time window into nodes (one per service or downstream resource) and edges (caller to callee), each annotated with request rate, average latency and error/fault percentages.
code
json · 21 lines{
"name": "checkout-api",
"id": "70de5b6f19ff9a0a",
"trace_id": "1-581cf771-a006649127e371903a2de979",
"start_time": 1478293361.271,
"end_time": 1478293361.449,
"http": {
"request": { "method": "POST", "url": "https://api.example.com/orders" },
"response": { "status": 200 }
},
"subsegments": [
{
"name": "Orders",
"id": "c5b0d1a1f9e2b3c4",
"start_time": 1478293361.280,
"end_time": 1478293361.402,
"namespace": "aws",
"aws": { "operation": "PutItem", "table_name": "Orders" }
}
]
}go deeper
Be able to say plainly that a segment is one service's record of a request and subsegments are the steps inside it, and that a shared trace id is what ties them together.
Explain the assembly mechanics: trace_id, id and parent_id, how subsegments get namespace aws or remote, and how the service map is recomputed from those relationships rather than stored.
Show that you read the map critically — that coverage gaps collapse two hops into one, that node statistics are aggregates over sampled traces, and that you go to individual traces before drawing conclusions.
Own the naming and coverage conventions across teams: consistent segment names so the map has stable nodes, and a policy on which services must be instrumented so the graph is trustworthy as an architecture document.
## The unit X-Ray actually stores X-Ray does not store "traces" as single objects that somebody writes. It stores **segments**, uploaded independently by each participant in a request, and stitches them together afterwards using identifiers they all carry. A **segment** is one service's account of its share of a request. It is a JSON document — the *segment document* — with a small set of required fields and a large set of optional ones: ```json { "name": "checkout-api", "id": "70de5b6f19ff9a0a", "trace_id": "1-581cf771-a006649127e371903a2de979", "start_time": 1478293361.271, "end_time": 1478293361.449, "http": { "request": { "method": "POST", "url": "https://api.example.com/orders" }, "response": { "status": 200 } }, "subsegments": [ ... ] } ``` `name` is what appears as the node label on the service map, so it should be the *service's* name, not the endpoint's. ## Subsegments A **subsegment** has the same shape but describes work *inside* the segment. Three kinds show up in practice: - **Remote subsegments** — an outbound call. The X-Ray SDKs patch common clients (an AWS SDK client, an HTTP client, a SQL driver) so each call opens a subsegment automatically. AWS SDK calls get `"namespace": "aws"` and carry the operation name; plain HTTP calls get `"namespace": "remote"`. - **Custom subsegments** — blocks you wrap by hand to time something the SDK cannot see, such as an expensive in-process transform. - **Subsegments created by AWS on your behalf** — for example the `Initialization`, `Invocation` and `Overhead` subsegments inside a Lambda function segment. The practical rule: a segment is *a service's* work; a subsegment is *a step within* that work. Subsegments are how you answer "where inside this 900 ms did the time go" without adding another service to the map. ## How the pieces find each other Three fields do the assembly: - `trace_id` — identical for every segment of one request. X-Ray's trace ids have a fixed shape: version, then the request start time as hex epoch seconds, then random hex, e.g. `1-581cf771-a006649127e371903a2de979`. The embedded timestamp is why X-Ray can bucket traces by time without reading the segments first. - `id` — a 64-bit hex identifier unique to this segment or subsegment. - `parent_id` — the `id` of whatever caused this work. A downstream service sets it to the caller's subsegment id, which it learned from the `X-Amzn-Trace-Id` header. Because each participant uploads independently, a trace is *eventually* complete: a segment can arrive late, and a segment can arrive `"in_progress": true` and be completed by a later upload. A trace whose caller never uploaded still renders — you just see an orphaned subtree. ## From segments to the service map The service map is a **derived view**, recomputed over the traces in the selected time range. X-Ray walks the parent/child relationships and produces: - **Nodes** — one per distinct service name, plus nodes for downstream resources inferred from `aws` and `remote` subsegments (a DynamoDB table, an S3 bucket, an external API). Clients that only send requests appear as a client node. - **Edges** — caller to callee, drawn from the parent relationships. - **Statistics per node and edge** — request rate, average latency, a latency distribution, and the share of responses that were errors (4xx), faults (5xx) or throttles (429). Those percentages are what colour the ring around each node. So the map answers "which hop is slow or failing", and clicking through to individual traces answers "why". You cannot make the map show something the segments never carried: if a service is not instrumented and does not appear in anyone's subsegments, it is simply absent, and its latency is silently attributed to whoever called it. ## What this means when you read one Two consequences trip people up. First, **the map is only as complete as your instrumentation** — an uninstrumented middle service makes two hops look like one slow hop. Second, **map latency is not user latency**: node statistics are aggregates over sampled traces in the window, so a rare slow path can be invisible on the map while being very visible to one customer. Open the traces, not just the graph.
- If a service in the middle of the call chain is not instrumented, what does the service map show?The map simply skips it. The caller's subsegment for the outbound HTTP call still exists, so you get an edge from the caller to whatever the uninstrumented service called — or to a generic remote node — and the missing service's own latency is folded into the caller's downstream time. The graph looks plausible and is quietly wrong, which is why coverage gaps matter more than instrumentation depth.
- Why would a trace show a subsegment marked in_progress that never completes?Segments can be uploaded in two parts: an in-progress record when work starts and a complete record when it ends. If the process crashed, was killed, or timed out before the second upload, the in-progress record is all X-Ray ever receives. It is a useful signal in itself — it usually points at a hard timeout or an out-of-memory kill rather than a slow dependency.
- What is the difference between an error, a fault and a throttle flag on a segment?They map to response classes. `error` is a client-side problem — a 4xx. `fault` is a server-side failure — a 5xx. `throttle` is the specific 429 case, broken out separately because rate limiting is a different remediation from a bug. The service map colours nodes by these ratios, so mislabelling them by hand makes the map lie about where the failure is.
A segment is one courier's log for the leg they carried the parcel; subsegments are the stops on that leg. Nobody writes the end-to-end journey — it is reconstructed from the tracking number every courier wrote down.
saying these in an interview costs you the question
- Thinking X-Ray stores a whole trace as one document
- Calling every downstream call a segment rather than a subsegment
- Believing the service map is configured rather than derived
- Assuming an absent service on the map means it was not called
- Reading map averages as end-user latency